Discover Machine Data Lake datasets in the Catalog

Use the Catalog to find raw tables, Splunk indexes, analytics tables, and federated datasets. Decide whether to query existing data, promote data from a raw table, or manage a dataset.

Use the Catalog guide for shared workflows

The Catalog is the shared Splunk Cloud Platform discovery space for datasets. For general guidance about opening the Catalog, filtering and sorting datasets, inspecting side-panel details, viewing fields, and starting a search, see Discovering data for your investigations using the Catalog, Catalog dataset details, and Catalog use cases.

Use this topic for the additional behavior you need when a dataset comes from Machine Data Lake.

Machine Data Lake information in the Catalog

Use the Datasets tab to find raw tables, promoted Splunk indexes, analytics tables, local Splunk indexes, and federated datasets. For Machine Data Lake work, identify whether the required data is already available in a queryable Splunk index or analytics table, or is available only in a raw table.

When data lands through a supported Machine Data Lake workflow, the raw table becomes available in the Catalog automatically to users who have permission to discover it. Catalog metadata can include the dataset name, event time span, source metadata, size, and lineage where available. Query, promotion, sharing, and management actions require their own permissions.

Writing files directly to the storage location behind Machine Data Lake does not add the data to Machine Data Lake. To make data appear in the Catalog and make it available for raw search or promotion, use a supported Machine Data Lake landing workflow.

The Catalog does not show every metadata update immediately. Dataset statistics and available field metadata refresh on a background schedule. Use raw search, preview, or event inspection to validate a recent test event; do not treat Catalog metadata refresh as event-level validation.

For recently promoted datasets, especially datasets created from large promotion jobs, default field-value filters such as source, sourcetype, host, and time might not find the promoted dataset immediately after the promotion job completes. Search by dataset name, open the promoted dataset from the promotion job details, or allow more time for metadata in the Catalog to refresh.

Find Machine Data Lake datasets

Use the filtering and sorting controls described in Find relevant datasets. For Machine Data Lake datasets, you can also narrow results with field-value filters for source, sourcetype, host, and time when that metadata is available.

Use the dataset type and dataset details to decide what the dataset represents.

  • Raw table: logs and events retained through a supported Machine Data Lake landing workflow. Raw tables are the source for Machine Data Lake promotion workflows.

    Analytics table: a promoted dataset for structured analysis of selected raw data.

  • Splunk index: a promoted dataset or local Splunk platform index that uses the supported Splunk search path.

  • Federated dataset: a dataset that remains in another supported system and is searched in place through a supported search path. Federated datasets are not Machine Data Lake raw tables.

Inspect Machine Data Lake details

Select a dataset row to open the details side panel. For general side-panel behavior, field inspection, and field-value sampling, see Catalog dataset details and Investigate dataset contents.

For Machine Data Lake raw tables, the side panel can include Search and Promote actions, event range, size, maximum rolling window, available field metadata, lineage where available, and outbound promotion jobs. For promoted datasets, the panel can show the source raw table, recent promotion jobs, and a supported Search action.

Promotion jobs tab

Use the Promotion jobs tab to review promotion jobs that you own or have permission to manage. The list can show the promoted dataset, promotion type, promotion mode, state, creation time, and available job actions.

Select a promotion job row to inspect state details, configuration, source dataset, promoted dataset, and additional job information. A promotion job tracks a promotion request and its state. It can create or update a promoted dataset, but the job is not itself a dataset.

Choose the next action

Search the Catalog before you create a promotion. If an existing Splunk index or analytics table contains the required data in a queryable format, use that dataset. If the data is available only in a raw table, or the existing format does not support the use case, choose a static promotion to a Splunk index or analytics table, or a streaming promotion to a Splunk index.