Get started with Machine Data Lake

Use this guide to onboard Machine Data Lake, land data, discover what landed, and promote the data you need for search and analytics.

What you can do with Machine Data Lake

Machine Data Lake lets you retain supported logs and events in economical object storage without deciding at ingest time whether every event must be indexed. Land data in a raw table, discover it in the Catalog, and promote the data required for indexed search or structured analytics when a use case arises. Eligible data can also be accessed through supported Open Sharing clients.

Use Machine Data Lake as a governed data lake for supported machine data. You can search and promote data in Splunk, and where supported, share selected data with external tools such as business intelligence, notebook, and machine learning platforms through Open Sharing. Land data only through supported Machine Data Lake workflows and manage access through the Catalog and role-based permissions. Data written directly to the underlying object storage does not become a Machine Data Lake raw table and is not available for supported search, promotion, Open Sharing, or governance workflows.

For more information, see Access control, roles, and capabilities, Manage Machine Data Lake datasets, and Machine Data Lake limits and caveats.

Promote only the data that needs richer field visibility, better search performance, dashboards, alerting, large-scale analysis, or artificial intelligence and machine learning workflows. Select the relevant time range and filters, and then choose a static promotion to an analytics table, a static promotion to a Splunk index, or a streaming promotion to a Splunk index.

Onboarding journey

Follow this workflow when you use Machine Data Lake for the first time: decide whether Machine Data Lake fits, prepare prerequisites and data sources, land and validate data, discover available data, create a promotion when needed, query or access data through a supported path, and govern and troubleshoot ongoing use.

1. Prepare prerequisites and data sources

Review the Machine Data Lake model, confirm prerequisites, configure the data management service account role, and prepare a supported data source. Choose a direct HTTP Event Collector (HEC), Universal Forwarder, or Heavy Forwarder path, or redirect an existing indexed flow through Edit data landing. Decide whether matching events land only in a raw table or also continue to a Splunk index through a supported dual-routing path.

More information: Machine Data Lake concepts, Prerequisites, Access control, roles, and capabilities, Prepare a data source

2. Land and validate data

Create or select a raw table, define conditions using source, sourcetype, or host, and preview the landing configuration before you save it. Send a uniquely identifiable test event. Verify the event in the raw table, verify any preserved index route separately, and confirm that non-matching events continue to the intended destination. If the results are unexpected, edit the landing configuration to restore the previous routing behavior.

Outcome: The uniquely identifiable test event is available through the supported raw-data validation path, and any configured index route contains its corresponding copy. Catalog metadata alone is not event-level validation.

More information: Onboard your first data source, Create a raw table, Inspect and edit raw tables

3. Discover data

Use the Catalog to find datasets that you have permission to discover. If an existing Splunk index or analytics table contains the data in a queryable format, use it. If the data exists only in a raw table or is not in the required format, create the appropriate promotion.

Outcome: You know which dataset contains the data and whether to query it directly or promote data from a raw table.

More information: Discover MDL datasets in the Catalog, Find MDL datasets in the Catalog, Query and access MDL data, Search MDL data from the Catalog

4. Promote data

Choose a destination based on the use case. Use static promotion to send a historical time range to a Splunk index or analytics table. Use streaming promotion to send matching events that arrive after the promotion job becomes active to a Splunk index. Streaming promotion applies only to new matching events and doesn't include data that is already stored in the raw table. Use a separate static promotion when you need historical data.

Outcome: The selected historical data or future matching events are available in the destination that you chose.

More information: , Promote data to a Splunk index, Promote data to an analytics table, Monitor and manage promotion jobs, Manage MDL datasets, MDL limits and caveats, Troubleshoot Machine Data Lake

How to use this guide

If you are new to Machine Data Lake, start with the overview and prerequisites, then follow the workflow tasks in order. If Machine Data Lake is already provisioned, start at the step that matches your current goal, such as creating a raw table, landing data, browsing the Catalog, or creating a promotion.

Each task topic gives the action to take, context for why the action matters, and the expected result or next step.