Machine Data Lake architecture and data flow
Understand how supported inputs route logs and events into a Machine Data Lake raw table, how you discover the data in the Catalog, and how promotion, sharing, query, and consumption paths differ.
Architecture at a glance
Logs and events enter through supported input configurations, including HTTP Event Collector (HEC), Universal Forwarder (UF), and Heavy Forwarder (HF). Data reaches a Machine Data Lake raw table either through a supported direct sender destination or through a data landing definition that selects events from an existing flow. Landed data becomes available automatically in the Catalog for discovery. After discovery, choose the appropriate promotion, query, or authorized external-access path. Data that lands in Machine Data Lake is automatically available in the Catalog for discovery. You can then promote or share selected data before using the appropriate query or consumption path.
Choose the routing topology before landing. Use Machine Data Lake-only routing when matching events do not need to continue to an existing Splunk index. Use dual routing that preserves an existing Splunk index path when current searches, dashboards, alerts, Enterprise Security content, or rollback plans must keep using that index while matching events also land in Machine Data Lake, where supported.
How the pieces fit together
Machine Data Lake separates routing and landing, discovery, promotion, query, external access, and lifecycle management. This separation lets you retain broad raw machine data, inspect what landed, and create only the search, analytics, or external-access target required by a specific use case.
- Inputs
- Supported HEC, Universal Forwarder, and Heavy Forwarder configurations send logs and events to the Machine Data Lake-enabled Splunk Cloud Platform environment. A supported sender can use a direct raw table destination, or its existing flow can be selected by a data landing definition.
- Route and land
- A supported direct sender destination or data landing definition determines which events reach the raw table. Supported landing-time processing affects the values written to the raw table. A dual-routing workflow preserves the existing Splunk index destination when required.
- Raw table
- A raw table stores landed data and provides the source for preview, Catalog discovery, raw search, promotion, sharing where supported, and retention controls.
- The Catalog
- The Catalog is the discovery stage. It shows Machine Data Lake raw tables and supported promoted datasets, metadata, fields, lineage, access, and promotion jobs. Landed data becomes available in the Catalog automatically, and discovery does not create another data copy.
- Promotion
- Promotion creates a derived promotion target from selected raw data. Promotion does not remove the source data from the raw table.
- Query and consume
- Use raw search for validation, Splunk index promotion for operational search, static promotion to an analytics table for structured analysis, Open Sharing for authorized external access where supported, or federated access for data that remains elsewhere.
How data moves through Machine Data Lake
-
Route and land data. Supported input configurations, including HEC, UF, and HF, send logs and events to the Machine Data Lake-enabled environment. The landing configuration selects matching events, applies supported landing-time processing, and writes matching events to the raw table. Non-matching events follow the configured input and routing destination.
- Retain raw data. Matching events are written to the Machine Data Lake raw table. The raw source remains available according to its own retention and deletion policy, independently of preserved index routes, promotion targets, or shared access paths.
-
Discover data. Data that lands in Machine Data Lake is automatically available in the Catalog. Use the Catalog to find raw and promoted datasets and inspect metadata, fields, lineage, access, and promotion jobs. Use preview or search for immediate event validation because Catalog metadata can lag behind recent landing.
-
Promote or share selected data. Use promotion-time processing to choose the promotion type, time range, filters, fields, and supported actions. Static promotion to a Splunk index, streaming promotion to a Splunk index, and static promotion to an analytics table create supported targets. Where supported, Open Sharing provides authorized read-only external access. The source raw data remains in the raw table according to raw table retention.
-
Query or consume data. Use raw search over the MDL raw table, Splunk search over a supported existing or promoted index, analytics search over a analytics table created by static promotion, or a Spark or ETL pipeline that accesses selected data through Open Sharing.
Where filters and transformations apply
| Stage | Where it happens | What it affects |
|---|---|---|
| Routing and dual routing | Supported input and landing routing configuration. | Whether matching events land only in Machine Data Lake, also continue to a preserved Splunk index path where supported, or follow another configured destination. |
| Landing-time processing | Raw table landing definition and related Ingest Processor processing. | Which events and values are written to the raw table. |
| Promotion-time processing | Promotion configuration and preview workflow. | Which subset of raw data and fields are written to the promotion target. |
| Open Sharing controls | Dataset sharing settings where supported. | Which authorized external consumers can access selected data and for how long. |
Storage and governance boundaries
Machine Data Lake stores raw data in object storage associated with your enabled Splunk Cloud Platform environment. Storage ownership, region, encryption, and key-management behavior depend on the provisioned configuration. Confirm these requirements before onboarding production data.
Writing data directly to object storage does not land the data in Machine Data Lake. To make data appear in the Catalog and make it available for raw search, promotion, sharing, or governance workflows, land the data through a supported Machine Data Lake workflow.