Choose an MDL data strategy

Choose how data enters Machine Data Lake, how existing Splunk index paths are preserved where needed, and how landed data is searched, promoted, shared, or queried.

Use this topic to decide whether Machine Data Lake (MDL) fits a data source and to choose the strategy for routing, retaining, searching, promoting, sharing, or querying the data. Complete this strategy before you prepare the source for onboarding.

Evaluate whether Machine Data Lake fits

Choose Machine Data Lake when you want to separate data landing from promotion and external consumption. Land the broad base of supported machine data once, inspect it in the Catalog, and promote only the data that needs a Splunk index or analytics table, and use Open Sharing when a supported external consumer needs read-only access.

Machine Data Lake works well when one or more of these conditions apply:

  • You need to retain high-volume logs and events, including cloud logs and network event data, without promoting every event immediately.

  • You do not yet know which part of the raw data becomes operationally important, but you want the option to inspect and promote it later.

  • You need a cost-aware path for historical investigation, selective promotion, or delayed promotion.

  • You need to support operational Splunk search and analytics-ready workflows from the same raw base.

  • You need to evaluate event ranges, field coverage, filters, and promotion volume before creating a promoted dataset.

Another path can be simpler when every event must be available immediately for continuous dashboards, alerts, and high-performance search, when the full dataset is already active and frequently searched, or when policy, residency, ownership, or governance requirements require the data to remain in another supported system.

Decision 1: Choose an ingest and routing strategy

Start with how matching events enter Machine Data Lake and whether any existing Splunk index path must continue to receive the same data.

Land only in Machine Data Lake
  • Use it when: You need to retain and evaluate broad supported machine data, and current workflows do not require the events in an existing Splunk index.

  • Search behavior: Data is available for supported raw search after it lands. Raw tables are storage-optimized, so raw search is best for validation and narrowing the next action.

  • Data and retention: Matching events land in the raw table without creating a separate Splunk index copy. Raw table retention controls how long the landed data remains available.

  • Permissions and trade-offs: Requires dataset access and supported landing permissions. This path can reduce unnecessary indexing, but repeated raw searches can still use compute resources.

Land in Machine Data Lake and preserve an existing Splunk index path (dual routing)
  • Use it when: Current searches, dashboards, alerts, Enterprise Security content, or rollback plans must continue using an existing Splunk index while matching events also land in Machine Data Lake, where supported.

  • Search behavior: The existing index keeps its current search behavior. Machine Data Lake provides raw search, promotion, and Open Sharing paths after landing for the landed copy.

  • Data and retention: Matching events can be written to more than one destination. Retention and lifecycle settings are managed separately for the raw table and the existing index.

  • Permissions and trade-offs: Requires the permissions and capacity needed to manage the existing route and the Machine Data Lake landing workflow. This path reduces migration risk but can increase storage and processing usage.

Use a direct Splunk index path
  • Use it when: Every event must be available immediately for continuous dashboards, alerts, scheduled searches, or high-performance operational search, and Machine Data Lake doesn't add meaningful value for the source.

  • Search behavior: The data uses the standard Splunk index search path and index-time behavior.

  • Data and retention: Events are indexed directly instead of landing in a Machine Data Lake raw table. Index retention controls availability.

  • Permissions and trade-offs: Uses existing Splunk index permissions, capacity, and cost model. You don't get the Machine Data Lake raw landing, discovery, and selective promotion and external-consumption workflow.

Keep data in another supported system
  • Use it when: Policy, residency, ownership, governance, or operational requirements require the data to remain outside your Machine Data Lake-enabled Splunk Cloud Platform environment.

  • Search behavior: Query behavior, latency, field behavior, and performance depend on the external system and supported federated search path.

  • Data and retention: The data remains in its existing system. Retention and lifecycle controls stay with that system.

  • Permissions and trade-offs: Requires access to the external source and any supported federation configuration. Use this path when in-place query is the goal.

Decision 2: Choose a post-landing consumption strategy

After data lands, choose the simplest consumption path that supports your use case. Don't promote or share broad raw data unless the workflow requires it.

Raw search
  • Use it when: You need to validate that events landed, inspect raw data, narrow a time range, or decide whether promotion is needed.

  • Search or query behavior: Raw search is available on supported raw tables, but it is slower and has more limited field behavior than promoted data.

  • Data movement and lifecycle: No additional data copy is created. Searches read the raw table within its retention period.

  • Permissions and trade-offs: Requires query access. Use raw search for targeted validation or exploration, not repeated high-performance operational search.

Static promotion to a Splunk index
  • Use it when: You need a bounded historical slice in a Splunk search-optimized target for faster repeated search, dashboards, scheduled searches, or comparison with other Splunk data.

  • Search or query behavior: After the promotion completes, searches use the supported Splunk index search path for the promoted dataset.

  • Data movement and lifecycle: The promoted data covers a fixed historical time range. Promotion creates a new promoted dataset and doesn't remove the source data from the raw table.

  • Permissions and trade-offs: Requires static promotion access and related service-account capabilities. Use the smallest time range and filters that satisfy the outcome.

Streaming promotion to a Splunk index
  • Use it when: Matching events must continue arriving in a Splunk search-optimized target for monitoring, alerting, dashboards, or high-performance search.

  • Search or query behavior: After setup, new matching events continue flowing into the promoted Splunk index path.

  • Data movement and lifecycle: This is continuous for matching events. Retention for the promoted destination is controlled by the configured rolling window and doesn't change the source raw table retention.

  • Permissions and trade-offs: Requires streaming promotion access and shared Ingest Processor capacity. Keep filters focused to control ongoing volume and usage.

Static promotion to an analytics table
  • Use it when: You need selected fields in a structured format for wide scans, business intelligence, notebooks, reporting, compliance review, or machine learning workflows.

  • Search or query behavior: After the promotion completes, the table supports structured analysis over selected fields. It is not intended for continuous monitoring or Splunk event search.

  • Data movement and lifecycle: The promoted table covers a fixed historical time range and selected fields. Promotion creates a new promoted dataset.

  • Permissions and trade-offs: Requires static promotion access. Select only the fields and time range that support the analysis.

Open Sharing
  • Use it when: A supported external consumer, such as a business intelligence tool, notebook, or machine learning platform, needs authorized read-only access to selected data.

  • Search or query behavior: External query behavior depends on the supported sharing mechanism and the consuming platform.

  • Data movement and lifecycle: Where supported, sharing provides read-only external access without a manual export. Access is limited by the shared dataset and configured expiration or revocation controls.

  • Permissions and trade-offs: Requires sharing or Open Sharing capability where supported. Share only the dataset and access duration that match the authorized use case.

Federated search
  • Use it when: The data must remain in another supported system and you need to search it in place from Splunk.

  • Search or query behavior: Query capabilities, latency, metadata, and performance depend on the federated source and supported search path.

  • Data movement and lifecycle: The data doesn't land in a Machine Data Lake raw table. Retention and lifecycle controls remain with the external system.

  • Permissions and trade-offs: Requires access to the federated source and supported federated search configuration. Use this path when data movement is not allowed or not useful.

Confirm custody and governance requirements

Before onboarding, confirm where the data must reside, who can view, query, promote, share, or manage the dataset, how long raw and promoted data must be retained, and which audit or governance controls apply. Machine Data Lake access depends on role-based access control, dataset access, sharing capabilities, and any configured attribute-based access control policies.

For more information, see Access control, roles, and capabilities, Raw tables and Retention and deletion lifecycle, and Machine Data Lake limits and caveats.

Complete your data strategy

Before you prepare the source for onboarding, confirm the data source, expected volume, retention expectation, search frequency, latency requirement, query shape, governance constraints, validation plan, rollback expectation, ingest and routing strategy, and post-landing consumption path.

If the strategy depends on preserving an existing Splunk index path while data also lands in Machine Data Lake, confirm support for that routing model with your administrator or implementation team before you change production data flows.

After you choose the data strategy, see Prepare for onboarding.