Review the crawler-discovered schema and partitions for an Amazon S3 dataset

Review, edit, and confirm the crawler-discovered schema and partitions for an Amazon S3 dataset that is backed by a Splunk-native data catalog, to ready the dataset for federated search.

When your Amazon S3 dataset is backed by a Splunk-native data catalog and you select Discover schema via crawler or Discover partitions via crawler during the Configure dataset step of dataset definition, selection of Create dataset on the Review step launches a crawler process that scans your dataset.
  • If you have selected Discover schema via crawler for your dataset, the crawler process discovers the dataset schema.

  • If you have selected Discover partitions via crawler for your dataset, the crawler process identifies the fields by which your dataset is partitioned.

The crawler process might take a few minutes to complete. If it completes without errors, your dataset will have a Status of Needs action.

What does Needs action mean? It means you need to review the schema and partition fields that were discovered by the crawler process, edit them as necessary, and confirm that you have inspected them, so you can use the new Amazon S3 dataset in federated searches.

  • A role on your Splunk Cloud Platform deployment with the edit_connections and edit_datasets capabilities. See Define roles on the Splunk platform with capabilities in the Splunk Cloud Platform Manage Users and Security manual.
  • You must have completed the definition of your Amazon S3 dataset and selected Create dataset on the Review page.

  • The Amazon S3 dataset must be backed by a Splunk-native data catalog, and you must have selected either Discover schema via crawler or Discover partitions via crawler (or both) on the Configure dataset step of the dataset definition.

  • On the Datasets list page, your Amazon S3 dataset must have a Status of Needs action.

  1. In the Data Management app, on the Datasets list page, select an Amazon S3 dataset with a Status of Needs action.

    You can use the filter on the Status column to quickly find Amazon S3 datasets that have the Needs action status:An icon that looks like a funnel. When the filter is active, the icon changes from an outline of a funnel to a fully shaded-in funnel.. Select the filter icon and choose the status value you want to filter on.

  2. In the right-hand sidebar, select Edit.
  3. If you selected Discover schema via crawler while defining this Amazon S3 dataset, go to the Schema section to review the schema fields that the crawler process discovered for your dataset and update the schema if necessary.
    Note: The following schema update options assume that you are using the Field list view for the schema. You can also edit, delete and add schema fields through the JSON view. Your input must match the JSON data schema (dataSchema). See JSON standards for the data and partition schemas.
    1. (Optional) Edit schema errors. You can change the Name and Data type of any given schema field, and you change the schema order so that it fits the actual schema in your dataset.
    2. (Optional) Delete fields that do not belong in the schema.
    3. (Optional) Replace fields that are missing from the schema by selecting Add field.
    4. Select I confirm that I have reviewed the schema.
  4. If you selected Discover partitions via crawler while defining this Amazon S3 dataset, scroll down to the Partitions section to review the partition fields that the crawler process discovered for your dataset and update them as necessary.
    1. If a partition field has an incorrect Data type, change it.
    2. Identify partition fields that are time fields, such as month or day. Select the edit icon, then select This is a time partition field and provide a Time format. For help with time format variables, see Using time variables in the SPL2 Search Manual.
    3. If there are time partition fields, ensure the correct Time zone is selected for them.
    4. Select I confirm that I have reviewed the partitions.
  5. Select Save to save your changes.
You have an Amazon S3 dataset backed by a Splunk-native data catalog with crawler-discovered schema or partition fields that is ready to be used in federated searches.
Now that this dataset is ready to be shared and searched, do these things: