Configure Microsoft Azure dataset details

Provide information about how Splunk software creates the Splunk-native data catalog that facilitates federated searches of your Microsoft Azure dataset.

To run federated searches over a Microsoft Azure dataset, Splunk software requires the dataset be backed by a data catalog that Splunk software creates for your dataset. In this step, you determine how this data catalog is created and managed.

You decide whether its schema is created manually or inferred automatically with a crawler. You optionally ensure whether the data catalog is automatically kept in sync with your dataset as it changes. You provide time field information if your data contains time-series data and you want to make use of time fields in your searches. And you provide partition field information as necessary to facilitate efficient federated searches.

  1. On the Configure dataset step of the Create dataset workflow, identify whether your data is stored in one of the following non-table formats: Parquet, CSV, or JSON.
  2. If your data is in CSV or JSON format, indicate whether the data is compressed with Gzip or is Uncompressed.
  3. (Optional) If this dataset is updated on an ongoing basis, and you want Splunk software to keep your Splunk-native data catalog in sync with the Microsoft Azure dataset it represents, set up the following things in the Microsoft Azure Portal:
    • An Azure Storage Queue that receives messages from downstream consumers.
    • An Azure Event Grid system topic scoped to your Azure Storage Account, with a subscription that forwards blob lifecycle events (created, deleted, and so on) to the Azure Storage Queue.

    For detailed setup instructions, see Ensure the Microsoft Azure dataset and its data catalog stay in synch with each other.

    When you set up the Azure Storage Queue, you can retrieve its URL and enter it into the Queue URL field.

    Note: Skip this step if your dataset is composed of historical data that is not subject to future updates, or if you are not interested in keeping your data catalog in sync with changes to your dataset.
  4. Indicate how you want to define the schema for your data catalog.
    • Select Define schema manually if you want to manually determine the columns in your dataset. You can use a Field list view or JSON view. If you select JSON view, your input must match the JSON data schema (dataSchema). For more information, see JSON standards for the data and partition schemas.
    • Select Discover schema via crawler if you want to have a crawler scan a sample set of files from your dataset to infer the overall schema. Use Number of files to scan to tell the crawler how many files it should scan.
      Note: The crawler can be applied only to data that is in Parquet, CSV, or JSON format. All files sampled by the crawler must follow the same schema. Inconsistent schemas across sampled files might result in a data catalog with incorrectly inferred fields in its schema.

      The crawler process will begin its scan after you reach the Review step of dataset definition and select Create. It might take a few minutes for the crawler process to complete. The Status value of the dataset tells you what to do next. See the table at the bottom of this topic for more information.

  5. (Optional) Select Define the time field if your dataset contains time-series data and you intend to use time-based filtering or SPL2 time functions when you run federated searches over it.

    If you select Define the time field, provide the Time field, Time format, and Unix time field.

    For more information, see Identify the time field in a Microsoft Azure dataset.

  6. Indicate whether your dataset is partitioned, and if so, whether its partitions follow Hive formatting. Answer Are your partitions Hive-compatible?
    • Yes: Decide whether you want to Define your partitions manually or let Splunk software Discover partitions via crawler.

      If you select Define partitions manually, you can use a Field list view or a JSON view. If you select JSON view, your input must match the partition schema (dataPartition.PartitionSchema). For more information, see JSON standards for the data and partition schemas.

      Note: If you have time partitions and you have selected Define partitions manually. See Identify time partitions in a Microsoft Azure dataset.

      If you select Discover partitions via crawler, a crawler process will scan your dataset when you reach the Review step of dataset definition and select Create dataset. The crawler process might take a few minutes to complete. The Status value of the dataset tells you what to do next. See the table at the bottom of this topic for more information.

    • No: Define partitions manually using a Field list view or a JSON view. If you have time partitions, ensure they are properly identified and defined. See Identify time partitions in a Microsoft Azure dataset.
    • I don't have partitions: Select Next.
  7. Select Next.
  8. On the Review page, review your dataset definition. If the details appear correct, select Create dataset to create your dataset.

Your Microsoft Azure dataset is created or is in the process of being created.

On the Datasets listing page you can see the Status of your Microsoft Azure dataset, and you can use that status value to guide your next actions regarding it.

Status Description Action
Ready The dataset is available for use in federated searches.
Processing The crawler process is running over the dataset.

If you have selected Discover schema via crawler or Discover partitions via crawler during the Configure dataset step of dataset definition, selection of Create dataset on the Review step causes the crawler process to initiate schema and partition field discovery for the dataset.

The crawler process might take a few minutes to complete.

Note: If more than 10 minutes pass and the crawler process is still in Processing status, a dataset setup error might be causing it to fail to complete. Review the current dataset configuration for errors such as an incorrect location path. Then delete the dataset that is stuck in Processing status and try to recreate it without errors.
Needs action The schema and partitions discovered by the crawler require review and confirmation.

Go to the Edit page for your Microsoft Azure dataset. Review the schema and partition fields that the crawler has discovered, make edits as necessary, and confirm that you have reviewed the discovered fields. See Review the crawler-discovered schema and partitions for a Microsoft Azure dataset.

Error An error occurred during dataset creation or processing. Review the configuration for your Microsoft Azure dataset, correct issues, and recreate the dataset if necessary.