Apply the dataset resource access policy to an AWS IAM role
Apply the Splunk-generated resource access policy to the IAM role associated with your connection to authenticate access to your Amazon S3 dataset.
To authenticate access to your Amazon S3 dataset, you must apply a Splunk-generated resource access policy to the IAM role associated with the dataset's connection. This resource access policy is generated by Splunk software. It controls the following:
- The resource access policy grants access to the Amazon S3 bucket that contains the Amazon S3 location you specified for the dataset in the Define dataset step. See Define an Amazon S3 dataset.
- If your Amazon S3 dataset is backed by an AWS Glue catalog table, the resource access policy grants access to the AWS Glue catalog that contains the Glue table. See Create an Amazon S3 dataset for federated search that is backed by an AWS Glue catalog table.
- If your Amazon S3 dataset is backed by a Splunk-native data catalog, and you have set up an SQS queue and event notification for the Amazon S3 bucket that contains your dataset, the resource access policy contains the SQS queue ARN that you supplied earlier in the dataset definition. This ARN enables the Splunk-native catalog to stay in synch with your dataset as its contents change. See Set up automated updates for Splunk-native data catalogs in AWS.
- If you are using server-side encryption (SSE-KMS) to encrypt data in the Amazon S3 bucket that contains your dataset or the AWS Glue catalog table that refers to it, the resource access policy allows Splunk software to access that data.
- Your user account on the Splunk Cloud Platform deployment must have a role with the
edit_connectionsandedit_datasetscapabilities. - Review the connection this dataset is associated with. Obtain the name of the IAM role that it uses for IAM role authentication. See Define an Amazon S3 dataset.
- You must have completed the Configure dataset step of the Create dataset workflow.
Your Amazon S3 dataset is created or is in the process of being created.
On the Datasets listing page you can see the Status of your Amazon S3 dataset, and you can use that status value to guide your next actions regarding it.
| Status | Description | Action |
|---|---|---|
| Ready | The dataset is available for use in federated searches. |
|
| Processing | The crawler process is running over the dataset. |
If your Amazon S3 dataset is backed by a Splunk-native data catalog and you have selected Discover schema via crawler or Discover partitions via crawler during the Configure dataset step of dataset definition, selection of Create dataset on the Review step causes the crawler process to initiate schema and parititon field discovery for the dataset. The crawler process might take a few minutes to complete.
Note: If more than 10 minutes pass and the crawler process is still in Processing status, a dataset setup error might be causing it to fail to complete. Review the current dataset configuration for errors such as an incorrect location path. Then delete the dataset that is stuck in Processing status and try to recreate it without errors.
|
| Needs action | The schema and partitions discovered by the crawler require review and confirmation. |
Go to the Edit page for your Amazon S3 dataset. Review the schema and partition fields that the crawler has discovered, make edits as necessary, and confirm that you have reviewed the discovered fields. See Review the crawler-discovered schema and partitions for an Amazon S3 dataset. |
| Error | An error occurred during dataset creation or processing. | Review the configuration for your Amazon S3 dataset, correct issues, and recreate the dataset if necessary. |