Create an Amazon S3 dataset for federated search that is backed by a Splunk-native data catalog

Set up a federated dataset with a Splunk-managed data catalog that Splunk software creates and maintains for you.

Set up an Amazon S3 dataset with a Splunk-native data catalog that Splunk software creates and maintains for you.
  1. On your Splunk Cloud Platform deployment, in the Data Management app, at the Configure dataset step of the Create dataset workflow, select I don't have a catalog.
  2. Identify whether your data is stored in one of the following non-table formats: Parquet, CSV, or JSON. If your data is in JSON or CSV format, indicate whether the data is compressed with Gzip, or is Uncompressed.
  3. (Optional) If this dataset is updated on an ongoing basis, and you want your Splunk-native data catalog to be updated automatically so it is consistent with the dataset it represents, answer Yes to Do you want to keep catalog in sync with dataset?.

    Set up an SQS queue and event notification for the Amazon S3 bucket that contains the dataset. Paste the ARN that you obtain when you set up the SQS queue into the SQS queue ARN field. For detailed instructions, see Set up automated updates for Splunk-native data catalogs in AWS.

    Note: You can skip this step if your dataset is composed of historical data that is not subject to future updates, or if you are not interested in keeping your data catalog in sync with changes to your dataset.
  4. Indicate how you want to define the schema for your data catalog.
    • Select Define schema manually if you want to manually determine the columns in your dataset. You can use a Field list view or JSON view. If you select JSON view, your input must match the JSON data schema (dataSchema). For more information, see JSON standards for the data and partition schemas.
    • Select Discover schema via crawler if you want a crawler process to scan a sample set of files from your dataset to discover the overall dataset schema. Use Number of files to scan to tell the crawler process how many files it should scan.
      Note: The crawler can be applied only to data that is in Parquet, CSV, or JSON format. All files sampled by the crawler must share the same schema. Inconsistent schemas across sampled files might result in a data catalog with incorrectly inferred fields in its schema.

      The crawler process will begin its scan after you reach the Review step of dataset definition and select Create dataset. It might take a few minutes for the crawler process to complete. After the crawler process completes with a Status of Needs action, review, edit, and confirm the crawler-discovered schema on the Edit page for the dataset. See Review the crawler-discovered schema and partitions for an Amazon S3 dataset.

    • Choose Select from pre-defined schema if your dataset follows a known schema such as the AWS CloudTrail log schema or the VPC flow log schema. When you apply a predefined schema to your dataset, that schema does not need to be manually defined or discovered by a crawler process.

      The data format you have selected for your dataset determines which known schema options you can select.

      • The AWS CloudTrail log predefined schema designation is available only when the dataset data format is JSON.

      • The VPC flow log predefined schema designation is available only when the dataset data format is Parquet or CSV.

  5. (Optional) Select Define the time field if your dataset contains time-series data and you intend to use time-based filtering or SPL2 time functions when you run federated searches over it.

    If you select Define the time field, fill out the Time settings: Time field, Time format, and Unix time field.

    For more information, see Identify the time field in an Amazon S3 dataset.

  6. Indicate whether your dataset is partitioned, and if so, whether its partitions follow Hive formatting. Answer Are your partitions Hive-compatible?
    • Yes: Decide whether you want to Define partitions manually or let Splunk software Discover partitions via crawler.

      If you select Define partitions manually, you can use a Field list view or a JSON view. If you select JSON view, your input must match the partition schema (dataPartition.PartitionSchema). For more information, see JSON standards for the data and partition schemas.

      Note: If you have time partitions and you have selected Define partitions manually, ensure they are properly identified and defined. See Identify time partitions.

      If you select Discover partitions via crawler, a crawler process will scan your dataset when you reach the Review step of dataset definition and select Create dataset. The crawler process might take a few minutes to complete. After the crawler process completes with a Status of Needs action, you review, edit, and confirm the crawler-discovered partitions on the Edit page for the dataset. This is also the point where you ensure that crawler-discovered time partitions are properly identified and defined. See Review the crawler-discovered schema and partitions for an Amazon S3 dataset.

    • No: Define partitions manually using a Field list view or a JSON view. If you have time partitions, ensure they are defined. See Identify time partitions.
    • I don't have partitions: Select Next.
  7. Select Next to move on to the Update policies step.
Authenticate access to your dataset by applying a Splunk-generated resource access policy to the AWS IAM role that is associated with the dataset's connection. See Apply the dataset resource access policy to an AWS IAM role.