Remove duplicate fields from pipelines

Remove duplicate events from Edge Processor pipelines on Splunk Enterprise.

Use the dedup command to remove duplicate events from Edge Processor pipelines on Splunk Enterprise.

The SPL2 dedup command removes events that contain an identical combination of values for the fields that you specify.

You can specify the number of duplicate events to keep for each value of a single field or for each combination of values among several fields.

Important: Verify that the dedup command is supported for Edge Processor pipelines in your target Splunk Enterprise release before you publish or use this procedure. The current Compatibility Quick Reference for SPL2 commands does not list dedup among the Edge Processor commands.

Deduplication overview

Plan deduplication for Edge Processor pipelines on Splunk Enterprise.

Removing duplicate events from an Edge Processor pipeline on Splunk Enterprise involves the following tasks:

  1. Identify your data source: Determine the pipeline input and the fields that cause duplication.

  2. Select a deduplication strategy: Choose between a visual UI configuration and custom SPL2 code.

  3. Define the scope: Specify the fields for the dedup command.

  4. Configure time constraints: Set the span and TTL to define how long the processor retains event information.

  5. Validate: Run the pipeline in Preview mode to verify the reduction in event volume.

How duplicates are identified

The dedup command identifies duplicate events by using a time range and an event-count range. Deduplication effectiveness depends on the runtime context, including batch, instance, and inter-batch processing, and on the memory and TTL configuration.

Steps

Configure an Edge Processor pipeline in Splunk Enterprise by using the Data Management app or custom SPL2 code.

Configure a pipeline to remove duplicate events by using the Data Management app

Complete the following steps to remove duplicate events from your pipeline.

  1. In the Data Management app on your Splunk Enterprise data management control plane, open the Pipelines page. Find the pipeline that you want to deduplicate, and select Edit.
  2. Select the plus icon next to Actions.
  3. Select Remove duplicates for.
  4. On the Remove duplicate field values page, set the deduplication parameters, and select Apply.
  5. Select Next to confirm the deduplicated data.
  6. Run a preview of your pipeline to verify your changes.
  7. Select Done to confirm your changes.

Configure a pipeline to remove duplicate events by using custom SPL2 code

Complete the following steps to remove duplicate events from your pipeline.

  1. In the Data Management app on your Splunk Enterprise data management control plane, open the Pipelines page. Find the pipeline that you want to deduplicate, and select Edit.
  2. In the pipeline editor, navigate to the fields that you want to deduplicate.
  3. Enter SPL2 code that uses the dedup command.
  4. Select Preview to review your changes.
  5. Save your changes.

Examples of deduplication pipelines

Examples of SPL2 pipeline code that uses the dedup command in Edge Processor pipelines on Splunk Enterprise.

Deduplicate by host within a batch

PYTHON
from $source | dedup host, batch_id()

Use a time interval to deduplicate events

PYTHON
from $source | eval field_with_batch_id = batch_id() | dedup host, field_with_batch_id, span(_time, 5m)

Use @maxmem as a runtime hint in memory-constrained environments

PYTHON
from $source | @maxmem('1GB') dedup host, batch_id()

See also