Remove duplicate fields from pipelines
Remove duplicate events from Edge Processor pipelines on Splunk Enterprise.
Use the dedup command to remove duplicate events from Edge Processor pipelines on Splunk Enterprise.
The SPL2 dedup command removes events that contain an identical combination of values for the fields that you specify.
You can specify the number of duplicate events to keep for each value of a single field or for each combination of values among several fields.
dedup command is supported for Edge Processor pipelines in your target Splunk Enterprise release before you publish or use this procedure. The current Compatibility Quick Reference for SPL2 commands does not list dedup among the Edge Processor commands.
Deduplication overview
Plan deduplication for Edge Processor pipelines on Splunk Enterprise.
Removing duplicate events from an Edge Processor pipeline on Splunk Enterprise involves the following tasks:
-
Identify your data source: Determine the pipeline input and the fields that cause duplication.
-
Select a deduplication strategy: Choose between a visual UI configuration and custom SPL2 code.
-
Define the scope: Specify the fields for the
dedupcommand. -
Configure time constraints: Set the span and TTL to define how long the processor retains event information.
-
Validate: Run the pipeline in Preview mode to verify the reduction in event volume.
How duplicates are identified
The dedup command identifies duplicate events by using a time range and an event-count range. Deduplication effectiveness depends on the runtime context, including batch, instance, and inter-batch processing, and on the memory and TTL configuration.
Steps
Configure an Edge Processor pipeline in Splunk Enterprise by using the Data Management app or custom SPL2 code.
Configure a pipeline to remove duplicate events by using the Data Management app
Complete the following steps to remove duplicate events from your pipeline.
- In the Data Management app on your Splunk Enterprise data management control plane, open the Pipelines page. Find the pipeline that you want to deduplicate, and select Edit.
- Select the plus icon next to Actions.
- Select Remove duplicates for.
- On the Remove duplicate field values page, set the deduplication parameters, and select Apply.
- Select Next to confirm the deduplicated data.
- Run a preview of your pipeline to verify your changes.
- Select Done to confirm your changes.
Configure a pipeline to remove duplicate events by using custom SPL2 code
Complete the following steps to remove duplicate events from your pipeline.
- In the Data Management app on your Splunk Enterprise data management control plane, open the Pipelines page. Find the pipeline that you want to deduplicate, and select Edit.
- In the pipeline editor, navigate to the fields that you want to deduplicate.
- Enter SPL2 code that uses the
dedupcommand. - Select Preview to review your changes.
- Save your changes.
Examples of deduplication pipelines
Examples of SPL2 pipeline code that uses the dedup command in Edge Processor pipelines on Splunk Enterprise.
Deduplicate by host within a batch
from $source | dedup host, batch_id()
Use a time interval to deduplicate events
from $source | eval field_with_batch_id = batch_id() | dedup host, field_with_batch_id, span(_time, 5m)
Use @maxmem as a runtime hint in memory-constrained environments
from $source | @maxmem('1GB') dedup host, batch_id()
See also
See also
Related SPL2 topics
For more information, see the following topics in the Splunk Enterprise SPL2 Search Reference: