10.2 Core Processors & Execution Order
Key Takeaways
Processors manipulate, enrich, filter, or batch telemetry data in memory as it flows between receivers and exporters.
The memory_limiter processor guards against Out-Of-Memory (OOM) crashes; the Collector's documented best practice is to make it the first processor in every pipeline so it refuses data before other processors allocate memory for it.
The batch processor groups telemetry into larger network payloads based on send_batch_size and timeout, maximizing compression and reducing backend RPC overhead; it should be placed last before exporters.
The attributes and resource processors modify telemetry metadata using actions such as insert, update, upsert, delete, hash (for PII redaction), and extract.
The k8sattributes processor enriches telemetry with Kubernetes pod, namespace, and node metadata by correlating incoming client IP addresses with the Kubernetes API server cache.
10.2 Core Processors & Execution Order
Quick Answer: Processors are the pipeline stages of the OpenTelemetry Collector that transform, filter, enrich, sample, and batch telemetry data in memory. Processors execute in the exact sequential order declared under
service.pipelines.<signal>.processors. The most important operational rule is processor ordering: thememory_limiterprocessor MUST ALWAYS BE FIRST in every pipeline to guard against Out-Of-Memory (OOM) fatal crashes, while thebatchprocessor should be placed LAST (immediately before exporters) to ensure optimal payload compression and reduce network RPC calls.
While receivers bring telemetry into the Collector and exporters send it to destination backends, Processors form the central computational engine of the OpenTelemetry Collector. Operating directly on in-memory pdata structures, processors allow platform engineers to enforce data governance, redact sensitive personally identifiable information (PII), drop high-volume synthetic health checks, enrich spans with infrastructure metadata, and bundle telemetry into efficient network batches.
Understanding processor functionality and the rules of processor sequencing is essential for running a stable Collector.
Core Processors
1. The memory_limiter Processor (First-Line Safeguard)
The memory_limiter processor is the primary operational safeguard preventing the Collector from crashing due to Out-Of-Memory (OOM) conditions. In containerized environments like Kubernetes, exceeding a container's memory limit causes the Linux kernel's OOM Killer to immediately terminate the process (exit code 137).
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
Why memory_limiter Exists
The Go runtime utilizes an automatic garbage collector (GC). Under heavy telemetry bursts, incoming data can allocate hundreds of megabytes of heap memory within milliseconds. Because Go's GC runs periodically rather than continuously, heap allocations can easily breach the container limit before a GC cycle completes. The memory_limiter constantly monitors the Go runtime heap and intervenes before fatal memory exhaustion occurs.
Configuration Parameters
check_interval: How often the processor measures memory. Its default is0s, so always set it explicitly;1sis the usual starting point, shorter for spiky traffic.limit_mib/limit_percentage: The hard limit. When memory rises above it, the processor additionally forces garbage collection.spike_limit_mib/spike_limit_percentage: The largest spike expected between two checks (default 20% of the limit). The soft limit islimit - spike_limit. Above the soft limit, the processor enters memory-limited mode and refuses new data by returning a retryable (non-permanent) error to the preceding component; receivers pass that back to clients, which should retry. Normal operation resumes when usage falls below the soft limit.- The Sizing Rule: On a 2 GiB container,
limit_percentage: 80(about 1.6 GiB) andspike_limit_percentage: 20(about 400 MiB) put the soft limit near 1.2 GiB. Refusals start there, leaving headroom to absorb spikes before the container's OOM limit. - Limits of the safeguard: The README warns that incoming data can consume memory in a receiver before the limiter can reject it, and that data is lost if the component before the limiter does not retry. The processor complements correct sizing and
GOMEMLIMIT; it does not replace them.
Important
Golden rule: place the memory_limiter processor first in every pipeline.
Processors execute in strict linear sequence. If memory_limiter is placed after memory-intensive processors (such as transform, k8sattributes, or batch), those preceding processors will allocate heap memory to deserialize, enrich, or buffer incoming telemetry bursts before the limiter ever evaluates the payload. By the time data reaches memory_limiter, the container has already breached its memory cgroup limit and been killed by the kernel!
2. The batch Processor (Network Efficiency & Compression)
Telemetry data arrives at the Collector as fragmented streams of individual spans, metrics, or log records. Transmitting each individual item over the network in its own HTTP or gRPC request generates massive TCP and TLS connection overhead, wastes network bandwidth, and overburdens downstream observability backends.
The batch processor buffers incoming data points in memory, grouping them into optimal batches before passing them to the exporter.
processors:
batch:
send_batch_size: 8192
timeout: 1s
send_batch_max_size: 10240
Configuration Parameters
send_batch_size(default8192): The number of spans, metric data points, or log records that triggers sending a batch regardless of the timeout. It is a trigger, not a cap on batch size.timeout(default200ms): The maximum time to wait before sending an incomplete batch. Under low traffic, this prevents telemetry from lingering in the Collector.send_batch_max_size(default0, meaning no limit): The upper limit on batch size (e.g.,10240), which must be at leastsend_batch_size. If incoming volume causes a batch to exceed this number, the processor splits the payload into multiple smaller batches. This prevents batches from exceeding downstream backend gRPC frame limits (default 4 MiB) or HTTP request entity size restrictions.
Execution Order Rule for batch
The batch processor should come after memory_limiter and any sampling or filtering processors, normally right before the exporters. This ensures that filtering, sampling, and attribute transformations have already occurred, allowing the batch processor to construct large, contiguous, final payloads that maximize gzip/zstd compression over the network.
3. The attributes & resource Processors
These processors modify metadata attached to telemetry items:
attributesProcessor: Operates on signal-level attributes (span attributes, log record attributes, metric data point attributes).resourceProcessor: Operates on entity-level attributes describing the producer of telemetry (service.name,service.version,host.name,cloud.region,deployment.environment.name).
processors:
resource:
attributes:
- key: deployment.environment.name
value: production
action: upsert
- key: host.id
action: delete
attributes/redact_pii:
actions:
- key: user.email
action: hash # Replaces email with SHA-1 hash for privacy compliance
- key: credit_card
action: delete
- key: app.client_ip
pattern: '^(?P<client_subnet>\d+\.\d+\.\d+)\.\d+$'
action: extract # Named group client_subnet becomes a new attribute
Supported Actions
| Action | Operational Behavior |
|---|---|
insert | Adds the attribute key and value only if the key does not already exist. Leaves existing values untouched. |
update | Modifies the attribute value only if the key already exists. Does nothing if the key is missing. |
upsert | Unconditionally sets the attribute. If the key exists, its value is overwritten; if missing, it is created. |
delete | Removes the specified attribute key and its value completely. |
hash | Replaces the attribute value with its SHA-1 hash (useful for pseudonymizing values such as user IDs or emails). |
extract | Uses a regular expression with named capture groups to parse an existing attribute and generate new attributes. |
4. The filter Processor
The filter processor drops unneeded spans, metrics, or log records from the pipeline based on deterministic boolean conditions or regular expressions. Discarding low-value telemetry early saves memory, CPU, network bandwidth, and downstream backend storage costs.
processors:
filter/drop_health_checks:
error_mode: ignore
traces:
span:
- 'attributes["url.path"] == "/healthz"'
- 'attributes["url.path"] == "/ready"'
logs:
log_record:
- 'severity_number < SEVERITY_NUMBER_INFO' # Drops TRACE and DEBUG logs
5. The k8sattributes Processor (Kubernetes Enrichment)
In containerized Kubernetes clusters, applications emit telemetry that typically lacks cluster-level metadata. The k8sattributes processor (available in the Contrib distribution) automatically enriches incoming telemetry with Kubernetes pod, namespace, and node metadata without requiring changes to application code.
processors:
k8sattributes:
auth_type: "serviceAccount"
passthrough: false
extract:
metadata:
- k8s.pod.name
- k8s.pod.uid
- k8s.namespace.name
- k8s.node.name
- k8s.deployment.name
labels:
- tag_name: app.label.version
key: app.kubernetes.io/version
pod_association:
- sources:
- from: connection # Extracts client IP from incoming TCP socket
How IP Correlation Works
- Informer Cache: The processor uses a Kubernetes client to watch the Kubernetes API server via Informers, maintaining an in-memory cache mapping Pod IP addresses to current Pod metadata, labels, and annotations.
- IP Extraction: When telemetry arrives, the processor identifies the source IP address (either from the incoming TCP connection context or from an existing resource attribute such as
k8s.pod.ip). - Metadata Injection: The processor matches the IP address against its in-memory pod table and injects the corresponding metadata (
k8s.pod.name,k8s.namespace.name,k8s.node.name, etc.) directly into the telemetry's resource attributes.
Pipeline Processor Ordering Rules (The Golden Sequence)
Processors execute strictly in the order they appear in the pipeline configuration. Violating the proper sequence introduces performance degradation, data corruption, or fatal process crashes.
CORRECT PIPELINE PROCESSOR SEQUENCE:
[Receivers]
|
v
1. memory_limiter <--- MUST BE FIRST: Drops/throttles data before memory is allocated
|
v
2. filter <--- Discard unneeded spans/logs early (saves downstream CPU & RAM)
|
v
3. transform / OTTL <--- Scrub PII, rewrite paths, normalize field structures
|
v
4. k8sattributes <--- Enrich retained telemetry with Kubernetes cluster metadata
|
v
5. batch <--- LAST: Groups processed items into large, compressed batches
|
v
[Exporters]
Analysis of Processor Sequencing Rules
| Pipeline Position | Processor Type | Operational Rationale | Consequence if Misplaced |
|---|---|---|---|
| 1st (Entry) | memory_limiter | Evaluates current Go heap allocations before any memory is allocated for mutations or buffers. | If placed later, heavy processors allocate memory for large incoming bursts, causing Kubernetes OOM kills before the limiter can act. |
| 2nd | filter / sampling | Drops high-frequency synthetic spans (e.g., /healthz) and debug logs immediately. | If placed after batch or k8sattributes, the Collector wastes CPU cycles enriching and batching data that is immediately thrown away. |
| 3rd | transform / attributes / resource | Modifies, standardizes, and hashes attributes on retained telemetry data. | If placed before filtering, the Collector burns CPU cycles masking PII on spans that will subsequently be filtered out. |
| 4th | k8sattributes | Queries local pod cache to inject cluster context onto filtered, sanitized telemetry. | If placed before filtering, IP correlation cache lookups run on high-volume synthetic health check spans, thrashing CPU. |
| Final (Exit) | batch | Buffers enriched, sanitized telemetry into optimal batches for network transmission. | If placed before filters or transforms, downstream processors fragment the assembled batches, resulting in small, inefficient network requests. |
A high-throughput OpenTelemetry Collector gateway deployed on a Kubernetes cluster with a 2 GiB memory limit repeatedly crashes with Out-Of-Memory (OOM) errors during morning traffic peaks. An engineer inspects the configuration and identifies the pipeline: processors: [k8sattributes, transform, batch, memory_limiter]. Why does this configuration cause the Collector to be terminated by the kernel, and what is the required remedy?
The batch processor timeout is too high, causing memory to leak; reducing the timeout from 1s to 100ms resolves the issue.
The transform processor requires a tail_sampling processor to precede it to reduce trace volume before OTTL statements execute.
The k8sattributes processor creates persistent TCP sockets to every pod in the cluster, exhausting operating system file descriptors.
The memory_limiter is placed last, meaning preceding processors allocate heap memory for incoming telemetry bursts before memory checks occur; placing memory_limiter first allows it to drop or throttle data before memory exhaustion.
An observability engineer notes that downstream APM network traffic from an OpenTelemetry Collector gateway consists of thousands of small, uncompressed network requests containing only 2 to 3 spans each, resulting in high network latency and excessive CPU utilization on the APM backend. The existing configuration shows: processors: [memory_limiter, batch, filter/drop_health]. How should the engineer restructure the processor pipeline to maximize network efficiency?
Move filter/drop_health before batch, and place batch immediately before the exporters with send_batch_size and timeout tuned for optimal aggregation.
Place batch ahead of memory_limiter so that telemetry is aggregated before memory thresholds are evaluated.
Remove the batch processor entirely and configure gzip compression directly on the OTLP receiver.
Declare multiple batch processors in parallel inside the receiver configuration block.
A platform team wants all microservices running inside a Kubernetes cluster to have their distributed traces enriched with k8s.pod.name, k8s.namespace.name, and k8s.node.name without altering application source code or SDK configurations. Which processor fulfills this requirement, and how does it determine which pod produced the telemetry?
The resourcedetection processor, which queries the cloud provider metadata API for host details.
The k8sattributes processor, which correlates the incoming TCP connection client IP address with its internal cache populated by watching the Kubernetes API server.
The resource processor, which parses the pod name from the span name using regular expression extract actions.
The transform processor, which executes reverse DNS lookups on every incoming span attribute.
Sections you finish are checked off in the contents.