9.2 Collector Configuration Structure & Pipelines
Key Takeaways
Collector YAML declares components under receivers, processors, exporters, connectors, and extensions, and activates them in the service block; each pipeline needs at least one receiver and one exporter.
Declaring a component under receivers, processors, or exporters merely configures an instance; it remains completely inactive until explicitly referenced in a pipeline under service.pipelines.
Pipelines are signal-specific (traces, metrics, logs) and support multiple named variants using forward-slash syntax (such as traces/sampling and traces/compliance).
Processor execution order within a pipeline is strictly linear and sequential: memory_limiter must be placed first to prevent OOM kills, followed by filters and transforms, with batch placed immediately before exporters.
Multiple pipelines operate independently and concurrently, allowing fan-in from multiple receivers and fan-out to multiple backend exporters without cross-signal data interference.
9.2 Collector Configuration Structure & Pipelines
Quick Answer: OpenTelemetry Collector configurations are defined in YAML using the top-level keys
receivers,processors,exporters,connectors(when used),extensions, andservice. Theserviceblock glues the configuration together by defining activeextensionsand signal-specificpipelines(traces,metrics,logs). Processors in a pipeline execute in the exact linear order listed. Critical rule:memory_limitershould execute first to guard against out-of-memory crashes, whilebatchshould execute after any sampling or filtering, right before the exporters.
The OpenTelemetry Collector relies on a purely declarative YAML configuration file. Rather than writing custom integration code or managing disparate agents, operators declare components, tune their parameters, and wire them together into independent data-processing pipelines. Understanding the hierarchical structure of this YAML document and mastering pipeline mechanics is central to operating a Collector.
The Top-Level YAML Configuration Blocks
A Collector configuration file is organized around these root-level keys (a sixth, connectors, appears when pipelines are joined by a connector; see Section 11.3):
receivers: # 1. Declares input protocols and network bindings
...
processors: # 2. Declares in-memory transformation, filtering, and batching
...
exporters: # 3. Declares destination endpoints, formats, and egress queues
...
extensions: # 4. Declares auxiliary runtime capabilities (health, profiling)
...
service: # 5. Activates extensions, telemetry, and wires components into pipelines
...
1. receivers
The receivers section configures how telemetry enters the Collector. Each receiver instance is defined using the component name or a named instance format: component_type[/custom_name].
receivers:
otlp: # Default instance of the otlpreceiver
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
otlp/internal: # Named instance of the otlpreceiver for internal traffic
protocols:
grpc:
endpoint: 127.0.0.1:5555
prometheus: # Scrape receiver for Prometheus metrics
config:
scrape_configs:
- job_name: 'collector-self'
scrape_interval: 15s
static_configs:
- targets: ['localhost:8888']
2. processors
The processors section defines in-memory operations applied to telemetry as it moves through the pipeline. Like receivers, multiple instances of the same processor type can be declared using named suffixes:
processors:
memory_limiter: # Must be configured to prevent OOM
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 20
batch: # Buffers data before sending to exporters
send_batch_size: 8192
timeout: 1s
send_batch_max_size: 10240
filter/drop_health: # Named filter processor to drop health probes
error_mode: ignore
traces:
span:
- 'attributes["url.path"] == "/healthz"'
3. exporters
The exporters section defines where processed telemetry is delivered, along with connection timeouts, TLS certificates, compression algorithms, and queuing mechanisms:
exporters:
otlp/primary: # Primary OTLP gRPC backend
endpoint: tempo.monitoring.svc.cluster.local:4317
tls:
insecure: true
sending_queue:
enabled: true
queue_size: 5000
retry_on_failure:
enabled: true
initial_interval: 5s
max_interval: 30s
debug: # Replaces legacy 'logging' exporter for troubleshooting
verbosity: detailed
4. extensions
The extensions section configures auxiliary components that provide server management, security, and diagnostics without processing telemetry batches:
extensions:
health_check:
endpoint: 0.0.0.0:13133
pprof:
endpoint: 0.0.0.0:1777
zpages:
endpoint: 0.0.0.0:55679
5. service (The Master Orchestrator)
The service section binds all declared components together into a working Collector. If a component is declared under receivers, processors, exporters, or extensions but is never referenced in the service block, it is completely ignored by the Collector runtime:
service:
extensions: [health_check, pprof, zpages] # Activates declared extensions
pipelines:
traces:
receivers: [otlp, otlp/internal]
processors: [memory_limiter, filter/drop_health, batch]
exporters: [otlp/primary, debug]
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, batch]
exporters: [otlp/primary]
telemetry: # Configures internal monitoring of the Collector itself
logs:
level: info
metrics:
readers:
- pull:
exporter:
prometheus:
host: 0.0.0.0
port: 8888
Multi-Signal Pipelines
A pipeline defines a unidirectional data processing pathway for a specific telemetry signal. Each pipeline connects one or more receivers, through a sequence of zero or more processors, to one or more exporters.
+-------------------------------------------------------------+
| PIPELINE |
| |
| [Receivers] ===> [Processor 1] ===> [Processor 2] ===> [Exporters]
| (Fan-In) (Sequential execution order) (Fan-Out)
+-------------------------------------------------------------+
Pipeline Signal Types and Naming Syntax
Pipelines are strictly typed by signal:
traces: Ingests, processes, and exports distributed trace spans.metrics: Ingests, processes, and exports gauge, counter, and histogram data points.logs: Ingests, processes, and exports structured log records.
When multiple pipelines are needed for the same signal (e.g., routing production traces differently from development traces), operators declare named pipelines using a forward slash:
traces/productiontraces/compliance_auditmetrics/internal
Fan-In and Fan-Out Mechanics
The Collector architecture natively supports complex routing topologies through fan-in and fan-out:
- Receiver Fan-In: A single pipeline can reference multiple receivers (
receivers: [otlp, zipkin, jaeger]). Telemetry arriving on any of these receivers is converted into internalpdataand fed into the shared processor chain. - Receiver Fan-Out: A single receiver can be referenced across multiple pipelines (
traces,metrics,logs). Each OTLP export request carries one signal, so theotlpreceiver hands trace requests to every traces pipeline that lists it and metric requests to the metrics pipelines. - Exporter Fan-Out: A single pipeline can reference multiple exporters (
exporters: [otlp/primary, otlp/backup, debug]). The pipeline hands every processed batch to each listed exporter, cloning the data when a consumer might modify it.
Pipeline Independence and Isolation
Pipelines do not share data: a transform in the traces pipeline never touches metrics. They are not perfectly isolated, though. A receiver shared by several pipelines waits for its downstream consumers, so backpressure in one pipeline can slow the shared receiver; separate receivers keep failure domains apart.
Processor Execution Order
Processors in a pipeline execute in the EXACT sequential order in which they are listed under service.pipelines.<signal>.processors. They do not execute concurrently, and they do not execute based on internal priority heuristics.
CORRECT PRODUCTION ORDER:
[Receivers]
|
v
1. memory_limiter <-- Drops/throttles before memory allocation; prevents OOM
|
v
2. filter <-- Discards unwanted data immediately (saves CPU/memory)
|
v
3. transform <-- Scrubs PII, normalizes attributes on retained data
|
v
4. k8sattributes <-- Enriches retained spans with pod/namespace metadata
|
v
5. batch <-- Assembles optimal network payloads right before sending
|
v
[Exporters]
The Operational Rules of Processor Sequencing
memory_limiterMUST ALWAYS BE FIRST: The memory limiter processor monitors the Go runtime heap. If available memory drops below configured safety limits, it drops data or signals backpressure to the receiver. Ifmemory_limiteris placed afterbatchortransform, the Collector will allocate large amounts of heap space parsing, mutating, and buffering payloads, triggering an operating system Out-Of-Memory (OOM) kill before the limiter ever evaluates the payload!filterandtransformBEFOREbatch: Filtering unwanted spans (such as high-frequency/healthzchecks) must happen before batching. Ifbatchis placed beforefilter, the Collector wastes CPU cycles and memory assembling large batches of data, only for the subsequent filter processor to tear those batches apart and discard half the items. This leads to inefficient, half-empty network payloads.batchNear the End (After Sampling and Filtering): Thebatchprocessor's README says to place it aftermemory_limiterand after any sampling processors, because batching should happen after data is dropped. Placing it just before the exporters lets outgoing requests carry full batches. (Newer Collector releases can also batch inside the exporter throughsending_queue::batch, see Section 10.3.)
Annotated Complete Production YAML Configuration
The following complete, production-ready configuration demonstrates the integration of all five configuration blocks across traces, metrics, and logs:
# ===================================================================
# 1. RECEIVERS: Ingest telemetry across multiple protocols
# ===================================================================
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
prometheus:
config:
scrape_configs:
- job_name: 'otel-collector'
scrape_interval: 15s
static_configs:
- targets: ['localhost:8888']
# ===================================================================
# 2. PROCESSORS: In-memory processing, ordering, and batching
# ===================================================================
processors:
# Rule: memory_limiter must always be the first processor!
memory_limiter:
check_interval: 500ms
limit_percentage: 80
spike_limit_percentage: 20
# Rule: filter early to reduce resource consumption
filter/drop_health:
error_mode: ignore
traces:
span:
- 'attributes["url.path"] == "/healthz"'
- 'attributes["url.path"] == "/ready"'
# Rule: enrich retained telemetry
resourcedetection:
detectors: [env, system]
timeout: 2s
# Rule: batch immediately before exporters
batch:
send_batch_size: 4096
timeout: 500ms
send_batch_max_size: 8192
# ===================================================================
# 3. EXPORTERS: Deliver processed pdata to destinations
# ===================================================================
exporters:
otlp/tempo:
endpoint: tempo.monitoring.svc.cluster.local:4317
tls:
insecure: true
sending_queue:
enabled: true
queue_size: 5000
retry_on_failure:
enabled: true
otlp/mimir:
endpoint: mimir.monitoring.svc.cluster.local:4317
tls:
insecure: true
debug:
verbosity: basic
# ===================================================================
# 4. EXTENSIONS: Auxiliary management and diagnostics
# ===================================================================
extensions:
health_check:
endpoint: 0.0.0.0:13133
pprof:
endpoint: 0.0.0.0:1777
zpages:
endpoint: 0.0.0.0:55679
# ===================================================================
# 5. SERVICE: Wires components into pipelines and activates runtime
# ===================================================================
service:
extensions: [health_check, pprof, zpages]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, filter/drop_health, resourcedetection, batch]
exporters: [otlp/tempo, debug]
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, resourcedetection, batch]
exporters: [otlp/mimir]
telemetry:
logs:
level: info
metrics:
readers:
- pull:
exporter:
prometheus:
host: 0.0.0.0
port: 8888
A DevOps engineer configures a new filter processor under the top-level processors block named filter/drop_ping to discard synthetic ping check traces. After restarting the OpenTelemetry Collector, ping check spans continue to be stored in the remote tracing backend, and the Collector logs show no errors during startup. What is the root cause of this behavior?
The filter processor only functions on metric streams and cannot evaluate trace span attributes.
The remote tracing backend automatically overrides and re-generates dropped synthetic spans.
The memory_limiter processor dynamically disabled the filter processor due to CPU resource starvation.
The filter/drop_ping processor was declared under the top-level processors block but was omitted from the service.pipelines.traces.processors list.
A high-throughput Collector instance in an enterprise Kubernetes cluster repeatedly crashes with Out-Of-Memory (OOM) errors during sudden traffic spikes. An investigation reveals the pipeline configuration: processors: [k8sattributes, batch, transform, memory_limiter]. How should the platform team restructure this processor list to stabilize the Collector?
Reorder the processors so that memory_limiter executes first, followed by transform, k8sattributes, and batch immediately before the exporters.
Place batch first, followed by memory_limiter, k8sattributes, and transform.
Move memory_limiter to execute between batch and the exporters to evaluate full batch allocations.
Increase the batch processor timeout to allow the operating system time to garbage-collect stale spans.
An infrastructure architect must configure an OpenTelemetry Collector gateway that ingests traces from internal microservices, dispatches high-priority payment spans to both an internal compliance audit store and an APM backend, while routing general application traces only to the APM backend. How can this routing architecture be achieved within a single Collector configuration?
Configure two separate OTLP receiver instances listening on different TCP port numbers for each microservice type.
Define multiple named traces pipelines (e.g., traces/audit and traces/apm) that share the same OTLP receiver but specify different processors and exporters.
Deploy a custom Go extension that dynamically rewrites pdata memory addresses at runtime.
Create multiple service blocks within the YAML configuration to isolate pipeline execution.
Sections you finish are checked off in the contents.