9.2 Collector Configuration Structure & Pipelines

Key Takeaways

  • Collector YAML declares components under receivers, processors, exporters, connectors, and extensions, and activates them in the service block; each pipeline needs at least one receiver and one exporter.

  • Declaring a component under receivers, processors, or exporters merely configures an instance; it remains completely inactive until explicitly referenced in a pipeline under service.pipelines.

  • Pipelines are signal-specific (traces, metrics, logs) and support multiple named variants using forward-slash syntax (such as traces/sampling and traces/compliance).

  • Processor execution order within a pipeline is strictly linear and sequential: memory_limiter must be placed first to prevent OOM kills, followed by filters and transforms, with batch placed immediately before exporters.

  • Multiple pipelines operate independently and concurrently, allowing fan-in from multiple receivers and fan-out to multiple backend exporters without cross-signal data interference.

Last updated: September 2026

9.2 Collector Configuration Structure & Pipelines

Quick Answer: OpenTelemetry Collector configurations are defined in YAML using the top-level keys receivers, processors, exporters, connectors (when used), extensions, and service. The service block glues the configuration together by defining active extensions and signal-specific pipelines (traces, metrics, logs). Processors in a pipeline execute in the exact linear order listed. Critical rule: memory_limiter should execute first to guard against out-of-memory crashes, while batch should execute after any sampling or filtering, right before the exporters.

The OpenTelemetry Collector relies on a purely declarative YAML configuration file. Rather than writing custom integration code or managing disparate agents, operators declare components, tune their parameters, and wire them together into independent data-processing pipelines. Understanding the hierarchical structure of this YAML document and mastering pipeline mechanics is central to operating a Collector.


The Top-Level YAML Configuration Blocks

A Collector configuration file is organized around these root-level keys (a sixth, connectors, appears when pipelines are joined by a connector; see Section 11.3):

receivers:   # 1. Declares input protocols and network bindings
  ...
processors:  # 2. Declares in-memory transformation, filtering, and batching
  ...
exporters:   # 3. Declares destination endpoints, formats, and egress queues
  ...
extensions:   # 4. Declares auxiliary runtime capabilities (health, profiling)
  ...
service:     # 5. Activates extensions, telemetry, and wires components into pipelines
  ...

1. receivers

The receivers section configures how telemetry enters the Collector. Each receiver instance is defined using the component name or a named instance format: component_type[/custom_name].

receivers:
  otlp: # Default instance of the otlpreceiver
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318
  otlp/internal: # Named instance of the otlpreceiver for internal traffic
    protocols:
      grpc:
        endpoint: 127.0.0.1:5555
  prometheus: # Scrape receiver for Prometheus metrics
    config:
      scrape_configs:
        - job_name: 'collector-self'
          scrape_interval: 15s
          static_configs:
            - targets: ['localhost:8888']

2. processors

The processors section defines in-memory operations applied to telemetry as it moves through the pipeline. Like receivers, multiple instances of the same processor type can be declared using named suffixes:

processors:
  memory_limiter: # Must be configured to prevent OOM
    check_interval: 1s
    limit_percentage: 75
    spike_limit_percentage: 20
  batch: # Buffers data before sending to exporters
    send_batch_size: 8192
    timeout: 1s
    send_batch_max_size: 10240
  filter/drop_health: # Named filter processor to drop health probes
    error_mode: ignore
    traces:
      span:
        - 'attributes["url.path"] == "/healthz"'

3. exporters

The exporters section defines where processed telemetry is delivered, along with connection timeouts, TLS certificates, compression algorithms, and queuing mechanisms:

exporters:
  otlp/primary: # Primary OTLP gRPC backend
    endpoint: tempo.monitoring.svc.cluster.local:4317
    tls:
      insecure: true
    sending_queue:
      enabled: true
      queue_size: 5000
    retry_on_failure:
      enabled: true
      initial_interval: 5s
      max_interval: 30s
  debug: # Replaces legacy 'logging' exporter for troubleshooting
    verbosity: detailed

4. extensions

The extensions section configures auxiliary components that provide server management, security, and diagnostics without processing telemetry batches:

extensions:
  health_check:
    endpoint: 0.0.0.0:13133
  pprof:
    endpoint: 0.0.0.0:1777
  zpages:
    endpoint: 0.0.0.0:55679

5. service (The Master Orchestrator)

The service section binds all declared components together into a working Collector. If a component is declared under receivers, processors, exporters, or extensions but is never referenced in the service block, it is completely ignored by the Collector runtime:

service:
  extensions: [health_check, pprof, zpages] # Activates declared extensions
  pipelines:
    traces:
      receivers: [otlp, otlp/internal]
      processors: [memory_limiter, filter/drop_health, batch]
      exporters: [otlp/primary, debug]
    metrics:
      receivers: [otlp, prometheus]
      processors: [memory_limiter, batch]
      exporters: [otlp/primary]
  telemetry: # Configures internal monitoring of the Collector itself
    logs:
      level: info
    metrics:
      readers:
        - pull:
            exporter:
              prometheus:
                host: 0.0.0.0
                port: 8888

Multi-Signal Pipelines

A pipeline defines a unidirectional data processing pathway for a specific telemetry signal. Each pipeline connects one or more receivers, through a sequence of zero or more processors, to one or more exporters.

+-------------------------------------------------------------+
|                          PIPELINE                           |
|                                                             |
|  [Receivers]  ===>  [Processor 1] ===> [Processor 2] ===> [Exporters]
|  (Fan-In)           (Sequential execution order)           (Fan-Out)
+-------------------------------------------------------------+

Pipeline Signal Types and Naming Syntax

Pipelines are strictly typed by signal:

  • traces: Ingests, processes, and exports distributed trace spans.
  • metrics: Ingests, processes, and exports gauge, counter, and histogram data points.
  • logs: Ingests, processes, and exports structured log records.

When multiple pipelines are needed for the same signal (e.g., routing production traces differently from development traces), operators declare named pipelines using a forward slash:

  • traces/production
  • traces/compliance_audit
  • metrics/internal

Fan-In and Fan-Out Mechanics

The Collector architecture natively supports complex routing topologies through fan-in and fan-out:

  1. Receiver Fan-In: A single pipeline can reference multiple receivers (receivers: [otlp, zipkin, jaeger]). Telemetry arriving on any of these receivers is converted into internal pdata and fed into the shared processor chain.
  2. Receiver Fan-Out: A single receiver can be referenced across multiple pipelines (traces, metrics, logs). Each OTLP export request carries one signal, so the otlp receiver hands trace requests to every traces pipeline that lists it and metric requests to the metrics pipelines.
  3. Exporter Fan-Out: A single pipeline can reference multiple exporters (exporters: [otlp/primary, otlp/backup, debug]). The pipeline hands every processed batch to each listed exporter, cloning the data when a consumer might modify it.

Pipeline Independence and Isolation

Pipelines do not share data: a transform in the traces pipeline never touches metrics. They are not perfectly isolated, though. A receiver shared by several pipelines waits for its downstream consumers, so backpressure in one pipeline can slow the shared receiver; separate receivers keep failure domains apart.


Processor Execution Order

Processors in a pipeline execute in the EXACT sequential order in which they are listed under service.pipelines.<signal>.processors. They do not execute concurrently, and they do not execute based on internal priority heuristics.

CORRECT PRODUCTION ORDER:
[Receivers] 
    |
    v
1. memory_limiter  <-- Drops/throttles before memory allocation; prevents OOM
    |
    v
2. filter          <-- Discards unwanted data immediately (saves CPU/memory)
    |
    v
3. transform       <-- Scrubs PII, normalizes attributes on retained data
    |
    v
4. k8sattributes   <-- Enriches retained spans with pod/namespace metadata
    |
    v
5. batch           <-- Assembles optimal network payloads right before sending
    |
    v
[Exporters]

The Operational Rules of Processor Sequencing

  1. memory_limiter MUST ALWAYS BE FIRST: The memory limiter processor monitors the Go runtime heap. If available memory drops below configured safety limits, it drops data or signals backpressure to the receiver. If memory_limiter is placed after batch or transform, the Collector will allocate large amounts of heap space parsing, mutating, and buffering payloads, triggering an operating system Out-Of-Memory (OOM) kill before the limiter ever evaluates the payload!
  2. filter and transform BEFORE batch: Filtering unwanted spans (such as high-frequency /healthz checks) must happen before batching. If batch is placed before filter, the Collector wastes CPU cycles and memory assembling large batches of data, only for the subsequent filter processor to tear those batches apart and discard half the items. This leads to inefficient, half-empty network payloads.
  3. batch Near the End (After Sampling and Filtering): The batch processor's README says to place it after memory_limiter and after any sampling processors, because batching should happen after data is dropped. Placing it just before the exporters lets outgoing requests carry full batches. (Newer Collector releases can also batch inside the exporter through sending_queue::batch, see Section 10.3.)

Annotated Complete Production YAML Configuration

The following complete, production-ready configuration demonstrates the integration of all five configuration blocks across traces, metrics, and logs:

# ===================================================================
# 1. RECEIVERS: Ingest telemetry across multiple protocols
# ===================================================================
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318
  prometheus:
    config:
      scrape_configs:
        - job_name: 'otel-collector'
          scrape_interval: 15s
          static_configs:
            - targets: ['localhost:8888']

# ===================================================================
# 2. PROCESSORS: In-memory processing, ordering, and batching
# ===================================================================
processors:
  # Rule: memory_limiter must always be the first processor!
  memory_limiter:
    check_interval: 500ms
    limit_percentage: 80
    spike_limit_percentage: 20

  # Rule: filter early to reduce resource consumption
  filter/drop_health:
    error_mode: ignore
    traces:
      span:
        - 'attributes["url.path"] == "/healthz"'
        - 'attributes["url.path"] == "/ready"'

  # Rule: enrich retained telemetry
  resourcedetection:
    detectors: [env, system]
    timeout: 2s

  # Rule: batch immediately before exporters
  batch:
    send_batch_size: 4096
    timeout: 500ms
    send_batch_max_size: 8192

# ===================================================================
# 3. EXPORTERS: Deliver processed pdata to destinations
# ===================================================================
exporters:
  otlp/tempo:
    endpoint: tempo.monitoring.svc.cluster.local:4317
    tls:
      insecure: true
    sending_queue:
      enabled: true
      queue_size: 5000
    retry_on_failure:
      enabled: true

  otlp/mimir:
    endpoint: mimir.monitoring.svc.cluster.local:4317
    tls:
      insecure: true

  debug:
    verbosity: basic

# ===================================================================
# 4. EXTENSIONS: Auxiliary management and diagnostics
# ===================================================================
extensions:
  health_check:
    endpoint: 0.0.0.0:13133
  pprof:
    endpoint: 0.0.0.0:1777
  zpages:
    endpoint: 0.0.0.0:55679

# ===================================================================
# 5. SERVICE: Wires components into pipelines and activates runtime
# ===================================================================
service:
  extensions: [health_check, pprof, zpages]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, filter/drop_health, resourcedetection, batch]
      exporters: [otlp/tempo, debug]
    metrics:
      receivers: [otlp, prometheus]
      processors: [memory_limiter, resourcedetection, batch]
      exporters: [otlp/mimir]
  telemetry:
    logs:
      level: info
    metrics:
      readers:
        - pull:
            exporter:
              prometheus:
                host: 0.0.0.0
                port: 8888
Loading diagram...
Collector Multi-Signal Pipelines and Sequential Processor Ordering
Test Your Knowledge

A DevOps engineer configures a new filter processor under the top-level processors block named filter/drop_ping to discard synthetic ping check traces. After restarting the OpenTelemetry Collector, ping check spans continue to be stored in the remote tracing backend, and the Collector logs show no errors during startup. What is the root cause of this behavior?

A

The filter processor only functions on metric streams and cannot evaluate trace span attributes.

B

The remote tracing backend automatically overrides and re-generates dropped synthetic spans.

C

The memory_limiter processor dynamically disabled the filter processor due to CPU resource starvation.

D

The filter/drop_ping processor was declared under the top-level processors block but was omitted from the service.pipelines.traces.processors list.

Test Your Knowledge

A high-throughput Collector instance in an enterprise Kubernetes cluster repeatedly crashes with Out-Of-Memory (OOM) errors during sudden traffic spikes. An investigation reveals the pipeline configuration: processors: [k8sattributes, batch, transform, memory_limiter]. How should the platform team restructure this processor list to stabilize the Collector?

A

Reorder the processors so that memory_limiter executes first, followed by transform, k8sattributes, and batch immediately before the exporters.

B

Place batch first, followed by memory_limiter, k8sattributes, and transform.

C

Move memory_limiter to execute between batch and the exporters to evaluate full batch allocations.

D

Increase the batch processor timeout to allow the operating system time to garbage-collect stale spans.

Test Your Knowledge

An infrastructure architect must configure an OpenTelemetry Collector gateway that ingests traces from internal microservices, dispatches high-priority payment spans to both an internal compliance audit store and an APM backend, while routing general application traces only to the APM backend. How can this routing architecture be achieved within a single Collector configuration?

A

Configure two separate OTLP receiver instances listening on different TCP port numbers for each microservice type.

B

Define multiple named traces pipelines (e.g., traces/audit and traces/apm) that share the same OTLP receiver but specify different processors and exporters.

C

Deploy a custom Go extension that dynamically rewrites pdata memory addresses at runtime.

D

Create multiple service blocks within the YAML configuration to isolate pipeline execution.

Sections you finish are checked off in the contents.