9.1 Collector Core Architecture & Component Lifecycle

Key Takeaways

  • The OpenTelemetry Collector is a high-performance, vendor-agnostic proxy process that receives, processes, and exports traces, metrics, and logs across disparate systems.

  • The Collector's internal architecture is structured around five component types: Receivers (push/pull ingress), Processors (in-memory transformation and batching), Exporters (egress translation), Extensions (auxiliary runtime capabilities), and Connectors (cross-pipeline bridges).

  • The internal pdata model is the Collector's in-memory representation of OTLP telemetry; it decouples ingestion protocols from destination formats.

  • The Collector starts components in dependency order (extensions, then exporters, processors, and receivers last) and shuts them down in reverse (receivers first, extensions last) so that in-flight data can drain; data held only in memory can still be lost if the process crashes.

  • The OpenTelemetry project publishes Core (otelcol), Contrib (otelcol-contrib), and Kubernetes (otelcol-k8s) Collector distributions, and the OpenTelemetry Collector Builder (OCB) compiles custom binaries with only the components you choose.

Last updated: September 2026

9.1 Collector Core Architecture & Component Lifecycle

Quick Answer: The OpenTelemetry Collector is a high-performance, vendor-agnostic proxy process that ingests, transforms, and dispatches telemetry data (traces, metrics, and logs). Its architecture consists of five core components: Receivers (ingest and convert telemetry into internal pdata), Processors (in-memory filtering, batching, and transformations), Exporters (translate pdata to external formats and send to backends), Extensions (auxiliary capabilities like health probes and profiling outside the data path), and Connectors (bridges between pipelines). The component lifecycle manages startup (extensions, exporters, processors, receivers) and graceful shutdown (stopping receivers first, flushing processors, draining exporters) so that in-flight data is not lost during a normal stop.

In modern cloud-native systems, microservices generate vast streams of traces, metrics, and logs. While application code can directly transmit telemetry to external storage backends using language SDKs, doing so introduces tight coupling to specific vendor APIs, burns application compute on network retries and compression, and creates security risks by exposing external backend credentials across application containers. The OpenTelemetry Collector solves these architectural challenges by providing a centralized, vendor-neutral proxy that abstracts telemetry ingestion, processing, and delivery.


What is the OpenTelemetry Collector?

The OpenTelemetry Collector is an out-of-process executable written in Go that acts as a telemetry router, transformation engine, and gateway. It receives telemetry from instrumented services, processes it according to declarative operational rules, and exports it to one or more observability backends (such as Prometheus, Jaeger, Elasticsearch, Google Cloud Trace, AWS X-Ray, Datadog, or custom OTLP stores).

+-----------------------+
| Instrumented Services |
| (Java, Go, Python...) |
+-----------------------+
           | (OTLP / Zipkin / Jaeger)
           v
+-------------------------------------------------------------+
|                  OpenTelemetry Collector                    |
|                                                             |
|  [Receivers] ---> [Processors] ---> [Exporters]             |
|       ^                                    |                |
|       |             [Extensions]           |                |
|       +----------- (Health, Pprof) --------+                |
+-------------------------------------------------------------+
           |                                    |
           v                                    v
+----------------------+             +----------------------+
| Prometheus / Metrics |             | Jaeger / Tracing APM |
+----------------------+             +----------------------+

Why Deploy a Collector Instead of Direct SDK Export?

Deploying an OpenTelemetry Collector provides several critical architectural benefits over direct-to-backend SDK export:

  1. Decoupling and Vendor Neutrality: Applications transmit telemetry using the standardized OpenTelemetry Protocol (OTLP). Changing an observability vendor or adding a secondary data lake requires updating a single Collector configuration file rather than recompiling, retesting, and redeploying dozens of application microservices.
  2. Offloading Compute and Network Overhead: In-memory batching, GZIP/ZSTD compression, retry backoffs, and TLS handshakes are offloaded from application runtimes into a dedicated proxy, keeping application containers lean and reducing tail latencies.
  3. Centralized Data Governance and PII Scrubbing: Security compliance rules—such as redacting personally identifiable information (PII), stripping credit card numbers, or masking authentication headers—can be enforced globally at the Collector layer before data ever exits the internal network perimeter.
  4. Metadata Enrichment: Collectors running as local agents or Kubernetes DaemonSets can automatically enrich incoming telemetry with host and container metadata (such as k8s.pod.name, k8s.namespace.name, host.id, and cloud availability zones) without application code needing access to the Kubernetes API.
  5. Advanced Sampling and Routing: Collectors can perform tail-based sampling, intelligent rate limiting, and multi-tenant routing that cannot be executed effectively within isolated application processes.

Core Architecture Components

The Collector processes telemetry through a modular pipeline architecture constructed from five fundamental component types:

1. Receivers

Receivers define how telemetry enters the Collector. A receiver can be push-based (listening on a network port for incoming payloads) or pull-based (actively scraping target endpoints on a configured schedule).

  • Push Receivers: Examples include otlpreceiver (listening on gRPC port 4317 and HTTP port 4318), zipkinreceiver (listening on port 9411), and jaegerreceiver (listening on gRPC port 14250 and HTTP port 14268). These receivers accept incoming network connections, deserialize the wire protocol, and parse the data.
  • Pull Receivers: Examples include prometheusreceiver (which scrapes Prometheus /metrics endpoints based on a scrape configuration) and host metrics receivers (which poll operating system /proc and kernel counters).
  • The Ingestion Invariant: Regardless of the incoming wire format (Zipkin JSON, Jaeger Thrift, Prometheus text format, or OTLP Protobuf), the receiver is strictly responsible for translating external payloads into the Collector's internal pdata (Pluggable Data) representation before passing it to the pipeline.

2. Processors

Processors sit between receivers and exporters, operating on in-memory pdata batches. Processors execute sequentially to transform, filter, enrich, sample, or buffer telemetry:

  • Batching (batch): Aggregates individual spans, metrics, or logs into larger batches based on byte size or time thresholds, dramatically improving network throughput and backend ingestion efficiency.
  • Memory Safety (memory_limiter): Constantly checks the Collector's Go heap allocations and drops data or rejects incoming requests when memory limits are approached, preventing out-of-memory (OOM) fatal crashes.
  • Transformation (transform): Uses the OpenTelemetry Transformation Language (OTTL) to modify span names, parse JSON bodies, set semantic attributes, or delete unwanted fields.
  • Filtering (filter): Evaluates boolean conditions to discard telemetry (e.g., dropping successful HTTP /healthz spans).
  • Enrichment (k8sattributes, resourcedetection): Detects environment variables or queries the Kubernetes API to append infrastructure metadata to the telemetry resource block.

3. Exporters

Exporters translate internal pdata structures into outgoing wire formats and dispatch them over the network to destination storage systems. Like receivers, exporters can be push-based or pull-based:

  • Push Exporters: Translate pdata into external protocols (such as OTLP gRPC/HTTP, Elasticsearch, Kafka, AWS CloudWatch, or Google Cloud Monitoring) and push payloads over the network. Push exporters feature configurable retry logic, exponential backoffs, and in-memory or persistent disk-backed sending queues.
  • Pull Exporters: Convert internal metrics into an endpoint that external scrapers can pull on demand (e.g., prometheusexporter hosting a /metrics scrape target, commonly configured on port 8889).

4. Extensions

Extensions provide optional auxiliary services that do not process telemetry directly. They operate outside the pipeline data path to provide management, diagnostic, and security capabilities:

  • Health Monitoring (health_check): Exposes an HTTP probe endpoint (default port 13133) indicating whether Collector components are operational.
  • Performance Diagnostics (pprof, zpages): Exposes Go runtime profiling on port 1777 and in-process diagnostic HTML pages on port 55679.
  • Security and Authentication (oidc, basicauth, bearertokenauth): Validates client authentication headers on incoming receiver connections or signs outgoing exporter requests.

5. Connectors

Introduced to eliminate architectural workarounds, Connectors act as both an exporter in one pipeline and a receiver in another pipeline. They bridge distinct signal pipelines without serializing data onto an external network loop:

  • spanmetrics Connector: Listens as an exporter at the end of a traces pipeline, aggregates request counts, error counts, and latency durations from spans, and emits those calculated measurements as metrics into the beginning of a metrics pipeline.
  • count Connector: Counts spans, metric data points, or log records passing through a pipeline and emits the counts as metrics.
  • routing and forward Connectors: routing sends data to different pipelines of the same signal based on OTTL conditions; forward simply passes data from one pipeline to another (it ships in the Core distribution).

The Internal pdata (Pluggable Data) Architecture

A critical design decision of the OpenTelemetry Collector is the use of pdata (go.opentelemetry.io/collector/pdata). In legacy telemetry proxies, components frequently converted data into generic maps (map[string]interface{}) or intermediate JSON strings. This created extreme garbage collection pressure, CPU serialization overhead, and memory churn under high throughput.

[Raw Ingress Wire Bytes] (Zipkin / Jaeger / OTLP)
           |
           v  (Receiver translates wire format to pdata)
+-------------------------------------------------------------+
|                   Internal pdata Model                      |
|                                                             |
|  ptrace.Traces   |   pmetric.Metrics   |   plog.Logs        |
|  - Wraps the generated OTLP protobuf structures             |
|  - Processors read and mutate data in place                 |
|  - Follows the OpenTelemetry data model exactly             |
+-------------------------------------------------------------+
           |
           v  (Exporter translates pdata to destination format)
[Destination Egress Bytes] (OTLP / Elasticsearch / Prometheus)

Why pdata Matters for Performance and Architecture

  1. In-Place Mutability: pdata structures allow processors to read and mutate attributes in place without serializing and deserializing payloads between pipeline steps (the pipeline clones data only when several consumers could otherwise modify the same batch).
  2. Protocol Decoupling (M×NM \times N Problem): Without pdata, supporting MM input formats and NN output backends would require writing M×NM \times N distinct protocol translation engines. With pdata, each receiver converts from its format into pdata (MM receivers), and each exporter converts from pdata into its target format (NN exporters), reducing system complexity to M+NM + N.
  3. No Conversion for OTLP: Because pdata wraps the OTLP protobuf structures, OTLP data enters and leaves the Collector without a translation step, which keeps CPU and garbage-collection cost low.

The Collector Component Lifecycle

The Collector runtime orchestrates startup and shutdown in a strict, sequential order to prevent telemetry loss during normal operation.

============================== STARTUP PHASE ==============================
1. Validate Config    --> Read YAML, verify syntax, check component factories
2. Start Extensions   --> Initialize health_check, auth, pprof, zpages
3. Start Exporters    --> Establish network sockets, connection pools, queues
4. Start Processors   --> Allocate in-memory buffers and limiter routines
5. Start Receivers    --> Open network listeners (4317, 4318); begin scraping

========================= RUNTIME PROCESSING PHASE ========================
Telemetry Ingestion   --> pdata pipelines stream batches; backpressure applies

============================= SHUTDOWN PHASE ==============================
1. Stop Receivers     --> Close network listeners; reject new ingress immediately
2. Flush Processors   --> Drain in-flight queues; force-flush pending batches
3. Drain Exporters    --> Transmit queued pdata to backends; close connections
4. Stop Extensions    --> Terminate health probes and profiling listeners

Startup Sequence Mechanics

When the Collector binary boots:

  1. Configuration Validation: The service reads the YAML configuration, checks all component factories, verifies that all referenced components exist in the binary, and validates configuration parameter types.
  2. Extensions Start First: Extensions (such as health_check, oidc, and pprof) initialize their runtime resources. This ensures that authentication mechanisms and liveness monitors are ready before network traffic begins.
  3. Exporters Start Second: Exporters establish network connections, configure authentication tokens, and initialize worker pools and sending queues. Destination backends must be reachable before data processing begins.
  4. Processors Start Third: Processors initialize internal memory counters, compile OTTL rules, and start background batch timers.
  5. Receivers Start Last: Receivers bind to operating system network ports (e.g., :4317, :4318) and start background scrape jobs. Starting receivers last guarantees that downstream pipelines are 100% operational before any client payload is accepted.

Graceful Shutdown Sequence Mechanics

When the Collector receives a termination signal (SIGTERM or SIGINT), it initiates a deterministic reverse-order shutdown:

  1. Receivers Stop First: Receivers immediately close network listening ports and cancel active scrape loops. This prevents new telemetry from entering the Collector and signals upstream clients (via TCP disconnect or HTTP 503) to fail over to another Collector instance.
  2. Processors Flush Second: Processors finish processing in-flight pdata batches. The batch processor immediately flushes all buffered items downstream, regardless of whether the configured batch size or timeout has been reached.
  3. Exporters Drain Third: Exporters drain their internal sending queues, dispatching all remaining batches to destination backends. Once queues are empty or the shutdown timeout expires, exporters close network sockets.
  4. Extensions Stop Last: Extensions terminate their endpoints, marking the health check as unhealthy and releasing diagnostic resources.

Core Distribution vs. Contrib Distribution vs. OCB

The OpenTelemetry project provides pre-compiled Collector binaries and a compilation utility to satisfy different operational requirements:

AttributeCore Distribution (otelcol)Contrib Distribution (otelcol-contrib)Custom Binary (builder / OCB)
MaintenanceOpenTelemetry Collector maintainers (collector-releases repository)Collector maintainers plus per-component code ownersYour organization / platform team
Component CountA curated set of about 27 componentsHundreds of receivers, processors, exporters, extensions, and connectorsExactly the components you select
Included Receiversotlp, hostmetrics, prometheus, jaeger, zipkin, kafka, nopEverything in Core plus many more (for example filelog, k8s_cluster, cloud-provider receivers)User-defined manifest list
Included Processorsbatch, memory_limiter, attributes, resource, span, filter, probabilistic_samplerCore's plus transform, k8sattributes, tail_sampling, resourcedetection, and moreUser-defined manifest list
Included Exportersotlp, otlphttp, debug, file, kafka, prometheus, prometheusremotewrite, zipkin, nopCore's plus vendor and storage exporters (for example loadbalancing, elasticsearch)User-defined manifest list
Extensions & Connectorshealth_check, pprof, zpages; forward connectorMany more, including file_storage, oidc, and the spanmetrics, count, and routing connectorsUser-defined manifest list
Binary FootprintSmallerLargestAs small as the component list allows
Vulnerability SurfaceSmaller than ContribBroadest (many third-party Go dependencies)Only the dependencies of the chosen components

OpenTelemetry Collector Builder (OCB)

In enterprise production environments, deploying the full Contrib binary often violates security and resource policies due to unnecessary third-party dependencies and larger container footprints. The OpenTelemetry Collector Builder (OCB) (builder) is an official CLI tool that compiles a tailored Collector binary containing only the exact components required.

Platform teams define a declarative builder-config.yaml manifest:

dist:
  name: otelcol-custom
  description: "Custom production OpenTelemetry Collector"
  output_path: ./dist
receivers:
  - gomod: go.opentelemetry.io/collector/receiver/otlpreceiver v0.110.0
processors:
  - gomod: go.opentelemetry.io/collector/processor/batchprocessor v0.110.0
  - gomod: go.opentelemetry.io/collector/processor/memorylimiterprocessor v0.110.0
  - gomod: github.com/open-telemetry/opentelemetry-collector-contrib/processor/transformprocessor v0.110.0
exporters:
  - gomod: go.opentelemetry.io/collector/exporter/otlpexporter v0.110.0
extensions:
  - gomod: github.com/open-telemetry/opentelemetry-collector-contrib/extension/healthcheckextension v0.110.0

Running builder --config=builder-config.yaml generates the Go source code for a main package, resolves the module dependencies, and compiles a binary containing only the listed components. Keep every component on the same Collector release version.

Loading diagram...
OpenTelemetry Collector Internal Architecture and Data Flow
Test Your Knowledge

A platform engineering team wants to decouple microservices from specific observability backends, eliminate client-side retry overhead, enforce centralized PII masking, and route telemetry to multiple storage destinations. Why does deploying an OpenTelemetry Collector provide a superior architectural design over direct-to-backend SDK export?

A

The Collector acts as a vendor-neutral proxy that abstracts data ingestion into an internal pluggable data model (pdata), offloading encoding, memory batching, retries, and multi-destination routing from application microservices.

B

The Collector replaces standard operating system network stacks with kernel-bypass sockets that accelerate network transmission by tenfold.

C

The Collector automatically instruments compiled application binaries at runtime without requiring any telemetry SDK or library dependencies.

D

The Collector converts all telemetry into proprietary compressed binary archives that eliminate the need for downstream observability backends.

Test Your Knowledge

During a scheduled Kubernetes rolling update, an OpenTelemetry Collector pod receives a SIGTERM termination signal. In what order must the Collector runtime execute its component lifecycle shutdown sequence to guarantee that in-flight telemetry is not lost?

A

Exporters immediately terminate connections, processors drop pending memory batches, extensions shut down, and receivers stop listening.

B

Receivers close network listeners and stop scraping, processors flush in-flight buffers and active batches, exporters drain sending queues to remote backends, and extensions terminate.

C

Processors halt processing immediately, receivers dump incoming network packets to local temporary files, extensions terminate, and exporters force-close connections.

D

Extensions terminate first, exporters immediately close network sockets, receivers reject new requests, and processors clear memory without flushing.

Test Your Knowledge

A cloud security architect requires an OpenTelemetry Collector deployment on edge nodes with strict memory limits (under 64MB) and zero tolerance for unused third-party dependencies with potential CVE vulnerabilities. The deployment requires only the OTLP receiver, the memory_limiter processor, the batch processor, the transform processor, and the OTLP exporter. Which distribution strategy should the engineering team adopt?

A

Deploy the official Core distribution (opentelemetry-collector), because it includes all core processors including the transform processor by default.

B

Deploy the official Contrib distribution (opentelemetry-collector-contrib) with all unused components disabled via environment variables.

C

Use the OpenTelemetry Collector Builder (OCB) to compile a custom minimal binary specifying only the exact required components in its manifest.

D

Deploy an uncompiled Go source repository directly inside the container and build the binary dynamically during container initialization.

Sections you finish are checked off in the contents.