13.1 Debugging Collector Pipelines & Telemetry Health

Key Takeaways

  • The OpenTelemetry Collector is self-observing, generating its own internal Prometheus metrics, structured logs, and distributed traces configured via the service.telemetry block.

  • Ingress pipeline health is measured via otelcol_receiver_accepted_* and otelcol_receiver_refused_* metrics; refused items indicate client decoding errors, bad formats, or upstream backpressure.

  • Processor metrics otelcol_processor_incoming_items and otelcol_processor_outgoing_items show how much each processor removes, while otelcol_process_runtime_heap_alloc_bytes shows heap pressure.

  • Egress metrics distinguish between queue overflow drops in the sending queue (otelcol_exporter_enqueue_failed_spans) and network transmission or retry exhaustion failures (otelcol_exporter_send_failed_spans).

  • The debug exporter (which replaced the removed logging exporter) with detailed verbosity shows full payloads, and the zpages extension (/debug/pipelinez, /debug/tracez) shows pipeline composition and internal spans.

Last updated: September 2026

13.1 Debugging Collector Pipelines & Telemetry Health

Quick Answer: The OpenTelemetry Collector provides comprehensive self-observability via the service.telemetry block, emitting internal metrics (by default a Prometheus endpoint at 127.0.0.1:8888/metrics), structured logs, and internal traces. To diagnose telemetry loss, think in three pipeline stages: Ingress (otelcol_receiver_accepted_* for successful items vs otelcol_receiver_refused_* for client format errors or backpressure), Processing (otelcol_processor_incoming_items versus otelcol_processor_outgoing_items for items removed by filters or samplers, and otelcol_process_runtime_heap_alloc_bytes for heap memory pressure), and Egress (otelcol_exporter_sent_* for successes, otelcol_exporter_enqueue_failed_spans for queue overflows during backend slowdowns, and otelcol_exporter_send_failed_spans for network or backend errors after retries). For live payload inspection, use the debug exporter with verbosity: detailed, and monitor internal component states with the zpages extension (/debug/pipelinez for pipeline composition and /debug/tracez for internal spans).

In modern production architectures, the OpenTelemetry Collector acts as the central nervous system for telemetry data. When traces disappear from APM dashboards, metrics flatline, or log streams lag behind real time, platform engineers face an urgent operational challenge: is the failure originating within the application instrumentations, the network transport layer, the storage backend, or the Collector pipeline itself? To troubleshoot these failures authoritatively, engineers must understand how the Collector monitors its own internal health and emits self-diagnostic signals.


Monitoring the OpenTelemetry Collector Itself

The OpenTelemetry Collector is designed with a fundamental operational philosophy: the observability pipeline must itself be fully observable. Rather than operating as an opaque black box, the Collector runtime generates its own three pillars of internal telemetry:

  1. Internal Metrics: Quantitative counters, gauges, and histograms measuring component throughput, batch sizes, queue depths, memory allocations, and dropped data points.
  2. Structured Logs: Contextual, leveled operational events (debug, info, warn, error) reporting component lifecycle state changes, connection terminations, and network errors.
  3. Internal Traces: Distributed tracing spans recording the Collector's internal processing latency, measuring the microsecond duration of receiver unmarshaling, processor batch evaluations, and exporter network transmissions.

By monitoring these internal signals, site reliability engineers (SREs) can detect backpressure, locate misconfigured pipeline components, and prevent catastrophic telemetry loss during traffic surges.


Configuring Collector Internal Telemetry (service.telemetry)

Internal self-observability is configured within the top-level service.telemetry block of the Collector configuration file. This section governs how the Collector exposes its operational logs, internal metrics, and internal distributed traces.

service:
  telemetry:
    logs:
      level: info            # Options: debug, info, warn, error
      development: false     # Enables development mode (more verbose DPanic handling)
      encoding: json         # Options: json, console
      output_paths: ["stdout"]
      error_output_paths: ["stderr"]
    metrics:
      level: normal          # Options: none, basic, normal (default), detailed
      readers:               # Without readers: Prometheus endpoint on 127.0.0.1:8888
        - pull:
            exporter:
              prometheus:
                host: 0.0.0.0
                port: 8888
    traces:
      # Experimental internal tracing of Collector operations
      processors: []
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp]

Metric Configuration & Verbosity Levels

By default, the Collector serves its metrics for Prometheus scraping at 127.0.0.1:8888/metrics, reachable only from inside the pod or host. To expose it, add a readers entry with host 0.0.0.0 (the older address setting is ignored since Collector v0.123.0), or push the metrics with a periodic OTLP reader. The service.telemetry.metrics.level setting controls how many internal series are emitted:

  • none: Disables all internal metric collection entirely.
  • basic: Essential service telemetry: receiver accepted/refused counts, exporter sent/failed/enqueue-failed counts and queue size, processor incoming/outgoing items, and process memory and CPU.
  • normal (Default): Adds standard indicators such as the batch processor's batch-size histogram and size- or timeout-triggered send counters.
  • detailed: Adds the most verbose series, such as HTTP and RPC client/server duration histograms and batch sizes in bytes. Use it for targeted debugging.

Logging Levels

The service.telemetry.logs.level controls operational verbosity:

  • debug: Logs verbose internal state transitions, incoming batch headers, and detailed retry evaluations.
  • info (Default): Logs component startup, pipeline registration, configuration changes, and graceful shutdown events.
  • warn: Logs transient network timeouts, rate limit responses from backends, and recoverable pipeline retries.
  • error: Logs unmarshaling failures, unrecoverable export drops, and component crash events.

Internal Collector Metrics

Knowing these metric names and their meanings makes troubleshooting fast. Metrics generated by Collector components share the prefix otelcol_ (a Prometheus endpoint may also add unit and _total suffixes). These metrics are categorized by pipeline phase: Ingress (Receiver), Processing (Processor), and Egress (Exporter).

1. Ingress (Receiver) Metrics

Receivers listen on network sockets or scrape endpoints to unmarshal incoming telemetry into internal OpenTelemetry data structures (pdata).

Metric NameMetric TypeOperational Meaning & Diagnostic Indicator
otelcol_receiver_accepted_spansCounterCumulative number of spans successfully decoded, parsed, and accepted into the Collector pipeline.
otelcol_receiver_accepted_metric_pointsCounterCumulative number of metric data points successfully received and accepted into the pipeline.
otelcol_receiver_accepted_log_recordsCounterCumulative number of log records successfully parsed and pushed downstream.
otelcol_receiver_refused_spansCounterNumber of spans rejected at ingress. Indicates client payload syntax errors, invalid protobuf, or receiver backpressure triggered by memory limits.
otelcol_receiver_refused_metric_pointsCounterNumber of metric data points rejected at ingress due to invalid structure or memory backpressure.
otelcol_receiver_refused_log_recordsCounterNumber of log records rejected at ingress. Spikes indicate upstream client issues or saturation.

2. Processing (Processor) Metrics

Processors sit between receivers and exporters to batch, sample, mutate, filter, and monitor telemetry in memory.

Metric NameMetric TypeOperational Meaning & Diagnostic Indicator
otelcol_processor_incoming_itemsCounterItems (spans, metric points, or log records) passed into each processor.
otelcol_processor_outgoing_itemsCounterItems each processor emitted. Outgoing lower than incoming at a filter or sampler is intentional reduction. Refusals by memory_limiter show up instead as otelcol_receiver_refused_*.
otelcol_process_runtime_heap_alloc_bytesGaugeCurrent Go runtime heap memory footprint in bytes. Essential for tracking memory limiter thresholds and preventing Linux OOM crashes.
otelcol_processor_batch_batch_send_sizeHistogramNumber of items in each batch the batch processor sends (normal level; a _bytes variant exists at detailed level).
otelcol_processor_batch_timeout_trigger_sendCounterNumber of batches pushed downstream due to timer expiry rather than reaching max batch size.

3. Egress (Exporter) Metrics

Exporters marshal batches into wire formats and transmit them to external observability backends over network sockets.

Metric NameMetric TypeOperational Meaning & Diagnostic Indicator
otelcol_exporter_sent_spansCounterNumber of spans successfully acknowledged by destination backends (HTTP 200/202, gRPC OK).
otelcol_exporter_sent_metric_pointsCounterNumber of metric data points successfully received and acknowledged by backends.
otelcol_exporter_sent_log_recordsCounterNumber of log records successfully transmitted to destination endpoints.
otelcol_exporter_enqueue_failed_spansCounterQueue overflow drops! Spans dropped because the exporter's in-memory sending_queue reached 100% capacity due to backend latency.
otelcol_exporter_enqueue_failed_metric_pointsCounterMetric points dropped due to saturated exporter sending queues.
otelcol_exporter_enqueue_failed_log_recordsCounterLog records dropped due to saturated exporter sending queues.
otelcol_exporter_send_failed_spansCounterTransmission drops! Spans permanently discarded after exhausting all retry_on_failure attempts (network timeouts, HTTP 4xx/5xx).
otelcol_exporter_send_failed_metric_pointsCounterMetric points discarded following unrecoverable network or remote backend failure.
otelcol_exporter_send_failed_log_recordsCounterLog records discarded following unrecoverable network or remote backend failure.
otelcol_exporter_queue_sizeGaugeCurrent number of batches waiting inside the exporter's sending queue.
otelcol_exporter_queue_capacityGaugeMaximum batch capacity configured for the exporter's sending queue.

Diagnostic Tools for Pipeline Debugging

When standard metrics alert engineers to dropped telemetry or misconfigured attributes, the Collector provides specialized diagnostic tools to inspect live payloads and pipeline state without interrupting traffic.

The debug Exporter (Superseding logging)

Historically, the Collector provided a logging exporter to print telemetry to stdout. It was deprecated in Collector v0.86.0, when the debug exporter was introduced, and removed in v0.111.0.

The debug exporter is designed for rapid pipeline troubleshooting. It prints uncompressed, decoded telemetry structures directly to the Collector's standard output stream:

exporters:
  debug:
    verbosity: detailed  # Options: basic, normal, detailed
    sampling_initial: 5  # Number of messages logged initially per second
    sampling_thereafter: 200 # Sampling rate thereafter to prevent log flooding

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp, debug]

Verbosity Levels in the debug Exporter

  • basic: Outputs a single summary line per batch, detailing the number of spans, metric points, or log records without showing attribute payloads.
  • normal: Outputs high-level metadata, including trace IDs, span IDs, span names, and instrumentation scope details.
  • detailed: Outputs the entire telemetry payload in human-readable format, including every resource attribute (service.name, host.id), span attribute (http.request.method, db.query.text), span events, exception stack traces, status codes, and links. Use detailed to verify whether OTTL attribute mutations, PII masking, or schema translations have succeeded.

The zpages Extension

Originating from Google's internal production debugging systems, the zpages extension provides live in-process HTML diagnostic pages hosted on port 55679 (http://localhost:55679/debug/...). It requires zero external storage backends and operates completely in-memory.

extensions:
  zpages:
    endpoint: 0.0.0.0:55679

service:
  extensions: [zpages]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlp]

Key zpages Diagnostic Endpoints

  • /debug/tracez: Inspects in-memory traces of the Collector's internal operations. It displays running active spans, latency percentiles across internal execution stages, and recent sampled error traces within the Collector runtime.
  • /debug/pipelinez: Lists each running pipeline with its type, whether it mutates data, and its receivers, processors, and exporters. It shows composition, not throughput; use the internal metrics for counts.
  • /debug/servicez, /debug/extensionz, and /debug/featurez: Build information, active extensions, and feature-gate status.

Other Diagnostic Extensions

  • health_check: Runs on port 13133 (http://localhost:13133/). Responds with HTTP 200 OK when all pipeline components are healthy, serving as Kubernetes liveness and readiness probes.
  • pprof: Runs on port 1777 (http://localhost:1777/debug/pprof/). Exposes Go runtime profiling endpoints for analyzing CPU profiles, heap memory allocations, mutex contention, and goroutine leaks.

Practical Runbook: Diagnosing Dropped Telemetry

When telemetry fails to reach destination backends, engineers must follow a systematic, metric-driven triage process to isolate the exact point of failure.

+-------------------------------------------------------------------------+
|                    Collector Telemetry Triage Runbook                   |
+-------------------------------------------------------------------------+
  Step 1: Ingress Triage (Receiver Metrics)
  - Check otelcol_receiver_refused_spans > 0
    * If YES: Client format error, unmarshaling bug, or memory backpressure.
    * If NO: Receiver accepted data; move to Step 2.
  
  Step 2: Processing Triage (Processor Metrics)
  - Compare otelcol_processor_incoming_items with outgoing_items
    * Gap at filter / tail_sampling: intentional reduction.
    * Receiver refusals + high heap_alloc_bytes: memory_limiter.
    * No gap anywhere: processors passed data; go to Step 3.
  
  Step 3: Egress Triage (Exporter Metrics)
  - Check otelcol_exporter_enqueue_failed_spans > 0
    * If YES: Backend is slow/rate-limiting; sending_queue overflowed.
  - Check otelcol_exporter_send_failed_spans > 0
    * If YES: Network timeout, DNS failure, 4xx/5xx after retry exhaustion.

Diagnostic Decision Matrix

Failure StagePrimary Metric IndicatorRoot CauseImmediate Action
Ingress Refusalotelcol_receiver_refused_* > 0Payload syntax error, mismatched TLS, or memory backpressureInspect receiver error logs; check memory_limiter thresholds; verify client protobuf format.
Memory Backpressureotelcol_receiver_refused_* spiking with high heap_alloc_bytesMemory limiter triggered; receivers return HTTP 429/503Scale out Collector replicas (HPA); increase container RAM limits; lower client ingestion rate.
Intentional Samplingotelcol_processor_outgoing_items below incoming_items at a filter or sampler, normal memoryExpected behavior of filter or tail-sampling processorsVerify sampling policy rules; confirm whether dropped spans match expected non-error routes.
Egress Queue Overflowotelcol_exporter_enqueue_failed_* > 0Downstream backend saturation, rate limits, or undersized queueIncrease sending_queue.queue_size; scale downstream backend; increase exporter worker concurrency (num_consumers).
Network Export Lossotelcol_exporter_send_failed_* > 0Network partition, expired auth token, or backend HTTP 5xxInspect exporter error logs; verify API credentials/mTLS certs; verify DNS resolution and egress firewall rules.
Loading diagram...
OpenTelemetry Collector Pipeline Telemetry Lifecycle and Metrics
Test Your Knowledge

A site reliability engineering team notices that several distributed traces are missing from their Jaeger backend following a traffic spike. In the Collector's Prometheus metrics endpoint (:8888/metrics), otelcol_receiver_accepted_spans is incrementing normally and otelcol_processor_outgoing_items matches otelcol_processor_incoming_items for every processor. However, otelcol_exporter_enqueue_failed_spans is spiking dramatically. What is the fundamental root cause of this telemetry loss?

A

Upstream microservices are sending malformed protobuf payloads rejected by the receiver

B

The memory_limiter processor has tripped due to excessive heap usage and is actively shedding spans

C

The network DNS resolver is returning NXDOMAIN errors for the external backend endpoint

D

The downstream backend is experiencing high latency or rate limits, causing the exporter's in-memory sending_queue to saturate and drop newly arriving batches

Test Your Knowledge

An observability engineer is configuring an OpenTelemetry Transformation Language (OTTL) statement inside the transform processor to mask customer credit card numbers in span attributes. Before deploying this pipeline to production, the engineer needs to verify the exact string mutations on stdout by inspecting the full attribute map, span events, and status codes of every span passing through the Collector. Which exporter and configuration setting should be used?

A

The debug exporter configured with verbosity set to detailed

B

The legacy logging exporter configured with log_level set to info

C

The zpages extension accessed via the /debug/pipelinez HTTP endpoint

D

The health_check extension with verbose logging enabled on port 13133

Test Your Knowledge

During a flash-sale event, client microservices begin receiving HTTP 503 (Service Unavailable) and gRPC Unavailable status codes when transmitting telemetry to an OpenTelemetry Collector gateway. Monitoring shows a sharp rise in otelcol_receiver_refused_spans while otelcol_process_runtime_heap_alloc_bytes is approaching the configured 80% threshold. What mechanism is triggering the receiver to refuse incoming spans?

A

The OTLP receiver certificate has expired, causing TLS handshakes to terminate abruptly

B

The memory_limiter processor has detected memory pressure and signaled the receiver to refuse incoming requests with backpressure to prevent a fatal process crash

C

The batch processor has filled its internal buffer and crashed the Go garbage collector thread

D

Upstream microservices are transmitting spans using an incompatible W3C traceparent version

Sections you finish are checked off in the contents.