13.1 Debugging Collector Pipelines & Telemetry Health
Key Takeaways
The OpenTelemetry Collector is self-observing, generating its own internal Prometheus metrics, structured logs, and distributed traces configured via the service.telemetry block.
Ingress pipeline health is measured via otelcol_receiver_accepted_* and otelcol_receiver_refused_* metrics; refused items indicate client decoding errors, bad formats, or upstream backpressure.
Processor metrics otelcol_processor_incoming_items and otelcol_processor_outgoing_items show how much each processor removes, while otelcol_process_runtime_heap_alloc_bytes shows heap pressure.
Egress metrics distinguish between queue overflow drops in the sending queue (otelcol_exporter_enqueue_failed_spans) and network transmission or retry exhaustion failures (otelcol_exporter_send_failed_spans).
The debug exporter (which replaced the removed logging exporter) with detailed verbosity shows full payloads, and the zpages extension (/debug/pipelinez, /debug/tracez) shows pipeline composition and internal spans.
13.1 Debugging Collector Pipelines & Telemetry Health
Quick Answer: The OpenTelemetry Collector provides comprehensive self-observability via the
service.telemetryblock, emitting internal metrics (by default a Prometheus endpoint at127.0.0.1:8888/metrics), structured logs, and internal traces. To diagnose telemetry loss, think in three pipeline stages: Ingress (otelcol_receiver_accepted_*for successful items vsotelcol_receiver_refused_*for client format errors or backpressure), Processing (otelcol_processor_incoming_itemsversusotelcol_processor_outgoing_itemsfor items removed by filters or samplers, andotelcol_process_runtime_heap_alloc_bytesfor heap memory pressure), and Egress (otelcol_exporter_sent_*for successes,otelcol_exporter_enqueue_failed_spansfor queue overflows during backend slowdowns, andotelcol_exporter_send_failed_spansfor network or backend errors after retries). For live payload inspection, use thedebugexporter withverbosity: detailed, and monitor internal component states with thezpagesextension (/debug/pipelinezfor pipeline composition and/debug/tracezfor internal spans).
In modern production architectures, the OpenTelemetry Collector acts as the central nervous system for telemetry data. When traces disappear from APM dashboards, metrics flatline, or log streams lag behind real time, platform engineers face an urgent operational challenge: is the failure originating within the application instrumentations, the network transport layer, the storage backend, or the Collector pipeline itself? To troubleshoot these failures authoritatively, engineers must understand how the Collector monitors its own internal health and emits self-diagnostic signals.
Monitoring the OpenTelemetry Collector Itself
The OpenTelemetry Collector is designed with a fundamental operational philosophy: the observability pipeline must itself be fully observable. Rather than operating as an opaque black box, the Collector runtime generates its own three pillars of internal telemetry:
- Internal Metrics: Quantitative counters, gauges, and histograms measuring component throughput, batch sizes, queue depths, memory allocations, and dropped data points.
- Structured Logs: Contextual, leveled operational events (
debug,info,warn,error) reporting component lifecycle state changes, connection terminations, and network errors. - Internal Traces: Distributed tracing spans recording the Collector's internal processing latency, measuring the microsecond duration of receiver unmarshaling, processor batch evaluations, and exporter network transmissions.
By monitoring these internal signals, site reliability engineers (SREs) can detect backpressure, locate misconfigured pipeline components, and prevent catastrophic telemetry loss during traffic surges.
Configuring Collector Internal Telemetry (service.telemetry)
Internal self-observability is configured within the top-level service.telemetry block of the Collector configuration file. This section governs how the Collector exposes its operational logs, internal metrics, and internal distributed traces.
service:
telemetry:
logs:
level: info # Options: debug, info, warn, error
development: false # Enables development mode (more verbose DPanic handling)
encoding: json # Options: json, console
output_paths: ["stdout"]
error_output_paths: ["stderr"]
metrics:
level: normal # Options: none, basic, normal (default), detailed
readers: # Without readers: Prometheus endpoint on 127.0.0.1:8888
- pull:
exporter:
prometheus:
host: 0.0.0.0
port: 8888
traces:
# Experimental internal tracing of Collector operations
processors: []
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp]
Metric Configuration & Verbosity Levels
By default, the Collector serves its metrics for Prometheus scraping at 127.0.0.1:8888/metrics, reachable only from inside the pod or host. To expose it, add a readers entry with host 0.0.0.0 (the older address setting is ignored since Collector v0.123.0), or push the metrics with a periodic OTLP reader. The service.telemetry.metrics.level setting controls how many internal series are emitted:
none: Disables all internal metric collection entirely.basic: Essential service telemetry: receiver accepted/refused counts, exporter sent/failed/enqueue-failed counts and queue size, processor incoming/outgoing items, and process memory and CPU.normal(Default): Adds standard indicators such as the batch processor's batch-size histogram and size- or timeout-triggered send counters.detailed: Adds the most verbose series, such as HTTP and RPC client/server duration histograms and batch sizes in bytes. Use it for targeted debugging.
Logging Levels
The service.telemetry.logs.level controls operational verbosity:
debug: Logs verbose internal state transitions, incoming batch headers, and detailed retry evaluations.info(Default): Logs component startup, pipeline registration, configuration changes, and graceful shutdown events.warn: Logs transient network timeouts, rate limit responses from backends, and recoverable pipeline retries.error: Logs unmarshaling failures, unrecoverable export drops, and component crash events.
Internal Collector Metrics
Knowing these metric names and their meanings makes troubleshooting fast. Metrics generated by Collector components share the prefix otelcol_ (a Prometheus endpoint may also add unit and _total suffixes). These metrics are categorized by pipeline phase: Ingress (Receiver), Processing (Processor), and Egress (Exporter).
1. Ingress (Receiver) Metrics
Receivers listen on network sockets or scrape endpoints to unmarshal incoming telemetry into internal OpenTelemetry data structures (pdata).
| Metric Name | Metric Type | Operational Meaning & Diagnostic Indicator |
|---|---|---|
otelcol_receiver_accepted_spans | Counter | Cumulative number of spans successfully decoded, parsed, and accepted into the Collector pipeline. |
otelcol_receiver_accepted_metric_points | Counter | Cumulative number of metric data points successfully received and accepted into the pipeline. |
otelcol_receiver_accepted_log_records | Counter | Cumulative number of log records successfully parsed and pushed downstream. |
otelcol_receiver_refused_spans | Counter | Number of spans rejected at ingress. Indicates client payload syntax errors, invalid protobuf, or receiver backpressure triggered by memory limits. |
otelcol_receiver_refused_metric_points | Counter | Number of metric data points rejected at ingress due to invalid structure or memory backpressure. |
otelcol_receiver_refused_log_records | Counter | Number of log records rejected at ingress. Spikes indicate upstream client issues or saturation. |
2. Processing (Processor) Metrics
Processors sit between receivers and exporters to batch, sample, mutate, filter, and monitor telemetry in memory.
| Metric Name | Metric Type | Operational Meaning & Diagnostic Indicator |
|---|---|---|
otelcol_processor_incoming_items | Counter | Items (spans, metric points, or log records) passed into each processor. |
otelcol_processor_outgoing_items | Counter | Items each processor emitted. Outgoing lower than incoming at a filter or sampler is intentional reduction. Refusals by memory_limiter show up instead as otelcol_receiver_refused_*. |
otelcol_process_runtime_heap_alloc_bytes | Gauge | Current Go runtime heap memory footprint in bytes. Essential for tracking memory limiter thresholds and preventing Linux OOM crashes. |
otelcol_processor_batch_batch_send_size | Histogram | Number of items in each batch the batch processor sends (normal level; a _bytes variant exists at detailed level). |
otelcol_processor_batch_timeout_trigger_send | Counter | Number of batches pushed downstream due to timer expiry rather than reaching max batch size. |
3. Egress (Exporter) Metrics
Exporters marshal batches into wire formats and transmit them to external observability backends over network sockets.
| Metric Name | Metric Type | Operational Meaning & Diagnostic Indicator |
|---|---|---|
otelcol_exporter_sent_spans | Counter | Number of spans successfully acknowledged by destination backends (HTTP 200/202, gRPC OK). |
otelcol_exporter_sent_metric_points | Counter | Number of metric data points successfully received and acknowledged by backends. |
otelcol_exporter_sent_log_records | Counter | Number of log records successfully transmitted to destination endpoints. |
otelcol_exporter_enqueue_failed_spans | Counter | Queue overflow drops! Spans dropped because the exporter's in-memory sending_queue reached 100% capacity due to backend latency. |
otelcol_exporter_enqueue_failed_metric_points | Counter | Metric points dropped due to saturated exporter sending queues. |
otelcol_exporter_enqueue_failed_log_records | Counter | Log records dropped due to saturated exporter sending queues. |
otelcol_exporter_send_failed_spans | Counter | Transmission drops! Spans permanently discarded after exhausting all retry_on_failure attempts (network timeouts, HTTP 4xx/5xx). |
otelcol_exporter_send_failed_metric_points | Counter | Metric points discarded following unrecoverable network or remote backend failure. |
otelcol_exporter_send_failed_log_records | Counter | Log records discarded following unrecoverable network or remote backend failure. |
otelcol_exporter_queue_size | Gauge | Current number of batches waiting inside the exporter's sending queue. |
otelcol_exporter_queue_capacity | Gauge | Maximum batch capacity configured for the exporter's sending queue. |
Diagnostic Tools for Pipeline Debugging
When standard metrics alert engineers to dropped telemetry or misconfigured attributes, the Collector provides specialized diagnostic tools to inspect live payloads and pipeline state without interrupting traffic.
The debug Exporter (Superseding logging)
Historically, the Collector provided a logging exporter to print telemetry to stdout. It was deprecated in Collector v0.86.0, when the debug exporter was introduced, and removed in v0.111.0.
The debug exporter is designed for rapid pipeline troubleshooting. It prints uncompressed, decoded telemetry structures directly to the Collector's standard output stream:
exporters:
debug:
verbosity: detailed # Options: basic, normal, detailed
sampling_initial: 5 # Number of messages logged initially per second
sampling_thereafter: 200 # Sampling rate thereafter to prevent log flooding
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlp, debug]
Verbosity Levels in the debug Exporter
basic: Outputs a single summary line per batch, detailing the number of spans, metric points, or log records without showing attribute payloads.normal: Outputs high-level metadata, including trace IDs, span IDs, span names, and instrumentation scope details.detailed: Outputs the entire telemetry payload in human-readable format, including every resource attribute (service.name,host.id), span attribute (http.request.method,db.query.text), span events, exception stack traces, status codes, and links. Usedetailedto verify whether OTTL attribute mutations, PII masking, or schema translations have succeeded.
The zpages Extension
Originating from Google's internal production debugging systems, the zpages extension provides live in-process HTML diagnostic pages hosted on port 55679 (http://localhost:55679/debug/...). It requires zero external storage backends and operates completely in-memory.
extensions:
zpages:
endpoint: 0.0.0.0:55679
service:
extensions: [zpages]
pipelines:
traces:
receivers: [otlp]
processors: [batch]
exporters: [otlp]
Key zpages Diagnostic Endpoints
/debug/tracez: Inspects in-memory traces of the Collector's internal operations. It displays running active spans, latency percentiles across internal execution stages, and recent sampled error traces within the Collector runtime./debug/pipelinez: Lists each running pipeline with its type, whether it mutates data, and its receivers, processors, and exporters. It shows composition, not throughput; use the internal metrics for counts./debug/servicez,/debug/extensionz, and/debug/featurez: Build information, active extensions, and feature-gate status.
Other Diagnostic Extensions
health_check: Runs on port13133(http://localhost:13133/). Responds with HTTP 200 OK when all pipeline components are healthy, serving as Kubernetes liveness and readiness probes.pprof: Runs on port1777(http://localhost:1777/debug/pprof/). Exposes Go runtime profiling endpoints for analyzing CPU profiles, heap memory allocations, mutex contention, and goroutine leaks.
Practical Runbook: Diagnosing Dropped Telemetry
When telemetry fails to reach destination backends, engineers must follow a systematic, metric-driven triage process to isolate the exact point of failure.
+-------------------------------------------------------------------------+
| Collector Telemetry Triage Runbook |
+-------------------------------------------------------------------------+
Step 1: Ingress Triage (Receiver Metrics)
- Check otelcol_receiver_refused_spans > 0
* If YES: Client format error, unmarshaling bug, or memory backpressure.
* If NO: Receiver accepted data; move to Step 2.
Step 2: Processing Triage (Processor Metrics)
- Compare otelcol_processor_incoming_items with outgoing_items
* Gap at filter / tail_sampling: intentional reduction.
* Receiver refusals + high heap_alloc_bytes: memory_limiter.
* No gap anywhere: processors passed data; go to Step 3.
Step 3: Egress Triage (Exporter Metrics)
- Check otelcol_exporter_enqueue_failed_spans > 0
* If YES: Backend is slow/rate-limiting; sending_queue overflowed.
- Check otelcol_exporter_send_failed_spans > 0
* If YES: Network timeout, DNS failure, 4xx/5xx after retry exhaustion.
Diagnostic Decision Matrix
| Failure Stage | Primary Metric Indicator | Root Cause | Immediate Action |
|---|---|---|---|
| Ingress Refusal | otelcol_receiver_refused_* > 0 | Payload syntax error, mismatched TLS, or memory backpressure | Inspect receiver error logs; check memory_limiter thresholds; verify client protobuf format. |
| Memory Backpressure | otelcol_receiver_refused_* spiking with high heap_alloc_bytes | Memory limiter triggered; receivers return HTTP 429/503 | Scale out Collector replicas (HPA); increase container RAM limits; lower client ingestion rate. |
| Intentional Sampling | otelcol_processor_outgoing_items below incoming_items at a filter or sampler, normal memory | Expected behavior of filter or tail-sampling processors | Verify sampling policy rules; confirm whether dropped spans match expected non-error routes. |
| Egress Queue Overflow | otelcol_exporter_enqueue_failed_* > 0 | Downstream backend saturation, rate limits, or undersized queue | Increase sending_queue.queue_size; scale downstream backend; increase exporter worker concurrency (num_consumers). |
| Network Export Loss | otelcol_exporter_send_failed_* > 0 | Network partition, expired auth token, or backend HTTP 5xx | Inspect exporter error logs; verify API credentials/mTLS certs; verify DNS resolution and egress firewall rules. |
A site reliability engineering team notices that several distributed traces are missing from their Jaeger backend following a traffic spike. In the Collector's Prometheus metrics endpoint (:8888/metrics), otelcol_receiver_accepted_spans is incrementing normally and otelcol_processor_outgoing_items matches otelcol_processor_incoming_items for every processor. However, otelcol_exporter_enqueue_failed_spans is spiking dramatically. What is the fundamental root cause of this telemetry loss?
Upstream microservices are sending malformed protobuf payloads rejected by the receiver
The memory_limiter processor has tripped due to excessive heap usage and is actively shedding spans
The network DNS resolver is returning NXDOMAIN errors for the external backend endpoint
The downstream backend is experiencing high latency or rate limits, causing the exporter's in-memory sending_queue to saturate and drop newly arriving batches
An observability engineer is configuring an OpenTelemetry Transformation Language (OTTL) statement inside the transform processor to mask customer credit card numbers in span attributes. Before deploying this pipeline to production, the engineer needs to verify the exact string mutations on stdout by inspecting the full attribute map, span events, and status codes of every span passing through the Collector. Which exporter and configuration setting should be used?
The debug exporter configured with verbosity set to detailed
The legacy logging exporter configured with log_level set to info
The zpages extension accessed via the /debug/pipelinez HTTP endpoint
The health_check extension with verbose logging enabled on port 13133
During a flash-sale event, client microservices begin receiving HTTP 503 (Service Unavailable) and gRPC Unavailable status codes when transmitting telemetry to an OpenTelemetry Collector gateway. Monitoring shows a sharp rise in otelcol_receiver_refused_spans while otelcol_process_runtime_heap_alloc_bytes is approaching the configured 80% threshold. What mechanism is triggering the receiver to refuse incoming spans?
The OTLP receiver certificate has expired, causing TLS handshakes to terminate abruptly
The memory_limiter processor has detected memory pressure and signaled the receiver to refuse incoming requests with backpressure to prevent a fatal process crash
The batch processor has filled its internal buffer and crashed the Go garbage collector thread
Upstream microservices are transmitting spans using an incompatible W3C traceparent version
Sections you finish are checked off in the contents.