10.3 Standard Exporters & Delivery Guarantees
Key Takeaways
Exporters serve as the egress gateway of the OpenTelemetry Collector, translating internal pdata into external wire protocols and transmitting data to storage backends.
The canonical otlp (gRPC on port 4317) and otlphttp (HTTP on port 4318) exporters provide standardized, high-performance telemetry delivery with configurable TLS, compression, and authentication.
The debug exporter writes telemetry to the Collector's console output with verbosity basic (default), normal, or detailed, and replaces the removed logging exporter.
Exporter resilience is powered by sending_queue (in-memory or persistent disk-backed FIFO buffering) and retry_on_failure (exponential backoff for transient errors like HTTP 429/503 and gRPC UNAVAILABLE).
With retries enabled, the Collector delivers queued data at least once across transient failures, which may create duplicates at the destination; permanent errors (HTTP 400, 401) are dropped immediately, and only a persistent queue protects queued data from a crash.
10.3 Standard Exporters & Delivery Guarantees
Quick Answer: Exporters are the egress components of the OpenTelemetry Collector that translate internal
pdatastructures into destination formats and transmit them over the network to external observability platforms. The canonical exporters areotlp(gRPC) andotlphttp(HTTP), whiledebug(which supersedes the deprecatedloggingexporter) outputs telemetry to stdout for local inspection. Exporter reliability relies on two foundational features:sending_queue(a FIFO queue that decouples ingestion from export latency, with optional disk-backed persistent storage) andretry_on_failure(exponential backoff for transient errors such as HTTP 429/503 and gRPCUNAVAILABLE). With retries enabled, delivery of queued data is at-least-once, so network retries can introduce duplicates at the backend, while non-retryable client errors (HTTP 400, 401) are dropped immediately to prevent queue poison pills; only a persistent queue survives a crash.
In an enterprise observability pipeline, telemetry must be routed reliably to disparate destinations: distributed tracing backends (such as Jaeger, Tempo, or AWS X-Ray), metric time-series databases (such as Prometheus, Mimir, or Google Cloud Monitoring), and log search clusters (such as Elasticsearch or OpenSearch). Exporters decouple internal Collector processing from the network transport requirements of these external systems.
Core Exporters
Know the canonical exporters, their key configuration parameters, verbosity levels, and operational behavior.
1. The otlp & otlphttp Exporters (The Canonical Standards)
The primary method for exporting telemetry from an OpenTelemetry Collector to downstream Collectors or OTLP-native observability platforms is via the OpenTelemetry Protocol exporters:
otlpExporter (gRPC): Transmits binary Protobuf payloads over HTTP/2 multiplexed gRPC connections (default destination port4317). It provides maximum network throughput, low CPU overhead, and persistent connection multiplexing.otlphttpExporter (HTTP): Transmits Protobuf or JSON payloads over standard HTTP/1.1 or HTTP/2 POST requests (default destination port4318, sending to paths/v1/traces,/v1/metrics, or/v1/logs). It is ideal when routing telemetry through corporate HTTP proxies, cloud Application Load Balancers (ALBs), or API gateways that do not support gRPC streaming.
exporters:
# Canonical OTLP gRPC exporter
otlp/backend:
endpoint: tempo.internal.net:4317
tls:
ca_file: /etc/otel/certs/ca.pem
cert_file: /etc/otel/certs/client.pem
key_file: /etc/otel/certs/client.key
compression: gzip # Supported: gzip, zstd, none
timeout: 10s
headers:
Authorization: "Bearer ${env:API_KEY}"
# Canonical OTLP HTTP exporter
otlphttp/external:
endpoint: https://otlp.monitoring.vendor.com
encoding: proto # Supported: proto, json
compression: gzip
timeout: 15s
2. The debug Exporter (Local Troubleshooting)
During development, staging verification, or production debugging, operators must inspect raw telemetry items entering or exiting a pipeline. The debug exporter prints telemetry payloads directly to the Collector's standard output (stdout) or standard error (stderr).
Important
The historical logging exporter was deprecated in Collector v0.86.0 in favor of the debug exporter and removed in v0.111.0. Use debug for console output.
exporters:
debug:
verbosity: detailed
sampling_initial: 5
sampling_thereafter: 100
Debug Verbosity Levels
basic(default): Emits concise one-line summaries displaying the count of processed spans, metric data points, or log records and their associated service names. Extremely lightweight; suitable for low-overhead production sanity checks.normal: Emits key metadata including span names, trace IDs, timestamps, and primary attributes. Provides a balance between visibility and log volume.detailed: Emits the complete, unabridged internalpdatastructure, printing every span event, link, attribute key-value pair, resource attribute block, and raw metric histogram bucket. Warning: Generates massive log volume; should only be enabled in testing or isolated staging environments.
3. The prometheus Exporter (Pull-Based Metric Egress)
While most exporters push data outward, the prometheus exporter hosts an HTTP server that exposes a /metrics scrape endpoint. External Prometheus or VictoriaMetrics servers can then scrape the Collector on a configured interval.
- Default Endpoint:
0.0.0.0:8889/metrics - Metric Mapping: Translates OpenTelemetry
pmetric.Metricsinto Prometheus exposition format. - Distinction: Do not confuse the
prometheusexporter (a pull server hosted on port 8889) with theprometheusremotewriteexporter (a push client that dispatches metrics via HTTP POST to a remote Prometheus Remote Write endpoint).
Exporter Resilience & Reliability Features
In enterprise networks, downstream observability backends experience network partitions, transient rate-limiting spikes, and maintenance outages. Without robust resilience mechanisms, transient backend disruptions would cause immediate telemetry loss or trigger cascading backpressure that stalls upstream microservices.
The OpenTelemetry Collector provides two tightly integrated resilience mechanisms on all standard push exporters: sending_queue and retry_on_failure.
exporters:
otlp/resilient:
endpoint: apm.internal.net:4317
sending_queue:
enabled: true
num_consumers: 16
queue_size: 10000
storage: file_storage # Optional: enables persistent disk-backed queuing
retry_on_failure:
enabled: true
initial_interval: 5s
max_interval: 30s
max_elapsed_time: 5m
1. sending_queue (In-Memory & Persistent Buffering)
The sending_queue decouples pipeline ingestion from exporter network transmission using a First-In, First-Out (FIFO) queue:
enabled: Activates the queuing mechanism (default:true).num_consumers: The number of concurrent worker goroutines that pull batches from the queue and execute network RPC calls to the destination backend (default:10). Increasing consumer workers improves export throughput over high-latency WAN connections.queue_size: The queue's capacity (default:1000, counted in requests, i.e. incoming batches, because the defaultsizerisrequests;itemsandbytessizers also exist). When the queue is full, new data is rejected (counted byotelcol_exporter_enqueue_failed_*), unlessblock_on_overflow: truemakes the caller wait.wait_for_result(defaultfalse): With the default, the pipeline reports success as soon as data is queued; the upstream client is not told whether the backend later accepts it.batch: Newer releases can batch inside the exporter (sending_queue::batch, disabled by default;batch: {}enables it with a 200 ms flush timeout and a minimum of 8192 items), an alternative to the batch processor.storage(Persistent Queuing): References a storage extension (such asfile_storage). When configured, queued batches are written to durable local disk storage rather than RAM. If the Collector pod crashes, runs out of memory, or restarts during a network partition, buffered telemetry is restored from disk upon boot and dispatched without data loss.
2. retry_on_failure (Exponential Backoff)
When an exporter attempts to deliver a batch and receives a network timeout or transient server error, retry_on_failure automatically schedules retries using an exponential backoff algorithm with jitter:
enabled: Activates automated retries (default:true).initial_interval: The duration to wait before executing the first retry attempt (default:5s).max_interval: The upper ceiling on the backoff duration (default:30s). Each subsequent retry increases the wait interval exponentially () plus random jitter to prevent thundering herd problems.max_elapsed_time: The maximum total cumulative time spent retrying a single batch (default:5m). If the backend remains unreachable aftermax_elapsed_timeexpires, the batch is permanently discarded, the failure is counted inotelcol_exporter_send_failed_spans(or the metric-point and log-record equivalents), and the worker moves to the next batch.
3. Retryable vs. Non-Retryable Errors (Poison Pill Avoidance)
Exporters sort backend HTTP and gRPC status codes into two groups:
| Error Classification | Status Codes / Conditions | Operational Behavior | Architectural Rationale |
|---|---|---|---|
| Retryable Errors (Transient) | HTTP 429 (Too Many Requests), HTTP 502 (Bad Gateway), HTTP 503 (Service Unavailable), HTTP 504 (Gateway Timeout); gRPC UNAVAILABLE (14), gRPC RESOURCE_EXHAUSTED (8) when the server indicates it can recover. | Exporter retains batch in sending_queue and retries with exponential backoff until max_elapsed_time is reached. | The backend is temporarily overloaded or experiencing a network hiccup; the payload is valid and will likely succeed once the backend recovers. |
| Non-Retryable Errors (Permanent) | HTTP 400 (Bad Request), HTTP 401 (Unauthorized), HTTP 403 (Forbidden), HTTP 404 (Not Found); gRPC INVALID_ARGUMENT (3), gRPC UNAUTHENTICATED (16), gRPC PERMISSION_DENIED (7). | Exporter immediately drops the batch, logs a fatal transmission error, and advances to the next batch. | The payload is malformed or credentials are invalid. Retrying will never succeed and would create a poison pill that blocks worker threads and exhausts queue memory. |
Caution
The poison-pill danger: If an exporter were to retry a 400 Bad Request or 401 Unauthorized error, the malformed batch would cycle through retry_on_failure indefinitely. This would pin worker consumer threads, fill the sending_queue to capacity, and trigger cascading backpressure that halts valid telemetry ingestion across the entire Collector! For this reason, all non-retryable errors are dropped immediately.
Delivery Guarantees: At-Least-Once vs. Best-Effort
Understanding the delivery guarantees provided by the OpenTelemetry Collector is vital for architecting resilient distributed systems:
At-Least-Once Delivery Semantics
With retry_on_failure enabled, the Collector's delivery to the backend is at-least-once for data it still holds: a batch leaves the queue only after a success or a permanent error. How far that guarantee extends depends on the queue:
- With the default
wait_for_result: false, the receiver acknowledges the client once the data is accepted into the pipeline and queued, not when the backend confirms it. An in-memory queue loses whatever it holds if the process crashes; only a persistent queue (storage: file_storage) carries queued data across a restart. - A batch is only removed from the queue after the downstream backend responds with a success code (
200 OKor gRPCOK) or a permanent error. - Duplicate Generation Scenario: Suppose the Collector transmits a batch of 5,000 spans to an APM backend. The backend successfully processes and writes the spans to its database, but a network router blips before the HTTP
200 OKresponse packet reaches the Collector. The Collector experiences a socket timeout, assumes delivery failed, and retries the batch. The backend receives the batch a second time, resulting in duplicate spans. - Handling Duplicates: Observability backends must handle deduplication using unique span IDs and trace IDs, or accept minor metric over-counting during network recovery events.
Best-Effort Delivery Semantics
Under certain failure modes, the Collector operates under Best-Effort delivery, where telemetry may be discarded to preserve system stability:
- When memory reaches the
memory_limiterthreshold and incoming data is dropped to prevent an OOM crash. - When the
sending_queuefills to maximum capacity (queue_sizeexceeded) during a prolonged outage. - When a batch exceeds
max_elapsed_timein the retry loop. - When non-retryable errors (HTTP
400/401) are dropped. - When the process crashes while data sits in an in-memory queue.
Note
Why exactly-once delivery is impractical in telemetry pipelines: Achieving strict Exactly-Once delivery across distributed network boundaries requires distributed transactional consensus (such as two-phase commit protocols or end-to-end distributed idempotency keys) across every client, proxy, and storage engine. In high-volume observability streams generating gigabytes of data per second, the compute and network latency costs of two-phase locking are prohibitively expensive. At-Least-Once delivery represents the optimal cloud-native engineering trade-off.
A network switch failure severs connectivity between an OpenTelemetry Collector gateway and a downstream distributed tracing backend for approximately three minutes. During this outage, microservices continue pushing traces to the Collector at normal rates. What exporter configuration guarantees that the Collector buffers incoming batches and automatically delivers them without loss once network connectivity is restored?
Enable the debug exporter with verbosity set to detailed so in-flight spans are captured in stdout.
Configure the prometheus exporter to mirror traces onto an HTTP scrape target.
Enable sending_queue with sufficient queue_size (or persistent file_storage) alongside retry_on_failure with max_elapsed_time set to at least 5 minutes.
Set the memory_limiter spike_limit_percentage to 90 to allow infinite in-memory trace buffering.
An OpenTelemetry Collector exporter dispatches batches of telemetry to a remote SaaS observability backend. A misconfigured microservice sends a corrupted span batch containing an invalid payload structure that causes the backend to respond with HTTP 400 (Bad Request). Simultaneously, another batch is rejected with HTTP 429 (Too Many Requests). How does the Collector exporter handle these two failure scenarios?
Both HTTP 400 and HTTP 429 errors are classified as transient and are retried indefinitely with exponential backoff.
HTTP 400 is retried with backoff, while HTTP 429 triggers an immediate container crash to alert operations.
HTTP 400 is dumped to an operating system swap partition, while HTTP 429 is converted into cumulative gauge metrics.
HTTP 429 is treated as a transient rate-limiting error and retried via retry_on_failure, while HTTP 400 is treated as a non-retryable client error and dropped immediately to prevent blocking the sending queue.
A site reliability engineer is troubleshooting an issue where trace spans emitted by an order service appear to be missing specific span attributes before reaching the production APM backend. The engineer wants to view the complete internal pdata structure of every span on a test Collector instance in real time without interrupting or altering telemetry delivery to the primary APM backend. How should the engineer configure the Collector?
Configure a debug exporter with verbosity: detailed, and include both the primary OTLP exporter and the debug exporter in the service.pipelines.traces.exporters list.
Replace the primary OTLP exporter with the legacy logging exporter using sampling_initial: 0.
Set the service.telemetry.logs.level parameter to detailed without configuring any exporter components.
Attach the zpages extension to the OTLP receiver to print incoming Protobuf network packets directly to stdout.
Sections you finish are checked off in the contents.