12.3 Resilience: Backpressure, Buffering & Persistent Queues

Key Takeaways

  • Telemetry backpressure propagates upstream: when backends throttle or fail, exporter queues fill, memory rises, and receivers refuse data with retryable errors (gRPC UNAVAILABLE or HTTP 503 when the memory limiter refuses it).

  • Default in-memory sending queues (sending_queue) provide zero-I/O buffering for short transient spikes, but introduce catastrophic data loss if the Collector process crashes or restarts during an extended backend outage.

  • The file_storage extension keeps a persistent sending queue in an on-disk key-value database (bbolt), so queued batches survive a Collector restart and drain when the backend returns.

  • Out-of-memory (OOM) fatal crashes are prevented by properly sizing the memory_limiter processor—setting limit_mib to ~80% and spike_limit_mib to ~20% of container memory limits—and configuring the Go runtime GOMEMLIMIT environment variable.

  • Combining durable file storage, conservative memory limits, and upstream client retry policies establishes an end-to-end resilient telemetry architecture capable of surviving network partitions and backend brownouts.

Last updated: September 2026

12.3 Resilience: Backpressure, Buffering & Persistent Queues

Quick Answer: When downstream observability backends experience outages or rate limits, the OpenTelemetry Collector protects itself and upstream systems through backpressure propagation: exporter queues fill, memory rises, and receivers return retryable errors (gRPC UNAVAILABLE or HTTP 503), prompting application SDKs to retry and, if their own queues fill, shed telemetry locally. While default in-memory sending queues offer zero I/O overhead, they lose their contents if the pod crashes or restarts; the file_storage extension backs a persistent queue on disk that survives container restarts. To prevent fatal out-of-memory (OOM) crashes, the memory_limiter processor must be placed first in every pipeline with limit_mib sized to 80% and spike_limit_mib to 20% of container limits, complemented by the Go runtime GOMEMLIMIT setting.

Observability pipelines operate under asynchronous, unpredictable conditions. Backend storage clusters (such as Prometheus TSDBs, OpenSearch clusters, or cloud APM platforms) undergo scheduled maintenance, hit API rate limits, or suffer transient network partitions. During downstream downtime, the OpenTelemetry Collector must act as a reliable shock absorber. If the Collector cannot manage backpressure, buffer incoming surges, and protect its memory boundaries, it risks crashing with Linux Out-Of-Memory (OOM) errors, corrupting data buffers, or cascading failures back into customer-facing applications.


Backpressure Propagation in Telemetry Pipelines

In an OpenTelemetry pipeline, telemetry data flows downstream from applications to storage backends. Conversely, backpressure travels upstream, flowing in reverse from the destination storage system back to the telemetry source:

=============================== TELEMETRY DATA FLOW (DOWNSTREAM) ===============================>
[App Microservice SDK] ──> [Receiver] ──> [Processor] ──> [Exporter Queue] ──> [Remote Backend]

<============================= BACKPRESSURE PROPAGATION (UPSTREAM) =============================
[SDK drops locally]   <── [HTTP 429]  <── [Halts Work] <── [Queue Full]   <── [Backend Outage]
[Protects App RAM]        [gRPC 14/8]     [Limiter Drops]  [Rejects Batches]   [HTTP 429 / 503]

The Step-by-Step Anatomy of a Downstream Failure

When a downstream observability backend slows down (e.g., returning HTTP 429 Too Many Requests due to rate limits) or suffers a hard outage (network timeout, HTTP 503 Service Unavailable):

  1. Exporter Sending Queue Fills: The exporter's retry mechanism kicks in with exponential backoff. In-flight batches cannot be drained, and new batches arriving from upstream processors accumulate in the exporter's sending_queue up to its configured capacity (queue_size).
  2. Exporter Rejects New Batches: Once the sending_queue is full, new data cannot be enqueued; it is counted in otelcol_exporter_enqueue_failed_* (unless block_on_overflow: true makes the caller wait for space).
  3. The Error Travels Back, Unless Something Absorbs It: Without an asynchronous stage in between, the enqueue error returns through the processors to the receiver. The batch processor, however, has already accepted the data and answered its caller, so a queue-full error after it becomes a drop rather than backpressure. This is one reason newer Collector releases can batch inside the exporter (sending_queue::batch) instead.
  4. Memory Pressure Triggers memory_limiter: As queued and buffered data pushes memory past the soft limit, the memory_limiter processor refuses new data with a retryable error (and above the hard limit it also forces garbage collection).
  5. Receivers Reject Ingress Requests: The OTLP receiver turns a refused request into a retryable status for the client:
    • OTLP/gRPC Receiver: a plain retryable error becomes gRPC 14 (UNAVAILABLE); an error that already carries 8 (RESOURCE_EXHAUSTED) keeps it.
    • OTLP/HTTP Receiver: UNAVAILABLE maps to HTTP 503 (Service Unavailable) and RESOURCE_EXHAUSTED to HTTP 429 (Too Many Requests).
  6. Application SDKs React: Upstream OpenTelemetry SDKs receive the rejection status codes:
    • The SDK's OTLP exporter retries retryable responses with exponential backoff while the BatchSpanProcessor keeps queuing new spans.
    • If the outage persists and the SDK's local in-memory buffer (max_queue_size, default 2048 spans) fills to capacity, the SDK drops telemetry locally at the application (SDKs that implement the self-observability conventions record this in otel.sdk.processor.span.processed with error.type=queue_full).

The Golden Rule of Observability Backpressure

Telemetry is an operational monitoring tool; it must never compromise the primary business function of the application. Upstream backpressure propagation guarantees that when catastrophic storage failures occur, telemetry is shed gracefully at the edges rather than allowing unchecked buffer growth to consume memory, trigger application container crashes, or lock up user request threads.


In-Memory Buffering vs. File-Backed Persistent Storage

The OpenTelemetry Collector provides two primary strategies for managing queued telemetry inside exporters: in-memory sending queues and file-backed persistent queues.

+-------------------------------------------------------------------------+
|               In-Memory Queue vs. File-Backed Persistent Queue          |
+-------------------------------------------------------------------------+
  IN-MEMORY SENDING QUEUE                 FILE-BACKED PERSISTENT QUEUE
  (sending_queue: storage: none)          (sending_queue: storage: file_storage)
  
  [RAM: Go Process Heap]                  [RAM Buffer] ──> [On-Disk Queue DB]
  - Blazing fast (zero disk I/O)          - Survives container crashes & OOMs
  - Microsecond latency                   - Survives Kubernetes node evictions
  - Zero storage volume config            - Spools gigabytes during outages
  -----------------------------------     -----------------------------------
  FATAL FLAW: Process crash or            RECOVERY: On boot, Collector reads
  pod restart permanently wipes           the queue from disk and drains all
  all buffered telemetry in RAM!          queued batches to restored backend!

1. In-Memory Sending Queue (sending_queue with default memory storage)

  • How It Operates: The exporter keeps a bounded queue of batches in the Go heap, sized by queue_size (default 1,000 requests).
  • Strengths: Ultra-high throughput, zero disk I/O latency, no requirement for Kubernetes PersistentVolumeClaims (PVCs) or stateful storage drivers.
  • The Critical Vulnerability: RAM is volatile. If the Collector container runs out of memory, is terminated by a Kubernetes rolling update, or experiences a worker node hardware reboot during an outage, every single telemetry batch buffered in memory is lost forever.

2. File-Backed Persistent Storage Queue (file_storage extension)

  • How It Operates: The file_storage extension (Contrib) stores data in an on-disk key-value database (bbolt) in its directory (default /var/lib/otelcol/file_storage), which can live on local SSD, a cloud block volume, or a Kubernetes PVC. With sending_queue::storage set, the queue lives only on disk; there is no in-memory queue.
  • The Persistence Lifecycle:
    1. When a batch enters the queue, it is serialized and written to the storage database before the enqueue call returns.
    2. Consumers read batches from storage and export them; a batch is deleted from storage after success or a permanent error.
    3. During an outage, batches accumulate on disk until queue_size batches are stored or the disk (or the extension's optional max_size) is full; after that, new data is rejected.
    4. Crash Recovery: If the Collector is killed or restarted, the new process opens the same directory and resumes exporting the stored batches. In Kubernetes this works only if the same volume is attached to the replacement pod.

Memory Safety & Out-Of-Memory (OOM) Prevention

Because the OpenTelemetry Collector is implemented in Go, managing memory boundaries requires understanding both the Go runtime Garbage Collector (GC) and the Collector's internal pipeline memory safeguards.

+-------------------------------------------------------------------------+
|               Collector Container Memory Layout (Example: 2 GiB)        |
+-------------------------------------------------------------------------+
  [ 0 MiB ]                                                    [ 2048 MiB ]
  ├────────────────────────────┼──────────────┼─────────────────────┤
  │ Normal Pipeline Operation  │ Soft Limit   │ Hard Container Limit│
  │ Heap Allocations & Batches │ Limiter Sheds│ (Kernel OOM Killer) │
  │                            │ (Spike Zone) │                     │
  ├────────────────────────────┼──────────────┼─────────────────────┤
  0 MiB                     1228 MiB       1638 MiB              2048 MiB
                            (Limit - Spike) (limit_mib)           (Container Max)

  1. Normal Ingestion: Heap stays well below 1228 MiB.
  2. Sudden Traffic Burst: Memory crosses 1228 MiB -> memory_limiter drops batches.
  3. Go GC Intervention: GOMEMLIMIT (about 1638 MiB) makes the garbage collector work harder near the limit.
  4. OOM Avoidance: Process stays below 2048 MiB container limit. SIGKILL avoided!

The Go Runtime Memory Dilemma

By default, Go's garbage collector uses a target percentage (GOGC=100), meaning GC triggers only when the heap doubles relative to reachable memory. During sudden bursts of telemetry (such as a cascading microservice outage generating massive trace spikes), memory allocations occur faster than the Go runtime schedules GC sweeps. The container's Resident Set Size (RSS) exceeds the Kubernetes memory limit, and the Linux kernel OOM Killer immediately terminates the container with SIGKILL (Exit Code 137).

The memory_limiter Processor

The memory_limiter processor is the primary line of defense against OOM terminations.

The Sizing Formula:

  • Let MM be the container's hard memory limit (e.g., 2048 MiB2048\text{ MiB} in Kubernetes resources.limits.memory).
  • limit_mib: The hard target limit for the Collector process. Must be set to approximately 80% of MM: limit_mib=2048 MiB×0.80≈1638 MiB\text{limit\_mib} = 2048\text{ MiB} \times 0.80 \approx 1638\text{ MiB}
  • spike_limit_mib: The safety buffer to absorb surges between measurement cycles. Must be set to approximately 20% of MM: spike_limit_mib=2048 MiB×0.20≈410 MiB\text{spike\_limit\_mib} = 2048\text{ MiB} \times 0.20 \approx 410\text{ MiB}
  • Soft Threshold Activation: When memory usage exceeds limit_mib - spike_limit_mib (1638−410=1228 MiB1638 - 410 = 1228\text{ MiB}), the processor refuses incoming data with retryable errors; if usage goes above limit_mib it also forces garbage collection.
  • check_interval: How often memory usage is checked (recommended: 1s for general workloads, 500ms for bursty workloads).

CRITICAL INVARIANT: The memory_limiter processor must always be placed as the FIRST processor in every pipeline sequence (processors: [memory_limiter, batch, ...]). If placed after the batch or transform processors, memory is allocated and consumed before the limiter can evaluate whether sufficient headroom exists to process the payload!

Tuning Go Runtime Environment Variables (GOMEMLIMIT & GOMAXPROCS)

In modern containerized deployments (Go 1.19+), platform engineers must configure two critical runtime environment variables:

  1. GOMEMLIMIT: A soft memory limit for the Go runtime. The memory_limiter README recommends setting it to about 80% of the Collector's hard memory limit (e.g., GOMEMLIMIT=1638MiB for a 2 GiB container); the garbage collector then runs more often as memory approaches the limit, reclaiming heap before the kernel's OOM killer acts.
  2. GOMAXPROCS: Since Go 1.25, the runtime derives the default GOMAXPROCS from the container's CPU limit on Linux. Binaries built with older Go versions used the host's core count, which could cause CPU-quota throttling in Kubernetes; they needed go.uber.org/automaxprocs or an explicit GOMAXPROCS matching the CPU limit.

Production Configuration: Resilient Collector with file_storage

The following complete configuration demonstrates production-grade resilience, combining the file_storage extension, properly sized memory_limiter, and a persistent OTLP exporter queue:

extensions:
  # Health probe for Kubernetes liveness/readiness
  health_check:
    endpoint: 0.0.0.0:13133

  # Durable on-disk storage (bbolt) for the persistent queue
  file_storage:
    directory: /var/lib/otelcol/file_storage
    create_directory: true
    timeout: 10s
    compaction:
      on_rebound: true
      directory: /var/lib/otelcol/file_storage/compaction

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  # CRITICAL: memory_limiter MUST be the first processor in every pipeline
  memory_limiter:
    check_interval: 1s
    limit_mib: 1638       # 80% of 2048 MiB container limit
    spike_limit_mib: 410   # 20% of 2048 MiB container limit

  batch:
    send_batch_size: 8192
    timeout: 1s

exporters:
  otlp:
    endpoint: remote-backend.observability.net:4317
    tls:
      insecure: false
      ca_file: /etc/ssl/certs/backend-ca.crt
    # Persistent sending queue stored through file_storage
    sending_queue:
      enabled: true
      storage: file_storage
      num_consumers: 16
      queue_size: 10000
    retry_on_failure:
      enabled: true
      initial_interval: 5s
      max_interval: 60s
      max_elapsed_time: 15m

service:
  extensions: [health_check, file_storage]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp]
  telemetry:
    logs:
      level: info
    metrics:
      readers:
        - pull:
            exporter:
              prometheus:
                host: 0.0.0.0
                port: 8888

Resilience Comparison Reference

Operational DimensionDefault In-Memory BufferFile-Backed Storage Buffer (file_storage)
Underlying MediumVolatile Go Process RAMDurable disk (bbolt database on a local disk or Persistent Volume)
I/O LatencySub-microsecond (RAM speed)Low millisecond (buffered disk writes)
Storage CapacityConstrained by container RAM (MiBs)Constrained only by attached disk volume (GiBs)
Behavior on Pod CrashTotal data loss for all queued itemsQueued batches kept on disk and resumed after restart
Behavior on Node EvictionTotal data loss when node terminatesPreserved only if the same volume is reattached to the new pod
Backend Outage ToleranceMinutes (until memory limit is hit)Hours to days (until disk volume fills)
Production RecommendationEphemeral dev/staging environmentsMission-critical enterprise production tiers
Loading diagram...
Backpressure Propagation and File Storage WAL Pipeline
Test Your Knowledge

During a major cloud provider network outage lasting 45 minutes, an external observability SaaS backend becomes completely unreachable. An OpenTelemetry Collector deployment is configured with default in-memory sending queues (sending_queue). Twenty minutes into the incident, the underlying Kubernetes worker node experiences a hardware failure, causing the Collector pod to be rescheduled on a new node. What is the impact on the telemetry data ingested during the first twenty minutes of the outage?

A

All buffered telemetry is permanently lost because standard in-memory sending queues reside in volatile process RAM that is destroyed when the container terminates

B

The telemetry is automatically recovered from the local Linux operating system page cache upon container restart

C

The upstream application SDKs detect the pod reschedule and retransmit all missing spans from their local disk buffers

D

The Kubernetes API server caches all dropped network packets and injects them into the new Collector pod

Test Your Knowledge

A platform engineer is configuring the memory_limiter processor for an OpenTelemetry Collector container with a Kubernetes memory limit of 2048 MiB (2 GiB). Under heavy load, the Collector previously experienced sudden traffic bursts that triggered the Linux kernel Out-Of-Memory (OOM) killer (exit code 137). Following OpenTelemetry best practices, how should limit_mib and spike_limit_mib be configured, and where must the processor be placed in the pipeline?

A

Set limit_mib to 2048 MiB and spike_limit_mib to 0 MiB, placed as the final processor right before exporters

B

Set limit_mib to approximately 80% of the container limit (1638 MiB) and spike_limit_mib to approximately 20% of the limit (410 MiB), placed as the very first processor in all active pipelines

C

Set limit_mib to 500 MiB and spike_limit_mib to 1500 MiB, placed between the batch and transform processors

D

Set limit_mib to 4096 MiB with swap enabled, placed only in trace pipelines

Test Your Knowledge

When an external observability backend experiences severe degradation and responds to outgoing OTLP export requests with HTTP 429 (Too Many Requests) or gRPC ResourceExhausted errors, how does backpressure propagate through a properly configured OpenTelemetry Collector pipeline to protect infrastructure resources?

A

The Collector immediately scales down its worker threads to zero and deletes all pending configuration files

B

The exporter converts incoming spans into uncompressed JSON logs and writes them to stdout until the backend recovers

C

The exporter sending queue fills to capacity, blocking processors and causing receivers to reject incoming client requests with HTTP 429 or gRPC Unavailable status codes, which prompts application SDKs to back off and buffer or shed data locally

D

Receivers silently swallow all incoming network traffic without acknowledging receipts so client applications assume data was delivered

Sections you finish are checked off in the contents.