12.2 Scaling, Load Balancing & High Availability

Key Takeaways

  • OpenTelemetry Collector workloads fall into two architectural categories: stateless pipelines (which scale linearly behind standard load balancers) and stateful pipelines (such as tail-based sampling, which require coordinated routing).

  • Traditional Layer 4 (TCP) and Layer 7 (HTTP) load balancers fail for tail-based sampling because spans belonging to the same distributed trace are scattered across different collector replicas, preventing any single collector from evaluating complete traces.

  • The loadbalancing exporter solves trace scatter by computing a consistent hash on the 128-bit TraceId of every span, deterministically routing all spans of a given trace to the identical downstream Gateway collector instance.

  • The loadbalancing exporter discovers downstream Collectors with the k8s resolver (EndpointSlices), the dns resolver (a headless Service's A records), the aws_cloud_map resolver, or a static list.

  • High availability in production requires deploying multiple Gateway replicas across availability zones with pod anti-affinity, configuring graceful rolling updates with readiness probes, and tuning Horizontal Pod Autoscalers (HPA) using CPU, memory, and custom collector metrics like otelcol_receiver_refused_spans.

Last updated: September 2026

12.2 Scaling, Load Balancing & High Availability

Quick Answer: Scaling an OpenTelemetry Collector fleet depends on whether the telemetry pipeline is stateless (such as batching, basic filtering, and OTLP export, which scale linearly behind standard Layer 4 or Layer 7 load balancers) or stateful (such as tail-based sampling and trace-to-metric aggregation, which require all spans of a trace to reach the identical collector instance). Because standard round-robin load balancers scatter spans across different pods, production architectures deploy the loadbalancing exporter on Tier 1 Agents to route spans using consistent hashing on TraceId. High availability requires multi-replica deployments with pod anti-affinity, graceful rolling updates, and Horizontal Pod Autoscalers (HPA) driven by refused-span metrics (otelcol_receiver_refused_spans).

As cloud-native architectures expand to thousands of microservices emitting millions of spans, metric data points, and log records per second, observability infrastructure must scale elastically without dropping critical telemetry. However, scaling an OpenTelemetry Collector fleet presents unique challenges that do not exist in conventional web application scaling. In particular, distributed traces consist of loosely coupled spans emitted across asynchronous microservices over time. Scaling collectors without understanding trace affinity can silently destroy data correlation and break sampling pipelines.


Stateless vs. Stateful Collector Workloads

To scale an OpenTelemetry Collector deployment effectively, platform engineers must categorize pipelines based on statefulness:

+-------------------------------------------------------------------------+
|                   Collector Pipeline Categorization                     |
+-------------------------------------------------------------------------+
  STATELESS PIPELINES                     STATEFUL PIPELINES
  - Ingestion (otlpreceiver)              - Tail-based sampling (tail_sampling)
  - Batching (batchprocessor)             - Trace-to-metrics (spanmetrics)
  - Memory limiting (memory_limiter)      - Metric deduplication & rollup
  - Stateless OTTL (transform)            - Sliding window rate limiting
  - Egress (otlpexporter)                 - Trace assembly & error detection
  -----------------------------------     ---------------------------------
  Scaling: Linear & trivial               Scaling: Requires consistent routing
  Load Balancer: Standard L4/L7           Load Balancer: loadbalancing exporter

Stateless Workloads

In a stateless pipeline, every telemetry record (an individual span, metric point, or log line) is processed independently. The components in the pipeline do not need to correlate records with previous or future records:

  • The batch processor groups whatever records arrive within a timeout window or byte size.
  • The memory_limiter processor monitors heap memory and applies backpressure.
  • The transform processor mutates attributes on individual records using OTTL.
  • The otlp exporter serializes and transmits batches.

Scaling Characteristics: Stateless collector pipelines scale linearly and horizontally. You can place NN collector replicas behind a standard Kubernetes ClusterIP Service or cloud load balancer. If traffic doubles, simply scale the replica count from 5 to 10. Any collector pod can receive any telemetry record from any microservice.

Stateful Workloads

In a stateful pipeline, components must accumulate records over a time window to perform cross-record analysis or aggregation:

  • tail_sampling Processor: Holds spans in memory for a configurable duration (e.g., 10–30 seconds) to determine whether the entire distributed trace contained an error (HTTP 5xx), an exception event, or an abnormal duration (latency > 2s). If an error occurred anywhere in the trace, all spans of that trace are sampled; otherwise, the trace is dropped.
  • spanmetrics Connector: Computes request rates, error rates, and duration histograms (RED metrics) from span streams, requiring continuous aggregation across spans.
  • Metric Deduplication: Combines overlapping metric streams from multiple agents.

The Scaling Invariant: Stateful workloads cannot be scaled using standard round-robin or least-connection load balancing! If spans belonging to Trace A are scattered across Pod 1, Pod 2, and Pod 3, no single pod has full visibility into the trace. Pod 1 might evaluate its sampling decision before Pod 3 processes the database error span, resulting in premature trace deletion.


The Load Balancing Challenge with Traces (The Scatter Problem)

In a distributed microservice architecture, a single user transaction traverses multiple independent services. Each service generates one or more spans that share the exact same 128-bit TraceId:

User Request
     │
     ▼
[Frontend Service]      --> generates Span 1 (TraceId: 4bf92f35...)
     │
     ▼
[Payment Service]       --> generates Span 2 (TraceId: 4bf92f35...)
     │
     ▼
[Database Service]      --> generates Span 3 (TraceId: 4bf92f35... - THROWS ERROR!)

Why Traditional Layer 4 (TCP) Load Balancers Fail

Layer 4 load balancers (such as Kubernetes ClusterIP or AWS Network Load Balancer) operate at the transport layer, routing TCP/UDP connections:

  • OpenTelemetry SDKs and agents use gRPC (HTTP/2) for high-performance OTLP transport.
  • gRPC maintains long-lived, persistent TCP connections over which thousands of multiplexed RPC streams are transmitted.
  • A Layer 4 load balancer balances only at connection establishment time. Once Frontend Service connects to Gateway Pod 1, all of Frontend Service's spans stream to Pod 1.
  • When Payment Service connects, the L4 balancer assigns it to Gateway Pod 2. All Payment Service spans stream to Pod 2.
  • Database Service connects to Gateway Pod 3. All Database Service spans stream to Pod 3.
  • The Failure Mode: Gateway Pod 1 evaluates Trace 4bf92f35... using only Frontend Service's span. Gateway Pod 1 sees a successful HTTP 200 response from the frontend and drops the trace, completely blind to the fact that Database Service threw a fatal SQL deadlock on Pod 3!

Why Traditional Layer 7 (HTTP) Load Balancers Fail

Layer 7 load balancers (such as Envoy, NGINX, or AWS ALB) understand HTTP/2 frames and distribute individual RPC requests across backend targets:

  • The load balancer round-robins individual ExportTraceServiceRequest RPC calls across Gateway Pods A, B, and C.
  • Span 1 lands on Pod A; Span 2 lands on Pod B; Span 3 lands on Pod C.
  • Every Gateway pod receives a broken, fragmented slice of every trace.
  • Every pod's tail-sampling cache times out waiting for missing spans, leading to corrupted sampling decisions, disjointed trace graphs, and severe memory bloat.

The loadbalancing Exporter (OpenTelemetry Collector Contrib)

The OpenTelemetry community developed the loadbalancing exporter specifically to solve the trace scatter problem in horizontally scaled environments. It sits on Tier 1 Node Agents or on an ingress routing Gateway layer, acting as a smart, telemetry-aware traffic dispatcher.

+-------------------------------------------------------------------------+
|              loadbalancing Exporter: Consistent Hashing                 |
+-------------------------------------------------------------------------+
  Incoming Spans from Microservices
  - Span A1 (TraceId: 4bf9...) ──────┐
  - Span B1 (TraceId: 4bf9...) ──────┼──> [loadbalancing exporter]
  - Span C1 (TraceId: 4bf9...) ──────┘         │
                                               │ Evaluates TraceId:
                                               │ hash("4bf9...") % Ring
                                               ▼
                                     Routes 100% of "4bf9..." spans
                                     to the EXACT SAME Gateway Pod!
                                               │
         ┌─────────────────────────────────────┼─────────────────────────┐
         ▼                                     ▼                         ▼
+─────────────────+                   +─────────────────+       +─────────────────+
|  Gateway Pod 1  |                   |  Gateway Pod 2  |       |  Gateway Pod 3  |
|  [Trace 4bf9...] |                  |  [Trace 9ca2...] |      |  [Trace 7ef1...] |
|  All spans for  |                   |                 |       |                 |
|  trace co-located|                  |                 |       |                 |
|  Tail sampling  |                   |                 |       |                 |
|  evaluates FULL |                   |                 |       |                 |
|  trace tree!    |                   |                 |       |                 |
+─────────────────+                   +─────────────────+       +─────────────────+

How It Works: Consistent Hashing on TraceId

  1. Batch Inspection: When a batch of spans arrives at the Tier 1 Agent, the loadbalancing exporter iterates through every span in the batch.
  2. Key Extraction: For each span, it extracts the 128-bit TraceId.
  3. Consistent Hashing: It hashes the routing key (the TraceId by default for traces) onto a consistent-hash ring built from the downstream Gateway endpoints. Every Tier 1 instance with the same configuration and endpoint list computes the same mapping.
  4. Batch Re-assembly and Dispatch: The exporter regroups spans destined for the same downstream Gateway instance into discrete OTLP gRPC batches and sends them directly to that specific pod's IP.
  5. Deterministic Invariant: Every span sharing the exact same TraceId—no matter which microservice emitted it, which worker node it ran on, or what millisecond it was generated—is deterministically routed to the identical downstream Gateway Collector instance.

The Value of Consistent Hashing During Cluster Rescaling

When downstream Gateway pods scale up (e.g., from 4 pods to 5 pods) or scale down due to node termination, a naive modulo hash (H mod NH \bmod N) would remap nearly 100% of all trace IDs, corrupting in-flight tail-sampling buffers across the entire fleet.

Because the loadbalancing exporter implements consistent hashing, adding or removing a Gateway replica remaps only about 1/N1/N of the hash space, so roughly (N−1)/N(N-1)/N of active traces stay pinned to their existing Gateway pods. Traces that do move during a scale event can still be split, so expect a small amount of disruption.

Downstream Routing Resolvers

The loadbalancing exporter discovers downstream Gateway instances with one of four resolvers:

  1. k8s (Kubernetes Resolver): Watches the EndpointSlice objects of a Service (given as name.namespace) and updates the ring as Gateway pods become ready or terminate. The Collector's service account needs get, list, and watch permissions on EndpointSlices.
  2. dns (DNS Resolver): Resolves a hostname, typically a headless Kubernetes Service such as otel-gateway-headless.monitoring.svc.cluster.local, to all of its IP addresses, re-checking every interval (default 5s).
  3. static (Static Resolver): Uses a fixed list of hostnames or IP addresses, suited for local development or fixed virtual machine clusters (an unavailable endpoint is not removed automatically).
  4. aws_cloud_map: Discovers instances registered in AWS Cloud Map.

The routing_key chooses what is hashed: traceID (the default for traces), service, resource, metric, streamID, attributes, or randomness.

Production Configuration Example: loadbalancing Exporter

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
    spike_limit_percentage: 20
  batch:
    send_batch_size: 8192
    timeout: 1s

exporters:
  loadbalancing:
    routing_key: "traceID"
    protocol:
      otlp:
        timeout: 5s
        tls:
          insecure: true
    resolver:
      k8s:
        service: otel-gateway-headless.observability
        ports:
          - 4317

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [loadbalancing]

High Availability (HA) Considerations in Kubernetes

Deploying a scalable OpenTelemetry Collector Gateway tier requires designing for node failures, zone outages, and rolling software updates:

1. Multi-Replica Redundancy & Pod Anti-Affinity

Never run a single Collector replica in production. Maintain a minimum of three replicas across separate availability zones. Configure Kubernetes podAntiAffinity so that the Kubernetes scheduler spreads Collector pods across distinct physical nodes and availability zones:

affinity:
  podAntiAffinity:
    preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 100
        podAffinityTerm:
          labelSelector:
            matchExpressions:
              - key: app.kubernetes.io/name
                operator: In
                values: ["opentelemetry-collector-gateway"]
          topologyKey: topology.kubernetes.io/zone
      - weight: 90
        podAffinityTerm:
          labelSelector:
            matchExpressions:
              - key: app.kubernetes.io/name
                operator: In
                values: ["opentelemetry-collector-gateway"]
          topologyKey: kubernetes.io/hostname

2. Pod Disruption Budgets (PDB)

To protect the Collector tier during planned Kubernetes node drains, cluster upgrades, or autoscaling events, define a PodDisruptionBudget:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: otel-gateway-pdb
  namespace: observability
spec:
  minAvailable: 2
  selector:
    matchLabels:
      app.kubernetes.io/name: opentelemetry-collector-gateway

3. Graceful Rolling Updates & Termination Handshakes

During a deployment update, Kubernetes sends a SIGTERM signal to terminating pods. If a Collector terminates abruptly, in-flight batches and network sockets are dropped:

  • Set terminationGracePeriodSeconds: 60 or higher in the pod specification.
  • When SIGTERM is received, the Collector executes its reverse lifecycle: receivers stop listening first, processors flush pending batches downstream, and exporters drain sending queues before process exit.
  • Configure Kubernetes Readiness Probes against the health_check extension on port 13133. Its default endpoint binds only to localhost, so set endpoint: 0.0.0.0:13133 (or the pod IP) for the kubelet to reach it. A pod only receives traffic once all internal pipelines are ready, and traffic is removed from the pod immediately upon termination.

4. Horizontal Pod Autoscaling (HPA) Triggers

Autoscaling Collector Gateways requires monitoring both infrastructure utilization and the Collector's internal metrics, served by default at http://127.0.0.1:8888/metrics (configure a readers entry with host 0.0.0.0 so a cluster Prometheus can scrape it):

Metric NameSourceOperational Meaning & Autoscaling Action
CPU UtilizationContainer RuntimeScale out when average CPU exceeds 70–75% to prevent thread scheduling delays and ingestion latency.
Memory UtilizationContainer RuntimeScale out when memory exceeds 75% to prevent memory_limiter drops.
otelcol_receiver_refused_spansCollector InternalCritical Custom Metric: Non-zero values indicate the Collector is rejecting spans due to backpressure or memory exhaustion. Must trigger rapid scale-out!
otelcol_receiver_refused_metric_pointsCollector InternalNon-zero values indicate metrics are actively being dropped at ingress.
otelcol_exporter_queue_sizeCollector InternalHigh queue depth relative to queue_capacity indicates downstream backend saturation or network bottlenecks.

HPA Autoscaling Policy Best Practice: Configure HPA with a rapid scale-up policy (e.g., doubling replicas within 15 seconds upon a traffic spike) and a conservative, stabilized scale-down window (e.g., 300 seconds) to prevent pod flapping and cache thrashing during transient dips in telemetry volume.

Loading diagram...
TraceId Consistent Hashing with the loadbalancing Exporter
Test Your Knowledge

A platform team deploys a tail-based sampling pipeline on a cluster of four OpenTelemetry Collector Gateway pods placed behind a standard Kubernetes ClusterIP Service. Although the tail-sampling processor is configured to retain 100% of traces that contain HTTP 500 errors, engineers notice that traces with server errors in downstream microservices are frequently dropped or appear with missing child spans. What is the fundamental reason standard Layer 4 or Layer 7 load balancers fail for tail-based sampling?

A

Kubernetes ClusterIP Services encrypt gRPC payloads, preventing the Collector from reading span attributes

B

Standard load balancers distribute requests or connections across pods without TraceId affinity, scattering spans of the same trace across different collector replicas so no single pod observes the complete trace

C

The tail-sampling processor can only operate on single-node VM architectures and crashes when scaled horizontally

D

Layer 7 load balancers automatically strip W3C traceparent headers from HTTP requests during proxying

Test Your Knowledge

To solve the trace-scattering problem in a horizontally scalable Two-Tier Collector architecture, an engineering team installs the loadbalancing exporter on Tier 1 Node Agents. How does the loadbalancing exporter ensure that tail-based sampling can evaluate complete traces across downstream Gateway Collector pods?

A

It broadcasts every incoming span to all downstream Gateway Collector replicas simultaneously using UDP multicast

B

It buffers an entire trace inside the Tier 1 Agent memory until the root span finishes, then sends the assembled trace in a single batch

C

It extracts the TraceId from every span and uses consistent hashing to deterministically route all spans sharing the same TraceId to the identical downstream Gateway replica

D

It converts all spans into Prometheus metrics before forwarding them through a round-robin proxy

Test Your Knowledge

A site reliability engineering team is setting up a Horizontal Pod Autoscaler (HPA) for an OpenTelemetry Collector Gateway deployment in Kubernetes. In addition to monitoring container CPU and memory utilization, which custom Prometheus metric exposed by the Collector's internal telemetry (:8888/metrics) should be used as a primary scaling metric to prevent telemetry data loss during traffic surges?

A

otelcol_process_uptime to measure how long the Go runtime has been executing

B

otelcol_processor_batch_metadata_cardinality to monitor the size of resource attributes

C

otelcol_exporter_sent_spans to measure the total number of successfully exported spans

D

otelcol_receiver_refused_spans (and otelcol_receiver_refused_metric_points) to detect when collectors are saturated and actively rejecting incoming telemetry

Sections you finish are checked off in the contents.