12.2 Scaling, Load Balancing & High Availability
Key Takeaways
OpenTelemetry Collector workloads fall into two architectural categories: stateless pipelines (which scale linearly behind standard load balancers) and stateful pipelines (such as tail-based sampling, which require coordinated routing).
Traditional Layer 4 (TCP) and Layer 7 (HTTP) load balancers fail for tail-based sampling because spans belonging to the same distributed trace are scattered across different collector replicas, preventing any single collector from evaluating complete traces.
The loadbalancing exporter solves trace scatter by computing a consistent hash on the 128-bit TraceId of every span, deterministically routing all spans of a given trace to the identical downstream Gateway collector instance.
The loadbalancing exporter discovers downstream Collectors with the k8s resolver (EndpointSlices), the dns resolver (a headless Service's A records), the aws_cloud_map resolver, or a static list.
High availability in production requires deploying multiple Gateway replicas across availability zones with pod anti-affinity, configuring graceful rolling updates with readiness probes, and tuning Horizontal Pod Autoscalers (HPA) using CPU, memory, and custom collector metrics like otelcol_receiver_refused_spans.
12.2 Scaling, Load Balancing & High Availability
Quick Answer: Scaling an OpenTelemetry Collector fleet depends on whether the telemetry pipeline is stateless (such as batching, basic filtering, and OTLP export, which scale linearly behind standard Layer 4 or Layer 7 load balancers) or stateful (such as tail-based sampling and trace-to-metric aggregation, which require all spans of a trace to reach the identical collector instance). Because standard round-robin load balancers scatter spans across different pods, production architectures deploy the
loadbalancingexporter on Tier 1 Agents to route spans using consistent hashing onTraceId. High availability requires multi-replica deployments with pod anti-affinity, graceful rolling updates, and Horizontal Pod Autoscalers (HPA) driven by refused-span metrics (otelcol_receiver_refused_spans).
As cloud-native architectures expand to thousands of microservices emitting millions of spans, metric data points, and log records per second, observability infrastructure must scale elastically without dropping critical telemetry. However, scaling an OpenTelemetry Collector fleet presents unique challenges that do not exist in conventional web application scaling. In particular, distributed traces consist of loosely coupled spans emitted across asynchronous microservices over time. Scaling collectors without understanding trace affinity can silently destroy data correlation and break sampling pipelines.
Stateless vs. Stateful Collector Workloads
To scale an OpenTelemetry Collector deployment effectively, platform engineers must categorize pipelines based on statefulness:
+-------------------------------------------------------------------------+
| Collector Pipeline Categorization |
+-------------------------------------------------------------------------+
STATELESS PIPELINES STATEFUL PIPELINES
- Ingestion (otlpreceiver) - Tail-based sampling (tail_sampling)
- Batching (batchprocessor) - Trace-to-metrics (spanmetrics)
- Memory limiting (memory_limiter) - Metric deduplication & rollup
- Stateless OTTL (transform) - Sliding window rate limiting
- Egress (otlpexporter) - Trace assembly & error detection
----------------------------------- ---------------------------------
Scaling: Linear & trivial Scaling: Requires consistent routing
Load Balancer: Standard L4/L7 Load Balancer: loadbalancing exporter
Stateless Workloads
In a stateless pipeline, every telemetry record (an individual span, metric point, or log line) is processed independently. The components in the pipeline do not need to correlate records with previous or future records:
- The
batchprocessor groups whatever records arrive within a timeout window or byte size. - The
memory_limiterprocessor monitors heap memory and applies backpressure. - The
transformprocessor mutates attributes on individual records using OTTL. - The
otlpexporter serializes and transmits batches.
Scaling Characteristics: Stateless collector pipelines scale linearly and horizontally. You can place collector replicas behind a standard Kubernetes ClusterIP Service or cloud load balancer. If traffic doubles, simply scale the replica count from 5 to 10. Any collector pod can receive any telemetry record from any microservice.
Stateful Workloads
In a stateful pipeline, components must accumulate records over a time window to perform cross-record analysis or aggregation:
tail_samplingProcessor: Holds spans in memory for a configurable duration (e.g., 10–30 seconds) to determine whether the entire distributed trace contained an error (HTTP 5xx), an exception event, or an abnormal duration (latency > 2s). If an error occurred anywhere in the trace, all spans of that trace are sampled; otherwise, the trace is dropped.spanmetricsConnector: Computes request rates, error rates, and duration histograms (RED metrics) from span streams, requiring continuous aggregation across spans.- Metric Deduplication: Combines overlapping metric streams from multiple agents.
The Scaling Invariant: Stateful workloads cannot be scaled using standard round-robin or least-connection load balancing! If spans belonging to Trace A are scattered across Pod 1, Pod 2, and Pod 3, no single pod has full visibility into the trace. Pod 1 might evaluate its sampling decision before Pod 3 processes the database error span, resulting in premature trace deletion.
The Load Balancing Challenge with Traces (The Scatter Problem)
In a distributed microservice architecture, a single user transaction traverses multiple independent services. Each service generates one or more spans that share the exact same 128-bit TraceId:
User Request
│
▼
[Frontend Service] --> generates Span 1 (TraceId: 4bf92f35...)
│
▼
[Payment Service] --> generates Span 2 (TraceId: 4bf92f35...)
│
▼
[Database Service] --> generates Span 3 (TraceId: 4bf92f35... - THROWS ERROR!)
Why Traditional Layer 4 (TCP) Load Balancers Fail
Layer 4 load balancers (such as Kubernetes ClusterIP or AWS Network Load Balancer) operate at the transport layer, routing TCP/UDP connections:
- OpenTelemetry SDKs and agents use gRPC (HTTP/2) for high-performance OTLP transport.
- gRPC maintains long-lived, persistent TCP connections over which thousands of multiplexed RPC streams are transmitted.
- A Layer 4 load balancer balances only at connection establishment time. Once Frontend Service connects to Gateway Pod 1, all of Frontend Service's spans stream to Pod 1.
- When Payment Service connects, the L4 balancer assigns it to Gateway Pod 2. All Payment Service spans stream to Pod 2.
- Database Service connects to Gateway Pod 3. All Database Service spans stream to Pod 3.
- The Failure Mode: Gateway Pod 1 evaluates Trace
4bf92f35...using only Frontend Service's span. Gateway Pod 1 sees a successful HTTP 200 response from the frontend and drops the trace, completely blind to the fact that Database Service threw a fatal SQL deadlock on Pod 3!
Why Traditional Layer 7 (HTTP) Load Balancers Fail
Layer 7 load balancers (such as Envoy, NGINX, or AWS ALB) understand HTTP/2 frames and distribute individual RPC requests across backend targets:
- The load balancer round-robins individual
ExportTraceServiceRequestRPC calls across Gateway Pods A, B, and C. - Span 1 lands on Pod A; Span 2 lands on Pod B; Span 3 lands on Pod C.
- Every Gateway pod receives a broken, fragmented slice of every trace.
- Every pod's tail-sampling cache times out waiting for missing spans, leading to corrupted sampling decisions, disjointed trace graphs, and severe memory bloat.
The loadbalancing Exporter (OpenTelemetry Collector Contrib)
The OpenTelemetry community developed the loadbalancing exporter specifically to solve the trace scatter problem in horizontally scaled environments. It sits on Tier 1 Node Agents or on an ingress routing Gateway layer, acting as a smart, telemetry-aware traffic dispatcher.
+-------------------------------------------------------------------------+
| loadbalancing Exporter: Consistent Hashing |
+-------------------------------------------------------------------------+
Incoming Spans from Microservices
- Span A1 (TraceId: 4bf9...) ──────┐
- Span B1 (TraceId: 4bf9...) ──────┼──> [loadbalancing exporter]
- Span C1 (TraceId: 4bf9...) ──────┘ │
│ Evaluates TraceId:
│ hash("4bf9...") % Ring
▼
Routes 100% of "4bf9..." spans
to the EXACT SAME Gateway Pod!
│
┌─────────────────────────────────────┼─────────────────────────┐
▼ ▼ ▼
+─────────────────+ +─────────────────+ +─────────────────+
| Gateway Pod 1 | | Gateway Pod 2 | | Gateway Pod 3 |
| [Trace 4bf9...] | | [Trace 9ca2...] | | [Trace 7ef1...] |
| All spans for | | | | |
| trace co-located| | | | |
| Tail sampling | | | | |
| evaluates FULL | | | | |
| trace tree! | | | | |
+─────────────────+ +─────────────────+ +─────────────────+
How It Works: Consistent Hashing on TraceId
- Batch Inspection: When a batch of spans arrives at the Tier 1 Agent, the
loadbalancingexporter iterates through every span in the batch. - Key Extraction: For each span, it extracts the 128-bit
TraceId. - Consistent Hashing: It hashes the routing key (the
TraceIdby default for traces) onto a consistent-hash ring built from the downstream Gateway endpoints. Every Tier 1 instance with the same configuration and endpoint list computes the same mapping. - Batch Re-assembly and Dispatch: The exporter regroups spans destined for the same downstream Gateway instance into discrete OTLP gRPC batches and sends them directly to that specific pod's IP.
- Deterministic Invariant: Every span sharing the exact same
TraceId—no matter which microservice emitted it, which worker node it ran on, or what millisecond it was generated—is deterministically routed to the identical downstream Gateway Collector instance.
The Value of Consistent Hashing During Cluster Rescaling
When downstream Gateway pods scale up (e.g., from 4 pods to 5 pods) or scale down due to node termination, a naive modulo hash () would remap nearly 100% of all trace IDs, corrupting in-flight tail-sampling buffers across the entire fleet.
Because the loadbalancing exporter implements consistent hashing, adding or removing a Gateway replica remaps only about of the hash space, so roughly of active traces stay pinned to their existing Gateway pods. Traces that do move during a scale event can still be split, so expect a small amount of disruption.
Downstream Routing Resolvers
The loadbalancing exporter discovers downstream Gateway instances with one of four resolvers:
k8s(Kubernetes Resolver): Watches theEndpointSliceobjects of a Service (given asname.namespace) and updates the ring as Gateway pods become ready or terminate. The Collector's service account needsget,list, andwatchpermissions on EndpointSlices.dns(DNS Resolver): Resolves a hostname, typically a headless Kubernetes Service such asotel-gateway-headless.monitoring.svc.cluster.local, to all of its IP addresses, re-checking everyinterval(default5s).static(Static Resolver): Uses a fixed list of hostnames or IP addresses, suited for local development or fixed virtual machine clusters (an unavailable endpoint is not removed automatically).aws_cloud_map: Discovers instances registered in AWS Cloud Map.
The routing_key chooses what is hashed: traceID (the default for traces), service, resource, metric, streamID, attributes, or randomness.
Production Configuration Example: loadbalancing Exporter
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch:
send_batch_size: 8192
timeout: 1s
exporters:
loadbalancing:
routing_key: "traceID"
protocol:
otlp:
timeout: 5s
tls:
insecure: true
resolver:
k8s:
service: otel-gateway-headless.observability
ports:
- 4317
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loadbalancing]
High Availability (HA) Considerations in Kubernetes
Deploying a scalable OpenTelemetry Collector Gateway tier requires designing for node failures, zone outages, and rolling software updates:
1. Multi-Replica Redundancy & Pod Anti-Affinity
Never run a single Collector replica in production. Maintain a minimum of three replicas across separate availability zones. Configure Kubernetes podAntiAffinity so that the Kubernetes scheduler spreads Collector pods across distinct physical nodes and availability zones:
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app.kubernetes.io/name
operator: In
values: ["opentelemetry-collector-gateway"]
topologyKey: topology.kubernetes.io/zone
- weight: 90
podAffinityTerm:
labelSelector:
matchExpressions:
- key: app.kubernetes.io/name
operator: In
values: ["opentelemetry-collector-gateway"]
topologyKey: kubernetes.io/hostname
2. Pod Disruption Budgets (PDB)
To protect the Collector tier during planned Kubernetes node drains, cluster upgrades, or autoscaling events, define a PodDisruptionBudget:
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: otel-gateway-pdb
namespace: observability
spec:
minAvailable: 2
selector:
matchLabels:
app.kubernetes.io/name: opentelemetry-collector-gateway
3. Graceful Rolling Updates & Termination Handshakes
During a deployment update, Kubernetes sends a SIGTERM signal to terminating pods. If a Collector terminates abruptly, in-flight batches and network sockets are dropped:
- Set
terminationGracePeriodSeconds: 60or higher in the pod specification. - When
SIGTERMis received, the Collector executes its reverse lifecycle: receivers stop listening first, processors flush pending batches downstream, and exporters drain sending queues before process exit. - Configure Kubernetes Readiness Probes against the
health_checkextension on port 13133. Its default endpoint binds only to localhost, so setendpoint: 0.0.0.0:13133(or the pod IP) for the kubelet to reach it. A pod only receives traffic once all internal pipelines are ready, and traffic is removed from the pod immediately upon termination.
4. Horizontal Pod Autoscaling (HPA) Triggers
Autoscaling Collector Gateways requires monitoring both infrastructure utilization and the Collector's internal metrics, served by default at http://127.0.0.1:8888/metrics (configure a readers entry with host 0.0.0.0 so a cluster Prometheus can scrape it):
| Metric Name | Source | Operational Meaning & Autoscaling Action |
|---|---|---|
| CPU Utilization | Container Runtime | Scale out when average CPU exceeds 70–75% to prevent thread scheduling delays and ingestion latency. |
| Memory Utilization | Container Runtime | Scale out when memory exceeds 75% to prevent memory_limiter drops. |
otelcol_receiver_refused_spans | Collector Internal | Critical Custom Metric: Non-zero values indicate the Collector is rejecting spans due to backpressure or memory exhaustion. Must trigger rapid scale-out! |
otelcol_receiver_refused_metric_points | Collector Internal | Non-zero values indicate metrics are actively being dropped at ingress. |
otelcol_exporter_queue_size | Collector Internal | High queue depth relative to queue_capacity indicates downstream backend saturation or network bottlenecks. |
HPA Autoscaling Policy Best Practice: Configure HPA with a rapid scale-up policy (e.g., doubling replicas within 15 seconds upon a traffic spike) and a conservative, stabilized scale-down window (e.g., 300 seconds) to prevent pod flapping and cache thrashing during transient dips in telemetry volume.
A platform team deploys a tail-based sampling pipeline on a cluster of four OpenTelemetry Collector Gateway pods placed behind a standard Kubernetes ClusterIP Service. Although the tail-sampling processor is configured to retain 100% of traces that contain HTTP 500 errors, engineers notice that traces with server errors in downstream microservices are frequently dropped or appear with missing child spans. What is the fundamental reason standard Layer 4 or Layer 7 load balancers fail for tail-based sampling?
Kubernetes ClusterIP Services encrypt gRPC payloads, preventing the Collector from reading span attributes
Standard load balancers distribute requests or connections across pods without TraceId affinity, scattering spans of the same trace across different collector replicas so no single pod observes the complete trace
The tail-sampling processor can only operate on single-node VM architectures and crashes when scaled horizontally
Layer 7 load balancers automatically strip W3C traceparent headers from HTTP requests during proxying
To solve the trace-scattering problem in a horizontally scalable Two-Tier Collector architecture, an engineering team installs the loadbalancing exporter on Tier 1 Node Agents. How does the loadbalancing exporter ensure that tail-based sampling can evaluate complete traces across downstream Gateway Collector pods?
It broadcasts every incoming span to all downstream Gateway Collector replicas simultaneously using UDP multicast
It buffers an entire trace inside the Tier 1 Agent memory until the root span finishes, then sends the assembled trace in a single batch
It extracts the TraceId from every span and uses consistent hashing to deterministically route all spans sharing the same TraceId to the identical downstream Gateway replica
It converts all spans into Prometheus metrics before forwarding them through a round-robin proxy
A site reliability engineering team is setting up a Horizontal Pod Autoscaler (HPA) for an OpenTelemetry Collector Gateway deployment in Kubernetes. In addition to monitoring container CPU and memory utilization, which custom Prometheus metric exposed by the Collector's internal telemetry (:8888/metrics) should be used as a primary scaling metric to prevent telemetry data loss during traffic surges?
otelcol_process_uptime to measure how long the Go runtime has been executing
otelcol_processor_batch_metadata_cardinality to monitor the size of resource attributes
otelcol_exporter_sent_spans to measure the total number of successfully exported spans
otelcol_receiver_refused_spans (and otelcol_receiver_refused_metric_points) to detect when collectors are saturated and actively rejecting incoming telemetry
Sections you finish are checked off in the contents.