7.1 Span Processors & Exporters
Key Takeaways
A SpanProcessor is the OpenTelemetry SDK component that intercepts span lifecycle events: onStart executes synchronously when a span begins to enrich attributes or context, while onEnd executes synchronously when a span completes to route the immutable span to buffer queues or exporters.
SimpleSpanProcessor hands each span to the exporter as soon as it ends; in many SDKs that puts export work on the request path, so it suits debugging, tests, and very low-volume or short-lived processes rather than busy services.
BatchSpanProcessor decouples application execution from telemetry export by buffering completed spans in a bounded queue and exporting them on a background worker thread, making it the recommended choice for production services.
The BatchSpanProcessor's key parameters are maxQueueSize (default 2048), scheduledDelayMillis (default 5000 ms), maxExportBatchSize (default 512, never larger than the queue size), and exportTimeoutMillis (default 30000 ms), set through the OTEL_BSP_* environment variables.
Applications must invoke TracerProvider.shutdown() or forceFlush() prior to process termination to gracefully drain buffered spans from memory and prevent data loss during container shutdown or Kubernetes pod eviction.
7.1 Span Processors & Exporters
Quick Answer: A SpanProcessor is the OpenTelemetry SDK component that hooks directly into the span lifecycle through two core callbacks:
onStart(span, parentContext)(which executes synchronously when a span is started to enrich or mutate span attributes) andonEnd(span)(which executes synchronously when a span ends to forward the completed span for batching or export). OpenTelemetry provides two standard implementations: SimpleSpanProcessor (which hands each span to the exporter as soon as it ends, suited to debugging, tests, and very low-volume processes) and BatchSpanProcessor (which queues spans in a bounded in-memory queue and exports them in batches on a background thread, the recommended choice for production services).
In the OpenTelemetry architecture, the Tracer API is decoupled from the Tracer SDK. While application code and instrumentation libraries invoke API methods—such as tracer.startSpan() and span.end()—they remain completely agnostic of how telemetry data is buffered, transformed, serialized, or transmitted over the network. The OpenTelemetry SDK pipeline bridges this gap. At the center of the trace SDK pipeline sit Span Processors and Span Exporters.
The Anatomy of a SpanProcessor
A SpanProcessor represents an extensible pipeline stage within the SDK. Registered with the TracerProvider, one or more span processors receive notifications whenever a span changes state. The OpenTelemetry specification defines these methods on the SpanProcessor interface (a further OnEnding hook, called just before the span becomes read-only, is still in Development status):
+-----------------------------------------------------------------------------------+
| SpanProcessor Interface |
+-----------------------------------------------------------------------------------+
| onStart(ReadWriteSpan span, Context parentContext) : void |
| --> Invoked synchronously when tracer.startSpan() is called |
| --> Span is still WRITABLE (name and attributes can be changed) |
+-----------------------------------------------------------------------------------+
| onEnd(ReadableSpan span) : void |
| --> Invoked synchronously when span.end() is called |
| --> Span is IMMUTABLE (read-only final representation) |
| --> Entry point for handoff to exporter or internal queue |
+-----------------------------------------------------------------------------------+
| forceFlush(timeoutMillis) : CompletableResultCode |
| --> Forces export of all spans that reached onEnd but are still buffered |
+-----------------------------------------------------------------------------------+
| shutdown(timeoutMillis) : CompletableResultCode |
| --> Flushes remaining buffered spans, releases worker threads and resources |
+-----------------------------------------------------------------------------------+
1. onStart(ReadWriteSpan span, Context parentContext)
The onStart method is invoked synchronously on the application caller's thread at the instant tracer.startSpan() or tracer.spanBuilder().startSpan() is executed. At this point, the span is represented as a ReadWriteSpan, meaning its attributes, name, and internal state are mutable.
Key characteristics of onStart:
- Attribute Enrichment: Allows custom processors to inspect the ambient
parentContext(such as distributed baggage or security tokens) and inject internal metadata (e.g.,tenant.id,datacenter.zone,thread.id) directly into the span attributes before any child spans are generated. - Execution Constraints: Because
onStartexecutes inline on the application thread, any logic insideonStartmust be non-blocking and highly performant. Blocking insideonStartdirectly delays the application's business logic. - Sampling Interaction:
onStartis only called for spans that the configuredSamplerhas decided to sample (or record). If a sampler drops a span completely (DROP),onStartis never invoked.
2. onEnd(ReadableSpan span)
The onEnd method is invoked synchronously on the application caller's thread at the instant span.end() is executed. At this milestone, the span's end timestamp is permanently captured, and the span transitions into an immutable ReadableSpan.
Key characteristics of onEnd:
- Immutability: Attributes, events, status codes, and timestamps cannot be added or modified after
onEndis triggered. - Pipeline Gateway:
onEndserves as the handoff point where completed spans are either immediately dispatched to an exporter or enqueued in memory for background batching. - Execution Constraints: Similar to
onStart,onEndexecutes on the application caller thread. If the processor performs synchronous network I/O or disk writes insideonEnd, the application thread blocks until that operation finishes.
3. forceFlush() and shutdown()
forceFlush(timeout): Instructs the processor to immediately export all spans that have completed (onEndinvoked) but remain buffered in memory. This is a blocking operation with a configurable timeout.shutdown(timeout): Shuts down the processor. It first performs a force-flush of all pending spans, halts background worker threads, releases network sockets, and flags the processor as inactive. Once shut down, subsequent invocations ofonStartandonEndare silently dropped.
SimpleSpanProcessor vs. BatchSpanProcessor
The OpenTelemetry SDK ships with two standard implementations of SpanProcessor. Selecting between them is one of the most critical operational decisions in application instrumentation.
| Architectural Dimension | SimpleSpanProcessor | BatchSpanProcessor |
|---|---|---|
| Export Timing | Immediate (synchronous on onEnd) | Delayed (asynchronous batch timer or buffer capacity) |
| Calling Thread Impact | Blocks caller thread during export network I/O | Non-blocking (enqueues to in-memory ring buffer) |
| Network Efficiency | Very low (1 RPC per span) | Very high (hundreds of spans per compressed RPC) |
| Memory Footprint | Near zero (no buffer queue) | Bounded in-memory buffer (max_queue_size) |
| Data Loss on Sudden Crash | Low (spans sent immediately upon completion) | Risk of dropping buffered spans if not gracefully flushed |
| Primary Use Cases | Serverless functions (AWS Lambda), local testing, CLI tools | Production microservices, high-throughput web APIs |
Deep-Dive: SimpleSpanProcessor
The SimpleSpanProcessor is designed for environments where background threading or memory buffering is undesirable or prohibited:
Application Thread: [End Span] ---> onEnd() ---> Exporter.export([span]) ---> [Network I/O Wait] ---> Return to App
- Synchronous Bottleneck: In a web application processing 1,000 requests per second, each request might generate 5 spans. If an exporter takes 20ms to serialize and transmit a span over HTTP to a collector,
SimpleSpanProcessoradds of artificial network latency to every single user request. - Serverless Considerations: In serverless execution environments like AWS Lambda, Google Cloud Functions, or Azure Functions, the execution environment can be frozen as soon as the request handler returns. If a background thread is still holding spans, they stay stranded until a later invocation (or are lost if the environment is destroyed). Handing each span to the exporter immediately avoids that queue, but the more common pattern is to keep the
BatchSpanProcessorand callforceFlush()on the provider before the handler returns.
Deep-Dive: BatchSpanProcessor
The BatchSpanProcessor is the industry standard for production web applications and microservices:
Application Thread: [End Span] ---> onEnd() ---> Enqueue to Ring Buffer (O(1)) ---> Return immediately
│
▼
[Bounded Memory Queue]
│
Background Worker Thread (Periodically)
│
▼
Exporter.export([Span1, Span2, ...])
- Asynchronous Decoupling: When
span.end()is invoked, the span is placed into an internal, thread-safe bounded queue (often implemented as a circular ring buffer or concurrent queue). This operation takes fractions of a microsecond ( lock-free enqueue). The application thread immediately returns to serving business logic. - Background Worker Thread: A dedicated daemon thread runs in the background. It wakes up either when a configurable timer expires (
scheduledDelayMillis) or when the accumulated spans reach a threshold batch size (maxExportBatchSize). It extracts the batch of spans and callsexporter.export(batch)in a single, multiplexed, compressed network call.
BatchSpanProcessor Tuning Parameters
Know the four tuning parameters of the BatchSpanProcessor, their default values, their environment variables, and how misconfiguring them affects reliability. The parameter names below are the specification's names; each language uses its own casing (Python, for example, uses schedule_delay_millis and export_timeout_millis).
| Parameter Name | OpenTelemetry Environment Variable | Canonical Default | Operational Definition & Constraints |
|---|---|---|---|
maxQueueSize | OTEL_BSP_MAX_QUEUE_SIZE | 2048 | Maximum number of completed spans held in the internal memory buffer awaiting export. If the queue is full, new spans are dropped. |
scheduledDelayMillis | OTEL_BSP_SCHEDULE_DELAY | 5000 (5s) | Interval in milliseconds between consecutive batch export attempts by the background worker thread. |
maxExportBatchSize | OTEL_BSP_MAX_EXPORT_BATCH_SIZE | 512 | Maximum number of spans exported in a single batch. Must be less than or equal to maxQueueSize. When a full batch is queued, it is exported without waiting for the delay. |
exportTimeoutMillis | OTEL_BSP_EXPORT_TIMEOUT | 30000 (30s) | Maximum time in milliseconds an exporter RPC is allowed to take before being aborted and cancelled. |
1. maxQueueSize (Default: 2048)
maxQueueSize defines the upper bound of the in-memory circular buffer. This bounding is vital: in the event of a downstream network partition or an overloaded OpenTelemetry Collector, spans cannot be exported. If the queue were unbounded, the application would rapidly experience memory leaks and crash with an Out-Of-Memory (OOM) error.
- Drop Semantics: When
maxQueueSizeis reached (e.g., 2048 spans in memory), any subsequent span arriving atonEnd()is immediately dropped. Drops appear in SDK diagnostic logs and, in SDKs that implement the SDK self-observability conventions, in theotel.sdk.processor.span.processedmetric witherror.typeset toqueue_full. - Production Sizing: In high-throughput services handling tens of thousands of requests per second, a burst of traffic can saturate 2048 spans in milliseconds. SREs frequently increase this value to
8192or16384while ensuring adequate heap allocation.
2. scheduledDelayMillis (Default: 5000ms / 5s)
This parameter governs the maximum delay before the background worker thread drains accumulated spans.
- Latency vs. Network Overhead: A 5000ms delay means spans may reside in process memory for up to 5 seconds before being exported. In real-time debugging scenarios or canary deployments, 5 seconds might feel sluggish. Reducing
scheduledDelayMillisto1000msor500mscauses spans to appear in backends much faster, at the cost of more frequent network calls and smaller, less efficient batches.
3. maxExportBatchSize (Default: 512)
This parameter limits the number of spans packed into a single payload sent to the exporter.
- Mandatory Constraint: The specification requires
maxExportBatchSizeto be less than or equal tomaxQueueSize. If a developer sets a batch size of 4096 while the queue holds 2048, the SDK rejects or clamps the value (behaviour varies by language). - Batch Efficiency: A larger batch size amortizes HTTP/gRPC header overhead and achieves superior gzip compression ratios across similar span attributes.
4. exportTimeoutMillis (Default: 30000ms / 30s)
If the telemetry collector or network hangs, the exporter RPC will not block indefinitely. After 30 seconds, the worker thread cancels the request, logs an exporter timeout error, and attempts to recover on the next scheduled cycle.
Graceful Shutdown and ForceFlush Mechanics
One of the most frequent causes of telemetry loss in Kubernetes environments is the failure to handle process termination lifecycles properly.
Why Process Termination Drops Spans
When an application using BatchSpanProcessor terminates:
- Spans generated during the final 5 seconds of execution still reside in the in-memory buffer queue.
- In-flight export network requests may be halfway through a TCP socket transmission.
- If the process terminates via
SIGKILLor exitsmain()abruptly, the operating system instantly destroys the process heap, permanently destroying all spans buffered in memory.
Kubernetes Pod Eviction & SIGTERM
In Kubernetes, when a pod is deleted, upgraded, or autoscaled down, the kubelet executes a graceful shutdown sequence:
- The kubelet removes the pod from service endpoints (stopping new ingress traffic).
- The kubelet sends a
SIGTERMsignal to PID 1 inside the container. - The container is granted a grace period governed by
terminationGracePeriodSeconds(default: 30 seconds). - If the process is still running when the grace period expires, the kubelet issues an uncatchable
SIGKILL.
To ensure zero telemetry loss, applications must catch SIGTERM and SIGINT, invoke TracerProvider.shutdown(), and allow the SDK to drain its queues before the application process exits.
import atexit
import signal
import sys
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
# Initialize TracerProvider
provider = TracerProvider()
# Configure OTLP Exporter
otlp_exporter = OTLPSpanExporter(endpoint="http://otel-collector:4317", insecure=True)
# Configure BatchSpanProcessor with tuned parameters
batch_processor = BatchSpanProcessor(
otlp_exporter,
max_queue_size=8192, # Increased buffer for traffic bursts
schedule_delay_millis=1000, # Flush every 1s for near-real-time visibility
max_export_batch_size=1024, # Efficient batching <= max_queue_size
export_timeout_millis=10000 # 10s export timeout
)
provider.add_span_processor(batch_processor)
trace.set_tracer_provider(provider)
# Graceful shutdown handler
def graceful_shutdown(signum=None, frame=None):
print("Flushing telemetry and shutting down TracerProvider...")
# shutdown() drains all BatchSpanProcessors and closes exporter sockets
provider.shutdown()
if signum is not None:
sys.exit(0)
# Register signal traps for Kubernetes SIGTERM and SIGINT
signal.signal(signal.SIGTERM, graceful_shutdown)
signal.signal(signal.SIGINT, graceful_shutdown)
atexit.register(provider.shutdown)
shutdown() vs. forceFlush()
- Use
forceFlush()when the application is remaining alive but must ensure that all spans completed up to this point have reached the exporter (e.g., at the end of an AWS Lambda invocation handler or after a crucial transactional milestone). - Use
shutdown()when the application runtime is permanently closing. It drains all queues, shuts down background threads, closes network connections, and permanently deactivates the SDK.
A high-throughput e-commerce checkout microservice experiences sporadic latency spikes of 150ms on p99 requests. Profiling reveals that the application thread handling checkout is synchronously waiting on socket network I/O to the telemetry collector whenever a span finishes. An inspection of the codebase indicates that the SDK configuration was recently refactored. Which configuration flaw is directly causing this latency degradation on the user-facing request thread?
The exporter was configured with OTLP/HTTP instead of OTLP/gRPC, causing HTTP keep-alive connection renegotiation on every span
The head-based sampler was set to ParentBased(AlwaysOn), forcing child spans to evaluate sampling decisions synchronously
The TracerProvider was configured with SimpleSpanProcessor instead of BatchSpanProcessor, causing synchronous network export on the caller thread during onEnd
The scheduledDelayMillis on the BatchSpanProcessor was set to 0, forcing background worker threads into an infinite busy loop
During sudden traffic spikes, an application's SDK logs show 'BatchSpanProcessor queue is full, dropping span'. Network bandwidth is plentiful, the Collector is lightly loaded, and exports succeed quickly. The configuration uses a queue size of 2048, a schedule delay of 5000 ms, and a batch size of 512. Which change most directly stops these burst-time drops?
Increase the export timeout from 30000 ms to 120000 ms so each export has more time to finish
Replace the BatchSpanProcessor with SimpleSpanProcessor so no queue exists to overflow
Reduce the maximum export batch size from 512 to 64 so that smaller payloads travel faster
Raise the maximum queue size so bursts can be absorbed, and raise the maximum export batch size so each export call drains more spans
A serverless microservice running as an AWS Lambda function processes incoming payment webhooks in under 80 milliseconds. When the function finishes executing, the Lambda runtime immediately freezes the execution environment. The operations team notices that spans generated during function execution never appear in the backend Jaeger dashboard, or only appear hours later during subsequent invocations. What is the root cause of this issue, and what is the proper architectural remedy?
The BatchSpanProcessor buffers spans in memory on a background worker thread that is frozen as soon as the Lambda handler returns; the function must call TracerProvider.forceFlush() or TracerProvider.shutdown() before exiting, or use SimpleSpanProcessor
AWS Lambda security groups block outbound UDP packets required by the OpenTelemetry Protocol (OTLP) specification
The W3C TraceContext propagator cannot serialize traceparent headers inside serverless runtimes because of ephemeral file systems
The OpenTelemetry API prevents spans from being ended when running on container virtualization runtimes
Sections you finish are checked off in the contents.