8.1 Head-Based Sampling & Built-in Samplers
Key Takeaways
Head-based sampling makes the irreversible sampling and recording decision at the ingress ('head') of a trace when the root span begins, before downstream execution or failure modes occur.
The sampling decision configures the lowest bit (0x01) of the TraceFlags byte in SpanContext, which is serialized into W3C traceparent headers and propagated downstream across network boundaries.
The OpenTelemetry SDK Sampler interface evaluates to one of three SamplingResult states: DROP (do not record, do not export), RECORD_ONLY (record attributes and events in memory without exporting), or RECORD_AND_SAMPLE (record in memory and export).
The classic built-in samplers are AlwaysOn, AlwaysOff, TraceIdRatioBased, and ParentBased (the SDK default is ParentBased with an AlwaysOn root); newer specification releases add ProbabilitySampler, AlwaysRecord, and composable samplers.
TraceIdRatioBased decides deterministically from the trace ID, but its exact algorithm was never standardized, so different SDKs can disagree; it is deprecated in favor of ProbabilitySampler, which compares 56 random trace-ID bits with a threshold.
8.1 Head-Based Sampling & Built-in Samplers
Quick Answer: Head-based sampling makes the sampling and recording decision at the very start ("head") of a trace when the root span is initiated, before child operations execute, latencies are measured, or errors are thrown. The decision sets the low-order bit (
0x01) ofTraceFlagswithin the immutableSpanContext, which is propagated across microservices via the W3Ctraceparentheader. OpenTelemetry SDKs provide three sampling result states (DROP,RECORD_ONLY,RECORD_AND_SAMPLE) and the classic built-in samplersAlwaysOn,AlwaysOff,TraceIdRatioBased, and the production-standardParentBasedwrapper.
In modern microservice ecosystems processing hundreds of thousands of transactions per second, capturing 100% of all distributed traces is economically and operationally prohibitive. Emitting, transmitting, serializing, indexing, and storing billions of spans consumes massive network egress bandwidth, exhausts collector CPU cycles, and inflates cloud storage invoices by tens of thousands of dollars each month. To maintain manageable overhead while preserving deep diagnostic observability, engineering teams must implement sampling strategies.
In the OpenTelemetry architecture, sampling falls into two distinct operational models: head-based sampling and tail-based sampling. This section delves into head-based sampling, the SDK-level decision contract, the W3C TraceFlags bitmask, and the built-in samplers specified by the OpenTelemetry standard.
Fundamentals of Head-Based Sampling: The Ingress Decision
Head-based sampling is defined by where and when the sampling decision occurs: at the ingress point ("head") of a distributed request, when the root span is initially created.
Incoming Request (HTTP / gRPC)
│
▼
┌──────────────────────────────────────────────┐
│ Microservice Ingress: Root Span Creation │
│ Sampler.shouldSample() is invoked │
│ │
│ Decision: [RECORD_AND_SAMPLE] (low bit = 1) │
│ or [DROP] (low bit = 0) │
└──────────────────────┬───────────────────────┘
│
┌──────────────┴──────────────┐
▼ ▼
[TraceFlags = 0x01] [TraceFlags = 0x00]
Span recorded & exported Span dropped / ignored
Context injected downstream Context injected downstream
The Operational Advantage: Maximum Ingress Efficiency
The primary advantage of head-based sampling is resource efficiency:
- Minimal CPU Overhead: The sampling algorithm evaluates a single mathematical or rule-based check once at the root span inception.
- Zero Downstream Wastage: When a trace is dropped at the head, downstream services that respect parent sampling do not waste CPU cycles serializing span JSON/Protobuf payloads or buffering spans in memory.
- Immediate Network Savings: Network bandwidth between microservices and observability collectors is conserved because unselected traces never leave application memory.
The Fundamental Limitation: Lack of Foresight
The fundamental limitation of head-based sampling is that the decision is made before the request executes:
- When a user clicks "Submit Order", the root span begins.
- At that exact microsecond, the SDK cannot know whether a downstream database connection will deadlock after 8 seconds.
- The SDK cannot know whether an inventory microservice will throw an unhandled
NullPointerException. - The SDK cannot know whether a payment gateway will return an HTTP 500 status code.
If the head sampler decides to drop the trace because a probabilistic 5% ratio was configured, that trace is gone forever. When the downstream database deadlocks 3 hops later, no trace exists to debug the incident. Overcoming this limitation requires understanding the SDK decision states and configuring the right samplers.
The Sampling Contract: The TraceFlags Low Bit
When an OpenTelemetry SDK starts a span, the Tracer delegates the decision to the configured Sampler implementation by invoking its shouldSample() method. The result of this decision is immutably stored in the span's SpanContext.
The Anatomy of TraceFlags
A SpanContext contains four core members:
TraceId(16 bytes / 32 hexadecimal characters)SpanId(8 bytes / 16 hexadecimal characters)TraceFlags(1 byte / 8-bit bitmap)TraceState(list of key-value pairs for vendor routing)
The TraceFlags byte is the standard cross-process vehicle for communicating sampling decisions. The sampling decision is the least significant bit (mask 0x01); W3C Trace Context Level 2 also defines the random flag (0x02), which newer probability samplers use:
TraceFlags Byte: [ 0 | 0 | 0 | 0 | 0 | 0 | 0 | S ]
│
└── S = Sampled Bit (Bit 0)
0 = Not Sampled (0x00)
1 = Sampled (0x01)
01(Sampled): Indicates that the trace was selected for recording and export. Downstream services extracting this context recognize that the parent was sampled and should record child spans accordingly.00(Not Sampled): Indicates that the trace was not selected for export. Downstream services extracting this context recognize that the parent was not sampled and drop child spans.
Wire Serialization in W3C traceparent
When context is propagated across network boundaries (such as HTTP headers), TraceFlags is serialized as the final 2-character hexadecimal field of the traceparent header:
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
│ │ │ │
Version TraceId (16 bytes) SpanId (8 bytes) TraceFlags (01 = Sampled)
If the low bit is 00 (traceparent: 00-4bf9...-00f0...-00), the downstream recipient extracts isSampled() == false.
The SDK Decision Model: SamplingResult States
In application code and custom SDK plugins, the Sampler interface returns a structured SamplingResult object containing a Decision enum, an optional set of span attributes to append, and an updated TraceState.
The OpenTelemetry specification defines exactly three SamplingDecision enum values:
┌────────────────────────────────────────────────────────────────────────┐
│ SamplingDecision Enum Values │
├────────────────────┬────────────────────┬──────────────────────────────┤
│ Decision State │ isRecording() │ Low Bit of TraceFlags (Wire) │
├────────────────────┼────────────────────┼──────────────────────────────┤
│ DROP │ false │ 0 (0x00) │
│ RECORD_ONLY │ true │ 0 (0x00) │
│ RECORD_AND_SAMPLE │ true │ 1 (0x01) │
└────────────────────┴────────────────────┴──────────────────────────────┘
1. DROP
- The span is completely discarded. Calling
span.isRecording()returnsfalse. - Any operations attempting to set attributes (
span.setAttribute()), add span events (span.addEvent()), or record exceptions are immediate no-ops, consuming zero memory or CPU cycles. - The span is never dispatched to registered
SpanProcessorinstances (such asBatchSpanProcessor) and is never exported. - The low bit of
TraceFlagsis0(0x00).
2. RECORD_ONLY
- The span records execution data in memory:
span.isRecording()returnstrue. - Attributes, events, links, and status changes are populated and processed by active
SpanProcessor.onStart()andSpanProcessor.onEnd()callbacks. - Crucial Detail: The low bit of
TraceFlagsis set to0(0x00). The span is not exported to the telemetry backend. - Primary Use Case: Generating real-time in-process metrics (e.g., span-to-metrics counters or latency histograms calculated by custom processors) across 100% of requests without incurring the massive network egress bandwidth and storage cost of exporting full trace graphs.
3. RECORD_AND_SAMPLE
- Full observability enabled:
span.isRecording()returnstrue. - All attributes, events, and status updates are recorded in memory.
- The low bit of
TraceFlagsis set to1(0x01). - The completed span is forwarded to registered
SpanProcessorsand queued for network transmission by exporters (e.g., OTLP gRPC/HTTP exporter).
Built-in OpenTelemetry Samplers
These four classic samplers are the ones every SDK exposes through OTEL_TRACES_SAMPLER:
Built-in Samplers
│
┌──────────────────┬──────────┴──────────┬──────────────────┐
▼ ▼ ▼ ▼
AlwaysOn AlwaysOff TraceIdRatioBased ParentBased
(100% Export) (0% Export) (Probabilistic N%) (Tree Wrapper)
1. AlwaysOn (Environment identifier: always_on)
- Behavior: Unconditionally returns
RECORD_AND_SAMPLEfor every span evaluated. - Configuration: Requires no parameters or arguments.
- Use Cases: Development environments, automated CI/CD integration testing, low-traffic critical microservices (e.g., an internal authentication authority processing 10 requests per minute), or regulatory audit compliance systems where missing a single transaction is unacceptable.
- Drawbacks: Disastrous in high-throughput production systems; leads to memory exhaustion, network saturation, and exorbitant storage bills.
2. AlwaysOff (Environment identifier: always_off)
- Behavior: Unconditionally returns
DROPfor every span evaluated. - Configuration: Requires no parameters or arguments.
- Use Cases: Completely disabling tracing in non-critical environments, emergency circuit breaking during severe collector outages, or silencing high-volume batch data pipelines where tracing is undesired.
3. TraceIdRatioBased(ratio) (Environment identifier: traceidratio)
- Behavior: Probabilistically samples a configurable percentage of root traces based on a floating-point ratio between
0.0(0%) and1.0(100%). For example, a ratio of0.05samples approximately 5% of traces. - The Algorithm: The decision must be deterministic: the sampler derives it from the
TraceId(commonly by comparing part of the ID, often its lower 64 bits, withratiotimes the maximum value) rather than from a random number generator. A sampler with a higher ratio must also sample every trace that a lower-ratio sampler would sample. - Compatibility Warning: The specification never fixed the exact algorithm, so SDKs in different languages (or different SDK versions) can reach different decisions for the same trace ID. The specification therefore recommends using it only as the root sampler inside
ParentBased. - Deprecation:
TraceIdRatioBasedis now deprecated in favor ofProbabilitySampler(Development status).ProbabilitySampleruses the W3C Level 2 randomness: it compares the trace ID's rightmost 56 bits (R) with a rejection threshold (T) derived from the ratio, samples when R ≥ T, and records the threshold asthin the OpenTelemetrytracestateentry, so every SDK makes the same decision. SDKs must keepTraceIdRatioBasedworking unchanged until at least January 1, 2027. - Critical Pitfall: Used alone on downstream services,
TraceIdRatioBasedignores the parent's sampled flag and decides independently. With different ratios, or with SDKs that hash differently, a child span can be dropped while its parent was sampled, creating broken traces.
4. ParentBased(root_sampler) (Environment identifiers: parentbased_always_on, parentbased_traceidratio, etc.)
- The Production Standard Wrapper: In production distributed systems, you should almost never configure
TraceIdRatioBasedorAlwaysOndirectly as your standalone sampler. You must wrap them insideParentBased! - Core Architectural Purpose:
ParentBasedis a composite sampler that respects the upstream caller's sampling decision for all child spans, guaranteeing that a distributed trace is captured as an unbroken, contiguous tree.
Incoming Span Creation
│
├───── Has Parent Context? ──────┐
│ (Yes) │ (No - Root Span)
▼ ▼
Is Parent Sampled? Delegate to root_sampler
├── YES (01) ──> SAMPLE CHILD (e.g., TraceIdRatioBased(0.10))
└── NO (00) ──> DROP CHILD
Two further samplers appear in the specification. JaegerRemoteSampler polls a Jaeger-compatible endpoint for per-service sampling strategies. AlwaysRecord wraps another sampler and turns its DROP decisions into RECORD_ONLY, so processors (for example span-to-metrics processors) see every span while only sampled spans are exported.
Configurable Delegate Paths in ParentBased
While ParentBased has intelligent defaults, the OpenTelemetry specification allows engineering teams to customize five distinct delegation scenarios:
| Delegate Hook | Default Sampler | Operational Behavior |
|---|---|---|
root | User-configured (e.g. TraceIdRatioBased(0.05)) | Invoked when no parent context exists (the span is the root of a new trace). |
remoteParentSampled | AlwaysOn | Invoked when a span has a remote parent (extracted from network headers) whose low bit was 01. Honors the upstream caller's decision to trace. |
remoteParentNotSampled | AlwaysOff | Invoked when a span has a remote parent whose low bit was 00. Avoids creating orphan child spans when upstream decided not to trace. |
localParentSampled | AlwaysOn | Invoked when a span has an in-process local parent span that was sampled. Guarantees in-process child spans are preserved. |
localParentNotSampled | AlwaysOff | Invoked when a span has an in-process local parent span that was dropped. Prevents generating child telemetry for dropped local operations. |
Built-in Samplers Architectural Reference
| Sampler Name | Standard Env String | Configuration Argument | Default Ingress Behavior | Downstream Child Behavior | Typical Production Use Case |
|---|---|---|---|---|---|
AlwaysOn | always_on | None | Always returns RECORD_AND_SAMPLE | Overrides parent; forces sampling | Dev, staging, CI/CD, low-traffic audit paths |
AlwaysOff | always_off | None | Always returns DROP | Overrides parent; forces drop | Emergency disabling, high-throughput batch jobs |
TraceIdRatioBased | traceidratio | Floating-point 0.0 to 1.0 | Deterministic decision from the TraceId (algorithm not standardized; deprecated) | Evaluates ratio independently of parent (risks fragmentation) | Rarely used standalone; serves as delegate inside ParentBased |
ParentBased(AlwaysOn) | parentbased_always_on | None | Samples 100% of new root traces | Strictly inherits parent sampling flag | Services that originate critical root flows and propagate downstream |
ParentBased(AlwaysOff) | parentbased_always_off | None | Drops 100% of new root traces | Strictly inherits parent sampling flag | Downstream internal microservices that should never initiate their own root traces |
ParentBased(TraceIdRatioBased) | parentbased_traceidratio | Floating-point 0.0 to 1.0 (e.g. 0.05) | Samples configured ratio at root ingress | Strictly inherits parent sampling flag | The universal production standard for cloud microservices |
Real-World Production Scenario: Preventing "Swiss Cheese" Traces
Consider an enterprise banking platform with three microservices in series:
API Gateway → Account Service → Ledger Service.
The Failure Mode: Unwrapped Ratio Sampling
The operations team configures all three services with standalone TraceIdRatioBased(0.10):
- An incoming customer deposit hits the
API Gateway. - The
API Gatewayevaluates its localTraceIdRatioBasedsampler. The hash lands under the 10% threshold:RECORD_AND_SAMPLE(TraceFlags01). - The
API Gatewaycalls theAccount Servicevia HTTP, passingtraceparent: 00-4bf9...-00f0...-01. - The
Account Servicealso uses a standaloneTraceIdRatioBased(0.10), so it ignores the parent's sampled flag and makes its own decision. It runs a different language SDK from the gateway, and because the ratio algorithm was never standardized, its hash of the sameTraceIddoes not match the gateway's: its decision is effectively independent. - The
Ledger Servicegoes further and uses a ratio of0.02. Even where hashing agrees, a 2% sampler drops most traces that a 10% sampler kept. - The result in the observability backend is a fragmented "swiss cheese" trace: an API Gateway root span exists, but downstream account and ledger spans are missing.
The Architectural Resolution
The team updates all services to use ParentBased(TraceIdRatioBased(0.10)):
- The
API Gatewayacts as the root boundary, sampling 10% of new transactions. - Whenever the
API Gatewaysamples a trace, bothAccount ServiceandLedger ServicedetectremoteParentSampledand execute their defaultAlwaysOndelegate. - 100% of the child spans for that transaction are captured, delivering an unbroken, fully correlated distributed execution DAG.
An API Gateway (Java) calls an Order Service (Go), which calls a Payment Service (Python). All three configure TraceIdRatioBased(0.10) directly as their sampler. The gateway samples about 10% of traces, but far fewer complete end-to-end traces reach the backend, and Payment Service spans are often missing. What change fixes the fragmentation?
Increase the ratio to 1.0 on the Payment Service only while leaving the gateway at 0.10
Match the BatchSpanProcessor queue size on the Order Service to the Payment Service's network buffer
Switch propagation from W3C Trace Context to the Zipkin B3 single header
Wrap the ratio sampler in ParentBased on every service, so downstream services follow the sampled flag they receive
A high-throughput payment processing service processes 50,000 transactions per second. The site reliability team wants to compute accurate in-process request rate and duration metrics across 100% of incoming transactions without overwhelming their distributed tracing backend with billions of exported spans. Which OpenTelemetry SamplingResult decision should a custom or configured sampler return to fulfill these operational criteria?
RECORD_ONLY, because it records span attributes and execution events in memory for in-process metric computation without exporting the span over the network
DROP, because dropping the span entirely forces the OpenTelemetry SDK to write summary telemetry directly to standard output
RECORD_AND_SAMPLE, because spans must be exported to an external collector before any metric instruments can process their attributes
ALWAYS_OFF, because disabling the TracerProvider redirects span memory allocation to the global MetricsProvider
An engineer notices that a Java service and a Go service, both configured with TraceIdRatioBased(0.20), sometimes disagree about whether the same trace is sampled. Which statement explains this and the direction the specification now recommends?
The sampler hashes service.name and deployment.environment.name, so services with different names always disagree
The specification requires a deterministic decision from the trace ID but never fixed the algorithm, so SDKs can differ; it recommends using the sampler only as the root of ParentBased and is replacing it with ProbabilitySampler, which uses the trace ID's 56 random bits
The sampler uses a per-process random number generator, so disagreement is intended and cannot be avoided
The sampler reads a sampling token from W3C Baggage, which the Go SDK does not propagate
Sections you finish are checked off in the contents.