3.4 Composability & Extension: Plugin Interfaces, Distributions & Contrib
Key Takeaways
The OpenTelemetry SDK exposes specified plugin interfaces, including Sampler, SpanProcessor, exporters, MetricReader, View, TextMapPropagator, resource detectors, and IdGenerator, so behaviour can be extended without forking.
Span processors run in registration order and receive a writable span in OnStart, so enrichment processors must be registered before the batch processor that exports.
Exporters receive read-only spans, and the specification says Export should not be called concurrently for one exporter instance, so enrichment belongs in a processor, not an exporter.
A distribution wraps upstream OpenTelemetry components as pure, plus, or minus packages; it is not a fork, so code written against the API keeps working.
The Collector is composed at build time from component modules listed in an OpenTelemetry Collector Builder manifest and wired together at run time by YAML pipelines.
3.4 Composability & Extension: Plugin Interfaces, Distributions & Contrib
Quick Answer: OpenTelemetry is built from interchangeable parts. The SDK exposes standard plugin interfaces, including samplers, span and log processors, exporters, metric readers, views, propagators, resource detectors, and ID generators, so teams can add behaviour without forking the project. Instrumentation libraries, vendor exporters, and extra propagators live in contrib repositories; vendors package upstream parts as distributions; and the Collector is composed at build time from receivers, processors, exporters, connectors, and extensions.
The "Composability and Extension" competency asks a practical question: when the default behaviour is not enough, where do you plug in your change? Choosing the right extension point keeps instrumentation portable. The wrong choice usually means patching the SDK, losing data, or hard-coding a vendor into application code.
Why OpenTelemetry Is Designed to Be Composed
Three design decisions make OpenTelemetry composable:
- API/SDK separation (Section 3.1). Libraries depend only on the API, so any SDK, whether the reference SDK or a vendor distribution, can be installed underneath without touching library code.
- Specified plugin interfaces. The specification defines the interfaces the SDK must expose (for example
Sampler,SpanProcessor,SpanExporter,MetricReader,TextMapPropagator). A custom component written against one of them works with the standard pipeline. - A common data model and protocol. Because every component speaks the same data model and OTLP, you can swap an exporter or add a Collector stage without changing the producers.
SDK Extension Points
| Extension point | Key operations | Typical custom use | Where it is registered |
|---|---|---|---|
Sampler | ShouldSample, GetDescription | Drop health-check traces at the head, or sample VIP tenants at 100% | TracerProvider (often as the root of ParentBased) |
SpanProcessor | OnStart, OnEnd, ForceFlush, Shutdown | Enrich spans at start, copy baggage, redact attributes, route to an exporter | TracerProvider, in registration order |
SpanExporter / MetricExporter / LogRecordExporter | Export(batch), ForceFlush, Shutdown | Send data to a backend that does not accept OTLP | Wrapped by a processor or a metric reader |
MetricReader | Collect, Shutdown | Push on a schedule or serve a pull endpoint | MeterProvider (several readers allowed) |
View | Instrument selection plus stream configuration | Rename a stream, drop attributes, change buckets or aggregation | MeterProvider |
LogRecordProcessor | OnEmit, ForceFlush, Shutdown | Enrich or filter log records before export | LoggerProvider |
TextMapPropagator | Inject, Extract, Fields | Support a proprietary header format | Global propagators (usually inside a composite) |
| Resource detector | Detect | Add platform metadata not covered by the standard detectors | Resource builder at startup |
IdGenerator | New trace ID, new span ID | Generate IDs that another system requires (the AWS X-Ray ID generator embeds a timestamp) | TracerProvider |
Several contract details from the specification matter when you write or combine these components:
- Processors run in registration order. Each
OnStartreceives a writable span, so an enrichment processor must be registered before the batch processor that exports. OnEndreceives a read-only span. An exporter cannot add attributes; enrichment belongs inOnStart(a newerOnEndinghook, still in Development status, allows a last change just before the span becomes read-only).Exportis serialized: the specification saysExportshould not be called concurrently for the same exporter instance, and the built-in processors follow that rule, so an exporter does not need its own locking aroundExport.ForceFlushandShutdowncascade: calling them on a provider calls them on every registered processor or reader, which in turn flush their exporters.
Example: A Custom Enrichment Processor
The following Python processor copies a tenant.id baggage entry onto every new span. It is registered before the batch processor so that the attribute exists when the span is exported:
from opentelemetry import baggage
from opentelemetry.sdk.trace import SpanProcessor, TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
class TenantSpanProcessor(SpanProcessor):
def on_start(self, span, parent_context=None):
tenant = baggage.get_baggage("tenant.id", parent_context)
if tenant is not None:
span.set_attribute("tenant.id", tenant)
provider = TracerProvider()
provider.add_span_processor(TenantSpanProcessor()) # enrichment first
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter())) # export last
You rarely need to write this one yourself: ready-made baggage span processors exist in the Java, JavaScript, Python, and Go contrib repositories. The example shows the pattern that every enrichment processor follows.
Core, Contrib, and Instrumentation Libraries
Each language has a core repository (API, SDK, OTLP exporter) and a contrib repository. Contrib holds the pieces that are optional or backend-specific:
- Instrumentation libraries for frameworks and clients (HTTP servers, database drivers, messaging clients).
- Extra propagators such as B3, Jaeger (deprecated), and AWS X-Ray.
- Resource detectors for cloud platforms and container runtimes.
- Vendor exporters, samplers such as the Jaeger remote sampler, and processors such as the baggage span processor.
Contrib components carry their own stability levels. The long-term goal is native instrumentation, where a library calls the OpenTelemetry API itself, so no separate instrumentation package is needed.
Distributions
A distribution is a customized package of an upstream OpenTelemetry component. It is a wrapper around upstream code, not a fork. The OpenTelemetry documentation describes three kinds:
| Type | What it does | Example |
|---|---|---|
| Pure | Same functionality as upstream, with packaging or default settings tuned for a backend | A vendor build of the Java agent that pre-sets the export endpoint |
| Plus | Upstream plus extra components that are not upstream | A Collector build that adds a vendor's own exporter |
| Minus | A subset of upstream, for smaller size or supportability | A Collector with only the receivers and exporters a company supports |
Because distributions keep the upstream APIs and data model, code instrumented against the OpenTelemetry API keeps working if you switch distributions.
Extending the Collector
The Collector is composable at build time. Every component (receiver, processor, exporter, extension, connector) is a Go module with a factory, and a binary contains only the components compiled into it. The OpenTelemetry Collector Builder (OCB, Section 9.1) turns a manifest listing those modules into a custom binary; a custom component is simply one more module in that manifest. At run time, the YAML configuration chooses which compiled components to instantiate and how to wire them into pipelines. Each component declares a stability level per signal (Development, Alpha, Beta, Stable, Deprecated, or Unmaintained), which you should check before relying on it in production.
Choosing the Right Extension Point
| Requirement | Extension point |
|---|---|
| Decide at trace start whether to record | Custom Sampler, usually wrapped in ParentBased |
| Add or remove span attributes before export | SpanProcessor.OnStart (or a Collector processor for fleet-wide rules) |
| Send data to a non-OTLP backend | Custom exporter, or OTLP to a Collector that has the backend's exporter |
| Support a proprietary trace header | Custom TextMapPropagator combined with tracecontext |
| Add platform metadata to every signal | Custom resource detector |
| Change a metric's buckets or drop an attribute | View |
| Add a backend-specific component to the Collector | Build a custom Collector with OCB |
A platform team wants every span to carry the tenant.id value that the API gateway places in W3C Baggage. Application code must not change, and the value must be present when spans are exported. Which extension point is the best fit?
A custom SpanProcessor whose OnStart copies the baggage entry onto the span, registered before the BatchSpanProcessor
A custom SpanExporter that adds tenant.id to each span just before serialization
A Collector attributes processor that reads the baggage header from incoming OTLP requests
A custom IdGenerator that encodes the tenant into every trace ID
An observability vendor ships a package that bundles the upstream OpenTelemetry Collector with its own backend exporter, which is not part of the upstream project, and changes no upstream code. How does the OpenTelemetry documentation classify this package?
A fork, because it contains code that upstream does not maintain
A minus distribution, because vendor packages always remove components
A plus distribution, because it adds components that are not in upstream
A pure distribution, because the upstream code is unchanged
A team exports traces through an AWS X-Ray pipeline that requires trace IDs whose first 4 bytes encode the start time as epoch seconds. The rest of the SDK pipeline should stay standard. Which SDK extension point addresses this requirement?
A custom Sampler that rewrites the trace ID when it makes the sampling decision
A custom TextMapPropagator that converts IDs only when injecting headers
A View that renames the trace ID field in exported data
An IdGenerator configured on the TracerProvider, such as the X-Ray ID generator from contrib
Sections you finish are checked off in the contents.