1.2 Core Telemetry Signals: Traces, Metrics, Logs & Baggage

Key Takeaways

  • The primary telemetry signals governed by OpenTelemetry are Distributed Traces, Metrics, and Logs, unified through standard context propagation and supplemented by runtime Baggage.

  • Distributed traces represent end-to-end transactional execution paths as directed acyclic graphs (DAGs) of spans, providing causal lineage, latency profiling, and error attribution across process boundaries.

  • Metrics offer numerically aggregated measurements over regular intervals, delivering high-speed anomaly detection and minimal storage consumption at the cost of high-cardinality debugging context.

  • Logs provide granular, timestamped execution detail for local component events, but require injected TraceId and SpanId context to avoid becoming unsearchable operational silos in distributed systems.

  • OpenTelemetry Baggage propagates arbitrary key-value metadata alongside traces across distributed boundaries, but is designed for runtime application context and is not automatically converted into telemetry attributes.

Last updated: September 2026

1.2 Core Telemetry Signals: Traces, Metrics, Logs & Baggage

Quick Answer: The core telemetry signals in OpenTelemetry are Distributed Traces, Metrics, and Logs, unified under a shared data model and resource schema. Distributed Traces map request paths across microservices using a DAG of spans; Metrics provide low-cost, numerically aggregated health trends over time; Logs supply detailed forensic records of discrete events; and Baggage propagates in-process and cross-process contextual metadata at runtime without automatically recording it as telemetry.

For years, the observability industry promoted the concept of the "Three Pillars of Observability": Traces, Metrics, and Logs. While conceptually useful, treating these signals as three isolated pillars led many engineering organizations to deploy three disconnected tools, three different agent daemons, and three incompatible data formats. OpenTelemetry redefines this model by treating traces, metrics, and logs not as independent silos, but as interconnected viewpoints of a single system execution fabric, bound together by context propagation and shared resource metadata.


Distributed Traces: Request Lifecycles Across Boundaries

A distributed trace records the complete execution path of a single transaction or request as it travels through a distributed architecture. In modern cloud applications, a single user click on a frontend web application or mobile device often cascades into tens or hundreds of distinct downstream network calls—spanning API gateways, service meshes, internal microservices, serverless functions, database queries, and asynchronous message brokers.

The Directed Acyclic Graph (DAG) and Spans

At a mathematical and structural level, a distributed trace is modeled as a Directed Acyclic Graph (DAG) of spans. A span represents a single contiguous unit of work or execution within a system.

Trace: [OrderCheckout - 120ms]
├── Span A: Frontend HTTP POST /orders (120ms) [Root Span]
│   ├── Span B: Auth Token Validation (15ms)
│   └── Span C: OrderService.CreateOrder (95ms)
│       ├── Span D: InventoryService.ReserveStock (30ms)
│       └── Span E: PaymentService.ChargeCard (55ms)
│           └── Span F: DB INSERT INTO payments (12ms)

Anatomy of an OpenTelemetry Span

Every span generated by the OpenTelemetry API carries a standardized set of properties defined in the OpenTelemetry specification:

  • Name — A concise, human-readable string identifying the operation (e.g., HTTP GET /users/{id} or SELECT FROM accounts). Span names must avoid dynamic high-cardinality values like raw IDs; those belong in attributes.
  • SpanContext — The immutable bundle of identifiers that uniquely identifies the span within a distributed trace. It contains:
    • TraceId (16 bytes / 32 hex characters) — Globally unique identifier shared by all spans belonging to the same transaction.
    • SpanId (8 bytes / 16 hex characters) — Unique identifier for this specific unit of work.
    • TraceFlags (8-bit field) — Controls delivery options, primarily the 1-bit sampled flag indicating whether the trace was selected for storage.
    • TraceState — Vendor-specific opaque routing and filtering metadata represented as a list of key-value pairs.
  • Parent SpanId — The SpanId of the parent operation. If a span has no parent (Parent SpanId is null or empty), it is designated as the Root Span of the trace.
  • Timestamps — High-precision Unix start time and end time, representing the exact duration of the unit of work.
  • Span Kind — Describes the relationship of the span to network and process boundaries:
    • SERVER — Synchronous handling of an incoming network request (e.g., HTTP server or gRPC handler).
    • CLIENT — Synchronous outbound network request to a downstream dependency.
    • PRODUCER — Asynchronous emission of a message to an event bus or message broker (e.g., Kafka or RabbitMQ).
    • CONSUMER — Asynchronous receipt and processing of a message from a queue.
    • INTERNAL — In-process operation that does not cross network boundaries (e.g., calculating a cryptographic hash or running a local algorithm).
  • Attributes — Key-value pairs containing structured contextual metadata (http.response.status_code = 200, db.system.name = "postgresql").
  • Events — In-span timestamped text annotations representing discrete, zero-duration informational milestones during the span's execution (e.g., cache_miss or acquired_lock).
  • Links — Causal references connecting this span to other spans in different traces, commonly used in batch processing or fan-out fan-in asynchronous patterns.
  • Status — Explicit health indicator: Unset (default), Ok (explicitly marked successful by developer), or Error (contains a failure or unhandled exception).

Metrics: Aggregate Numeric Measurements Over Time

A metric is a numerically aggregated measurement captured over regular time intervals. Unlike traces, which preserve detailed request-level execution paths, metrics discard individual transaction identities in favor of statistical summaries (counts, rates, sums, averages, percentiles).

Metric Instruments in OpenTelemetry

The OpenTelemetry Metrics API defines seven instruments, grouped by how they are called (synchronous or asynchronous) and by what they measure:

  1. Counter (Synchronous) — Monotonically increasing value (can only increment, or reset to zero on restart). Example: a custom shop.orders.completed counter.
  2. Asynchronous Counter (Observable Counter) — Monotonically increasing total read by a callback at collection time. Example: process.cpu.time.
  3. UpDownCounter (Synchronous) — Value that can increment or decrement over time. Example: http.server.active_requests.
  4. Asynchronous UpDownCounter (Observable UpDownCounter) — Additive up-and-down value reported by a callback. Example: jvm.memory.used.
  5. Histogram (Synchronous) — Records a distribution of values; the default SDK aggregation is an explicit-bucket histogram. Example: http.server.request.duration.
  6. Gauge (Synchronous) — Records a non-additive current value when a change event occurs, for example a CPU fan speed reported through an on-change listener.
  7. Asynchronous Gauge (Observable Gauge) — Reads a non-additive current value through a callback. Example: system.memory.utilization.

Strengths and Trade-offs

  • Strengths: Metrics have a minimal, highly predictable storage footprint. Querying 30 days of metric aggregates executes in milliseconds regardless of whether the system processed 10,000 or 100,000,000 requests. This makes metrics the gold standard for real-time alerting, dashboards, and automated scaling (such as Kubernetes Horizontal Pod Autoscalers).
  • Limitations: Metrics lack execution context. A metric can alert you that the 99th percentile latency of your service spiked from 50ms to 4,000ms, but it cannot tell you why, nor can it reveal which specific customer transaction suffered the delay.

Logs: Granular Execution Records & The Logging Bridge

A log is an append-only, timestamped record of an event that occurred within a software component. Traditionally emitted as unstructured plaintext via standard output (stdout) or system files, modern logs are formatted as structured JSON or serialized OTLP entities.

The OpenTelemetry Logs Data Model

Unlike traces and metrics, which were authored from greenfield specifications in OpenTelemetry, logging had decades of entrenched historical tooling (Log4j, Logback, SLF4J, Zap, Winston, Python standard logging, Fluentd, Logstash). Rather than forcing developers to discard these established libraries, OpenTelemetry introduced the OpenTelemetry Logs Data Model and a Logging Bridge API.

The OpenTelemetry log data model defines standard top-level fields for every log record:

  • Timestamp — The exact time the event occurred.
  • ObservedTimestamp — The time the log record was ingested by an OpenTelemetry Collector or SDK.
  • SeverityNumber and SeverityText — Numerical representation (1-24) mapped to standard text levels (TRACE, DEBUG, INFO, WARN, ERROR, FATAL).
  • Body — The actual log message payload, which can be a simple string, a JSON map, or a byte array.
  • Attributes — Key-value metadata specific to the local event context.
  • Resource — Standard metadata identifying the originating process, container, host, and cloud provider.
  • InstrumentationScope — The logger (library or module) that emitted the record.
  • EventName — Set only on events; a log record with a non-empty EventName is an OpenTelemetry event.
  • TraceId, SpanId, and TraceFlags — Trace linkage filled in automatically by the OpenTelemetry logging bridge from the active context.

The Role of the Logging Bridge

OpenTelemetry does not ask developers to rewrite existing log statements. Developers keep their favorite logging library (such as Zap in Go, Logback in Java, or Winston in Node.js). The Logs API exists mainly for the authors of these bridges; the current specification also allows instrumentation libraries and applications to call it directly, which is how structured events are usually emitted. An OpenTelemetry Log Appender / Bridge is attached to the existing logger. This bridge automatically intercepts log statements, extracts the active TraceId and SpanId from the thread-local context, appends standard resource attributes, and exports the log records via the OpenTelemetry Protocol (OTLP).


Baggage: Distributed Runtime Context Propagation

Baggage represents contextual key-value pairs that are propagated across process and network boundaries alongside distributed traces. Implemented according to the W3C Baggage specification, baggage values are carried across HTTP boundaries in the baggage header:

GET /checkout HTTP/1.1
Host: api.example.com
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
baggage: userId=alice123,customerTier=enterprise,datacenter=us-east-1

Key Concept: Baggage vs. Telemetry Attributes

A common misconception concerns the precise operational role of Baggage:

Baggage is for runtime application context propagation. Baggage is NOT automatically recorded as telemetry attributes!

When a service puts a key-value pair into Baggage (such as customerTier=enterprise), that value is propagated through the distributed context across downstream services. Any downstream microservice can inspect the Baggage and use that value for runtime decisions—such as dynamic traffic routing, feature-flag toggling, multi-tenant database partitioning, or rate-limiting prioritization.

However, the OpenTelemetry SDK will NEVER automatically attach Baggage key-values to trace spans, metric datapoints, or log records by default. There are two critical reasons for this architectural decision:

  1. Performance and Storage Overhead — Automatically converting all Baggage keys into span attributes would cause uncontrolled data bloat across every downstream span.
  2. Security and Data Privacy — Baggage travels in plaintext over HTTP network headers. If a developer sets sensitive customer data in Baggage, automatically dumping it into public telemetry stores would violate privacy compliance (GDPR, HIPAA, PCI-DSS).

If an engineering team wants a Baggage value to appear on a span or log record, application code must explicitly read the key from the active context and add it as a span attribute, or the team must register an opt-in baggage span processor (published in the Java, JavaScript, Python, and Go contrib repositories) that copies selected entries onto each span as it starts. The Collector cannot do this later: baggage travels in request headers between services and is not part of the OTLP data that SDKs export.


Comparison: The Four Signals Across Dimensions

SignalPrimary PurposeData Volume & CostQuery LatencyStorage ProfileCorrelation Link
Distributed TracesMap end-to-end request flows, isolate latency bottlenecks, attribute distributed errorsModerate to High (mitigated by sampling strategies)Fast for specific TraceId; moderate for multi-attribute scansColumnar or document stores with TTL retentionSpans linked via TraceId, SpanId, and Parent SpanId
MetricsHigh-level system health monitoring, real-time alerting, autoscaling triggersLow and predictable (constant across request volume)Millisecond query times over months of dataTime Series Databases (TSDB) with rollups and downsamplingLinked to traces via Exemplars (attaching TraceId to datapoint)
Structured LogsDeep forensic context, stack traces, local component state trackingHigh to Massive (linear with request and error rate)High resource consumption on unindexed text queriesCompressed document/search indexes (Elasticsearch, Loki)Linked to traces via injected TraceId and SpanId fields
BaggageRuntime cross-process context propagation for business logic and routingMinimal (carried transiently in network headers)In-memory lookup from context (instantaneous)Ephemeral (not stored; discarded after request completion)Propagated alongside W3C TraceContext headers
Loading diagram...
Unified OpenTelemetry Data Model
Test Your Knowledge

A backend engineer configures an API gateway to inject 'tenant_tier = enterprise' into OpenTelemetry Baggage. However, when inspecting trace spans for downstream microservices in the distributed tracing backend, the 'tenant_tier' attribute is nowhere to be found. Why did this occur?

A

Baggage can only propagate numeric integer values across network boundaries

B

Downstream OpenTelemetry Collector pipelines automatically strip all HTTP headers starting with the letter b

C

Baggage is deprecated in the OpenTelemetry specification and has been superseded by Span Links

D

Baggage propagates context across process boundaries at runtime but is not automatically recorded as span attributes

Test Your Knowledge

An operations team must implement an automated scaling policy and real-time paging alerts for an API handling 300,000 requests per second. Which telemetry signal is best suited to serve as the foundation for these automated systems?

A

Metrics, because their numerical aggregation over fixed intervals provides minimal storage overhead and instantaneous query evaluation

B

Unstructured logs, because full-text search engines provide the fastest percentile calculations under heavy load

C

Distributed traces with 100% head-based sampling, because every individual request must be parsed to detect load changes

D

Baggage key-value pairs, because they reside in HTTP headers and evaluate scaling conditions directly in the network proxy

Test Your Knowledge

When debugging an unexpected server error in a microservice, what two contextual identifiers should be injected into the structured log records to allow an engineer to immediately navigate between the log statement and the corresponding distributed trace?

A

ProcessId and ThreadId

B

TraceId and SpanId

C

HostName and ContainerIP

D

InstrumentName and MeterVersion

Sections you finish are checked off in the contents.