3.3 Span Events, Attributes & Status Codes

Key Takeaways

  • Span attributes are typed key-value pairs; primitives and homogeneous arrays are the most widely supported, and the current specification also allows byte arrays, maps, and heterogeneous arrays (AnyValue).

  • High-cardinality attributes (such as user IDs or transaction hashes) are encouraged on spans because tracing backends index them as discrete event records rather than continuous time series.

  • Span events represent instantaneous, zero-duration milestones within a span's lifecycle, containing a name, timestamp, and optional attributes.

  • Calling span.recordException(err) records an annotated exception event but does NOT automatically change the span's status code to Error.

  • OpenTelemetry span status follows a three-state model (Unset, Ok, Error), where marking an operation as a failure requires explicitly calling span.setStatus(StatusCode.ERROR, description).

Last updated: September 2026

3.3 Span Events, Attributes & Status Codes

While structural identifiers (TraceID, SpanID, SpanKind) establish the topology of a distributed trace, the diagnostic value of tracing depends on contextual metadata attached to individual spans. OpenTelemetry provides three standardized mechanisms for enriching spans: Attributes, Span Events, and Status Codes.


Span Attributes: Typed Contextual Metadata

Span Attributes are key-value pairs applied directly to a span to describe the properties, configuration, and environment of the timed operation.

Permitted Attribute Data Types

For years the specification limited attribute values to primitives and homogeneous arrays, and most instrumentation still uses only those. The current specification (v1.61.0, September 2026) defines an attribute value as any AnyValue:

  • Primitive Types: string, boolean, int64 (signed 64-bit integer), and double (IEEE 754 floating-point number).
  • Homogeneous Arrays: arrays whose elements share one primitive type ([]string, []bool, []int64, []double).
  • Byte arrays, arrays of AnyValue, maps of string to AnyValue (arbitrarily nested, similar to a JSON object), and an empty value where the language supports one.

Other rules from the same section of the specification:

  • Keys must be non-empty strings, and case is preserved: user.id and user.ID are two different attributes.
  • Empty strings, zeros, and empty arrays are meaningful values and must be passed on to processors and exporters.
  • APIs should warn that arrays and maps may cost more than primitive values, and support for the complex types still varies by language and backend (non-OTLP exporters may flatten them into JSON strings).

Practical rule: Prefer primitive values that match a semantic convention. Reach for maps or nested arrays only when the structure genuinely matters and your SDK and backend both handle it.

Cardinality: Tracing vs. Metrics

A critical conceptual distinction is how cardinality is handled across signals:

  • In Metrics: Adding high-cardinality attributes (such as user.id, order.id, ip.address, or session.token) is a catastrophic anti-pattern. Metric databases (like Prometheus or Mimir) generate an independent time-series stream for every unique combination of label values. High cardinality causes "cardinality explosion," exhausting memory and crashing time-series databases.
  • In Distributed Tracing: High-cardinality attributes are actively encouraged! Spans are not stored as continuous time series; they are stored as discrete, indexed event documents in columnar or search backends (e.g., Elasticsearch, ClickHouse, Jaeger, Tempo). Setting user.id = "usr_98412" or payment.transaction_id = "tx_44129" on a span allows engineers to search for the exact trace corresponding to a specific customer complaint.

Semantic Conventions

Attributes should follow the official OpenTelemetry Semantic Conventions using dot-separated hierarchical namespaces:

  • http.request.method: e.g., "POST"
  • http.response.status_code: e.g., 200
  • db.system.name: e.g., "postgresql"
  • db.query.text: e.g., "SELECT * FROM customers WHERE id = ?"
  • rpc.method: e.g., "orders.v1.OrderService/Submit"

SDK Attribute Limits

To prevent rogue loops from consuming excessive memory, the OpenTelemetry SDK enforces configurable limits via standard environment variables:

  • OTEL_SPAN_ATTRIBUTE_VALUE_LENGTH_LIMIT: Maximum length of a string attribute value (default: no limit; longer strings are truncated when a limit is set).
  • OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT: Maximum number of attributes per span (default 128). Attributes added beyond this limit are dropped, and the SDK increments the span's dropped_attributes_count field.
  • OTEL_SPAN_EVENT_COUNT_LIMIT and OTEL_SPAN_LINK_COUNT_LIMIT: Maximum events and links per span (default 128 each). The general OTEL_ATTRIBUTE_COUNT_LIMIT (default 128) and OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT apply when no signal-specific limit is set.

Span Events: Instantaneous Milestones

A Span Event is an in-span timestamped annotation that represents an instantaneous milestone or state transition occurring during the execution of a span. Unlike spans, events have no duration (Duration = 0).

Anatomy of a Span Event

Every span event consists of three components:

  1. Event Name: A low-cardinality string identifying the milestone (e.g., "cache_miss", "connection_acquired", "payload_parsed").
  2. Timestamp: High-precision Unix epoch timestamp recording the exact nanosecond the milestone occurred.
  3. Event Attributes: An optional dictionary of typed key-value pairs providing specific contextual details about the milestone.
Span: [Process Payment] -----------------------------------------------------> (Total: 150ms)
  |-- [T+10ms] Event: "tls_handshake_completed" { cipher: "TLS_AES_256_GCM_SHA384" }
  |-- [T+45ms] Event: "fraud_score_evaluated" { score: 12, passed: true }
  `-- [T+135ms] Event: "gateway_response_buffered" { bytes_received: 4096 }

Architectural Decision: Span Event vs. Child Span

When should an engineer create a child span versus recording a span event?

CriteriaUse a Child SpanUse a Span Event
Execution DurationThe operation takes measurable time with a distinct beginning and ending (e.g., executing a database query, calling an external microservice).The milestone is instantaneous (zero duration) or duration measurement is irrelevant (e.g., socket connected, cache miss, exception thrown).
ConcurrencyThe operation executes asynchronously or concurrently with other tasks.The milestone occurs strictly in-line with the span's current execution thread.
Context & HierarchyThe sub-task needs to be passed down the call chain or needs its own child spans.The metadata is purely an internal annotation strictly scoped to the parent span.

Span Events vs. Structured Logs

While span events resemble structured log messages, their lifecycles differ fundamentally:

  • Span Events are strictly bound to the span lifecycle: If the span is dropped by a head-based sampler (e.g., TraceIdRatioBased at 1%), all events inside that span are immediately discarded and never exported.
  • OpenTelemetry Logs are independent first-class signals: Logs can exist with or without an active trace context and travel through independent log pipelines, processors, and exporters.

Recording Exceptions on Spans

Handling runtime errors and exceptions is a primary use case for distributed tracing. OpenTelemetry provides a dedicated API method for capturing errors: span.recordException(err).

What recordException Actually Does

When span.recordException(err) is called, the SDK does not alter the span's duration or status. Instead, it creates a standardized Span Event with the reserved name "exception" and populates standard semantic attributes:

  • exception.type: The fully qualified class or type name of the error (e.g., "java.net.ConnectException", "*pgconn.PgError", "KeyError").
  • exception.message: The human-readable error message string (e.g., "connection refused by peer: 10.0.4.12:5432").
  • exception.stacktrace: The complete formatted stack trace string showing the exact code location of the failure.
  • exception.escaped: Deprecated in current semantic conventions. It once flagged exceptions escaping the span; the guidance now is simply not to record exceptions that were handled inside the span's scope.

The Critical Rule: Exceptions vs. Status Codes

CRITICAL RULE: Calling span.recordException(err) DOES NOT automatically set the span's status code to Error! This is one of the most common instrumentation mistakes.

Why does OpenTelemetry keep exception recording separate from status code setting? Because in real-world software, many exceptions are non-fatal, expected, or successfully handled:

  • A service encounters a transient network timeout, catches the exception, and successfully executes on a second retry attempt. The overall span succeeded!
  • A cache lookup throws a KeyNotFoundException, but the application catches it and falls back to querying the primary database. The business operation succeeded!

If the SDK automatically marked a span as failed every time an exception occurred, tracing backends would be flooded with false-positive error alerts. Therefore, the developer must explicitly set the status code to Error if the operation actually failed.


Span Status Codes: The Three-State Model

OpenTelemetry defines a strict three-state model for classifying the execution outcome of a span via the StatusCode enumeration:

                +------------------+
                | StatusCode.UNSET |
                |  (Default State) |
                +------------------+
                    /          \
                   /            \
                  v              v
       +------------------+    +------------------+
       | StatusCode.ERROR |    |  StatusCode.OK   |
       | (Explicit Fail)  |    | (Explicit Success|
       +------------------+    |  Overrides Error)| 
                  \            +------------------+
                   \             ^
                    \___________/

1. StatusCode.UNSET (Default)

  • Every span begins in the UNSET state.
  • UNSET indicates that the operation completed without an explicit developer determination of success or failure.
  • Crucial Rule: Backends treat UNSET as a successful operation. Instrumentation libraries normally leave successful spans UNSET and set ERROR on failure; setting OK is reserved for an explicit decision by the application or operator.

2. StatusCode.OK

  • Explicitly stamps the operation as successful.
  • Setting OK is generally reserved for situations where a span recovered from an earlier partial failure (e.g., an error occurred and was recorded, but a fallback strategy succeeded).
  • Precedence Rule: Setting OK overrides any previous ERROR status.

3. StatusCode.ERROR

  • Marks the operation as failed, indicating that the span did not fulfill its primary technical or contractual objective.
  • Should be accompanied by a human-readable Status Description explaining the cause of the failure (e.g., "failed to commit payment: card expired"). A description is only meaningful with ERROR; it is ignored for OK and UNSET.
  • Tracing visualization tools render ERROR spans in bright red and use them to calculate service error rate SLOs.

Status Transition and Precedence Rules

  • UNSET can transition to ERROR or OK.
  • ERROR can transition to OK (if an explicit recovery step succeeded).
  • Once a span is set to OK, it cannot be changed back to ERROR or UNSET. OK is final.

Complete Production Code Walkthrough

The following code demonstrates the complete canonical workflow for creating a span, populating semantic attributes, recording milestone events, capturing exceptions, and explicitly setting the status code upon failure:

from opentelemetry import trace
from opentelemetry.trace import StatusCode

tracer = trace.get_tracer("com.example.billing.payment", "1.0.0")

def process_customer_payment(order_id: str, amount_cents: int, customer_id: str):
    # 1. Start span with low-cardinality operation name
    with tracer.start_as_current_span("process_payment", kind=trace.SpanKind.SERVER) as span:
        # 2. Add high-cardinality business attributes
        span.set_attribute("order.id", order_id)
        span.set_attribute("customer.id", customer_id)
        span.set_attribute("payment.amount_cents", amount_cents)
        span.set_attribute("payment.currency", "USD")

        try:
            # 3. Add a zero-duration milestone span event
            span.add_event(
                name="fraud_check_started",
                attributes={"fraud.engine.version": "2.4.1"}
            )
            
            # Execute business logic (which may throw an exception)
            charge_credit_card(customer_id, amount_cents)
            
            # 4. Add completion milestone
            span.add_event("payment_authorized")
            
            # Note: We do NOT need to set StatusCode.OK here; 
            # leaving it UNSET defaults to success in backends.
            
        except PaymentGatewayException as err:
            # 5. Record the exception event (creates 'exception' event with stacktrace)
            span.record_exception(err)
            
            # 6. EXPLICITLY set the status to ERROR with a descriptive message
            # Without this line, the span remains UNSET and renders as SUCCESS!
            span.set_status(
                StatusCode.ERROR, 
                description=f"Payment failed: {str(err)}"
            )
            
            # Re-raise to ensure proper HTTP error response to caller
            raise
Loading diagram...
Span Lifecycle, Events, and Status Code Evaluation
Test Your Knowledge

A backend engineer instruments an order processing microservice using the OpenTelemetry API. When a database connectivity timeout occurs, the service enters its catch block, executes span.recordException(err), logs the failure, and returns an HTTP 500 error to the client. However, when SREs review the trace in their visualization backend, the order processing span is rendered as green (successful). What is the root cause of this issue?

A

The engineer recorded the exception but failed to explicitly call span.setStatus(StatusCode.ERROR), leaving the span in the default StatusCode.UNSET state.

B

The OpenTelemetry SDK automatically suppresses exception events whenever the span duration is less than the configured network timeout limit.

C

The database exception was recorded as a span event rather than being passed into the LoggerProvider pipeline as an independent log record.

D

The trace backend automatically overrides span status to StatusCode.OK unless the span was configured with SpanKind.CLIENT.

Test Your Knowledge

An observability architect reviews attribute designs for a billing service instrumented with a current OpenTelemetry SDK. Which statement matches the current specification's rules for attributes?

A

Attribute values may only be strings, so numeric amounts must be converted to text before they are recorded

B

Attribute keys are case-insensitive, so user.ID and user.id always resolve to the same attribute

C

Empty strings and zero values carry no information, so the SDK discards them before export

D

Primitive values and homogeneous arrays are the most widely supported, and the specification now also permits maps and heterogeneous arrays, which may cost more and may be flattened by some backends

Test Your Knowledge

A distributed caching service handles user preference queries. When a requested key is absent from the in-memory Redis cluster, the application encounters an expected cache miss, records a milestone event, queries the persistent database, updates the cache, and successfully returns the requested preference to the client. How should the Redis lookup span's status and events be configured?

A

Set the Redis span status to StatusCode.ERROR with the description 'Cache miss' because the primary lookup operation failed to locate the key.

B

Call span.recordException() with a CacheMissException and set the status to StatusCode.ERROR to alert operators to low cache hit ratios.

C

Record a span event named 'cache_miss' on the span and allow its status to remain StatusCode.UNSET, as the miss is an expected operational branch.

D

Throw a fatal runtime exception from the Redis client to force the batch processor to flag the trace in the Collector pipeline.

Sections you finish are checked off in the contents.