3.3 Span Events, Attributes & Status Codes
Key Takeaways
Span attributes are typed key-value pairs; primitives and homogeneous arrays are the most widely supported, and the current specification also allows byte arrays, maps, and heterogeneous arrays (AnyValue).
High-cardinality attributes (such as user IDs or transaction hashes) are encouraged on spans because tracing backends index them as discrete event records rather than continuous time series.
Span events represent instantaneous, zero-duration milestones within a span's lifecycle, containing a name, timestamp, and optional attributes.
Calling span.recordException(err) records an annotated exception event but does NOT automatically change the span's status code to Error.
OpenTelemetry span status follows a three-state model (Unset, Ok, Error), where marking an operation as a failure requires explicitly calling span.setStatus(StatusCode.ERROR, description).
3.3 Span Events, Attributes & Status Codes
While structural identifiers (TraceID, SpanID, SpanKind) establish the topology of a distributed trace, the diagnostic value of tracing depends on contextual metadata attached to individual spans. OpenTelemetry provides three standardized mechanisms for enriching spans: Attributes, Span Events, and Status Codes.
Span Attributes: Typed Contextual Metadata
Span Attributes are key-value pairs applied directly to a span to describe the properties, configuration, and environment of the timed operation.
Permitted Attribute Data Types
For years the specification limited attribute values to primitives and homogeneous arrays, and most instrumentation still uses only those. The current specification (v1.61.0, September 2026) defines an attribute value as any AnyValue:
- Primitive Types:
string,boolean,int64(signed 64-bit integer), anddouble(IEEE 754 floating-point number). - Homogeneous Arrays: arrays whose elements share one primitive type (
[]string,[]bool,[]int64,[]double). - Byte arrays, arrays of AnyValue, maps of string to AnyValue (arbitrarily nested, similar to a JSON object), and an empty value where the language supports one.
Other rules from the same section of the specification:
- Keys must be non-empty strings, and case is preserved:
user.idanduser.IDare two different attributes. - Empty strings, zeros, and empty arrays are meaningful values and must be passed on to processors and exporters.
- APIs should warn that arrays and maps may cost more than primitive values, and support for the complex types still varies by language and backend (non-OTLP exporters may flatten them into JSON strings).
Practical rule: Prefer primitive values that match a semantic convention. Reach for maps or nested arrays only when the structure genuinely matters and your SDK and backend both handle it.
Cardinality: Tracing vs. Metrics
A critical conceptual distinction is how cardinality is handled across signals:
- In Metrics: Adding high-cardinality attributes (such as
user.id,order.id,ip.address, orsession.token) is a catastrophic anti-pattern. Metric databases (like Prometheus or Mimir) generate an independent time-series stream for every unique combination of label values. High cardinality causes "cardinality explosion," exhausting memory and crashing time-series databases. - In Distributed Tracing: High-cardinality attributes are actively encouraged! Spans are not stored as continuous time series; they are stored as discrete, indexed event documents in columnar or search backends (e.g., Elasticsearch, ClickHouse, Jaeger, Tempo). Setting
user.id = "usr_98412"orpayment.transaction_id = "tx_44129"on a span allows engineers to search for the exact trace corresponding to a specific customer complaint.
Semantic Conventions
Attributes should follow the official OpenTelemetry Semantic Conventions using dot-separated hierarchical namespaces:
http.request.method: e.g.,"POST"http.response.status_code: e.g.,200db.system.name: e.g.,"postgresql"db.query.text: e.g.,"SELECT * FROM customers WHERE id = ?"rpc.method: e.g.,"orders.v1.OrderService/Submit"
SDK Attribute Limits
To prevent rogue loops from consuming excessive memory, the OpenTelemetry SDK enforces configurable limits via standard environment variables:
OTEL_SPAN_ATTRIBUTE_VALUE_LENGTH_LIMIT: Maximum length of a string attribute value (default: no limit; longer strings are truncated when a limit is set).OTEL_SPAN_ATTRIBUTE_COUNT_LIMIT: Maximum number of attributes per span (default 128). Attributes added beyond this limit are dropped, and the SDK increments the span'sdropped_attributes_countfield.OTEL_SPAN_EVENT_COUNT_LIMITandOTEL_SPAN_LINK_COUNT_LIMIT: Maximum events and links per span (default 128 each). The generalOTEL_ATTRIBUTE_COUNT_LIMIT(default 128) andOTEL_ATTRIBUTE_VALUE_LENGTH_LIMITapply when no signal-specific limit is set.
Span Events: Instantaneous Milestones
A Span Event is an in-span timestamped annotation that represents an instantaneous milestone or state transition occurring during the execution of a span. Unlike spans, events have no duration (Duration = 0).
Anatomy of a Span Event
Every span event consists of three components:
- Event Name: A low-cardinality string identifying the milestone (e.g.,
"cache_miss","connection_acquired","payload_parsed"). - Timestamp: High-precision Unix epoch timestamp recording the exact nanosecond the milestone occurred.
- Event Attributes: An optional dictionary of typed key-value pairs providing specific contextual details about the milestone.
Span: [Process Payment] -----------------------------------------------------> (Total: 150ms)
|-- [T+10ms] Event: "tls_handshake_completed" { cipher: "TLS_AES_256_GCM_SHA384" }
|-- [T+45ms] Event: "fraud_score_evaluated" { score: 12, passed: true }
`-- [T+135ms] Event: "gateway_response_buffered" { bytes_received: 4096 }
Architectural Decision: Span Event vs. Child Span
When should an engineer create a child span versus recording a span event?
| Criteria | Use a Child Span | Use a Span Event |
|---|---|---|
| Execution Duration | The operation takes measurable time with a distinct beginning and ending (e.g., executing a database query, calling an external microservice). | The milestone is instantaneous (zero duration) or duration measurement is irrelevant (e.g., socket connected, cache miss, exception thrown). |
| Concurrency | The operation executes asynchronously or concurrently with other tasks. | The milestone occurs strictly in-line with the span's current execution thread. |
| Context & Hierarchy | The sub-task needs to be passed down the call chain or needs its own child spans. | The metadata is purely an internal annotation strictly scoped to the parent span. |
Span Events vs. Structured Logs
While span events resemble structured log messages, their lifecycles differ fundamentally:
- Span Events are strictly bound to the span lifecycle: If the span is dropped by a head-based sampler (e.g.,
TraceIdRatioBasedat 1%), all events inside that span are immediately discarded and never exported. - OpenTelemetry Logs are independent first-class signals: Logs can exist with or without an active trace context and travel through independent log pipelines, processors, and exporters.
Recording Exceptions on Spans
Handling runtime errors and exceptions is a primary use case for distributed tracing. OpenTelemetry provides a dedicated API method for capturing errors: span.recordException(err).
What recordException Actually Does
When span.recordException(err) is called, the SDK does not alter the span's duration or status. Instead, it creates a standardized Span Event with the reserved name "exception" and populates standard semantic attributes:
exception.type: The fully qualified class or type name of the error (e.g.,"java.net.ConnectException","*pgconn.PgError","KeyError").exception.message: The human-readable error message string (e.g.,"connection refused by peer: 10.0.4.12:5432").exception.stacktrace: The complete formatted stack trace string showing the exact code location of the failure.exception.escaped: Deprecated in current semantic conventions. It once flagged exceptions escaping the span; the guidance now is simply not to record exceptions that were handled inside the span's scope.
The Critical Rule: Exceptions vs. Status Codes
CRITICAL RULE: Calling
span.recordException(err)DOES NOT automatically set the span's status code toError! This is one of the most common instrumentation mistakes.
Why does OpenTelemetry keep exception recording separate from status code setting? Because in real-world software, many exceptions are non-fatal, expected, or successfully handled:
- A service encounters a transient network timeout, catches the exception, and successfully executes on a second retry attempt. The overall span succeeded!
- A cache lookup throws a
KeyNotFoundException, but the application catches it and falls back to querying the primary database. The business operation succeeded!
If the SDK automatically marked a span as failed every time an exception occurred, tracing backends would be flooded with false-positive error alerts. Therefore, the developer must explicitly set the status code to Error if the operation actually failed.
Span Status Codes: The Three-State Model
OpenTelemetry defines a strict three-state model for classifying the execution outcome of a span via the StatusCode enumeration:
+------------------+
| StatusCode.UNSET |
| (Default State) |
+------------------+
/ \
/ \
v v
+------------------+ +------------------+
| StatusCode.ERROR | | StatusCode.OK |
| (Explicit Fail) | | (Explicit Success|
+------------------+ | Overrides Error)|
\ +------------------+
\ ^
\___________/
1. StatusCode.UNSET (Default)
- Every span begins in the
UNSETstate. UNSETindicates that the operation completed without an explicit developer determination of success or failure.- Crucial Rule: Backends treat
UNSETas a successful operation. Instrumentation libraries normally leave successful spansUNSETand setERRORon failure; settingOKis reserved for an explicit decision by the application or operator.
2. StatusCode.OK
- Explicitly stamps the operation as successful.
- Setting
OKis generally reserved for situations where a span recovered from an earlier partial failure (e.g., an error occurred and was recorded, but a fallback strategy succeeded). - Precedence Rule: Setting
OKoverrides any previousERRORstatus.
3. StatusCode.ERROR
- Marks the operation as failed, indicating that the span did not fulfill its primary technical or contractual objective.
- Should be accompanied by a human-readable Status Description explaining the cause of the failure (e.g.,
"failed to commit payment: card expired"). A description is only meaningful withERROR; it is ignored forOKandUNSET. - Tracing visualization tools render
ERRORspans in bright red and use them to calculate service error rate SLOs.
Status Transition and Precedence Rules
UNSETcan transition toERRORorOK.ERRORcan transition toOK(if an explicit recovery step succeeded).- Once a span is set to
OK, it cannot be changed back toERRORorUNSET.OKis final.
Complete Production Code Walkthrough
The following code demonstrates the complete canonical workflow for creating a span, populating semantic attributes, recording milestone events, capturing exceptions, and explicitly setting the status code upon failure:
from opentelemetry import trace
from opentelemetry.trace import StatusCode
tracer = trace.get_tracer("com.example.billing.payment", "1.0.0")
def process_customer_payment(order_id: str, amount_cents: int, customer_id: str):
# 1. Start span with low-cardinality operation name
with tracer.start_as_current_span("process_payment", kind=trace.SpanKind.SERVER) as span:
# 2. Add high-cardinality business attributes
span.set_attribute("order.id", order_id)
span.set_attribute("customer.id", customer_id)
span.set_attribute("payment.amount_cents", amount_cents)
span.set_attribute("payment.currency", "USD")
try:
# 3. Add a zero-duration milestone span event
span.add_event(
name="fraud_check_started",
attributes={"fraud.engine.version": "2.4.1"}
)
# Execute business logic (which may throw an exception)
charge_credit_card(customer_id, amount_cents)
# 4. Add completion milestone
span.add_event("payment_authorized")
# Note: We do NOT need to set StatusCode.OK here;
# leaving it UNSET defaults to success in backends.
except PaymentGatewayException as err:
# 5. Record the exception event (creates 'exception' event with stacktrace)
span.record_exception(err)
# 6. EXPLICITLY set the status to ERROR with a descriptive message
# Without this line, the span remains UNSET and renders as SUCCESS!
span.set_status(
StatusCode.ERROR,
description=f"Payment failed: {str(err)}"
)
# Re-raise to ensure proper HTTP error response to caller
raise
A backend engineer instruments an order processing microservice using the OpenTelemetry API. When a database connectivity timeout occurs, the service enters its catch block, executes span.recordException(err), logs the failure, and returns an HTTP 500 error to the client. However, when SREs review the trace in their visualization backend, the order processing span is rendered as green (successful). What is the root cause of this issue?
The engineer recorded the exception but failed to explicitly call span.setStatus(StatusCode.ERROR), leaving the span in the default StatusCode.UNSET state.
The OpenTelemetry SDK automatically suppresses exception events whenever the span duration is less than the configured network timeout limit.
The database exception was recorded as a span event rather than being passed into the LoggerProvider pipeline as an independent log record.
The trace backend automatically overrides span status to StatusCode.OK unless the span was configured with SpanKind.CLIENT.
An observability architect reviews attribute designs for a billing service instrumented with a current OpenTelemetry SDK. Which statement matches the current specification's rules for attributes?
Attribute values may only be strings, so numeric amounts must be converted to text before they are recorded
Attribute keys are case-insensitive, so user.ID and user.id always resolve to the same attribute
Empty strings and zero values carry no information, so the SDK discards them before export
Primitive values and homogeneous arrays are the most widely supported, and the specification now also permits maps and heterogeneous arrays, which may cost more and may be flattened by some backends
A distributed caching service handles user preference queries. When a requested key is absent from the in-memory Redis cluster, the application encounters an expected cache miss, records a milestone event, queries the persistent database, updates the cache, and successfully returns the requested preference to the client. How should the Redis lookup span's status and events be configured?
Set the Redis span status to StatusCode.ERROR with the description 'Cache miss' because the primary lookup operation failed to locate the key.
Call span.recordException() with a CacheMissException and set the status to StatusCode.ERROR to alert operators to low cache hit ratios.
Record a span event named 'cache_miss' on the span and allow its status to remain StatusCode.UNSET, as the miss is an expected operational branch.
Throw a fatal runtime exception from the Redis client to force the batch processor to flag the trace in the Collector pipeline.
Sections you finish are checked off in the contents.