6.1 OpenTelemetry Logs Data Model & Architecture
Key Takeaways
Logs are the most ubiquitous yet historically fragmented telemetry signal, predating modern observability by decades across incompatible formats such as Syslog, stdout streams, and proprietary logging frameworks.
The OpenTelemetry Log Record data model defines Timestamp, ObservedTimestamp, TraceId, SpanId, TraceFlags, SeverityText, SeverityNumber, Body, Attributes, and EventName, plus the Resource and InstrumentationScope the record belongs to.
SeverityNumber provides a 24-point normalized integer scale partitioned into six standard tiers (TRACE 1–4, DEBUG 5–8, INFO 9–12, WARN 13–16, ERROR 17–20, FATAL 21–24), enabling cross-language severity querying across polyglot microservice architectures.
The Log Record Body is an AnyValue, so it can hold plain strings, numbers, booleans, byte arrays, arrays, or deeply nested maps without losing structure.
OpenTelemetry categorizes logs into legacy unstructured text streams, structured JSON records, and direct OpenTelemetry event logs, unifying all three into the telemetry graph through common resource attribution and trace context correlation.
6.1 OpenTelemetry Logs Data Model & Architecture
Quick Answer: The OpenTelemetry Logs Data Model defines a universal, vendor-neutral structure for representing log records across distributed systems. Unlike legacy logging where logs were isolated strings written to disk or stdout, an OpenTelemetry Log Record standardizes precise event and observed timestamps, distributed trace correlation identifiers (
TraceId,SpanId,TraceFlags), a 24-point normalized numerical severity scale (SeverityNumber), extensible key-value attributes, and structuredAnyValuebody payloads, all contextualized by sharedResourcemetadata.
Logs are the oldest and most widespread telemetry signal in software engineering. Decades before distributed tracing or standardized metric time-series emerged, developers diagnosed malfunctioning applications using print statements, local text files, and operating system syslog daemons. However, this longevity has made logging the most challenging signal to standardize across modern enterprise environments.
The Evolution of Logs in OpenTelemetry
When distributed tracing and metric collection emerged in the cloud-native ecosystem, projects like OpenTracing, OpenMetrics, and OpenCensus established clean, modern specifications unburdened by legacy technical debt. Traces and metrics were adopted in greenfield architectures or standardized through modern client libraries. Logs, by contrast, presented a fundamentally different problem:
- Decades of Entrenched Ecosystems — Organizations possess hundreds of microservices, third-party libraries, and legacy monolithic applications already integrated with language-specific logging frameworks (such as Log4j, Logback, and SLF4J in Java; Winston and Pino in Node.js;
loggingand Structlog in Python; Zap and Zerolog in Go; and Serilog in .NET). - Incompatible Transport and Storage Standards — Infrastructure components rely on Syslog (RFC 3164 and RFC 5424), Windows Event Logs, journald, or plain text streams routed through heterogeneous log shippers like Fluentd, Fluent Bit, Logstash, Promtail, and Vector.
- Arbitrary Severity and Field Naming — One framework logs errors as
"CRITICAL"with integer level 50, another logs"FATAL", while Syslog definesEmergency(severity 0) andAlert(severity 1). Timestamps vary from Unix epoch milliseconds to microsecond ISO-8601 strings or locale-dependent textual dates.
The OpenTelemetry Strategy for Logs
Recognizing that asking the global software industry to rewrite millions of existing logging statements was impossible, the OpenTelemetry project adopted a pragmatic, non-invasive design strategy:
- Do not replace existing logging APIs — Developers should not abandon their idiomatic logging frameworks.
- Define a unified universal data model — Create a standardized in-memory and on-the-wire representation for log records that bridges existing frameworks into modern observability pipelines.
- Correlate logs with traces and metrics — Transform logs from isolated textual records into fully integrated nodes within the telemetry graph by embedding distributed trace context (
TraceId,SpanId) and sharedResourceattribution (service.name,k8s.pod.name).
The OpenTelemetry Log Record Data Model
The OpenTelemetry specification defines a Log Record as a discrete event at a specific point in time. All of its fields are optional:
| Specification Field | Data Type | Requirement | Operational Definition & Usage |
|---|---|---|---|
Timestamp | Unix Nanoseconds (uint64) | Optional | Exact point in time when the event occurred at the source application runtime. Expressed as nanoseconds since Unix epoch (1970-01-01T00:00:00Z). |
ObservedTimestamp | Unix Nanoseconds (uint64) | Optional | When OpenTelemetry code first observed the event. For records created through the SDK or a log bridge it is normally the generation time; for external logs read by the Collector it is the read time. |
TraceId | 16-Byte Hex String (32 Hex Chars) | Optional | Distributed trace identifier correlating the log record with the active execution trace. Matches the W3C Trace Context trace-id. |
SpanId | 8-Byte Hex String (16 Hex Chars) | Optional | Distributed span identifier correlating the log record with the specific active span during which the log was emitted. Matches the W3C Trace Context parent-id. |
TraceFlags | 8-Bit Bitmap (byte) | Optional | W3C trace flags representing options such as whether the parent trace was sampled (01). Allows storage engines to prioritize logs belonging to sampled traces. |
SeverityNumber | Integer (1–24) | Optional | Standardized numerical severity level mapping the log to one of 24 canonical tiers, providing cross-language querying consistency. |
SeverityText | String | Optional | Original severity string as emitted by the source logging framework (e.g., "WARN", "CRITICAL", "Information", "Notice"). Preserves native fidelity. |
Body | AnyValue | Optional | The primary payload or log message. Can be a plain string, numeric value, boolean, binary blob, homogeneous array, or structured key-value map. |
Attributes | Key-Value Map (map[string]AnyValue) | Optional | Extensible contextual attributes providing operational metadata (e.g., exception.type, exception.stacktrace, db.query.text, http.response.status_code). |
EventName | String | Optional | Identifies the class of an Event; any log record with a non-empty EventName is an event (e.g., browser.page_view). |
InstrumentationScope | Scope (name, version, attributes) | Associated Scope | The logger that emitted the record, such as the logger name bridged from Log4j. |
Resource | Entity Metadata (Resource) | Associated Scope | The entity producing the telemetry (e.g., service.name, service.version, k8s.namespace.name, host.id). Attached at the batch or scope level. |
Detailed Analysis of Critical Fields
1. Timestamp vs. ObservedTimestamp
A critical distinction is the relationship between Timestamp and ObservedTimestamp:
Timestamprepresents the event generation time recorded by the application runtime. If an application executeslogger.info("User logged in"), the runtime captures the system clock at that microsecond.ObservedTimestamprepresents the moment OpenTelemetry code first observed the event. For a record created through the OpenTelemetry SDK or a log bridge, that is normally the generation time, so it usually equalsTimestamp. For a log collected from outside OpenTelemetry, for example a file line read by the Collector'sfilelogreceiver, it is the time the receiver read the line.
The two timestamps diverge when logs are collected late. Suppose devices or legacy hosts write plain log files while offline and the files are only read by a Collector hours later. Timestamp (parsed from each log line) reflects when the event occurred, while ObservedTimestamp reflects when the Collector read it. Distinguishing these two timestamps prevents index corruption, supports accurate out-of-order log analysis, and ensures that pipeline lag can be measured precisely (ObservedTimestamp - Timestamp = Telemetry Lag).
2. Trace and Span Correlation Identifiers
OpenTelemetry logs are not isolated silos. When an application logs a message inside an active distributed span, the logging integration automatically extracts TraceId, SpanId, and TraceFlags from the ambient context. This enables bi-directional telemetry correlation:
- Trace-to-Log Navigation: An engineer viewing a slow span in a distributed tracing UI can click a single button to view all log messages emitted during that exact span's execution lifecycle.
- Log-to-Trace Navigation: An engineer searching error logs in a central log management dashboard can immediately jump from an unhandled exception log directly into the end-to-end distributed trace DAG (Directed Acyclic Graph) to see what upstream caller triggered the failure.
3. The Body as an AnyValue
Legacy logging assumed log bodies were plain text strings (e.g., "Order 4092 failed for customer 108"). Modern logging engines emit deeply nested structured JSON. OpenTelemetry accommodates both paradigms without compromise by defining Body as an AnyValue type:
stringValue: Simple human-readable string messages.int64Value/doubleValue/boolValue: Primitive numeric or boolean payloads.bytesValue: Raw binary data payloads.arrayValue: Ordered collections ofAnyValueitems.kvlistValue: Nested key-value maps, enabling fully structured, tree-like JSON payloads to be ingested natively without converting them into serialized strings.
Categorization of Logs in OpenTelemetry
OpenTelemetry accommodates three distinct categories of logs across enterprise architectures:
- Legacy Unstructured Logs — Applications write plain, unformatted text strings directly to local disk files, syslog, or container stdout/stderr (e.g.,
"2026-09-29 10:14:22 [main] ERROR - Failed to connect to DB"). These logs lack explicit field boundaries, requiring downstream collectors to employ regex, grok, or delimiter parsers to extract timestamps, severities, and messages. - Structured JSON Logs — Applications utilize modern logging libraries configured to output serialized JSON objects to stdout or local log collectors (e.g.,
{"time":"2026-09-29T10:14:22Z","level":"error","msg":"Failed to connect to DB","db.name":"users"}). Parsers can readily extract these fields into OpenTelemetry attributes and body structures without fragile regex pattern matching. - Direct OpenTelemetry Event Logs — Applications emit native OpenTelemetry Log Records directly through an OpenTelemetry Log Appender or SDK Logger. These records contain native strongly-typed attributes, pre-populated trace context identifiers, normalized
SeverityNumbervalues, and attachedResourcemetadata before ever leaving the application process.
SeverityNumber Mapping and Cross-Language Normalization
In polyglot organizations, microservices written in different languages use disparate severity vocabularies. A Python service may emit CRITICAL, a Go service emits PANIC, a Java service emits FATAL, and an edge router emits Syslog level 1 (Alert). Querying across all services for severe failures traditionally required maintaining complex, error-prone regular expression filters.
OpenTelemetry solves this by introducing a 24-point numerical severity scale (SeverityNumber). The 24 numbers are partitioned into six canonical tiers, with four numerical increments per tier to accommodate framework-specific sub-levels (e.g., DEBUG, DEBUG2, DEBUG3). The mappings below follow the example table in the specification's data-model appendix (Zap's DPanic and Panic map to ERROR2 and ERROR3):
| SeverityNumber Range | OpenTelemetry Level | Syslog (RFC 5424) | Java (SLF4J / Log4j2) | Python logging | Description & Operational Semantics |
|---|---|---|---|---|---|
| 1–4 | TRACE (1=TRACE, 2=TRACE2, 3=TRACE3, 4=TRACE4) | — | TRACE | — (no TRACE level) | Fine-grained informational events for detailed application execution tracing. |
| 5–8 | DEBUG (5=DEBUG, 6=DEBUG2, 7=DEBUG3, 8=DEBUG4) | 7 (Debug) | DEBUG | DEBUG (10) | Diagnostic information useful for developers and support during debugging. |
| 9–12 | INFO (9=INFO, 10=INFO2, 11=INFO3, 12=INFO4) | 6 (Informational) = INFO; 5 (Notice) = INFO2 | INFO | INFO (20) | Normal operational milestones indicating regular business progression. |
| 13–16 | WARN (13=WARN, 14=WARN2, 15=WARN3, 16=WARN4) | 4 (Warning) | WARN | WARNING (30) | Potentially harmful situations or unexpected conditions that do not halt operation. |
| 17–20 | ERROR (17=ERROR, 18=ERROR2, 19=ERROR3, 20=ERROR4) | 3 (Error) = ERROR; 2 (Critical) = ERROR2; 1 (Alert) = ERROR3 | ERROR | ERROR (40) | Error events that impair specific operations but allow application execution to continue. |
| 21–24 | FATAL (21=FATAL, 22=FATAL2, 23=FATAL3, 24=FATAL4) | 0 (Emergency) = FATAL | FATAL (Log4j2) / ERROR (SLF4J) | CRITICAL (50) | Severe catastrophic conditions causing system abort, panic, or unrecoverable failure. |
Cross-Framework Querying in Practice
With SeverityNumber normalization, a centralized monitoring alerting rule or SRE dashboard query does not need to enumerate language-specific strings. To capture all critical production failures across a 50-service polyglot cluster, the alerting system executes a single mathematical condition:
SeverityNumber >= 17
This simple numerical filter instantly captures every ERROR from Java, ERROR and CRITICAL from Python, Error and Panic from Go, and Error, Critical, Alert, and Emergency from Syslog appliances. Concurrently, the original framework-level string is preserved in SeverityText (e.g., "CRITICAL"), ensuring that developers inspecting individual log records in search UIs see the exact familiar terms emitted by their source code.
Practical Operational Scenario: Resolving Ingestion Lag and Severity Drift
Consider an international e-commerce platform operating an asynchronous payment settlement service on an isolated legacy host. The settlement worker writes plain-text log files locally, and during a three-hour network partition nothing ships those files anywhere.
When network connectivity recovers at 04:00 UTC:
- A recovery job copies the 500,000-line log file to a central log host, where an OpenTelemetry Collector's
filelogreceiver reads it. - The receiver sets each record's
ObservedTimestampto04:00:xx UTC, the moment OpenTelemetry code first observed the line. - The receiver's timestamp parsing sets
Timestampfrom the time written in each line, so it correctly shows the execution time between01:00 UTCand04:00 UTC. - When on-call engineers receive an alert regarding customer transaction failures, querying by
Timestampallows them to reconstruct the exact timeline of database transaction aborts during the 01:00–04:00 window. - Simultaneously, comparing
ObservedTimestampagainstTimestampalerts the infrastructure team to the 3-hour telemetry ingestion lag, prompting automated checks on network partition recovery. - Furthermore, because the payment worker is written in Python (logging
CRITICALfor database deadlocks) while the inventory service is written in Java (loggingFATALfor stock reservation lockouts), querying forSeverityNumber >= 21immediately isolates all catastrophic settlement failures across both services in a single correlated dashboard.
During an extended network outage, a fleet of payment terminals keeps writing transaction logs to local plain-text files. After reconnecting, the files are uploaded to a central host, where an OpenTelemetry Collector's filelog receiver parses them hours later. An engineer needs to reconstruct the order in which events actually happened on the terminals. Which Log Record field should the query sort by, and why?
ObservedTimestamp, because it records the exact instant the Collector read and parsed each line
TraceFlags, because setting the sampled flag to 01 forces the backend to reorder logs chronologically
SeverityText, because logging libraries add a monotonic sequence counter to the severity string after network loss
Timestamp, because it holds the time the event occurred on the terminal (parsed from each line), whereas ObservedTimestamp holds when the Collector read the line
An enterprise observability team is constructing a centralized alerting pipeline across a polyglot microservice architecture composed of Java (using Log4j2), Python (using standard logging), and Go services, as well as legacy network switches emitting Syslog messages. The team wants a single alert definition to notify on-call engineers of all error-level and catastrophic fatal failures without writing complex regular expressions for language-specific severity strings like CRITICAL, FATAL, Error, and Emergency. How does the OpenTelemetry Logs Data Model address this requirement?
By mapping all incoming logs onto a standardized numerical scale (SeverityNumber) from 1 to 24, allowing alert rules to filter for SeverityNumber >= 17 to capture all ERROR and FATAL tiers uniformly across all runtimes
By permanently overwriting the original source severity string with a single fixed uppercase string in the SeverityText field, discarding the legacy logger level name permanently
By automatically converting all log records with severity levels above informational into synthetic distributed trace spans with an Error StatusCode
By forcing all programming language SDKs to deprecate language-specific logging levels and accept only a 3-tier enum consisting of INFO, WARNING, and ERROR
An application developer inspects a distributed trace in an observability UI and clicks on an HTTP span. The UI displays three contextual log records detailing a database connection failure that occurred during the span's execution. However, the developer did not write code to manually append trace or span IDs to the log messages. How did OpenTelemetry achieve this correlation between the log records and the active distributed trace?
The OpenTelemetry Collector used machine learning pattern matching on the log body text to associate the log with the most recent distributed trace
The OpenTelemetry logging bridge inspected the ambient in-process context at the moment of log emission and automatically populated the Log Record's TraceId, SpanId, and TraceFlags fields
The logging bridge created a brand-new distributed trace for every individual log statement and linked the original trace using a Baggage header key-value pair
The database driver injected its internal socket descriptor into the Resource attributes of both the span and the log record to bind them in storage
Sections you finish are checked off in the contents.