1.3 Telemetry Analysis, SLOs & High Cardinality
Key Takeaways
Site Reliability Engineering (SRE) utilizes Service Level Indicators (SLIs) to measure actual system performance, Service Level Objectives (SLOs) to set target reliability goals, and Error Budgets to balance feature velocity against system stability.
High cardinality occurs when metric dimensions contain an immense number of unique values, causing a combinatorial explosion of distinct time series that exhausts Time Series Database (TSDB) memory and degrades query performance.
While metrics require bounded, low-cardinality dimensions, distributed traces and structured logs are architecturally designed to handle high-cardinality metadata such as User IDs, Order IDs, and Tenant IDs.
OpenTelemetry Exemplars bridge metrics and traces by attaching specific Trace IDs and Span IDs directly to metric histogram buckets and counter measurements, enabling instant transition from aggregated anomaly to root-cause trace.
Telemetry cost and fidelity optimization demands deliberate governance: leveraging tail-based sampling in the Collector, pre-aggregating high-volume metrics, filtering redundant debug logs, and enforcing tiered data retention.
1.3 Telemetry Analysis, SLOs & High Cardinality
Quick Answer: Telemetry analysis enables modern Site Reliability Engineering (SRE) by measuring Service Level Indicators (SLIs) against Service Level Objectives (SLOs) to govern Error Budgets. A major challenge in telemetry management is high cardinality—an explosion of unique label combinations that can crash metric time series databases. OpenTelemetry resolves this through Exemplars, which attach trace identifiers to aggregate metric buckets, allowing engineers to jump from high-level SLO breaches directly to granular high-cardinality distributed traces.
Collecting telemetry is merely the operational prerequisite; the ultimate goal is telemetry analysis—deriving actionable insights that maintain system reliability, guide engineering investments, and minimize downtime. In modern Site Reliability Engineering (SRE), this is accomplished through formal reliability objectives, rigorous metric hygiene, and intelligent correlation between aggregated trends and granular execution traces.
Reliability Engineering: SLIs, SLOs, and Error Budgets
Rather than aiming for an unrealistic target of "100% uptime" (which paralyzes deployment velocity and creates unsustainable engineering costs), modern organizations adopt the Site Reliability Engineering framework popularized by Google. This framework centers on three core constructs:
1. Service Level Indicator (SLI)
An SLI is a carefully defined quantitative measure of some aspect of the level of service provided. It represents the ratio of "good events" to "total events":
In telemetry pipelines, SLIs are computed directly from OpenTelemetry metrics or distributed traces. Common SLIs include:
- Availability SLI — The percentage of HTTP server requests returning a status code less than 500:
- Latency SLI — The percentage of requests completing in under a defined duration threshold (such as ), calculated from an OpenTelemetry explicit bucket histogram (
http.server.request.duration):
2. Service Level Objective (SLO)
An SLO is a formal target reliability level for an SLI over a rolling compliance window (typically 7, 30, or 90 days), agreed upon by engineering, product, and business stakeholders. For example:
- "99.9% of user authentication requests must succeed with HTTP status < 500 over any rolling 30-day window."
- "95% of database query operations must complete in under 50ms over a rolling 7-day period."
3. Service Level Agreement (SLA)
An SLA is a formal legal or business contract with external customers that specifies financial or contractual penalties (such as billing credits or refunds) if the service fails to meet agreed thresholds. SLOs are always set stricter than SLAs. If your customer SLA promises 99.5% availability, your internal engineering SLO should be set to at least 99.9% to provide a safety margin before contractual penalties apply.
4. Error Budget
The Error Budget is the complement of the SLO:
For a service with a 99.9% availability SLO over a 30-day window, the allowable failure rate is 0.1%. If the service receives 10,000,000 requests per month, the engineering team can experience up to 10,000 failed requests before exhausting the error budget.
The error budget acts as an organizational governor balancing feature velocity against stability:
- Surplus Budget — When the error budget is healthy, developers can aggressively deploy new features, perform architectural refactoring, and conduct chaos engineering experiments.
- Exhausted Budget — When the error budget is depleted due to outages or regressions, feature deployments freeze. Engineering resources are redirected entirely toward reliability improvements, test automation, and infrastructure hardening.
| Reliability Construct | Core Definition | Formula / Implementation | Role in Incident Management |
|---|---|---|---|
| Service Level Indicator (SLI) | Quantitative metric of actual service behavior | Ratio of valid good events over total evaluated events | Real-time detection of performance or availability drops |
| Service Level Objective (SLO) | Target reliability percentage over a rolling window | Agreed target (such as 99.9% over rolling 30 days) | Internal engineering standard governing team priorities |
| Service Level Agreement (SLA) | Contractual commitment to external clients | Contract terms with financial penalties or credits | Business boundary; set looser than internal SLO |
| Error Budget | Allowable failure margin during the SLO window | 100% minus the target SLO percentage | Release governor: freezes deployments when exhausted |
The High Cardinality Challenge in Metrics
In mathematics and database theory, cardinality refers to the number of unique elements in a particular set. In observability metrics, cardinality refers to the total number of unique time series generated by the Cartesian product of all tag or attribute dimensions attached to a metric instrument.
The Combinatorial Explosion
Consider an OpenTelemetry counter metric http.server.requests measuring incoming web traffic. Suppose an engineer adds several standard operational attributes:
environment(production, staging) = 2 valuesregion(us-east, us-west, eu-central, ap-southeast) = 4 valueshttp.request.method(GET, POST, PUT, DELETE) = 4 valueshttp.response.status_code(200, 201, 400, 401, 403, 404, 500, 502, 503) = 9 values
The total number of active time series created in the Time Series Database (TSDB) is:
This is completely safe and represents low cardinality. A modern TSDB (such as Prometheus, Thanos, Cortex, or M3DB) can easily handle hundreds of thousands of concurrent time series.
Now suppose a developer unknowingly adds user_id as an attribute to this same metric to track per-user traffic, and the application has 2,000,000 active users. The calculation becomes:
This is a catastrophic cardinality explosion. Every unique time series requires in-memory indexing, inverted label index updating, disk blocks, and chunk compaction. Within minutes, the TSDB will experience severe RAM exhaustion, trigger Out-Of-Memory (OOM) kernel kills, thrash the disk cache, and render metric querying completely inoperable for the entire engineering organization.
Golden Rule of Metric Cardinality
- Allowed in Metrics (Bounded Low Cardinality): Status codes, HTTP methods, regions, service names, deployment environments, error classes (enum-like sets with fewer than a few hundred possible values).
- Prohibited in Metrics (Unbounded High Cardinality): User IDs, email addresses, order IDs, UUIDs, container IDs, client IP addresses, credit card hashes, or arbitrary search query strings.
High Cardinality in Traces and Logs
If high-cardinality attributes are strictly prohibited in metrics, where do they belong? They belong in distributed traces and structured logs!
The fundamental architectural difference lies in the underlying storage models:
- Metrics Engine (TSDB): Assumes that every unique combination of labels defines a continuous, infinite time series stream that must be tracked, indexed in memory, and scraped at continuous intervals.
- Trace and Log Stores (Document/Columnar Stores): Treat data as discrete, append-only events. Storing
user_id = "user_98721"ororder_id = "ord_55210"on an OpenTelemetry span attribute simply adds a string value to that individual span's JSON/Protobuf payload in an object store or columnar index (such as ClickHouse, Tempo, Jaeger, or Elasticsearch). It does not create a new infinite time series stream.
Therefore, high-cardinality contextual metadata should always be placed on span attributes or structured log records, where it powers needle-in-a-haystack debugging without threatening database stability.
| Telemetry Signal | Cardinality Suitability | Storage Architecture | Example Safe Attributes | Dangerous Prohibited Attributes |
|---|---|---|---|---|
| Metrics | Low cardinality only (strictly bounded) | In-memory time series index (TSDB) | http.response.status_code, http.request.method, region | user_id, email, order_id, uuid |
| Distributed Traces | High cardinality supported | Document/columnar event store with sampling | user_id, order_id, tenant_id, cart_id | Dynamic span names (names must stay generic) |
| Structured Logs | High cardinality supported | Inverted text index and columnar log store | user_id, stack_trace, query_param | High-frequency unbatched debug logs |
Exemplars: The Dynamic Bridge Between Metrics and Traces
While metrics are ideal for detecting aggregate SLO anomalies and traces are ideal for deep high-cardinality debugging, an operational gap historically separated them: When a metric histogram shows an anomalous latency spike, how do you find the exact trace that caused it?
Historically, an engineer had to look at the graph, note the timestamp (e.g., 14:23:15 UTC), switch to a trace search tool, filter for spans around that timestamp, and manually guess which trace matched the spike.
OpenTelemetry solves this problem through Exemplars.
How Exemplars Work
An Exemplar is a concrete sample data point attached directly to a metric measurement. When a synchronous instrument (such as a Histogram or Counter) records a value, the SDK first applies the exemplar filter. The default, OTEL_METRICS_EXEMPLAR_FILTER=trace_based, makes a measurement eligible only when it is recorded inside a sampled span (always_on and always_off are the other standard values). An eligible measurement goes to the Exemplar Reservoir, which keeps a small sample of the value, its timestamp, and:
TraceIdSpanId- Optional filtered attributes
Metric: http.server.request.duration (Histogram Bucket: 1000ms - 2500ms)
└── Aggregated Bucket Count: 42
└── Attached Exemplar:
├── Value: 1842ms
├── Timestamp: 2026-09-29T14:23:15.102Z
├── TraceId: 4bf92f3577b34da6a3ce929d0e0e4736
└── SpanId: 00f067aa0ba902b7
In modern observability visualization tools (such as Grafana), exemplars are rendered as clickable dots overlaid directly on the metric time series graph. An engineer observing a sudden spike in the 99th percentile latency can click directly on the outlier exemplar dot, which immediately opens the full distributed trace DAG showing the exact microservice, database query, and high-cardinality user_id that triggered the latency.
Telemetry Cost vs. Fidelity Optimization
In high-throughput distributed systems processing billions of events daily, capturing 100% of all telemetry with full attributes is financially and technically impossible. Telemetry engineering requires balancing diagnostic fidelity against network, compute, and storage costs.
Optimization Strategies
- Head-Based Sampling (at the SDK) — The sampling decision is made at the very beginning of a trace lifecycle (at the root span) before the full transaction executes. Examples include probabilistic sampling (e.g., sample 5% of all traces) or rate-limiting sampling (sample at most 100 traces per second). While head-based sampling is lightweight and consumes minimal memory, it risks dropping rare, critical error traces.
- Tail-Based Sampling (at the Collector) — The sampling decision is deferred until the entire distributed trace completes. The OpenTelemetry Collector buffers spans in memory across all services. When the trace finishes, a tail-sampling processor evaluates rules: retain 100% of traces containing HTTP 5xx errors, retain 100% of traces exceeding latency thresholds (such as ), and sample down routine 200 OK traces to 0.1%. This keeps the diagnostically valuable traces while sharply cutting stored volume; the actual savings depend on the traffic mix and the policies you write.
- Metric Volume Reduction — SDK Views (Section 5.3) can drop high-cardinality attributes before aggregation, and intermediate Collectors can filter unneeded metric streams, strip attributes, or convert temporality before forwarding, so the backend stores fewer series.
- Log Filtering and OTTL Transformation — High-volume debug logs can be filtered out at the Collector level in production environments. Repetitive access logs can be parsed using the OpenTelemetry Transformation Language (OTTL) to generate aggregate counter metrics, dropping the raw log lines entirely.
- Tiered Storage Architecture — Configure hot storage (in-memory or fast NVMe SSD) for the most recent 7 days of high-fidelity data, warm object storage (S3/GCS) for 30-90 days of downsampled metrics and sampled traces, and cold compressed archives for compliance audit logs.
Practical Case Study: Debugging a Multi-Tenant SaaS Latency Spike
To understand how these concepts converge in real-world operations, consider a multi-tenant B2B analytics platform where enterprise tenants share common computing infrastructure.
Incident Timeline
- SLO Burn Rate Alert Triggers — An automated alert alerts the on-call engineer that the 30-day 99th percentile latency SLO (set at ) is burning through its error budget at 14x the normal rate.
- Metric Inspection — The SRE checks the dashboard for
http.server.request.duration. A clear bifurcated latency distribution appears: the median (p50) remains steady at 45ms, but the tail (p99) has surged to 4,800ms. - Exemplar Navigation — The SRE highlights the outlier points in the elevated histogram bucket and clicks on an attached Exemplar.
- Trace DAG Drill-Down — The UI opens the exemplar trace. The root span
/api/v2/analytics/reportstook 4,812ms. Following the parent-child span hierarchy immediately reveals that an internal child span callingReportQueryService.ExecuteSQLconsumed 4,650ms. - High-Cardinality Attribute Analysis — Inspecting the span's semantic attributes reveals:
tenant.id = "tenant_enterprise_772"query.complexity = "unbounded_aggregation"db.rows_scanned = 48500000
- Resolution — The engineer identifies that a single enterprise tenant initiated an unindexed report query scanning 48 million rows across a 5-year date range. The team isolates the query to a dedicated read-replica, applies an index migration, and enforces query date partitioning, restoring the SLO and protecting the error budget from total exhaustion.
Without exemplars and correlated high-cardinality trace attributes, isolating this single tenant across hundreds of microservices would have required hours of manual log parsing.
A developer adds 'user_email' as an attribute to an OpenTelemetry counter metric tracking successful user logins. Within two hours of deployment, the cluster's Prometheus monitoring backend crashes due to out-of-memory errors. What is the fundamental cause of this failure?
OpenTelemetry metrics APIs explicitly reject string attribute values and throw fatal runtime exceptions
Prometheus does not support counter metric instruments generated by OpenTelemetry SDKs
The immense number of unique email addresses caused a high-cardinality combinatorial explosion of time series that exhausted TSDB memory
The OpenTelemetry Collector corrupted the OTLP wire format due to missing W3C Baggage headers
An SRE observes a sudden latency regression on a service level indicator (SLI) metric dashboard. Which OpenTelemetry feature allows the engineer to click directly on the elevated latency histogram bucket and immediately open the corresponding distributed trace?
W3C Baggage propagation headers
Tail-based sampling policies
Asynchronous metric gauge callbacks
Exemplars
A high-throughput financial microservice processes 250 million transactions daily. Storing 100% of all trace spans in persistent storage is cost-prohibitive, but the engineering team cannot afford to lose visibility into rare 5xx errors or requests with severe tail latency. Which telemetry strategy best satisfies both requirements?
Deploying tail-based sampling in the OpenTelemetry Collector to buffer traces and retain 100% of error and high-latency traces while downsampling routine 200 OK traces
Storing all distributed traces inside a high-cardinality time series database metric instrument
Disabling W3C context propagation across microservices to truncate span sizes
Converting all distributed traces into unstructured plaintext log statements
Sections you finish are checked off in the contents.