5.3 Metric Aggregations & Temporality

Key Takeaways

  • OpenTelemetry SDK aggregations transform raw instrument measurements into time-series data points; default aggregations include Sum for Counters, ExplicitBucketHistogram for Histograms, and LastValue for Gauges.

  • Explicit Bucket Histograms group values into fixed numerical boundary arrays, providing deterministic bucket counts and compatibility with Prometheus, but require upfront tuning to avoid boundary clumping.

  • Exponential Bucket Histograms dynamically allocate buckets using an exponential power base (22−scale2^{2^{-\text{scale}}}), delivering bounded relative error across wide dynamic ranges with compact memory usage.

  • Temporality defines the measurement timeframe: Delta temporality exports changes (Δ\Delta) since the last collection cycle (common in StatsD and cloud backends), while Cumulative temporality exports running totals since process start (required by Prometheus).

  • OpenTelemetry SDK Views allow operators to alter metric streams without modifying code, enabling attribute filtering to prevent cardinality explosions, bucket boundary tuning, and aggregation type overrides.

Last updated: September 2026

5.3 Metric Aggregations & Temporality

Quick Answer: In the OpenTelemetry SDK, Aggregations define how individual raw measurements are mathematically combined into exported time-series points (such as Sum, LastValue, or ExplicitBucketHistogram). Aggregation Temporality defines the time window over which measurements accumulate: Delta Temporality reports the change (Δ\Delta) that occurred strictly since the preceding export cycle (resetting to zero afterwards), whereas Cumulative Temporality reports the continuous running total accumulated since process startup. Views allow developers and operators to customize aggregations, drop dangerous high-cardinality attributes, and tune histogram buckets without modifying application source code.

When an application invokes counter.add(1) or histogram.record(0.045), the OpenTelemetry API does not immediately serialize and transmit each individual measurement over the network. Transmitting every raw measurement would saturate network interfaces and overwhelm backend time-series databases.

Instead, the OpenTelemetry SDK acts as an in-memory aggregation engine. It accumulates raw measurements across regular collection intervals, compresses them into statistical aggregates, applies temporality rules, and exports the resulting data points.


Metric Aggregations in the OpenTelemetry SDK

An Aggregation is the concrete mathematical operation applied by the SDK to convert a stream of raw measurements into an aggregated metric data point. The OpenTelemetry SDK specification defines the following aggregations:

Aggregation TypeDescriptionDefault Applied Instrument
DefaultSelects the aggregation for the instrument kind (the mappings below).All instruments unless a View overrides it
SumCalculates the arithmetic sum of all recorded numeric values.Counter, UpDownCounter, ObservableCounter, ObservableUpDownCounter
LastValueRetains only the most recently observed value, discarding prior readings.Gauge and ObservableGauge
ExplicitBucketHistogramAllocates recorded values into fixed, pre-configured upper bound intervals.Histogram
Base2ExponentialBucketHistogramDynamically allocates values into exponential scale buckets (22−scale2^{2^{-\text{scale}}}).Alternative for Histogram (select with a View or OTEL_EXPORTER_OTLP_METRICS_DEFAULT_HISTOGRAM_AGGREGATION=base2_exponential_bucket_histogram)
DropDiscards all incoming measurements for the matched instrument, emitting zero data points.Used via Views to silence unwanted metrics

Histogram Aggregations: Explicit vs. Exponential Buckets

Because histograms capture statistical distributions, their aggregation mechanism directly dictates memory consumption, network bandwidth, and percentile query accuracy.

1. Explicit Bucket Histograms

An Explicit Bucket Histogram uses an immutable, ordered array of boundary values that partition the continuous numeric spectrum into discrete intervals. The SDK's default boundaries are [0, 5, 10, 25, 50, 75, 100, 250, 500, 750, 1000, 2500, 5000, 7500, 10000], which suit millisecond values. Seconds-based semantic-convention histograms such as http.server.request.duration therefore carry an advisory boundary list, [0.005, 0.01, 0.025, 0.05, 0.075, 0.1, 0.25, 0.5, 0.75, 1, 2.5, 5, 7.5, 10], that instrumentation passes to the SDK.

When a measurement is recorded (e.g., 0.042s), the SDK finds the lowest boundary greater than or equal to the measurement (here, 0.05) and increments that bucket's counter.

The Data Model Emitted

For each collection window, an explicit bucket histogram emits:

  • The total count of all recorded samples (NN)
  • The arithmetic sum of all recorded samples
  • The observed min and max values
  • An array of bucket counts corresponding to each boundary interval, plus an implicit (+∞)(+\infty) bucket.

Advantages and Critical Limitations

  • Advantage: Predictable, constant memory footprint; universal native compatibility with Prometheus TSDB exposition format.
  • Limitation (Boundary Selection Dilemma): Operators must guess the expected distribution of values upfront. If boundaries are spaced too widely, percentile calculations suffer severe interpolation error. If a production outage causes latencies to spike from 100ms to 45 seconds, all degraded requests clump into the single +Inf bucket, completely blinding on-call engineers to the actual tail latency severity.

2. Exponential Bucket Histograms (High-Resolution Base-2)

To overcome the boundary selection dilemma, OpenTelemetry introduced Exponential Bucket Histograms (standardized in OTLP and adopted by modern observability engines). Instead of static manual boundaries, bucket boundaries are calculated mathematically using an exponential power formula:

boundary=2i⋅2−scale\text{boundary} = 2^{i \cdot 2^{-\text{scale}}}

where ii is the integer bucket index, and scale\text{scale} is an integer between −10-10 and +20+20. The SDK defaults are MaxSize = 160 buckets per positive or negative range and MaxScale = 20.

Scale = 0:  Boundaries double at each index:  ..., 0.5, 1, 2, 4, 8, 16, 32, ...
Scale = 1:  Boundaries grow by sqrt(2):      ..., 1, 1.414, 2, 2.828, 4, ...
Scale = 3:  Boundaries grow by 2^(1/8):      ~9% resolution across all scales

Dynamic Downscaling & Relative Error Guarantees

  • Fixed Relative Error: An exponential histogram provides a guaranteed relative error bound (e.g., within 5% of true value) across the entire dynamic range—whether measuring a 10-microsecond database ping or a 60-second batch import.
  • Dynamic Downscaling: If an application observes measurements spanning an unexpectedly wide range that exceeds the maximum configured bucket limit (e.g., 160 buckets), the OpenTelemetry SDK automatically downscales the histogram (reducing scale from 4 to 3, then 2). It merges adjacent bucket pairs in memory without losing accumulated counts, sums, or min/max bounds.

Aggregation Temporality: Delta vs. Cumulative

A core metrics concept is Aggregation Temporality. Temporality defines the time window over which accumulated metric measurements are calculated before being reported to an observability backend.

Measurement Events:   [+5]         [+10]        [+8]
Time Intervals:     0 ─── T1 ─────── T2 ──────── T3

Delta Temporality:     [ 5 ]        [ 10 ]       [ 8 ]   (Reports difference since last scrape)

Cumulative:            [ 5 ]        [ 15 ]       [ 23 ]  (Reports running sum since process start)

1. Delta Temporality

In Delta Temporality, each exported metric data point represents strictly the numeric change (Δ\Delta) that occurred during the time interval since the immediately preceding collection cycle ([Tn−1,Tn][T_{n-1}, T_n]).

Operational Mechanics

  • At the end of each collection cycle, the SDK exports the delta value accumulated during that window.
  • Immediately after export, the SDK's internal accumulator is reset to zero.
  • The start_time_unix_nano timestamp on the exported data point updates on every cycle to reflect the start of the current collection window.

Strengths & Weaknesses

  • Strengths:
    • Highly memory efficient: Ephemeral time-series (e.g., a burst of requests with a unique attribute) can be evicted from SDK memory after export, preventing memory leaks.
    • Resilient to container restarts: Process restarts do not create confusing counter reset jumps for downstream rate calculators.
  • Weaknesses:
    • Vulnerability to Packet Loss: If an export packet or scrape request is permanently dropped due to network instability, that measurement slice is permanently lost. The backend has no way to reconstruct the missed values, resulting in an irreversible undercount.
  • Target Ecosystems: StatsD-style systems and push-based backends that prefer or require delta data.

2. Cumulative Temporality

In Cumulative Temporality, each exported metric data point represents the continuous running total accumulated from the process startup time (T0T_0) up to the current collection timestamp (TnT_n).

Operational Mechanics

  • The internal accumulator is never reset to zero during the lifetime of the process.
  • The start_time_unix_nano timestamp remains permanently fixed at the initial startup timestamp of the process or instrument.
  • The exported value continuously climbs (for monotonic counters) or fluctuates (for UpDownCounters).

Strengths & Weaknesses

  • Strengths:
    • Network Resilience: If network partitioning causes one or two collection cycles to fail, zero data is lost. The next successful scrape carries the full cumulative running total from process start, allowing the backend to smoothly calculate rates.
  • Weaknesses:
    • Memory Overhead: The SDK must retain state for every unique attribute set ever encountered throughout the entire lifetime of the process, which can lead to memory exhaustion under high cardinality.
    • Reset Handling Complexity: When a process restarts, the cumulative value drops from a large number back to zero. Backend query engines must implement complex heuristics (like Prometheus rate() and increase()) to detect and handle counter resets.
  • Target Ecosystems: Prometheus (pull-based scraping model).

Temporality Preferences Across Protocols & Backends

Different metrics protocols and observability backends mandate specific temporality models. The OpenTelemetry SDK provides configurable temporality preferences on metric readers and exporters:

Observability Backend / ProtocolRequired / Preferred TemporalityConsequence of Wrong Temporality
Prometheus (pull scrape or Prometheus exporter)CumulativePrometheus counters and histograms are cumulative, and its rate() and increase() functions expect running totals; Prometheus exporters and pull readers always use cumulative.
Push backends that prefer deltaDeltaSet the exporter's temporality preference to delta, or convert in the Collector.
OTLP exporter (default)Cumulative (configurable)OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE accepts cumulative (default), delta (delta for Counter, Asynchronous Counter, and Histogram), or lowmemory (delta only for synchronous Counter and Histogram).

Exam Tip: Remember that Prometheus requires Cumulative temporality. Many cloud-native push backends prefer Delta temporality. In the Collector, the cumulativetodelta processor converts cumulative streams to delta, and the deltatocumulative processor does the reverse.


OpenTelemetry SDK Views: Declarative Metric Customization

A View is an OpenTelemetry SDK mechanism that allows operators, platform engineers, and developers to customize how metric streams are processed, aggregated, and exported without modifying a single line of application source code.

Application Code (API):     Meter.create_histogram("http.server.request.duration")
                                      │
                                      ▼
OpenTelemetry SDK View:     [ Match: "http.server.request.duration" ]
                            [ Action 1: Filter attributes -> Keep only http.route, http.response.status_code ]
                            [ Action 2: Custom Buckets -> [0.01, 0.05, 0.1, 0.5, 1.0, 5.0] ]
                                      │
                                      ▼
Exported Metric Stream:     Clean, low-cardinality, optimally bucketed time-series

Core Capabilities of SDK Views

OpenTelemetry Views provide five primary capabilities:

  1. Cardinality Defense (Attribute Filtering): In production, developers sometimes inadvertently attach high-cardinality attributes to metrics—such as user.id, customer.email, or raw URLs containing query strings. This creates millions of unique time-series that crash the backend TSDB. An SDK View can match the metric and define an attribute allow-list (the attribute_keys stream setting), automatically stripping dangerous attributes before aggregation.
  2. Customizing Histogram Bucket Boundaries: The SDK's default histogram boundaries may not suit specific domains. For an in-memory Redis cache client, latency ranges between 100 microseconds and 5 milliseconds; the default buckets (extending to 10 seconds) provide zero visibility. A View can supply custom boundaries (e.g., [0.0001, 0.0005, 0.001, 0.002, 0.005]) for that specific instrument.
  3. Changing the Aggregation Type: A View can override the default aggregation—for instance, switching a Histogram from ExplicitBucketHistogram to ExponentialBucketHistogram, or assigning the Drop aggregation to completely suppress noisy third-party telemetry.
  4. Renaming Streams and Changing Descriptions: Views can rename the output stream (e.g., http.server.request.duration to legacy.http.latency during a migration) and change its description. A View cannot change the unit.
  5. Setting a Cardinality Limit: A View can set aggregation_cardinality_limit for a stream. Without one, the SDK uses the reader's default or 2,000 series per metric; measurements for new attribute sets beyond the limit are folded into one overflow series tagged otel.metric.overflow=true.

Practical Code: Configuring a View in the SDK

The following Python example demonstrates configuring a View within the MeterProvider to drop unwanted attributes and apply custom latency buckets:

from opentelemetry.sdk.metrics import MeterProvider
from opentelemetry.sdk.metrics.view import View, ExplicitBucketHistogramAggregation

# Define a View that matches http.server.request.duration
custom_http_view = View(
    instrument_name="http.server.request.duration",
    # 1. Cardinality Defense: Retain only necessary low-cardinality attributes
    attribute_keys={"http.request.method", "http.response.status_code", "http.route"},
    # 2. Custom Explicit Buckets optimized for sub-second web transactions
    aggregation=ExplicitBucketHistogramAggregation(
        boundaries=[0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5]
    )
)

# Register the View during MeterProvider initialization
meter_provider = MeterProvider(
    views=[custom_http_view]
)
Loading diagram...
Cumulative vs Delta Measurement Timeline
Test Your Knowledge

A platform engineering team is streaming high-volume payment metrics to an OTLP-based observability backend using Delta temporality. A brief network partition causes two consecutive metric export packets to be permanently dropped. What is the impact on the backend's metric calculations?

A

The metric data for the dropped collection intervals is permanently lost, causing an irreversible undercount in the backend's cumulative total calculations

B

The backend automatically queries the SDK's local persistent disk buffer to replay the missing delta frames

C

The next successful packet will automatically include the values from the dropped intervals because Delta instruments retain unacknowledged frames

D

The SDK terminates the application process due to an unacknowledged TCP export timeout

Test Your Knowledge

A production microservice experiences severe memory pressure in its Prometheus TSDB backend caused by a runaway metric dimension: a developer added user_id as an attribute on http.server.request.duration. The infrastructure team wants to drop the user_id attribute across all running services without modifying source code or refactoring application instrumentation. How can this be accomplished within the OpenTelemetry SDK?

A

By switching all metrics from synchronous Histograms to Asynchronous Counters

B

By configuring an OpenTelemetry SDK View that defines an attribute filter to drop user_id from the target metric

C

By altering the W3C traceparent header to truncate the baggage payload

D

By forcing the SDK into Exponential Bucket Histogram mode

Test Your Knowledge

An SRE team is designing high-precision latency tracking for a financial trading API with latencies ranging from 50 microseconds to 10 seconds. Using explicit bucket histograms results in either massive memory bloat (thousands of buckets) or poor resolution in critical percentiles. Which OpenTelemetry aggregation addresses this challenge by dynamically scaling buckets based on an exponential power function?

A

Drop Aggregation

B

LastValue Aggregation

C

Exponential Bucket Histogram Aggregation

D

Sum Aggregation with Delta Temporality

Sections you finish are checked off in the contents.