5.2 Asynchronous (Observable) Metric Instruments

Key Takeaways

  • Asynchronous (Observable) instruments operate on a pull/callback model, where registered callbacks are executed periodically by the SDK during metric collection cycles rather than inline per request.

  • Because asynchronous callbacks execute on an SDK background collection thread, they are NOT associated with an active distributed execution Context and cannot capture inline TraceId/SpanId exemplars.

  • The three asynchronous instruments are ObservableCounter (monotonic cumulative totals like CPU execution time), ObservableUpDownCounter (non-monotonic counts like active threads or open file descriptors), and ObservableGauge (instantaneous non-additive states like temperature or memory usage percentage).

  • Batch Callbacks enable registering a single callback function across multiple observable instruments, allowing expensive underlying subsystem statistics (such as /proc/diskstats) to be scraped once and recorded atomically.

  • Choosing between synchronous and asynchronous instruments depends primarily on the event origin: synchronous for high-frequency internal application events, and asynchronous for periodic polling of external subsystem states.

Last updated: September 2026

5.2 Asynchronous (Observable) Metric Instruments

Quick Answer: Asynchronous (Observable) metric instruments—also known as callback or gauge-like instruments—are not called inline during application transactions. Instead, the OpenTelemetry SDK invokes user-registered callback functions periodically at every collection/scrape interval to report the current state of a system. Asynchronous instruments execute outside the request lifecycle, meaning they are NOT associated with an active distributed trace Context and cannot automatically record trace exemplars. The three asynchronous types are ObservableCounter, ObservableUpDownCounter, and ObservableGauge.

While synchronous instruments excel at capturing measurements triggered directly by application business logic (such as an incoming HTTP call or a completed database transaction), many critical operational metrics originate outside the immediate request flow. Operating system metrics, hardware telemetry, kernel counters, thread pool internals, and external cache statistics exist as standing states rather than discrete request-level events.

Attempting to measure these states synchronously—such as querying the Linux kernel's /proc/stat on every incoming HTTP request—would introduce unacceptable CPU and I/O overhead to production services. To solve this problem, OpenTelemetry provides Asynchronous (Observable) Instruments.


The Asynchronous Invocation Model

Asynchronous instruments operate on an inverted control flow often described as a pull model or callback model:

Application Startup:
  └── meter.create_observable_gauge("system.memory.usage", callbacks=[observe_memory])

Runtime (No Overhead on Request Threads):
  ├── Request 1 Processed ──> (No metric code executed)
  ├── Request 2 Processed ──> (No metric code executed)
  └── Request N Processed ──> (No metric code executed)

Periodic Collection Cycle (default every 60 s):
  └── OpenTelemetry SDK MetricReader Scraper
        ├── Invokes observe_memory(observer)
        │     ├── Reads OS memory statistics
        │     └── observer.observe(used_bytes, attributes)
        └── Exports gathered measurements via OTLP

The Role of Registered Callbacks

When defining an observable instrument, the developer does not receive an object with an add() or record() method. Instead, the developer supplies one or more callback functions during instrument creation (or registers callbacks after instantiation).

These callbacks remain dormant while the application executes requests. When the configured OpenTelemetry SDK MetricReader (such as a periodic exporting metric reader or a Prometheus scrape endpoint) initiates a collection cycle, the SDK invokes the registered callbacks. The callback inspects the target resource, reads its current state, and submits observations to an Observer object.

Context Isolation: No Ambient TraceContext

A foundational concept is the execution context of asynchronous callbacks:

Asynchronous instrument callbacks are NEVER associated with an active distributed execution Context!

When a synchronous instrument executes, it runs on the application thread handling a specific request, allowing the SDK to access Context.current() and capture the active TraceId and SpanId as an exemplar. In contrast, asynchronous callbacks execute on an independent SDK timer or exporter thread. At the instant the callback runs, there is no active HTTP request, no active RPC span, and no distributed trace context in scope.

Consequently, asynchronous instruments cannot record exemplars or link measurements directly to individual distributed traces. They provide macro-level system state observations, not request-correlated traces.


The Three Asynchronous Instrument Types

The OpenTelemetry Metrics API provides three asynchronous instrument primitives:

                         Asynchronous Instruments
                                    │
         ┌──────────────────────────┼──────────────────────────┐
         ▼                          ▼                          ▼
 ObservableCounter          ObservableUpDownCounter        ObservableGauge
   (Monotonic)                  (Non-Monotonic)             (Instantaneous)
 Cumulative Total              Fluctuating State             Non-Additive

1. ObservableCounter (Asynchronous Counter)

An ObservableCounter is an asynchronous instrument that reports a monotonically increasing cumulative value via a callback. Like its synchronous counterpart, it represents a sum that only ever increases (or resets to zero upon system reboot).

Operational Characteristics

  • Callback Reporting: The callback observes the cumulative total reached at the instant of the collection scrape.
  • Monotonicity: Each observed value must be greater than or equal to the previously observed value under the same attribute set (unless a system reboot or counter reset occurred).
  • Standard Use Cases:
    • Operating system CPU execution time in seconds obtained from /proc/stat (process.cpu.time)
    • Cumulative network octets or packets received by a network interface since OS boot
    • Cumulative page faults handled by the virtual memory subsystem
    • Cumulative garbage collection collection time reported by a runtime JMX/V8 engine

2. ObservableUpDownCounter (Asynchronous UpDownCounter)

An ObservableUpDownCounter is an asynchronous instrument that reports non-monotonic values via a callback, where the reported value can rise and fall over time, and where values are additive across entities.

Operational Characteristics

  • Additive Property: The reported values can be meaningfully summed across multiple instances, nodes, or processes. For example, summing the active thread counts of 10 microservice pods yields the total active threads across the service fleet.
  • Standard Use Cases:
    • Current number of JVM threads (jvm.thread.count)
    • Process open file descriptor count (process.unix.file_descriptor.count)
    • Total allocated heap memory bytes currently retained by an application
    • Total number of worker jobs currently queued in an external Celery or Redis queue

3. ObservableGauge (Asynchronous Gauge)

An ObservableGauge is an asynchronous instrument that records non-additive instantaneous values via a callback.

The Critical Concept: Non-Additive Values

The fundamental distinction between an ObservableUpDownCounter and an ObservableGauge lies in additivity:

  • If you have two application servers where Server A has 50 active threads and Server B has 70 active threads, summing them equals 120 total active threads. Thread count is additive, making it an ObservableUpDownCounter.
  • If Server A has a CPU temperature of 75°C and Server B has a CPU temperature of 65°C, adding them together (75+65=140∘C75 + 65 = 140^\circ\text{C}) is mathematically meaningless. You can calculate the average (70°C), minimum, or maximum, but you cannot sum them. Therefore, temperature is non-additive, making it an ObservableGauge.

Standard Use Cases

  • Current CPU core temperature in Celsius
  • Memory utilization percentage (0.0% to 100.0%)
  • Instantaneous filesystem usage fraction (system.filesystem.utilization)
  • Current room humidity or atmospheric pressure in an edge IoT facility
  • Remaining battery percentage

Batch Callbacks (Multi-Instrument Callbacks)

In real-world infrastructure monitoring, querying system status often produces multiple related metrics from a single operational source. For example, reading Linux /proc/diskstats provides read operations, write operations, read bytes, write bytes, and I/O time simultaneously.

If an engineer registers five separate Observable instruments with five independent callbacks, the SDK will invoke five separate file system operations on /proc/diskstats. This creates two major problems:

  1. Redundant CPU and I/O Overhead: Parsing the same file five times per scrape cycle.
  2. Temporal State Skew: Measurements are captured at slightly different millisecond offsets, resulting in inconsistencies between related metrics.

The Batch Callback Solution

OpenTelemetry solves this via Batch Callbacks. A batch callback allows an application to register a single callback associated with a list of multiple observable instruments. When the collection cycle executes, the callback runs exactly once, retrieves all metrics from the underlying subsystem, and observes values across all target instruments simultaneously:

// Go exposes multi-instrument callbacks through Meter.RegisterCallback.
// (Java offers meter.batchCallback; Python's API has only per-instrument callbacks.)
diskIO, _ := meter.Int64ObservableCounter("system.disk.io",
    metric.WithUnit("By"), metric.WithDescription("Disk bytes transferred"))
diskOps, _ := meter.Int64ObservableCounter("system.disk.operations",
    metric.WithUnit("{operation}"), metric.WithDescription("Disk operations"))

// One callback reads /proc/diskstats once and observes both instruments.
_, err := meter.RegisterCallback(func(ctx context.Context, o metric.Observer) error {
    for _, dev := range readDiskStats() { // parses /proc/diskstats once per collection
        read := metric.WithAttributes(
            attribute.String("system.device", dev.Name),
            attribute.String("disk.io.direction", "read"))
        write := metric.WithAttributes(
            attribute.String("system.device", dev.Name),
            attribute.String("disk.io.direction", "write"))
        o.ObserveInt64(diskIO, dev.ReadBytes, read)
        o.ObserveInt64(diskIO, dev.WriteBytes, write)
        o.ObserveInt64(diskOps, dev.Reads, read)
        o.ObserveInt64(diskOps, dev.Writes, write)
    }
    return nil
}, diskIO, diskOps)
if err != nil {
    log.Fatal(err)
}

Comparison: Synchronous vs. Asynchronous Instruments

Architectural DimensionSynchronous InstrumentsAsynchronous (Observable) Instruments
Invocation MethodInline push (add(), record()) on calling threadPeriodic pull callback (observe()) on SDK thread
Trigger PointExact moment an application event occursEach collection: the reader's interval (60 s by default for the periodic reader) or a scrape
Trace Context BindingYES: Automatically captures active TraceId and SpanIdNO: Executes outside request context; no ambient span
Exemplar SupportFully supported (links metric outliers to trace spans)Not supported (no trace context available)
Primary Overhead LocationExecution path of live user requests (must be non-blocking)Scrape interval execution (background timer thread)
Primary Target DomainRequest counts, latency distributions, inline queue deltasOS stats, CPU times, thread pools, memory levels, temperatures

Complete Master Table: All Seven OpenTelemetry Instruments

The Metrics API defines seven instruments:

Instrument NameSynchronicityMonotonic?Additive?Core API MethodTrace Context?Canonical Production Example
CounterSynchronousYES (≥0\ge 0)YESadd(value, attrs)YESshop.orders.completed (custom)
UpDownCounterSynchronousNO (+/−+/-)YESadd(value, attrs)YEShttp.server.active_requests
HistogramSynchronousN/A (Distribution)YESrecord(value, attrs)YEShttp.server.request.duration
GaugeSynchronousNONOrecord(value, attrs)YESa fan-speed change listener
ObservableCounterAsynchronousYES (≥0\ge 0)YESCallback observe()NOprocess.cpu.time
ObservableUpDownCounterAsynchronousNO (+/−+/-)YESCallback observe()NOprocess.unix.file_descriptor.count
ObservableGaugeAsynchronousNO (+/−+/-)NOCallback observe()NOsystem.memory.utilization
Loading diagram...
Synchronous Push vs Asynchronous Pull Invocation Models
Test Your Knowledge

An operations team needs to collect the current room ambient temperature in a data center and the instantaneous CPU throttling percentage of a server every 30 seconds. These values represent non-additive state measurements where summing them across instances or intervals is meaningless. Which OpenTelemetry instrument is designed specifically for this use case?

A

Synchronous Counter

B

ObservableGauge (Asynchronous Gauge)

C

ObservableCounter

D

UpDownCounter

Test Your Knowledge

A software architect wants to record a custom metric that associates every measurement with the currently active distributed trace (TraceId and SpanId) to support exemplar correlation in their APM backend. Why must the architect use a synchronous metric instrument rather than an asynchronous (observable) instrument?

A

Asynchronous instruments only support integer values, which cannot store 128-bit trace identifiers

B

Asynchronous callbacks are only supported in compiled languages like Go and C++, not dynamic runtimes

C

Synchronous instruments execute inline within the active execution thread and bind to the current Context, whereas asynchronous callbacks execute on an independent scheduler thread outside any request context

D

Asynchronous instruments automatically drop all attributes and metadata during SDK export

Test Your Knowledge

An engineer is instrumenting a storage subsystem where reading Linux /proc/diskstats incurs measurable I/O overhead. The subsystem needs to update three distinct metrics: system.disk.read_bytes, system.disk.write_bytes, and system.disk.operations. What OpenTelemetry mechanism should be used to read /proc/diskstats exactly once per collection cycle and update all three instruments atomically?

A

A Synchronous Histogram with multi-value records

B

Three independent synchronous UpDownCounters wrapped in a mutex

C

A custom Collector processor that executes bash scripts

D

An OpenTelemetry Batch Callback registered with the Meter for multiple observable instruments

Sections you finish are checked off in the contents.