5.1 Synchronous Metric Instruments
Key Takeaways
Synchronous instruments (Counter, UpDownCounter, Histogram, and Gauge) are invoked inline during application execution, capturing measurements at the instant an event occurs and binding to the active execution Context.
The Counter instrument measures strictly non-negative (), monotonically increasing cumulative totals via add(value, attributes) for events such as HTTP requests, bytes transferred, and cache hits.
The UpDownCounter instrument supports non-monotonic values with both positive and negative deltas via add(value, attributes), making it ideal for tracking active concurrent requests, queue sizes, and connection pool availability.
The Histogram instrument samples the statistical distribution of values over time using record(value, attributes), enabling computation of counts, sums, min/max, and percentiles (p50, p95, p99) for latencies and payload sizes.
OpenTelemetry metric naming follows lowercase dot-separated notation with standardized UCUM units (e.g., http.server.request.duration in s, messaging.publish.messages in {message}).
5.1 Synchronous Metric Instruments
Quick Answer: Synchronous metric instruments are invoked inline as part of an application's live request execution flow. They record measurements at the exact instant an event occurs and automatically bind to the active distributed execution Context (capturing
TraceIdandSpanIdfor exemplar linking). The OpenTelemetry Metrics API defines four synchronous instruments: Counter (monotonically increasing non-negative values), UpDownCounter (non-monotonic values accepting positive and negative deltas), Histogram (distributions of values such as request durations and payload sizes), and Gauge (non-additive current values recorded when they change).
In distributed cloud-native architectures, metrics provide the primary telemetry signal for real-time alerting, health monitoring, service level objective (SLO) tracking, and automated capacity scaling. While distributed traces offer deep transactional lineage through directed acyclic graphs of spans, metrics aggregate numeric observations over time into highly compressed, query-efficient time-series data.
Within the OpenTelemetry specification, the Metrics API is decoupled from the Metrics SDK. Application developers author telemetry against the abstract API, while the SDK manages accumulators, aggregation algorithms, views, temporality, and export pipelines. Central to the Metrics API is the concept of metric instruments. These instruments are categorized along two primary architectural dimensions:
- Invocation Model: Synchronous (inline push) versus Asynchronous (callback pull).
- Mathematical Nature: Monotonic counters, non-monotonic counters, statistical histograms, and point-in-time gauges.
What are Synchronous Metric Instruments?
Synchronous metric instruments are invoked directly inline by application code or middleware during the execution of a business operation or network transaction. When a user checks out an order, an HTTP server dispatches an endpoint handler, or a worker thread dequeues a message from a Kafka partition, synchronous instruments are called at that exact execution milestone to record a numeric measurement.
The Inline Execution Lifecycle
Unlike asynchronous instruments that rely on scheduled background pollers, synchronous instruments execute on the application's active calling thread:
Application Thread: [Incoming HTTP Request]
│
├── Active Span Created: TraceId=4bf92f... SpanId=00f067...
├── Business Logic Execution
│ ├── database_queries_counter.add(1, attributes)
│ └── active_requests_pool.add(1, attributes)
├── HTTP Response Formatted (200 OK, 45ms)
│ ├── active_requests_pool.add(-1, attributes)
│ ├── http_duration_histogram.record(0.045, attributes)
│ └── http_requests_total.add(1, attributes)
└── Request Completed (Active Span Closed)
Execution Context Association & Exemplars
A critical distinguishing property of all synchronous instruments is their Context association. Whenever application code calls a synchronous instrument, the OpenTelemetry SDK inspects the runtime environment's ambient execution context (Context.current()).
Because the call occurs inline within the request lifecycle, the SDK automatically extracts the active TraceId and SpanId. Compliant OpenTelemetry SDKs can attach this contextual trace linkage to the metric data point as an Exemplar. An exemplar is a concrete trace identifier associated with a specific metric measurement (such as an extreme latency outlier recorded in a histogram bucket). This allows site reliability engineers inspecting a Prometheus or Grafana latency spike to jump directly into the exact distributed trace that caused the degradation with a single click.
Performance & Thread Safety Guarantees
Because synchronous instruments execute on high-frequency application paths (potentially invoked tens of thousands of times per second per container), SDK implementations are engineered for:
- Non-blocking execution: Instrument calls do not initiate network I/O, disk writes, or remote procedure calls.
- Atomic concurrency: Measurements update in-memory thread-safe lock-free accumulators (such as atomic long/double registers or striped adders).
- Low-allocation hot paths: Mature SDKs optimize attribute handling to limit garbage-collection pressure during recording (some offer bound instruments or pre-built attribute sets for this).
The Four Synchronous Instrument Types
The OpenTelemetry Metrics API provides four synchronous instruments, each designed for a distinct mathematical and operational requirement:
Synchronous Instruments
┌──────────────────┬───────────────┴──┬──────────────────┐
▼ ▼ ▼ ▼
Counter UpDownCounter Histogram Gauge
(Monotonic) (Non-Monotonic) (Distribution) (Current value)
add(v >= 0) add(+/- v) record(v) record(v)
1. Counter (Synchronous Monotonic Sum)
A Counter is a synchronous instrument that records values that can only increase over time (or be reset to zero when the hosting process restarts). Mathematically, a Counter represents a monotonically increasing cumulative sum.
Operational Rules & Method Signature
- Method:
add(value, attributes) - Mathematical Constraint: The
valueparameter MUST be non-negative (). Passing a negative value is a violation of the OpenTelemetry specification. - SDK Error Handling: If an application passes a negative number to
Counter.add(), compliant OpenTelemetry SDKs do not crash the application; instead, they reject or drop the negative increment, retain the current cumulative sum, and optionally record an internal diagnostic warning. - Data Types: Supported in both integer (
LongCounter) and floating-point (DoubleCounter) variants.
Standard Use Cases
- Completed business transactions: a custom
shop.orders.completedcounter with unit{order} - Retries or failures by type: a custom
payment.retriescounter with anerror.typeattribute - Bytes uploaded by the application: a custom
upload.transferredcounter with unitBy - Cache lookups split by outcome: a custom
cache.lookupscounter with acache.resultattribute (hitormiss)
Note that the semantic conventions do not define a separate HTTP request counter: request counts come from the count field of the http.server.request.duration histogram.
2. UpDownCounter (Synchronous Non-Monotonic Sum)
An UpDownCounter is a synchronous instrument that records values that can both increase and decrease over time. It represents a non-monotonic cumulative sum where discrete positive and negative delta changes occur inline.
Operational Rules & Method Signature
- Method:
add(value, attributes) - Mathematical Constraint: The
valueparameter accepts both positive and negative values (). A positive value increments the running total; a negative value decrements the running total. - Data Types: Supported in both integer (
LongUpDownCounter) and floating-point (DoubleUpDownCounter) variants.
Critical Distinction: UpDownCounter vs. Gauge
A common point of confusion is distinguishing between an UpDownCounter and a Gauge:
- An UpDownCounter records relative changes (: ) to an additive total as events happen in application code.
- A Gauge records an absolute, non-additive value (such as a temperature or a fan speed). OpenTelemetry offers it in two forms: a synchronous
Gaugewhoserecord()is called when a change event fires, and anAsynchronous Gaugewhose callback reads the current value at collection time.
Standard Use Cases
- Currently active HTTP requests on a server, the semantic-convention metric
http.server.active_requests(increment by 1 when the handler begins, decrement by 1 when it returns). - Items currently waiting in an in-memory application work queue (increment by 1 on enqueue, decrement by 1 on dequeue).
- Connections in a client connection pool:
db.client.connection.count, split bydb.client.connection.state(usedoridle). - Active WebSocket sessions or connected client sockets.
3. Histogram (Statistical Distribution)
A Histogram is a synchronous instrument that samples the statistical distribution of numeric values over time. Rather than merely summing values or tracking an instantaneous count, a Histogram captures individual measurements and allocates them into configured intervals (buckets) or dynamic sketch structures.
Operational Rules & Method Signature
- Method:
record(value, attributes) - Mathematical Behavior: Each invocation of
record()captures a discrete sample. The SDK automatically computes and tracks:- The count of recorded samples ().
- The arithmetic sum of all recorded samples ().
- The statistical minimum () and maximum () values observed within the aggregation window.
- The distribution of samples across pre-defined explicit bucket boundaries or dynamic exponential bucket scales.
- Data Types: Supported in integer and floating-point representations (e.g., recording fractional seconds such as
0.0125seconds).
Why Histograms Are Essential for SLOs
A standard Counter or UpDownCounter cannot provide statistical percentiles. If a service processes 1,000 requests where 990 requests complete in 5 milliseconds and 10 requests hang for 30 seconds, an average latency metric calculates a misleading figure of ~305 milliseconds. By recording individual durations into a Histogram, the observability backend can accurately calculate the median (p50: 5ms), the 95th percentile (p95: 5ms), and the critical tail latency (p99: 30s), enabling precise Service Level Objective (SLO) error budget alerting.
Standard Use Cases
- Inbound HTTP request duration:
http.server.request.duration(measured in secondss) - Outbound database query duration:
db.client.operation.duration - Request and response payload body sizes:
http.server.request.body.size(measured in bytesBy) - Inbound RPC call duration:
rpc.server.call.duration(seconds)
4. Gauge (Synchronous Non-Additive Value)
A synchronous Gauge records the current value of something that cannot meaningfully be summed across instances, at the moment the application learns about a change.
- Method:
record(value, attributes) - Typical trigger: a change-event subscription, for example
fanSpeedSensor.onChange(v -> fanGauge.record(v)). - Default aggregation: Last Value.
- When to prefer the asynchronous form instead: if the value is exposed through a getter you can read at any time, register an Asynchronous Gauge callback rather than recording on every change.
OpenTelemetry Metric Naming & Unit Conventions
To ensure consistency across polyglot microservices, the API defines a name syntax and the semantic conventions add naming guidance:
Metric Naming Rules
- API syntax: An instrument name starts with a letter and continues with letters, digits,
_,.,-, or/, up to 255 characters, and is case-insensitive. The semantic conventions add that names SHOULD be lowercase. - Hierarchical Dot Notation: Names follow a dot-separated namespace hierarchy, for example:
http.server.request.durationjvm.memory.useddb.client.connection.count
- Pluralization: Names are not pluralized unless the metric counts discrete things measured in a non-unit such as
{operation}(system.disk.operationsis plural,http.server.request.durationis not). UpDownCounters are never pluralized (system.process.count, notsystem.processes) and should not end in_total. - No Unit in the Metric Name: A core naming rule is that units are not included in the name when the unit is carried in instrument metadata.
- Incorrect:
http.server.request.duration_secondsorjvm.memory.used_bytes - Correct: Metric name
http.server.request.durationwith the instrument's unit field set to"s"; metric namejvm.memory.usedwith the unit field set to"By".
- Incorrect:
Standard UCUM Unit Annotations
OpenTelemetry standardizes on the Unified Code for Units of Measure (UCUM) standard for instrument unit fields:
| Measured Dimension | Canonical Unit String | Description & Examples |
|---|---|---|
| Time (Seconds) | s | Seconds. Canonical default in modern OpenTelemetry semantic conventions. |
| Time (Milliseconds) | ms | Milliseconds (used in legacy or high-precision local systems). |
| Data Volume | By | Bytes (capital B, lowercase y). Megabytes: MBy, Kilobytes: kBy. |
| Dimensionless Count | {entity} | Arbitrary entity counts enclosed in curly brackets: {request}, {message}, {error}, {connection}. |
| Ratio / Fraction | 1 | Dimensionless ratio representing fractions (e.g., 0.0 to 1.0) or percentage. |
Decision Framework: Selecting Synchronous Instruments
When instrumenting application code, engineers can walk through these questions to select the correct instrument:
| Architectural Question | Decision Path | Selected Instrument |
|---|---|---|
| Does the metric track latency, payload size, or a statistical distribution where percentiles (p50, p99) are required? | YES | Histogram (record()) |
| Is the measurement a cumulative sum where the value only ever increases and cannot decrease? | YES | Counter (add(val >= 0)) |
| Does the measurement represent a dynamic total that can increase and decrease inline as events occur? | YES | UpDownCounter (add(+/- val)) |
| Is it a non-additive current value (temperature, fan speed) reported when a change event fires? | YES | Gauge (record()) |
| Is the metric read periodically by polling a value or an external system rather than inline during a request? | YES | Not Synchronous! Use an Asynchronous (Observable) instrument |
Practical Code Implementation
The following example demonstrates how to obtain a Meter from an initialized OpenTelemetry MeterProvider and instantiate three of the synchronous instruments in application code:
import time
from opentelemetry import metrics
# 1. Obtain a named Meter from the global MeterProvider
meter = metrics.get_meter(
name="com.example.checkout.service",
version="1.4.0"
)
# 2. Instantiate a Synchronous Monotonic Counter
# Measures cumulative completed orders (strictly non-negative)
orders_counter = meter.create_counter(
name="order.completed.count",
unit="{order}",
description="Cumulative total of successfully completed orders"
)
# 3. Instantiate a Synchronous UpDownCounter
# Measures active in-flight checkout transactions (increments and decrements)
active_checkouts_gauge = meter.create_up_down_counter(
name="checkout.active.requests",
unit="{request}",
description="Number of concurrent checkout requests currently in flight"
)
# 4. Instantiate a Synchronous Histogram
# Measures the distribution of payment processing latencies in seconds
payment_duration_histogram = meter.create_histogram(
name="payment.processing.duration",
unit="s",
description="Statistical distribution of payment gateway processing time"
)
def process_checkout(order_id: str, payment_method: str):
common_attrs = {
"payment.method": payment_method,
"service.environment": "production"
}
# Increment in-flight active checkout count inline (+1)
active_checkouts_gauge.add(1, common_attrs)
start_time = time.monotonic()
try:
# Execute transactional checkout logic
execute_payment(order_id)
# Record successful order completion (+1)
orders_counter.add(1, common_attrs)
finally:
duration_seconds = time.monotonic() - start_time
# Record payment duration distribution into the Histogram
payment_duration_histogram.record(duration_seconds, common_attrs)
# Decrement in-flight active checkout count inline (-1)
active_checkouts_gauge.add(-1, common_attrs)
In this workflow, every measurement occurs inline. If an active distributed trace span exists in context during process_checkout, the SDK attaches the active TraceId to the recorded duration as an exemplar, enabling seamless trace-metric correlation.
An engineer needs to track the number of active, currently running database queries in an asynchronous connection pool. As queries begin, the metric should increment, and as queries complete or timeout, the metric should decrement inline. Which OpenTelemetry instrument must be selected?
Counter
Histogram
UpDownCounter
ObservableGauge
A site reliability engineer observes that an application developer created a metric using meter.create_counter("http.server.request.duration") and called counter.add(latency_ms) at the end of each HTTP request. What is the fundamental flaw in this metric instrument selection?
Counters only record integer values, preventing floating-point latency measurements
The OpenTelemetry API rejects calls to create_counter when the metric name ends in duration
Counters automatically reset to zero after each HTTP request, wiping out historical duration records
Counters aggregate data into a single monotonic sum rather than capturing the statistical distribution, percentiles, and variance of request durations
A microservice developer attempts to decrement an OpenTelemetry Counter by calling request_counter.add(-5) during a batch retry rollback. According to the OpenTelemetry specification, how should a compliant SDK handle this negative value passed to a synchronous Counter?
The SDK rejects or ignores negative values passed to a Counter because Counter measurements must be strictly non-negative ()
The SDK converts the Counter into an UpDownCounter automatically at runtime
The SDK throws an unhandled fatal runtime exception that halts the calling process
The SDK accepts the negative value and subtracts 5 from the cumulative total
Sections you finish are checked off in the contents.