10.1 Standard Receivers & Ingestion Protocols

Key Takeaways

  • Receivers serve as the ingress gateway of the OpenTelemetry Collector, responsible for accepting or scraping external telemetry data and translating it into the unified internal pdata representation.

  • The canonical otlp receiver natively supports both gRPC (default port 4317) and HTTP (default port 4318), featuring configuration controls for TLS, maximum message sizes, connection limits, and CORS origins.

  • The prometheus receiver embeds the Prometheus scrape engine to poll target /metrics endpoints, translating Prometheus text and OpenMetrics formats into OpenTelemetry cumulative metrics.

  • The hostmetrics receiver directly captures host-level hardware and operating system telemetry across CPU, disk, memory, filesystem, and network scrapers without requiring external exporters like node_exporter.

  • Push ingestion models (OTLP, Zipkin, Jaeger) require open ingress ports and client-side retry handling, whereas pull models (Prometheus scraping) require target discovery and network reachability from the Collector to the targets.

Last updated: September 2026

10.1 Standard Receivers & Ingestion Protocols

Quick Answer: Receivers are the ingress components of the OpenTelemetry Collector that accept or scrape telemetry from external systems and convert it into the Collector's internal, zero-copy pdata (Pluggable Data) representation. The canonical receiver is otlp, which natively supports both gRPC (port 4317) and HTTP (port 4318) with TLS, CORS, and message-size tuning. Other foundational receivers include prometheus (pull-based scraping of Prometheus exposition endpoints), hostmetrics (direct operating system hardware counters), filelog (disk log tailing with multiline and regex/JSON parsing), and legacy protocol bridges like jaeger and zipkin. Push models place ingress ports on the Collector and rely on client-side backpressure, whereas pull models require target discovery and scheduled scraping loops.

In distributed observability architectures, telemetry arrives from heterogeneous sources: microservices emitting OpenTelemetry Protocol (OTLP) payloads over gRPC, legacy services instrumented with Jaeger or Zipkin agents, third-party infrastructure exporting Prometheus /metrics endpoints, and system daemons writing unstructured logs to local files. The OpenTelemetry Collector abstracts this complexity through Receivers.

Receivers sit at the entry boundary of the Collector. Regardless of whether data is pushed over a network socket or pulled from an HTTP endpoint, the receiver's sole architectural responsibility is to decode the external wire protocol, validate the payload structure, and convert the telemetry into the Collector's unified internal representation: pdata (ptrace.Traces, pmetric.Metrics, and plog.Logs). Once translated into pdata, downstream processors and exporters operate uniformly without needing any knowledge of the originating network format.


Core Receivers

Know the standard receivers, their default network ports, key configuration parameters, and operational use cases.

1. The otlp Receiver (The Canonical Standard)

The otlp receiver is the primary, vendor-neutral ingress component of the OpenTelemetry Collector. It provides native support for the OpenTelemetry Protocol over two transport mechanisms: gRPC and HTTP.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
        max_recv_msg_size_mib: 16
        max_concurrent_streams: 1024
        tls:
          cert_file: /etc/otel/certs/server.crt
          key_file: /etc/otel/certs/server.key
          client_ca_file: /etc/otel/certs/ca.crt
          client_auth: require_and_verify
      http:
        endpoint: 0.0.0.0:4318
        cors:
          allowed_origins:
            - "https://app.example.com"
            - "https://*.example.com"
          allowed_headers:
            - "*"
          max_age: 7200

Key OTLP Protocol Specifications

  • gRPC Ingress (protocols.grpc):

    • Default Port: 4317
    • Transport: HTTP/2 multiplexed streams with binary Protocol Buffers (Protobuf) serialization.
    • max_recv_msg_size_mib: Sets the maximum allowed incoming gRPC message size in MiB (default: 4 MiB). In high-throughput production clusters where microservices send large trace batches or dense log buffers, the default 4 MiB limit can cause gRPC RESOURCE_EXHAUSTED errors. Increasing this threshold (e.g., to 16 or 32 MiB) prevents dropped batches.
    • tls: Configures Transport Layer Security for encrypted transport and mutual TLS (mTLS) authentication. When client_auth is set to require_and_verify, client microservices must present a trusted X.509 certificate signed by the specified client_ca_file.
  • HTTP Ingress (protocols.http):

    • Default Port: 4318
    • Endpoints: Telemetry signals are routed to standardized URL paths:
      • Traces: POST http://<host>:4318/v1/traces
      • Metrics: POST http://<host>:4318/v1/metrics
      • Logs: POST http://<host>:4318/v1/logs
    • Encodings: Supports both binary Protobuf (Content-Type: application/x-protobuf) and JSON (Content-Type: application/json).
    • CORS (cors): Cross-Origin Resource Sharing settings (allowed_origins, allowed_headers, max_age). This is essential when telemetry is emitted directly from client-side single-page web applications (SPAs) or browser SDKs to prevent web browsers from blocking outbound telemetry calls.

2. The prometheus Receiver (Pull-Based Metric Scraper)

Rather than waiting for applications to push metrics, the prometheus receiver acts as an active scraper. It embeds the official Prometheus scraping engine directly into the Collector, polling external HTTP endpoints that expose metrics in Prometheus text format or OpenMetrics format.

receivers:
  prometheus:
    config:
      scrape_configs:
        - job_name: 'kubernetes-pods'
          scrape_interval: 15s
          scrape_timeout: 10s
          metrics_path: /metrics
          scheme: http
          static_configs:
            - targets: ['api-service.prod:8080', 'payment-service.prod:8081']
          relabel_configs:
            - source_labels: [__address__]
              target_label: instance

Prometheus Ingestion Mechanics

  • Scrape Configuration (scrape_configs): Leverages identical syntax to standard Prometheus YAML configurations, including job_name, scrape_interval, scrape_timeout, static_configs, kubernetes_sd_configs, and relabeling rules (relabel_configs).
  • Data Model Conversion: The receiver maps Prometheus metric types into OpenTelemetry pmetric.Metrics:
    • Prometheus Counter →\rightarrow OpenTelemetry Cumulative Sum (monotonic).
    • Prometheus Gauge →\rightarrow OpenTelemetry Gauge.
    • Prometheus Histogram →\rightarrow OpenTelemetry Cumulative Histogram with explicit bucket bounds.
    • Prometheus Summary →\rightarrow OpenTelemetry Summary.
  • Internal Synthetic Metrics: Generates standard Prometheus synthetic time series such as up (indicating whether the scrape succeeded, 1 or 0) and scrape_duration_seconds.

3. The hostmetrics Receiver (Host & Hardware Telemetry)

When the OpenTelemetry Collector runs as a local agent or daemon on physical bare-metal servers, virtual machines, or Kubernetes nodes, the hostmetrics receiver collects hardware and operating system counters without requiring external daemons (such as Prometheus node_exporter or Windows windows_exporter).

receivers:
  hostmetrics:
    collection_interval: 30s
    scrapers:
      cpu:
      memory:
      disk:
      filesystem:
      network:
      load:
      paging:
      processes:

Standard Hostmetrics Scrapers

  • cpu: Tracks CPU utilization states (system.cpu.time across user, system, idle, iowait, steal).
  • memory: Captures RAM distribution (system.memory.usage across used, free, buffered, cached).
  • disk & filesystem: Reports disk I/O operations, read/write throughput, filesystem capacity, and inode utilization.
  • network: Reports interface packet rates, network bandwidth (system.network.io), errors, and drop counters.
  • load: Measures operating system load averages across 1-minute, 5-minute, and 15-minute intervals.
  • processes / process: Reports aggregate process counts and per-process CPU and memory utilization.

4. The filelog Receiver (Disk Log Ingestion & Parsing)

The filelog receiver tails, parses, and ingests log files directly from disk storage. It is the primary component for capturing container stdout/stderr files (e.g., in /var/log/pods/) and legacy flat-file application logs.

receivers:
  filelog:
    include:
      - /var/log/apps/*.log
    exclude:
      - /var/log/apps/*.gz
    start_at: end
    storage: file_storage # Checkpoints file read offsets across restarts
    multiline:
      line_start_pattern: '^[0-9]{4}-[0-9]{2}-[0-9]{2}T'
    operators:
      - type: regex_parser
        regex: '^(?P<time>\S+)\s+(?P<severity>\w+)\s+(?P<message>.*)$'
        timestamp:
          parse_from: attributes.time
          layout: '%Y-%m-%dT%H:%M:%S.%LZ'
        severity:
          parse_from: attributes.severity

Key Filelog Capabilities

  • File Discovery & Tracking: Supports glob expressions (include, exclude) and tracks read positions. Offsets survive a restart only when storage points at a storage extension such as file_storage, which must also be declared under extensions and listed in service.extensions; without it, a restarted receiver falls back to start_at and can skip or re-read lines.
  • Multiline Assembly (multiline): Solves the fractured log problem by reassembling multiline exceptions (such as Java stack traces or Python error dumps). The line_start_pattern defines a regular expression indicating the first line of a new log entry; all subsequent non-matching lines are appended to the active log record.
  • Internal Operator Pipeline: Executes a sequential chain of operators to parse JSON (json_parser), extract regex tokens (regex_parser), parse event timestamps (timestamp), and normalize severity levels (severity).

5. Legacy Protocol Receivers (Migration Bridges)

During organizational migration from legacy distributed tracing systems to OpenTelemetry, the Collector provides drop-in protocol receivers that allow existing application agents to transmit data without code changes:

  • jaeger Receiver: Listens on gRPC port 14250 (Jaeger gRPC), HTTP port 14268 (Jaeger Thrift over HTTP), and UDP ports 6831/6832 (Jaeger Thrift Compact/Binary).
  • zipkin Receiver: Listens on HTTP port 9411, serving /api/v1/spans and /api/v2/spans for Zipkin JSON, Thrift, and Protobuf formats.

Push vs. Pull Ingestion Models

A central design consideration when configuring Collector ingress is the distinction between push-based and pull-based telemetry ingestion. Each architectural pattern presents distinct trade-offs across network discovery, firewall traversal, and backpressure management.

Operational DimensionPush Ingestion (e.g., OTLP, Zipkin, Jaeger)Pull Ingestion (e.g., Prometheus Scraper, Hostmetrics)
InitiatorClient service or upstream agent initiates the TCP connection.Collector initiates the HTTP/TCP connection to target endpoints.
Network DirectionInbound into the Collector (Client -> Collector).Outbound from the Collector (Collector -> Target Endpoint).
Firewall ConfigurationOpen ingress ports on the Collector; clients require outbound egress.Collector requires network reachability into target application subnets.
Service DiscoveryResides on the client side (DNS, L4/L7 load balancer, or service mesh).Resides in the Collector (Kubernetes API, Consul, DNS SRV, static IPs).
Backpressure HandlingExplicit: the Collector returns a retryable error (gRPC UNAVAILABLE / HTTP 503, or RESOURCE_EXHAUSTED / HTTP 429).Implicit: Scraper paces scrapes or drops slow/timed-out scrape intervals.
Load CharacteristicsUnpredictable: Traffic spikes at microservices create sudden Collector load bursts.Predictable: Regular scraping intervals ensure deterministic ingest rates.
EphemeralityIdeal for short-lived batch jobs, serverless functions, and lambdas.Poor for ephemeral jobs; endpoints may terminate before a scrape cycle runs.

Architectural Trade-offs in Practice

  1. Firewall and Security Traversal: Push architectures simplify network security when Collectors are deployed as centralized cluster gateways. Microservices push telemetry outward to a shared endpoint. Conversely, pull architectures require the Collector to establish network connections into secure, internal application subnets, which often requires complex network peering and firewall holes.
  2. Backpressure Propagation: In push pipelines, when a Collector's memory reaches capacity or downstream backends fail, the Collector pushes back directly onto clients by returning retryable errors; when the memory_limiter refuses data, the OTLP receiver answers with gRPC UNAVAILABLE or HTTP 503. The client SDK buffers or drops data locally. In pull pipelines, backpressure cannot be exerted on the application; if the Collector is overloaded, it skips scrape cycles, resulting in coarse-grained metric gaps.

Complete Production Receivers YAML Configuration

The following configuration demonstrates a complete, production-grade receivers block combining OTLP, Prometheus, Hostmetrics, and Filelog ingress:

receivers:
  # Standard OTLP ingress for traces, metrics, and logs
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
        max_recv_msg_size_mib: 16
        keepalive:
          server_parameters:
            max_connection_idle: 11s
            max_connection_age: 30s
      http:
        endpoint: 0.0.0.0:4318
        cors:
          allowed_origins:
            - "https://*.internal.company.com"
          allowed_headers:
            - "*"

  # Pull-based infrastructure metrics scraper
  prometheus:
    config:
      scrape_configs:
        - job_name: 'collector-internal'
          scrape_interval: 15s
          static_configs:
            - targets: ['localhost:8888']

  # Local host operating system telemetry
  hostmetrics:
    collection_interval: 15s
    scrapers:
      cpu:
      memory:
      disk:
      filesystem:
      network:
      load:

  # File log collection with multiline parsing and offset tracking
  filelog:
    include:
      - /var/log/containers/*.log
    exclude:
      - /var/log/containers/*-debug.log
    start_at: end
    storage: file_storage
    multiline:
      line_start_pattern: '^\d{4}-\d{2}-\d{2}T'
    operators:
      - type: json_parser
        id: parse_docker_json
        timestamp:
          parse_from: attributes.time
          layout: '%Y-%m-%dT%H:%M:%S.%LZ'
Loading diagram...
Push vs Pull Ingestion Flow into Internal pdata Representation
Test Your Knowledge

An enterprise platform engineering team is designing a centralized OpenTelemetry Collector gateway. The gateway must accept distributed trace spans from backend microservices using high-performance gRPC, receive client-side traces from single-page web applications running in user browsers, support incoming gRPC batch payloads of up to 12 MiB, and allow cross-origin requests from the company's web domain. How should the team configure the receivers block in the Collector YAML?

A

Configure the otlp receiver with both protocols.grpc (specifying endpoint 0.0.0.0:4317 and max_recv_msg_size_mib: 12) and protocols.http (specifying endpoint 0.0.0.0:4318 with cors.allowed_origins configured for the web domain).

B

Configure two separate receiver components: the otlp receiver listening on port 4317 for gRPC and the jaeger receiver listening on port 4318 for browser HTTP requests.

C

Deploy the prometheus receiver configured with a custom gRPC hook on port 4317 and enable the CORS extension in the service block.

D

Deploy an external reverse proxy container in front of the Collector because the otlp receiver cannot host both gRPC and HTTP endpoints simultaneously within a single configuration block.

Test Your Knowledge

A system administrator needs to monitor Linux operating system performance (CPU utilization, RAM consumption, disk I/O operations, and network bandwidth) across a fleet of virtual machines running an OpenTelemetry Collector agent. The team lead wants to avoid installing and maintaining standalone third-party monitoring daemons such as the Prometheus node_exporter. Which native Collector component should the administrator deploy to fulfill this requirement?

A

Configure the filelog receiver to tail and parse /proc/stat, /proc/meminfo, and /proc/net/dev every 15 seconds.

B

Configure the hostmetrics receiver with scrapers including cpu, memory, disk, filesystem, and network running at a configured collection_interval.

C

Configure the otlp receiver in pull mode pointing directly to the Linux kernel sysfs socket.

D

Configure the prometheus receiver with a static scrape job targeting localhost port 9100 without running an exporter process.

Test Your Knowledge

A site reliability engineer is configuring log ingestion for a legacy enterprise Java application running on a Kubernetes cluster. The application writes logs to local disk files, producing multi-line stack traces whenever unhandled exceptions occur. In addition, the Collector pod is frequently rescheduled during autoscaling events. Which receiver configuration correctly reconstructs full stack traces into unified log records and guarantees that log reading resumes without re-ingesting previous data after pod restarts?

A

Use the otlp receiver configured with an in-memory batch processor and a multiline extension.

B

Use the prometheus receiver configured to scrape the log file directory using standard HTTP get requests.

C

Use the filelog receiver configured with a multiline operator using line_start_pattern to identify entry boundaries, combined with a persistent storage extension to checkpoint file read offsets.

D

Use the jaeger receiver configured with Thrift UDP ingestion and a regex_parser operator.

Sections you finish are checked off in the contents.