10.1 Standard Receivers & Ingestion Protocols
Key Takeaways
Receivers serve as the ingress gateway of the OpenTelemetry Collector, responsible for accepting or scraping external telemetry data and translating it into the unified internal pdata representation.
The canonical otlp receiver natively supports both gRPC (default port 4317) and HTTP (default port 4318), featuring configuration controls for TLS, maximum message sizes, connection limits, and CORS origins.
The prometheus receiver embeds the Prometheus scrape engine to poll target /metrics endpoints, translating Prometheus text and OpenMetrics formats into OpenTelemetry cumulative metrics.
The hostmetrics receiver directly captures host-level hardware and operating system telemetry across CPU, disk, memory, filesystem, and network scrapers without requiring external exporters like node_exporter.
Push ingestion models (OTLP, Zipkin, Jaeger) require open ingress ports and client-side retry handling, whereas pull models (Prometheus scraping) require target discovery and network reachability from the Collector to the targets.
10.1 Standard Receivers & Ingestion Protocols
Quick Answer: Receivers are the ingress components of the OpenTelemetry Collector that accept or scrape telemetry from external systems and convert it into the Collector's internal, zero-copy
pdata(Pluggable Data) representation. The canonical receiver isotlp, which natively supports both gRPC (port4317) and HTTP (port4318) with TLS, CORS, and message-size tuning. Other foundational receivers includeprometheus(pull-based scraping of Prometheus exposition endpoints),hostmetrics(direct operating system hardware counters),filelog(disk log tailing with multiline and regex/JSON parsing), and legacy protocol bridges likejaegerandzipkin. Push models place ingress ports on the Collector and rely on client-side backpressure, whereas pull models require target discovery and scheduled scraping loops.
In distributed observability architectures, telemetry arrives from heterogeneous sources: microservices emitting OpenTelemetry Protocol (OTLP) payloads over gRPC, legacy services instrumented with Jaeger or Zipkin agents, third-party infrastructure exporting Prometheus /metrics endpoints, and system daemons writing unstructured logs to local files. The OpenTelemetry Collector abstracts this complexity through Receivers.
Receivers sit at the entry boundary of the Collector. Regardless of whether data is pushed over a network socket or pulled from an HTTP endpoint, the receiver's sole architectural responsibility is to decode the external wire protocol, validate the payload structure, and convert the telemetry into the Collector's unified internal representation: pdata (ptrace.Traces, pmetric.Metrics, and plog.Logs). Once translated into pdata, downstream processors and exporters operate uniformly without needing any knowledge of the originating network format.
Core Receivers
Know the standard receivers, their default network ports, key configuration parameters, and operational use cases.
1. The otlp Receiver (The Canonical Standard)
The otlp receiver is the primary, vendor-neutral ingress component of the OpenTelemetry Collector. It provides native support for the OpenTelemetry Protocol over two transport mechanisms: gRPC and HTTP.
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
max_recv_msg_size_mib: 16
max_concurrent_streams: 1024
tls:
cert_file: /etc/otel/certs/server.crt
key_file: /etc/otel/certs/server.key
client_ca_file: /etc/otel/certs/ca.crt
client_auth: require_and_verify
http:
endpoint: 0.0.0.0:4318
cors:
allowed_origins:
- "https://app.example.com"
- "https://*.example.com"
allowed_headers:
- "*"
max_age: 7200
Key OTLP Protocol Specifications
-
gRPC Ingress (
protocols.grpc):- Default Port:
4317 - Transport: HTTP/2 multiplexed streams with binary Protocol Buffers (Protobuf) serialization.
max_recv_msg_size_mib: Sets the maximum allowed incoming gRPC message size in MiB (default:4 MiB). In high-throughput production clusters where microservices send large trace batches or dense log buffers, the default 4 MiB limit can cause gRPCRESOURCE_EXHAUSTEDerrors. Increasing this threshold (e.g., to16or32MiB) prevents dropped batches.tls: Configures Transport Layer Security for encrypted transport and mutual TLS (mTLS) authentication. Whenclient_authis set torequire_and_verify, client microservices must present a trusted X.509 certificate signed by the specifiedclient_ca_file.
- Default Port:
-
HTTP Ingress (
protocols.http):- Default Port:
4318 - Endpoints: Telemetry signals are routed to standardized URL paths:
- Traces:
POST http://<host>:4318/v1/traces - Metrics:
POST http://<host>:4318/v1/metrics - Logs:
POST http://<host>:4318/v1/logs
- Traces:
- Encodings: Supports both binary Protobuf (
Content-Type: application/x-protobuf) and JSON (Content-Type: application/json). - CORS (
cors): Cross-Origin Resource Sharing settings (allowed_origins,allowed_headers,max_age). This is essential when telemetry is emitted directly from client-side single-page web applications (SPAs) or browser SDKs to prevent web browsers from blocking outbound telemetry calls.
- Default Port:
2. The prometheus Receiver (Pull-Based Metric Scraper)
Rather than waiting for applications to push metrics, the prometheus receiver acts as an active scraper. It embeds the official Prometheus scraping engine directly into the Collector, polling external HTTP endpoints that expose metrics in Prometheus text format or OpenMetrics format.
receivers:
prometheus:
config:
scrape_configs:
- job_name: 'kubernetes-pods'
scrape_interval: 15s
scrape_timeout: 10s
metrics_path: /metrics
scheme: http
static_configs:
- targets: ['api-service.prod:8080', 'payment-service.prod:8081']
relabel_configs:
- source_labels: [__address__]
target_label: instance
Prometheus Ingestion Mechanics
- Scrape Configuration (
scrape_configs): Leverages identical syntax to standard Prometheus YAML configurations, includingjob_name,scrape_interval,scrape_timeout,static_configs,kubernetes_sd_configs, and relabeling rules (relabel_configs). - Data Model Conversion: The receiver maps Prometheus metric types into OpenTelemetry
pmetric.Metrics:- Prometheus Counter OpenTelemetry Cumulative Sum (monotonic).
- Prometheus Gauge OpenTelemetry Gauge.
- Prometheus Histogram OpenTelemetry Cumulative Histogram with explicit bucket bounds.
- Prometheus Summary OpenTelemetry Summary.
- Internal Synthetic Metrics: Generates standard Prometheus synthetic time series such as
up(indicating whether the scrape succeeded,1or0) andscrape_duration_seconds.
3. The hostmetrics Receiver (Host & Hardware Telemetry)
When the OpenTelemetry Collector runs as a local agent or daemon on physical bare-metal servers, virtual machines, or Kubernetes nodes, the hostmetrics receiver collects hardware and operating system counters without requiring external daemons (such as Prometheus node_exporter or Windows windows_exporter).
receivers:
hostmetrics:
collection_interval: 30s
scrapers:
cpu:
memory:
disk:
filesystem:
network:
load:
paging:
processes:
Standard Hostmetrics Scrapers
cpu: Tracks CPU utilization states (system.cpu.timeacrossuser,system,idle,iowait,steal).memory: Captures RAM distribution (system.memory.usageacrossused,free,buffered,cached).disk&filesystem: Reports disk I/O operations, read/write throughput, filesystem capacity, and inode utilization.network: Reports interface packet rates, network bandwidth (system.network.io), errors, and drop counters.load: Measures operating system load averages across 1-minute, 5-minute, and 15-minute intervals.processes/process: Reports aggregate process counts and per-process CPU and memory utilization.
4. The filelog Receiver (Disk Log Ingestion & Parsing)
The filelog receiver tails, parses, and ingests log files directly from disk storage. It is the primary component for capturing container stdout/stderr files (e.g., in /var/log/pods/) and legacy flat-file application logs.
receivers:
filelog:
include:
- /var/log/apps/*.log
exclude:
- /var/log/apps/*.gz
start_at: end
storage: file_storage # Checkpoints file read offsets across restarts
multiline:
line_start_pattern: '^[0-9]{4}-[0-9]{2}-[0-9]{2}T'
operators:
- type: regex_parser
regex: '^(?P<time>\S+)\s+(?P<severity>\w+)\s+(?P<message>.*)$'
timestamp:
parse_from: attributes.time
layout: '%Y-%m-%dT%H:%M:%S.%LZ'
severity:
parse_from: attributes.severity
Key Filelog Capabilities
- File Discovery & Tracking: Supports glob expressions (
include,exclude) and tracks read positions. Offsets survive a restart only whenstoragepoints at a storage extension such asfile_storage, which must also be declared underextensionsand listed inservice.extensions; without it, a restarted receiver falls back tostart_atand can skip or re-read lines. - Multiline Assembly (
multiline): Solves the fractured log problem by reassembling multiline exceptions (such as Java stack traces or Python error dumps). Theline_start_patterndefines a regular expression indicating the first line of a new log entry; all subsequent non-matching lines are appended to the active log record. - Internal Operator Pipeline: Executes a sequential chain of operators to parse JSON (
json_parser), extract regex tokens (regex_parser), parse event timestamps (timestamp), and normalize severity levels (severity).
5. Legacy Protocol Receivers (Migration Bridges)
During organizational migration from legacy distributed tracing systems to OpenTelemetry, the Collector provides drop-in protocol receivers that allow existing application agents to transmit data without code changes:
jaegerReceiver: Listens on gRPC port14250(Jaeger gRPC), HTTP port14268(Jaeger Thrift over HTTP), and UDP ports6831/6832(Jaeger Thrift Compact/Binary).zipkinReceiver: Listens on HTTP port9411, serving/api/v1/spansand/api/v2/spansfor Zipkin JSON, Thrift, and Protobuf formats.
Push vs. Pull Ingestion Models
A central design consideration when configuring Collector ingress is the distinction between push-based and pull-based telemetry ingestion. Each architectural pattern presents distinct trade-offs across network discovery, firewall traversal, and backpressure management.
| Operational Dimension | Push Ingestion (e.g., OTLP, Zipkin, Jaeger) | Pull Ingestion (e.g., Prometheus Scraper, Hostmetrics) |
|---|---|---|
| Initiator | Client service or upstream agent initiates the TCP connection. | Collector initiates the HTTP/TCP connection to target endpoints. |
| Network Direction | Inbound into the Collector (Client -> Collector). | Outbound from the Collector (Collector -> Target Endpoint). |
| Firewall Configuration | Open ingress ports on the Collector; clients require outbound egress. | Collector requires network reachability into target application subnets. |
| Service Discovery | Resides on the client side (DNS, L4/L7 load balancer, or service mesh). | Resides in the Collector (Kubernetes API, Consul, DNS SRV, static IPs). |
| Backpressure Handling | Explicit: the Collector returns a retryable error (gRPC UNAVAILABLE / HTTP 503, or RESOURCE_EXHAUSTED / HTTP 429). | Implicit: Scraper paces scrapes or drops slow/timed-out scrape intervals. |
| Load Characteristics | Unpredictable: Traffic spikes at microservices create sudden Collector load bursts. | Predictable: Regular scraping intervals ensure deterministic ingest rates. |
| Ephemerality | Ideal for short-lived batch jobs, serverless functions, and lambdas. | Poor for ephemeral jobs; endpoints may terminate before a scrape cycle runs. |
Architectural Trade-offs in Practice
- Firewall and Security Traversal: Push architectures simplify network security when Collectors are deployed as centralized cluster gateways. Microservices push telemetry outward to a shared endpoint. Conversely, pull architectures require the Collector to establish network connections into secure, internal application subnets, which often requires complex network peering and firewall holes.
- Backpressure Propagation: In push pipelines, when a Collector's memory reaches capacity or downstream backends fail, the Collector pushes back directly onto clients by returning retryable errors; when the
memory_limiterrefuses data, the OTLP receiver answers with gRPCUNAVAILABLEor HTTP503. The client SDK buffers or drops data locally. In pull pipelines, backpressure cannot be exerted on the application; if the Collector is overloaded, it skips scrape cycles, resulting in coarse-grained metric gaps.
Complete Production Receivers YAML Configuration
The following configuration demonstrates a complete, production-grade receivers block combining OTLP, Prometheus, Hostmetrics, and Filelog ingress:
receivers:
# Standard OTLP ingress for traces, metrics, and logs
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
max_recv_msg_size_mib: 16
keepalive:
server_parameters:
max_connection_idle: 11s
max_connection_age: 30s
http:
endpoint: 0.0.0.0:4318
cors:
allowed_origins:
- "https://*.internal.company.com"
allowed_headers:
- "*"
# Pull-based infrastructure metrics scraper
prometheus:
config:
scrape_configs:
- job_name: 'collector-internal'
scrape_interval: 15s
static_configs:
- targets: ['localhost:8888']
# Local host operating system telemetry
hostmetrics:
collection_interval: 15s
scrapers:
cpu:
memory:
disk:
filesystem:
network:
load:
# File log collection with multiline parsing and offset tracking
filelog:
include:
- /var/log/containers/*.log
exclude:
- /var/log/containers/*-debug.log
start_at: end
storage: file_storage
multiline:
line_start_pattern: '^\d{4}-\d{2}-\d{2}T'
operators:
- type: json_parser
id: parse_docker_json
timestamp:
parse_from: attributes.time
layout: '%Y-%m-%dT%H:%M:%S.%LZ'
An enterprise platform engineering team is designing a centralized OpenTelemetry Collector gateway. The gateway must accept distributed trace spans from backend microservices using high-performance gRPC, receive client-side traces from single-page web applications running in user browsers, support incoming gRPC batch payloads of up to 12 MiB, and allow cross-origin requests from the company's web domain. How should the team configure the receivers block in the Collector YAML?
Configure the otlp receiver with both protocols.grpc (specifying endpoint 0.0.0.0:4317 and max_recv_msg_size_mib: 12) and protocols.http (specifying endpoint 0.0.0.0:4318 with cors.allowed_origins configured for the web domain).
Configure two separate receiver components: the otlp receiver listening on port 4317 for gRPC and the jaeger receiver listening on port 4318 for browser HTTP requests.
Deploy the prometheus receiver configured with a custom gRPC hook on port 4317 and enable the CORS extension in the service block.
Deploy an external reverse proxy container in front of the Collector because the otlp receiver cannot host both gRPC and HTTP endpoints simultaneously within a single configuration block.
A system administrator needs to monitor Linux operating system performance (CPU utilization, RAM consumption, disk I/O operations, and network bandwidth) across a fleet of virtual machines running an OpenTelemetry Collector agent. The team lead wants to avoid installing and maintaining standalone third-party monitoring daemons such as the Prometheus node_exporter. Which native Collector component should the administrator deploy to fulfill this requirement?
Configure the filelog receiver to tail and parse /proc/stat, /proc/meminfo, and /proc/net/dev every 15 seconds.
Configure the hostmetrics receiver with scrapers including cpu, memory, disk, filesystem, and network running at a configured collection_interval.
Configure the otlp receiver in pull mode pointing directly to the Linux kernel sysfs socket.
Configure the prometheus receiver with a static scrape job targeting localhost port 9100 without running an exporter process.
A site reliability engineer is configuring log ingestion for a legacy enterprise Java application running on a Kubernetes cluster. The application writes logs to local disk files, producing multi-line stack traces whenever unhandled exceptions occur. In addition, the Collector pod is frequently rescheduled during autoscaling events. Which receiver configuration correctly reconstructs full stack traces into unified log records and guarantees that log reading resumes without re-ingesting previous data after pod restarts?
Use the otlp receiver configured with an in-memory batch processor and a multiline extension.
Use the prometheus receiver configured to scrape the log file directory using standard HTTP get requests.
Use the filelog receiver configured with a multiline operator using line_start_pattern to identify entry boundaries, combined with a persistent storage extension to checkpoint file read offsets.
Use the jaeger receiver configured with Thrift UDP ingestion and a regex_parser operator.
Sections you finish are checked off in the contents.