12.1 Agent vs Gateway Deployment Topologies
Key Takeaways
OpenTelemetry Collector architectures span three primary deployment patterns: the local Agent pattern (DaemonSet or sidecar), the centralized Gateway pattern (standalone Deployment), and the production-standard Hybrid Two-Tier architecture.
The Agent pattern runs on the local compute host (one DaemonSet pod per node, sidecar per pod, or local systemd process) to receive telemetry over a one-hop local connection (the node's host IP, or localhost for sidecars), scrape container logs from /var/log/pods, capture host metrics via /proc, and enrich data using the k8sattributes processor.
The Gateway pattern centralizes telemetry egress behind an internal load balancer, creating a hardened security perimeter that consolidates external backend credentials, enforces enterprise data governance (OTTL), and runs stateful pipelines like tail-based sampling.
The Hybrid Two-Tier architecture combines Tier 1 Node Agents for immediate local offloading and pod metadata enrichment with Tier 2 Central Gateways for sampling, PII redaction, secret isolation, and multi-backend routing.
A sidecar deployment provides absolute tenant isolation and works in serverless environments like AWS Fargate, but multiplies memory consumption and operational management overhead across every single application pod.
12.1 Agent vs Gateway Deployment Topologies
Quick Answer: The OpenTelemetry Collector supports three primary deployment patterns: the Agent pattern (deployed as a Kubernetes DaemonSet, pod sidecar, or host systemd daemon for low-latency local offload, host metrics collection, and local container log scraping), the Gateway pattern (a centralized, horizontally scalable Kubernetes Deployment behind a load balancer that isolates backend credentials, enforces egress security, and runs compute-heavy OTTL transformations and tail-based sampling), and the Hybrid Two-Tier pattern (the production standard, pairing Tier 1 Node Agents for local enrichment with Tier 2 Central Gateways for sampling, PII redaction, and multi-backend routing).
In enterprise software architectures, telemetry collection is an operational side-channel: it must capture exhaustive runtime data across traces, metrics, and logs without degrading application performance, consuming excessive compute, or exposing sensitive credentials. While language-level OpenTelemetry SDKs can theoretically transmit telemetry directly from microservices to remote software-as-a-service (SaaS) or on-premises backends, this direct-export approach couples application runtimes to specific backend endpoints, requires distributing secret API tokens to hundreds of application containers, and burns application CPU on compression, retries, and network socket management.
Deploying an OpenTelemetry Collector decouples instrumentation from telemetry transport. However, designing where and how the Collector runs is one of the most critical architectural decisions in any observability rollout. This section analyzes the spectrum of Collector deployment patterns, comparing their operational characteristics, security perimeters, and resource profiles.
The Spectrum of Collector Deployment Patterns
Collector deployment models exist along a continuum between extreme localization (running right beside the application code) and extreme centralization (running as a shared platform infrastructure service):
+-------------------------------------------------------------------------+
| Collector Deployment Architecture Spectrum |
+-------------------------------------------------------------------------+
Local Agent (Sidecar) --> Local Agent (DaemonSet) --> Central Gateway
[Per-Pod Isolation] [Per-Node Aggregation] [Cluster-Wide Scale]
- Zero cross-app impact - Hostmetrics & logs - Central secrets
- High memory footprint - Low overhead per node - Tail-based sampling
- Simple serverless fit - Local k8sattributes - Single egress point
In modern cloud-native systems, these patterns resolve into three primary production models:
- The Agent Pattern: Deployed locally on the compute node or pod alongside the workload.
- The Gateway Pattern: Deployed as a centralized, scalable service tier shared across workloads.
- The Hybrid Two-Tier Pattern: A layered combination where local agents offload telemetry to centralized gateways.
The Agent Pattern (Node DaemonSet & Pod Sidecar)
In the Agent pattern, the Collector process runs in immediate physical or virtual proximity to the instrumented application. In Kubernetes, this is implemented either as a DaemonSet (one Collector pod per Kubernetes worker node) or as a Sidecar container (one Collector container inside each application pod sharing the localhost network namespace). On virtual machines or bare-metal servers, it runs as a local systemd daemon.
+-------------------------------------------------------------------+
| Kubernetes Worker Node |
| |
| +----------------------+ +----------------------+ |
| | App Pod 1 (OTel) | | App Pod 2 (OTel) | |
| +----------------------+ +----------------------+ |
| | | |
| | OTLP (HostIP:4318) | OTLP (HostIP:4317) |
| v v |
| +-------------------------------------------------------------+ |
| | Tier 1 Node Agent (Collector DaemonSet Pod) | |
| | | |
| | - hostmetricsreceiver <-- mounts /proc and /sys | |
| | - filelogreceiver <-- tails /var/log/pods | |
| | - k8sattributes <-- resolves client IP to pod meta | |
| +-------------------------------------------------------------+ |
+-------------------------------------------------------------------+
|
v (OTLP over cluster network)
[Tier 2 Gateway or Backends]
Operational Responsibilities of the Agent
- Immediate Telemetry Offload: A sidecar is reached at
localhost:4317(gRPC) orlocalhost:4318(HTTP). A DaemonSet agent is reached at the node's IP, which the pod learns through the Kubernetes Downward API (status.hostIP), with the agent exposing ahostPort. Either way the hop is short and local, so SDK export queues (BatchSpanProcessor) stay small. - Host Metrics Collection: Because the DaemonSet agent runs with node-level access, it mounts host kernel filesystems (
/procand/sys) to run thehostmetricsreceiver. It collects CPU utilization, memory pressure, disk I/O, network interface counters, and filesystem capacity without needing remote polling credentials. - Local Container Log Scraping: The DaemonSet agent mounts
/var/log/podsfrom the host. Using thefilelogreceiver, it tails, parses, and checkpoints stdout and stderr log streams generated by all container runtimes on that node. - Local Resource Enrichment: Using the
k8sattributesprocessor, the agent inspects the source IP address of incoming TCP connections, maps that IP to the local node's pod registry, and automatically enriches every incoming span, metric, and log record withk8s.pod.name,k8s.namespace.name,k8s.node.name, andk8s.pod.uid. When the processor is configured with a node filter (filter: node_from_env_var: KUBE_NODE_NAME), each agent watches only the pods on its own node, which keeps load on the Kubernetes API server low.
Strengths of the Agent Pattern
- Negligible Application Overhead: Applications do not spend CPU cycles on compression algorithms (GZIP, ZSTD), TLS session negotiation, or lengthy retry loops. The application delegates all network I/O to the local agent.
- Unmatched Infrastructure Context: Running directly on the host node enables unified correlation between application telemetry, container stdout logs, and hardware performance metrics (
hostmetrics). - Local Buffer Resilience: If the wider cluster network or external internet experiences a transient interruption, the local agent continues buffering telemetry locally on the node without stalling application execution.
Limitations of the Agent Pattern
- Coupled Resource Scaling: In a DaemonSet deployment, the Collector's memory and CPU allocations are tied to worker node counts rather than telemetry data volume. High-traffic nodes running telemetry-heavy services may experience memory pressure, while low-traffic nodes waste allocated memory limits.
- Sidecar Resource Bloat: If deployed as a sidecar across 2,000 application pods, the cluster runs 2,000 independent Collector processes, each with its own baseline memory and CPU, which adds up to a large amount of duplicated overhead.
- Broad Attack Surface for Secrets: If the Agent must export directly to an external vendor, secret API tokens and TLS client certificates must be distributed to every node or sidecar across the cluster, violating security best practices.
- Inability to Perform Cluster-Wide State Operations: Local agents have visibility only into telemetry emitted by workloads running on their specific node. They cannot perform cluster-wide tail-based sampling or cross-service metric aggregations because spans of a distributed trace are scattered across multiple nodes.
The Gateway Pattern (Centralized Standalone Tier)
In the Gateway pattern, the OpenTelemetry Collector is deployed as a centralized, standalone service tier. In Kubernetes, this is managed as a standard horizontal Deployment fronted by a Kubernetes Service (ClusterIP), an internal Layer 4 Network Load Balancer (NLB), or an Envoy/Ingress gateway.
[App Pods / Microservices]
[Distributed across Nodes]
│
│ OTLP (gRPC / HTTP)
▼
+───────────────────────────────────────────────────────────+
| Internal Load Balancer / VIP |
+───────────────────────────────────────────────────────────+
│
├──────────────────────┬──────────────────────┐
▼ ▼ ▼
+──────────────────────+ +──────────────────────+ +──────────────────────+
| Gateway Replica 1 | | Gateway Replica 2 | | Gateway Replica 3 |
| (Deployment Pod) | | (Deployment Pod) | | (Deployment Pod) |
| | | | | |
| - Central Secrets | | - Central Secrets | | - Central Secrets |
| - OTTL PII Redaction| | - OTTL PII Redaction| | - OTTL PII Redaction|
| - Tail Sampling | | - Tail Sampling | | - Tail Sampling |
| - Heavy Batching | | - Heavy Batching | | - Heavy Batching |
+──────────────────────+ +──────────────────────+ +──────────────────────+
│ │ │
└──────────────────────┼──────────────────────┘
│ Single Hardened Egress
▼
+─────────────────────────────────+
| Observability Backends / SaaS |
| (Prometheus, Jaeger, Cloud APM) |
+─────────────────────────────────+
Operational Responsibilities of the Gateway
- Centralized Credential Management: Sensitive observability backend credentials—such as Datadog API keys, Dynatrace tokens, Google Cloud service account JSON keys, and AWS IAM roles—are mounted exclusively into the Gateway deployment. Application developers and node-level DaemonSets never have access to backend secrets.
- Egress Traffic Consolidation: The Gateway tier acts as a single, hardened network exit point. Security teams can restrict outbound firewall and NAT gateway rules so that only the Gateway tier's CIDR blocks or service accounts are permitted to send outbound traffic to external internet endpoints over port 443.
- Data Governance & PII Scrubbing: The Gateway executes centralized OpenTelemetry Transformation Language (OTTL) rules. It strips Authorization headers, masks credit card patterns, redacts email addresses, drops debug spans from production namespaces, and normalizes semantic conventions across all teams.
- Stateful Pipeline Execution: Because the Gateway runs as a consolidated cluster, it can host stateful processors such as the
tail_samplingprocessor (evaluating full trace error states), thespanmetricsconnector (calculating golden signals from spans), and metric deduplication. - Batching and Egress Optimization: Gateways consolidate telemetry from hundreds of upstream services, grouping data into large, highly compressed batches (ZSTD or GZIP). This minimizes outbound HTTP/gRPC request overhead, optimizes network MTU utilization, and significantly reduces SaaS ingestion bills.
Strengths of the Gateway Pattern
- Strict Security Perimeter: Zero credential leakage into developer namespaces or node agents; centralized audit logging and strict egress filtering.
- Independent Elastic Scaling: The Gateway deployment scales horizontally via the Horizontal Pod Autoscaler (HPA) based strictly on telemetry ingress throughput (CPU, memory, refused spans), completely decoupled from the number of Kubernetes worker nodes.
- Central Policy Control: Observability platform teams can modify retention policies, change vendor destinations, add secondary cold-storage S3 exports, or adjust sampling ratios in a single Gateway configuration without restarting application pods.
Limitations of the Gateway Pattern
- Additional Network Hop: Telemetry travels from application pods over the internal cluster network to reach the Gateway pods, introducing a network hop and consuming inter-pod bandwidth.
- Application Buffering Risk: If the central Gateway cluster suffers an outage or network partition, application SDKs sending directly to the Gateway must buffer telemetry in memory. If SDK queues fill up, applications will drop telemetry locally.
- No Direct Host Visibility: A central Gateway cannot inspect
/procor/syson remote application worker nodes and cannot tail local container log files from/var/log/pods.
The Hybrid Two-Tier Pattern (Production Best Practice)
In large-scale production Kubernetes and multi-cloud environments, the Hybrid Two-Tier Pattern is a common choice; the OpenTelemetry documentation describes combining agents with gateways in exactly this way. It combines the low-latency, node-level visibility of the Agent pattern with the centralized security, governance, and stateful capabilities of the Gateway pattern.
============================== TIER 1: NODE AGENT ==============================
Workload Pods (App SDKs) ----> HostIP:4317 -------> Collector DaemonSet
- hostmetrics (/proc)
- filelog (/var/log/pods)
- k8sattributes (pod IP)
- minimal memory buffer
│
============================== INTER-TIER ROUTING ==============================
▼
Internal Load Balancing Exporter
(Consistent Hashing on TraceId)
│
============================== TIER 2: CENTRAL GATEWAY =========================
▼
Collector Gateway Deployment
- Tail-based sampling
- OTTL PII redaction
- Secret credentials & mTLS
- Multi-destination routing
│
============================== EGRESS DESTINATIONS ============================
▼
[Prometheus] [Tempo/Jaeger] [Elastic/S3]
Division of Labor in a Two-Tier Architecture
| Functional Responsibility | Tier 1: Node Agent (DaemonSet) | Tier 2: Central Gateway (Deployment) |
|---|---|---|
| Primary Ingress Source | Application SDKs over the node's host IP (or localhost for sidecars) | Tier 1 Agents over OTLP gRPC |
| Hardware & OS Metrics | Scrapes /proc via hostmetricsreceiver | None (does not scrape host kernels) |
| Container Log Scraping | Tails /var/log/pods via filelogreceiver | Ingests parsed logs from Tier 1 |
| Kubernetes Metadata | Enriches via k8sattributes using local pod IP | Normalizes and validates attributes |
| Data Redaction & PII | Basic sanitization (optional) | Enterprise OTTL regex masking and scrubbing |
| Credential Management | Zero external secrets; authenticates only to Tier 2 | Holds SaaS API tokens, OAuth, and mTLS certificates |
| Sampling Strategy | Head-based sampling or forwards 100% | Stateful tail-based sampling (tail_sampling) |
| Scaling Trigger | Scales with node count (1 per node) | Scales with telemetry ingestion volume (HPA) |
| Outbound Egress | Internal cluster network to Tier 2 | External network to cloud/on-prem backends |
Why Two-Tier is the Gold Standard
- Minimal Client Buffering: Applications offload to the Tier 1 node agent over a short node-local hop, keeping SDK queues small.
- Perfect Infrastructure Enrichment: The Tier 1 agent injects exact node, pod, and container metadata at the point of origin, where network mapping is fast and unambiguous.
- Rock-Solid Security: All sensitive vendor API keys reside solely within the Tier 2 Gateway deployment. Compromising an application pod or worker node does not reveal production observability secrets.
- Massive Cost Optimization: The Tier 2 Gateway performs global tail-based sampling (e.g., dropping 90% of boring HTTP 200 traces while preserving 100% of HTTP 500 errors) and aggregates spans into metrics before sending data to expensive SaaS vendors.
Comparison Matrix: Collector Deployment Topologies
| Architectural Attribute | Agent: DaemonSet | Agent: Sidecar | Gateway: Deployment | Hybrid: Two-Tier |
|---|---|---|---|---|
| Kubernetes Resource | DaemonSet | Container in Pod spec | Deployment + Service | DaemonSet + Deployment |
| Deployment Density | 1 per worker node | 1 per application pod | replicas per cluster | 1 per node + gateways |
| Telemetry Ingress Path | HostIP (or localhost with hostNetwork) | localhost:4317 | Cluster IP / Internal LB | HostIP then internal LB |
| Host Metrics & Log Tail | Supported natively | Not supported | Not supported | Supported natively (Tier 1) |
| k8s Metadata Enrichment | Highly efficient | Limited / Inefficient | High Kube-API load | Highly efficient (Tier 1) |
| Credential Isolation | Poor (secrets on all nodes) | Disastrous (in app pod) | Exceptional (centralized) | Exceptional (Tier 2 only) |
| Tail-Based Sampling | Not supported | Not supported | Supported (with routing) | Supported natively (Tier 2) |
| Resource Overhead | Low to moderate | Extreme (multiplied) | Highly elastic | Balanced and cost-effective |
| Failure Domain | Node-level impact | Pod-level impact | Gateway tier impact | Resilient multi-tier isolation |
| Operational Complexity | Low | Moderate to high | Moderate | Moderate to high |
Summary of Architectural Guidelines
- Use the Agent (DaemonSet) Pattern when you need unified host metrics, container log tailing, and local pod attribute enrichment in a self-hosted Kubernetes cluster with homogeneous workloads.
- Use the Agent (Sidecar) Pattern when deploying to serverless container platforms (such as AWS Fargate, Google Cloud Run, or Azure Container Apps) where node-level access is prohibited, or where strict multi-tenant network isolation requires zero shared processes.
- Use the Gateway (Deployment) Pattern when your primary goal is centralizing egress network traffic, securely managing vendor API tokens, enforcing enterprise data governance with OTTL, and running stateful tail-based sampling.
- Use the Hybrid Two-Tier Pattern for enterprise production environments to combine low-latency local offloading and metadata enrichment with centralized egress security, elastic scaling, and intelligent data reduction.
An infrastructure engineering team is designing a telemetry collection architecture for a multi-tenant Kubernetes cluster with 400 worker nodes running thousands of ephemeral microservice pods. The team needs to collect node-level hardware metrics from /proc and /sys, tail container stdout/stderr log files directly from /var/log/pods, enrich all distributed traces with Kubernetes pod and namespace metadata, and ensure application pods offload telemetry over a node-local connection without buffering telemetry across external networks. Which OpenTelemetry Collector deployment pattern is purpose-built to satisfy these requirements?
A centralized Gateway deployment fronted by an external network load balancer
A sidecar Collector container deployed inside every single application pod sharing local storage
A Kubernetes DaemonSet running one Collector agent per node with local volume mounts and the k8sattributes processor
A serverless function triggered asynchronously by Kubernetes API audit event streams
An enterprise security compliance policy mandates that SaaS observability API keys and private mTLS certificates must never be mounted into application pod namespaces or distributed across hundreds of worker nodes. Additionally, the security operations center requires all telemetry leaving the private cloud network to traverse a single, audited egress proxy that scrubs sensitive customer PII using OpenTelemetry Transformation Language (OTTL) rules. Which Collector deployment pattern should the platform team deploy to enforce these security controls?
Sidecar Collector containers running in application pods with secrets injected via environment variables
A Kubernetes DaemonSet where every node agent holds the SaaS API tokens in local memory
Direct telemetry export from application SDKs using hardcoded API credentials
A centralized Gateway deployment running as a dedicated Kubernetes Deployment with secrets mounted exclusively in its secure namespace
An enterprise is modernizing its observability pipeline across a hybrid cloud infrastructure consisting of thousands of microservices. The architecture team needs to minimize memory and CPU overhead within application runtimes, ensure automatic Kubernetes metadata enrichment, implement cluster-wide tail-based sampling, and strictly isolate external vendor credentials. Which deployment pattern represents the production industry standard to achieve all of these operational goals?
The Hybrid Two-Tier architecture combining Tier 1 node-level DaemonSet agents for local offload and enrichment with Tier 2 centralized Gateway deployments for sampling, transformation, and credential management
The standalone Sidecar pattern with individual Collector containers embedded in every application pod performing independent tail-based sampling
The direct-to-backend SDK export model with client-side batching and authentication tokens compiled into application binaries
A single monolithic Collector instance running on a dedicated bare-metal server ingesting all cluster telemetry over public internet endpoints
Sections you finish are checked off in the contents.