2.2 Telemetry Ingestion Architecture & Data Pipelines
Key Takeaways
Unified Elastic Agent replaces standalone Beats shippers by consolidating endpoint security, log collection, and system metrics into a single binary centrally managed through Fleet Server.
Fleet Server operates as the control plane for Elastic Agents, managing policy distribution, integration updates, and agent status checks over an outbound HTTPS long-poll connection on port 8220.
Logstash provides an intermediate pipeline layer featuring in-memory or persistent disk queues and dead-letter queues, making it essential for complex multi-source ETL, heavy regex parsing, and protocol buffering.
Elasticsearch Ingest Pipelines execute pre-indexing processors directly on ingest nodes, standardizing raw logs into ECS fields using dissect, grok, set, rename, and geoip processors.
Pipeline error handling strategies, including processor-level and pipeline-level on_failure blocks, prevent event loss during parsing anomalies and preserve raw telemetry in event.original.
Ingestion pipelines form the critical acquisition and transformation backbone of any Security Information and Event Management (SIEM) deployment. If raw telemetry is delayed, dropped due to queue saturation, or improperly mapped, downstream detection rules fail silently. Modern Elastic SIEM deployments employ a multi-tiered ingestion architecture that balances lightweight endpoint shipping, centralized policy orchestration, resilient intermediate buffering, and high-performance pre-indexing transformations.
1. The Telemetry Collection Ecosystem: Beats vs. Elastic Agent
Historically, the Elastic Stack utilized specialized, standalone shippers known as Beats. While legacy environments still run Beats, modern deployments center on the Unified Elastic Agent managed by Fleet.
Standalone Beats Shippers
Beats are lightweight, single-purpose data forwarders written in Go. They install directly on servers and network devices to capture specific telemetry streams:
- Filebeat: High-throughput log forwarding. Employs a harvest registry to track file read offsets and guarantees at-least-once delivery with backpressure handling. Includes pre-packaged modules for common security data sources (Suricata, Zeek, Cisco ASA/FTD, Palo Alto Networks, Okta, AWS CloudTrail) that package pre-configured Ingest Pipelines, ECS field mappings, and Kibana dashboards.
- Winlogbeat: Windows Event Log collection engine. Interacts directly with the native Windows Event Log API to read standard channels (Security, System, Application) and specialized diagnostic channels (Microsoft-Windows-Sysmon/Operational, Microsoft-Windows-PowerShell/Operational, Microsoft-Windows-Windows Defender/Operational). Supports channel filtering, event ID whitelisting, and localized event XML parsing.
- Packetbeat: Passive network packet analyzer. Uses libpcap or Npcap to capture wire traffic on endpoints or network taps without proxying connections. Dissects protocols including DNS, HTTP, TLS (extracting SNI, cipher suites, certificate fingerprints), SMB, and flow statistics, outputting normalized ECS network documents.
- Auditbeat: Host audit framework and integrity monitoring agent for Linux and macOS. Interfaces with the Linux
auditdkernel module to capture process execution, syscalls, and user logins. Incorporates a File Integrity Monitoring (FIM) module that computes SHA-256 hashes and monitors metadata changes on sensitive directories (/etc,/bin,/usr/bin).
Unified Elastic Agent
While Beats are highly efficient, maintaining separate configuration files (filebeat.yml, winlogbeat.yml, auditbeat.yml) and orchestrating binary upgrades across tens of thousands of endpoints creates significant operational overhead. The Elastic Agent unifies this ecosystem into a single binary.
Under the hood, Elastic Agent acts as a managed supervisor daemon that controls sub-processes for log collection, metrics, OSquery inspection, and Elastic Defend (endpoint security and EDR capabilities). Rather than configuring local YAML files, administrators manage Elastic Agent entirely through Kibana using centralized Agent Policies.
2. Fleet Server Architecture and Communication Model
Fleet Server is the centralized control plane that orchestrates Elastic Agents across an enterprise infrastructure. It runs as the Fleet Server integration inside an Elastic Agent, either self-managed or hosted by Elastic Cloud.
+-------------------------------------------------------------+
| Kibana UI |
| (Agent Policy Creation, Integration Config) |
+------------------------------+------------------------------+
|
v (REST API)
+-------------------------------------------------------------+
| Elasticsearch Cluster |
| (Stores .fleet-* Indices & Telemetry) |
+------------------------------+------------------------------+
^
| (Internal API)
+------------------------------v------------------------------+
| Fleet Server |
| (HTTPS Control Plane, Port 8220) |
+------------------------------+------------------------------+
^
| Outbound HTTPS Long-Polling
+--------------------+--------------------+
| |
+---------+---------+ +---------+---------+
| Elastic Agent | | Elastic Agent |
| (Windows Endpoint)| | (Linux Server) |
+-------------------+ +-------------------+
Communication Mechanism
- Outbound-Only Connectivity: Elastic Agents initiate an outbound HTTPS connection to Fleet Server (default port
8220). Endpoints never listen on inbound network ports, eliminating firewall attack surface on managed hosts. - Long-Polling Check-In: Agents maintain an active HTTP long-polling connection with Fleet Server. Fleet Server holds the request open. When a security engineer modifies an Agent Policy in Kibana (such as enabling Sysmon ingestion or adjusting Elastic Defend malware blocking), Fleet Server pushes the configuration delta over the established connection within seconds.
- Status Heartbeats: Agents check in regularly with their health (healthy, degraded or failed components, updating) and integration status, which Fleet uses to show statuses such as Healthy, Unhealthy, Updating and Offline, which records them in the
.fleet-*system indices. - Stateless Scalability: Fleet Server is completely stateless. Multiple Fleet Server instances can be deployed behind a standard enterprise load balancer (e.g., F5, HAProxy, AWS ALB). If a Fleet Server instance crashes, connected agents seamlessly reconnect to alternate instances without interruption.
- Enrollment Tokens and Scoped API Keys: Agents enroll into Fleet using an enrollment token associated with a specific Agent Policy. Upon successful validation, Fleet Server issues unique Elasticsearch API keys to each agent with least-privilege permissions constrained strictly to the data streams required by its assigned integrations.
3. Logstash: Ingestion Buffering and Complex ETL
While Elastic Agent can stream telemetry directly into Elasticsearch, enterprise SOCs often position Logstash as an intermediate aggregation and processing layer between edge shippers and the Elasticsearch cluster.
When to Deploy Logstash in SecOps
- Surge Protection and Buffering: During security incidents (such as automated brute-force attacks or malware propagation), event volumes can spike by 500% to 1000%. If data nodes experience indexing backpressure, directly connected agents may buffer data locally or drop events. Logstash provides Persistent Queues (PQ) (
queue.type: persisted), buffering incoming telemetry on local disk to ensure zero event loss while Elasticsearch recovers. - Multi-Source Enrichment: Logstash pipelines can enrich events by executing external queries against relational databases (
jdbc_streaming), Memcached key-value stores, or threat intelligence platforms during transit. - Central Protocol Aggregation: Logstash can act as a central listener for many inputs (syslog over UDP/TCP, Beats, Kafka, HTTP) and route events to several destinations. Elastic Agent can also receive syslog and NetFlow directly through integrations such as Custom UDP/TCP logs and NetFlow Records. Logstash is therefore a choice for central buffering and complex routing, not a requirement for these protocols.
- Dead-Letter Queues (DLQ): When Elasticsearch rejects an index request due to mapping conflicts (such as an IP field receiving a string that cannot be parsed), Logstash can write the rejected document to an on-disk dead-letter queue (
dead_letter_queue.enable: true), allowing engineers to resolve the mapping conflict and replay the event.
4. Elasticsearch Ingest Pipelines and Key Processors
Ingest Pipelines execute pre-indexing document transformations directly on Elasticsearch nodes assigned the ingest role. Pipelines consist of an ordered series of processors that execute sequentially before a document is written into a Lucene segment.
Critical Ingest Processors for SecOps
| Processor | Operational Mechanism | Primary Security Telemetry Use Case |
|---|---|---|
dissect | High-speed delimiter-based string parsing without regular expressions | Parsing structured firewall logs (CEF, CSV, space-delimited syslog) |
grok | Regular expression pattern matching using Oniguruma regex syntax | Extracting fields from complex, variable-length, unstructured log messages |
set | Assigns static values or copies dynamic values between fields | Setting event.dataset, event.module, or organizational environment tags |
rename | Renames source fields to target fields | Aligning proprietary vendor field names with Elastic Common Schema (ECS) |
drop | Drops documents matching specific conditional criteria | Discarding benign, high-volume noise (e.g., health-check pings, internal scans) |
geoip | Looks up geographic coordinates, country, and ASN from IP addresses | Enriching public source/destination IPs using MaxMind GeoLite2 databases |
date | Parses string timestamps into ISO 8601 format and assigns to @timestamp | Aligning divergent vendor log timestamps with cluster-wide @timestamp |
fingerprint | Generates a cryptographic hash (SHA-256) across specified fields | Creating deterministic document _id values to prevent duplicate indexing |
enrich | Matches fields against internal Elasticsearch indices at ingest time | Appending CMDB asset tags, vulnerability scores, or threat intel indicators |
dissect vs. grok: Performance Trade-offs
A fundamental optimization in security data pipelines is choosing between dissect and grok:
dissect(Delimiter Parsing): Operates by matching fixed string delimiters. It does not compile regular expressions or perform backtracking. As a result,dissectis typically much faster thangrokand consumes less CPU. It should always be used when log formats have a consistent structure (such as Palo Alto Networks CSV logs or standard CEF headers).grok(Regex Parsing): Necessary when log lines have variable structures, nested optional fields, or unstructured text (such as custom application audit logs or Linux daemon messages). Because regex backtracking can cause severe CPU saturation under high ingestion volumes, grok patterns must be carefully bounded (avoiding greedy.*patterns).
{
"description": "Example Security Ingest Pipeline with ECS Normalization",
"processors": [
{
"dissect": {
"field": "message",
"pattern": "%{syslog_timestamp} %{host.name} %{process.name}[%{process.pid}]: %{event_message}",
"ignore_failure": false
}
},
{
"date": {
"field": "syslog_timestamp",
"target_field": "@timestamp",
"formats": ["MMM d HH:mm:ss", "MMM dd HH:mm:ss"],
"timezone": "UTC"
}
},
{
"geoip": {
"field": "source.ip",
"target_field": "source.geo",
"ignore_missing": true
}
}
]
}
5. Pipeline Error Handling and Ingestion Failure Troubleshooting
In production SIEM pipelines, malformed logs, schema mismatches, and unexpected characters will inevitably trigger processor exceptions. If unhandled, the pipeline fails, the document is not indexed, and the client receives an error for that document. The security event is dropped unless the shipper retries or stores it elsewhere.
Handling Failures with on_failure
Elasticsearch Ingest Pipelines support two levels of on_failure error handling:
- Processor-Level
on_failure: Attached directly to an individual processor. If that specific processor errors (e.g., aconvertprocessor fails because a port field contains non-numeric text), the processor-level fallback executes without failing the rest of the pipeline. - Pipeline-Level
on_failure: Defined as a global array of processors at the root of the pipeline. If any processor in the pipeline fails and lacks a processor-level fallback, execution immediately jumps to the pipeline-levelon_failureblock.
Best Practice SecOps Error-Handling Pattern
A resilient security ingest pipeline must:
- Capture the exact error string into
error.message. - Append a failure tag (such as
_dissectfailureor_grokparsefailure) totags. - Retain the unparsed raw log payload in
event.original. - Route the failed document into an error index or dead-letter data stream (e.g.,
logs-ingest_failures-default) rather than discarding it.
{
"description": "Ingest Pipeline with Robust on_failure Catch-All",
"processors": [
{
"grok": {
"field": "message",
"patterns": ["%{COMMON_PATTERN}"]
}
}
],
"on_failure": [
{
"set": {
"field": "error.message",
"value": "{{ _ingest.on_failure_message }}"
}
},
{
"append": {
"field": "tags",
"value": ["_pipeline_failure", "_grokparsefailure"]
}
},
{
"set": {
"field": "_index",
"value": "logs-pipeline_error-default"
}
}
]
}
Debugging with the Simulate API
Security engineers can validate pipeline logic, benchmark processor latency, and test error handling using the Ingest Simulate API (POST _ingest/pipeline/<pipeline-id>/_simulate) with real log payloads before deploying pipelines to production.
6. Architectural Comparison: Agent vs. Beats vs. Logstash
| Evaluation Dimension | Unified Elastic Agent | Standalone Beats | Logstash |
|---|---|---|---|
| Primary Role | Unified endpoint & server collector | Lightweight edge shipper | Intermediate pipeline & buffer |
| Management Model | Centralized via Fleet Server & Kibana | Distributed local YAML configs | Centralized via Pipeline Management or local files |
| Policy Distribution | Real-time push via HTTPS long-poll | Manual deployment / Ansible / Puppet | CI/CD pipeline or Kibana Central Management |
| Endpoint Protection | Integrated native EDR (Elastic Defend) | None (Requires separate EDR) | None (Server-side pipeline only) |
| Intermediate Buffering | Local disk spooling (limited) | Local disk spooling (limited) | Full Persistent Queue (PQ) on disk |
| Transformation Engine | Delegates to Ingest Pipelines | Delegates to Ingest Pipelines | Rich native filter plugins (Ruby, Grok, Dissect) |
| Resource Footprint | Single binary supervising several components | Lightweight, one process per shipper | JVM-based; usually the heaviest of the three |
| Primary SOC Scenario | Standard workstations, servers, cloud VMs | Legacy devices, specialized air-gapped hosts | High-volume syslog aggregator, multi-SIEM routing |
A security engineer is configuring an Ingest Pipeline to parse high-volume firewall connection logs arriving at 30,000 events per second. The firewall log format is strictly structured with fixed delimiter-separated tokens. Which processor should be selected to minimize CPU overhead on the ingest nodes?
Use the dissect processor because it extracts fields using fixed string delimiter matching without the computational overhead of regular expression compilation and backtracking.
Use the grok processor because regular expressions automatically optimize string allocations during multi-threaded parsing.
Use the script processor with custom Painless code to split strings on space characters and assign array indices.
Use the enrich processor to match each delimiter token directly against an internal reference table in Elasticsearch.
An enterprise Ingest Pipeline encounters an unhandled parsing error when a web application sends an unexpected log structure, causing a grok processor to fail. By default, what occurs to the incoming event, and how should a SOC engineer configure the pipeline to preserve raw forensic data?
The document is silently indexed with unparsed fields omitted; no configuration modification is required.
The document is cached in the coordinating node's heap memory and retried automatically every 60 seconds until parsing succeeds.
The Elasticsearch cluster marks the data node as degraded and places all active indices into read-only mode.
The document is not indexed and the client receives an ingest error; configure an on_failure block to capture error.message, tag the event, preserve event.original, and route to a fallback stream.
In a centralized enterprise deployment using Fleet and Elastic Agent, how do thousands of distributed Elastic Agents receive updated integration policies and configuration changes?
Fleet Server initiates inbound SSH connections to each managed endpoint on TCP port 22 to push updated YAML configuration files.
Elastic Agents maintain an outbound HTTPS long-polling connection to Fleet Server on port 8220, allowing Fleet Server to push configuration deltas within seconds.
Elastic Agents query the Elasticsearch master node directly via the _cluster/state API every 5 seconds to detect Kibana policy modifications.
Fleet Server broadcasts UDP multicast packets across the local subnet to trigger simultaneous configuration reloads across all connected agents.
Sections you finish are checked off in the contents.