2.1 Elastic Stack Architecture for Security Operations

Key Takeaways

  • Elasticsearch operates as a distributed cluster where dedicated node roles (master, data tiers, ingest, coordinating, and machine learning) isolate workloads to ensure ingestion stability and low-latency SIEM investigations.

  • The data tiering architecture (Hot, Warm, Cold, and Frozen) maps telemetry age and query frequency to physical storage media ranging from NVMe SSDs to cloud object storage.

  • Lucene internals rely on immutable segments containing inverted indices for full-text keyword search, doc values for aggregations and sorting, and BKD trees for numeric, date, and IP range queries.

  • Elastic's sizing guidance targets shards of 10 GB to 50 GB, respects the default limit of 1,000 non-frozen shards per node, and keeps JVM heap at no more than 50% of RAM and below the roughly 32 GB compressed-oops threshold.

  • Cluster health states reflect shard availability: Green indicates all primary and replica shards are allocated, Yellow denotes unallocated replica shards with zero data loss, and Red warns of unallocated primary shards requiring immediate SOC triage.

Last updated: September 2026

Modern Security Operations Centers (SOCs) rely on the Elastic Stack as a centralized Security Information and Event Management (SIEM) platform capable of ingesting tens of thousands of events per second while simultaneously evaluating complex threat detection rules and supporting interactive forensic investigations. Understanding the distributed architecture of Elasticsearch, how data is indexed and queried across specialized node roles, and how storage tiers align with compliance retention policies is essential for both the Elastic Certified SIEM Analyst exam and day-to-day SOC administration.

Note

Elastic's objective here is to describe basic Stack architecture. For the exam, be able to name the components and how data flows between them. Elastic Agent (or Beats) collects data, and Fleet manages agents centrally. Optional Logstash transforms and routes data, Elasticsearch stores and searches it, and Kibana (including the Security app) is the user interface. The node-role, storage and sizing detail below explains why searches and rules behave as they do. It is background knowledge rather than a separate exam objective.


1. Distributed Node Roles in a SOC Cluster

An Elasticsearch cluster is a collection of one or more connected node instances sharing cluster state and data storage. In high-volume SIEM deployments, running homogeneous nodes where every instance performs all tasks leads to severe resource contention. Under sustained ingestion surges or intensive threat hunting queries, unspecialized nodes frequently exhaust memory and drop incoming telemetry. Production SOC clusters therefore implement dedicated, specialized node roles.

Node RoleConfiguration KeyPrimary SecOps ResponsibilityResource Profile
Master-Eligiblenode.roles: [ master ]Cluster state management, shard allocation, index metadata, node quorum electionHigh CPU availability, moderate RAM, low I/O
Data (Hot)node.roles: [ data_hot, data_content ]Active telemetry ingestion, real-time detection rule execution, interactive triage queriesHigh-speed NVMe/SSD, high RAM, high CPU
Data (Warm)node.roles: [ data_warm ]Read-only telemetry storage, segment force-merging, secondary forensic queriesHigh-density SSD or high-throughput HDD, moderate RAM
Data (Cold)node.roles: [ data_cold ]Historical investigation data, fully mounted searchable snapshotsLocal disk holding a full copy of snapshot data; snapshot kept in object storage
Data (Frozen)node.roles: [ data_frozen ]Long-term compliance retention, partially mounted searchable snapshotsCloud object storage (S3/GCS/Azure Blob) plus a shared local cache
Ingestnode.roles: [ ingest ]Execution of Ingest Pipelines (grok, dissect, geoip, rename, set) prior to indexingHigh CPU, moderate RAM, no persistent shard storage
Coordinatingnode.roles: [ ] (empty list)Client HTTP termination, scatter-gather query routing, search result reductionHigh network bandwidth, high CPU for aggregation reduction
Machine Learningnode.roles: [ ml ]Anomaly detection jobs, trained model inference (DGA detection, process risk scoring)Dedicated multi-core CPU, high RAM, isolated from indexing

Master-Eligible Nodes

Master-eligible nodes maintain the authoritative cluster state, which includes the global mapping definitions, index settings, active shard routing tables, and the list of joined nodes. The active master is elected by Elasticsearch's cluster coordination subsystem, which requires a quorum of master-eligible nodes (the cluster.initial_master_nodes setting is used only to bootstrap a brand-new cluster). Dedicated master nodes do not handle client search requests or store data shards. In a SOC cluster ingesting terabytes daily, keeping master nodes isolated ensures that massive indexing spikes or heavy aggregation queries cannot starve the master node of CPU or cause it to drop out of the cluster, preventing catastrophic split-brain scenarios or cluster freezes.

Data Nodes and Tiered Storage

Data nodes hold the physical shards containing indexed security documents. Rather than treating all data nodes identically, Elasticsearch organizes data nodes into distinct storage tiers (hot, warm, cold, and frozen). This ensures hardware resources match the lifecycle value of security telemetry.

Ingest Nodes

Ingest nodes execute Ingest Pipelines before documents are written into Lucene segments. They perform CPU-intensive parsing tasks such as regular expression extraction (grok), delimiter splitting (dissect), IP geolocation enrichment (geoip), and ECS field renaming. By routing telemetry through dedicated ingest nodes, SOC engineers prevent ingest transformation overhead from degrading query latency on data nodes.

Coordinating (Client) Nodes

Coordinating-only nodes accept incoming HTTP traffic from Kibana, Elastic Agents, and external API integrations. When a security analyst executes a Kibana Query Language (KQL) or ES|QL query across 90 days of firewall logs, the coordinating node coordinates the two-phase search process: scattering the query across all relevant data shards in the cluster, gathering the partial hits, executing the final aggregation reduce phase in memory, and returning the structured JSON response to Kibana.

Machine Learning Nodes

Elastic Security includes unsupervised machine learning for anomaly detection (such as rare process execution, anomalous DNS tunneling, or abnormal user authentication spikes) and supervised models for real-time inference. Machine learning jobs perform intensive mathematical matrix operations. Placing ML workloads on dedicated nodes prevents CPU starvation on data nodes running scheduled detection rules.


2. Data Tiers for Security Telemetry

Enterprise security compliance frameworks (such as PCI DSS, SOC 2, and ISO 27001) mandate retaining audit logs, authentication records, and network telemetry for one to seven years. However, active threat detection and SOC alert triage focus predominantly on the most recent 24 to 72 hours of data. Elastic data tiers solve this economic and architectural imbalance by systematically shifting indices across four distinct tiers.

Data TierStorage MediaIndex StateRelative Query SpeedExample SecOps Retention & Use Case
HotNVMe or high-performance SSDRead/Write (Active write index)FastestReal-time detection engine rules, active alert triage, streaming telemetry (0–7 days)
WarmHigh-capacity SSD or fast HDDRead-only (often force-merged)FastSecondary investigations, Timeline correlation, lateral movement hunting (7–30 days)
ColdLocal disk with a full copy of snapshot data (snapshot in S3/GCS/Azure)Read-only (Fully mounted snapshot)SlowerQuarterly compliance audits, historical forensic verification (30–90 days)
FrozenCloud object storage (S3/GCS/Azure Blob) + shared local cacheRead-only (Partially mounted snapshot)SlowestAnnual compliance retention, retrospective threat hunting, legal hold (90–365+ days)

The retention windows in the last column are example policy choices, not Elastic defaults. You set them in your own ILM policies.

Hot Tier Architecture

The Hot tier handles all active document indexing. Every Data Stream writes exclusively to an active backing index located on a data_hot node. Because indexing requires continuous disk writes, memory flushes, and background segment merging, Hot nodes require high-throughput NVMe storage and fast multi-core CPUs. The Elastic Security Detection Engine evaluates scheduled rules against Hot indices every few minutes.

Warm Tier Architecture

Once a backing index satisfies its rollover conditions (such as reaching 50 GB in primary shard size or 30 days of age), Index Lifecycle Management (ILM) migrates the index to the Warm tier. The index is marked read-only, preventing any further write operations. A background forcemerge action consolidates smaller Lucene segments into a single segment, dramatically reducing memory overhead and optimizing search speed for incident responders investigating multi-week campaigns.

Cold and Frozen Tiers with Searchable Snapshots

The Cold and Frozen tiers decouple compute from storage by leveraging Searchable Snapshots. Instead of retaining local primary and replica shard copies across expensive storage arrays, Elasticsearch takes a snapshot of the rolled-over index and stores it in cloud object storage (such as AWS S3, Google Cloud Storage, or Azure Blob Storage). In the Cold tier, indices are fully mounted: each node keeps a complete local copy of the snapshot's shard data, and no replica is needed because the snapshot repository provides resilience. In the Frozen tier, indices are partially mounted: nodes keep only a shared local cache of recently used data, and other data is fetched from the object store when a search needs it. This allows a SOC to retain a full year of raw proxy, endpoint, and firewall telemetry at a fraction of traditional storage costs, while keeping the data immediately searchable via KQL and ES|QL.


3. Lucene Storage Internals in a Security Context

At the storage engine layer, Elasticsearch builds upon Apache Lucene. Understanding Lucene data structures explains why certain queries execute instantaneously while others cause severe cluster memory pressure.

Inverted Index

The inverted index is the core data structure for full-text search and keyword matching. It tokenizes text fields and constructs an alphabetical dictionary of distinct terms, mapping each term to a sorted posting list of document IDs containing that term. When an analyst searches for a suspicious command-line utility (such as process.name: "powershell.exe"), Elasticsearch does not scan millions of log lines; it performs a direct lookup in the inverted index posting list, locating the matching documents in microseconds.

Doc Values

While the inverted index excels at finding documents matching a term, it is inefficient for sorting, aggregations, and script execution. Lucene solves this with Doc Values: a column-oriented on-disk data structure generated at index time. Doc values map document IDs to their column values. When a Kibana Lens dashboard calculates the top 10 destination IP addresses or plots authentication failure histograms over time, Elasticsearch reads directly from doc values. Doc values are stored on disk and rely on the operating system filesystem cache, keeping JVM heap utilization minimal.

BKD Trees

For numeric types (long, integer, float), date fields (@timestamp), and IP address types (ip), Elasticsearch utilizes Block K-d (BKD) trees. A BKD tree is a multi-dimensional binary search tree optimized for range evaluations and geometric lookups. When an analyst queries for connections within a specific subnet (such as source.ip: "192.168.1.0/24") or filters events across a 4-hour window (@timestamp >= "now-4h"), BKD trees allow Elasticsearch to identify matching document ranges with minimal I/O operations.

Lucene Segments and the Merge Process

An Elasticsearch index is divided into shards, and each shard is physically comprised of multiple immutable Lucene segments. When documents are ingested, they are initially written to an in-memory indexing buffer. Periodically (by default every 1 second, or when the buffer fills), Elasticsearch performs a refresh, writing the buffer into a new immutable Lucene segment on disk. Because segments are immutable:

  1. Read operations require no concurrency locking, allowing parallel search threads.
  2. Inverted indices and BKD trees never suffer from on-disk fragmentation.
  3. Updates and deletions do not modify existing data; updates write a new document and mark the old one as deleted in a live-documents bitset.

As hundreds of small segments accumulate, Elasticsearch executes background segment merges, combining multiple segments into a larger segment and permanently purging tombstones. During the ILM Warm transition, administrators execute a forcemerge to combine all segments into max_num_segments: 1, ensuring optimal query performance and reclaiming deleted document space.


4. SOC Cluster Sizing, Sharding, and JVM Memory Management

Improper cluster sizing and misconfigured JVM memory are the primary root causes of SIEM stability failures during security incidents.

Primary Shard Sizing Benchmark

Elastic's sizing guidance is to aim for shards of 10 GB to 50 GB (and up to about 200 million documents). The default ILM policies used by integrations roll over at 50 GB per primary shard or 30 days.

  • Over-sharding Risk (many shards far below 10 GB): Creating hundreds of tiny shards (such as hourly shards or separate indices for low-volume devices) consumes excessive JVM heap memory. Each Lucene segment maintains term dictionaries, doc values metadata, and file pointers in heap. Furthermore, coordinating nodes must manage separate search threads for every shard queried, causing high CPU overhead and query queue saturation.
  • Oversized Shard Risk (> 50 GB per shard): Shards much larger than 50 GB impede cluster resilience. If a node fails, rebalancing a 150 GB shard across the network causes prolonged I/O bottlenecks. In addition, background segment merges on massive shards can trigger high disk latency and indexing throttling.

Shard Limits and Heap Overhead

Older guidance quoted a rule of thumb of about 20 shards per GB of heap. Elastic's current 8.x sizing guide no longer uses that ratio. Instead it tells you to:

  • Stay within the cluster shard limit, which by default allows 1,000 non-frozen shards per node (cluster.max_shards_per_node) and 3,000 frozen shards per dedicated frozen node.
  • Allow enough heap for mapped fields and per-index overhead. Large numbers of mapped fields across many indices consume heap on every data node.
  • Give master-eligible nodes at least 1 GB of heap per 3,000 indices.

For a SIEM, the practical lesson is the same: avoid thousands of tiny indices, reuse ECS mappings rather than inventing new fields, and let data-stream rollover create reasonably sized shards.

JVM Heap Allocation and Compressed OOPs

Elasticsearch runs on the Java Virtual Machine (JVM). Correct heap configuration is critical:

  1. Keep JVM heap at no more than 50% of physical host RAM. Elasticsearch 8.x sizes the heap automatically by default. If you set it manually, leave at least half of the RAM for the operating system filesystem cache. The filesystem cache holds Lucene doc values and inverted index structures, enabling sub-second search without reading from physical disk.
  2. Keep heap below the compressed-oops threshold (near 32 GB). Elastic notes that about 26 GB is safe on most systems. In modern 64-bit systems, JVM pointers require 64 bits. However, when heap is configured below approximately 32 GB, the JVM enables Compressed Object Pointers (UseCompressedOops), encoding 64-bit object references into 32-bit offsets. If heap is configured above the threshold (for example 33 GB or 64 GB), Compressed OOPs is disabled, expanding every pointer from 4 bytes to 8 bytes. A heap just above the threshold can provide less usable memory than one just below it while significantly increasing garbage collection (GC) pause durations and degrading CPU L1/L2 cache hit rates.

5. Cluster Health Indicators and Search Optimization

Cluster Health Statuses

The cluster health API (GET _cluster/health) returns one of three discrete states:

  • Green: All primary shards and replica shards are fully allocated and operational across active nodes. The cluster possesses full redundancy.
  • Yellow: All primary shards are active and allocated, guaranteeing zero data loss and full read/write availability, but at least one replica shard cannot be allocated. This commonly occurs in single-node dev environments, following planned node reboots, or when rack-awareness rules cannot find an alternate fault domain.
  • Red: At least one primary shard is unallocated on the cluster. Ingest requests to affected indices fail, and forensic searches return partial data. A Red status demands immediate SOC investigation (checking unassigned shard reasons via GET _cluster/allocation/explain).

Filter Context vs. Query Context in SecOps

When creating detection rules or querying via KQL, understanding Lucene query context is critical:

  • Query Context (must, should): Evaluates whether a document matches and calculates a relevance score (_score) representing match precision. Scoring requires intensive floating-point CPU calculations.
  • Filter Context (filter, must_not): Evaluates a binary "yes/no" condition without computing relevance scores. Elasticsearch automatically caches filter bitsets in memory. For security monitoring—where an event either originated from source.ip: "10.0.0.1" or did not—queries should always operate in filter context to maximize throughput and leverage node-level query caches.
Loading diagram...
Elastic Stack Distributed Architecture for Security Operations
Test Your Knowledge

A security operations team observes that their production Elasticsearch cluster health indicator has transitioned from Green to Yellow following a scheduled rolling restart of data nodes. Cluster ingestion continues without interruption, and scheduled detection rules execute normally. What is the operational state of the cluster?

A

At least one primary shard has failed to allocate, causing forensic search queries across recent data to return partial results.

B

All primary shards are successfully allocated and active, but at least one replica shard cannot currently be allocated to an active node.

C

The active master node has lost quorum, forcing the cluster into a read-only state to prevent split-brain index corruption.

D

Ingestion queues have saturated available memory on the ingest nodes, initiating backpressure throttling on connecting Elastic Agents.

Test Your Knowledge

A SIEM architect is designing a multi-node Elasticsearch deployment to ingest 1.5 TB of daily Windows Sysmon and network connection telemetry. Which configuration adheres to Elastic best practices for shard sizing and JVM memory management on the data nodes?

A

Allocate 64 GB of JVM heap per node to maximize memory buffers and configure 5 GB primary shards for maximum parallel query execution.

B

Allocate 80% of host RAM to JVM heap and configure 100 GB primary shards to minimize the number of open Lucene segment files.

C

Keep heap at no more than 50% of host RAM and below the roughly 32 GB compressed-oops threshold, and aim for shards between 10 GB and 50 GB.

D

Allocate a fixed 16 GB heap regardless of host RAM and configure exactly one primary shard per calendar month regardless of ingested volume.

Test Your Knowledge

A financial enterprise must retain security audit logs and firewall traffic records for 365 days to comply with regulatory standards. Security analysts rarely query data older than 90 days, but when compliance audits occur, they must run direct KQL and ES|QL queries across the full year without manually rehydrating or re-indexing archived files. Which storage tier configuration satisfies these requirements at the lowest operational storage cost?

A

Retain all 365 days of telemetry in the Hot tier backed by NVMe SSDs to guarantee sub-second audit query performance.

B

Retain telemetry older than 90 days in the Warm tier on high-capacity spinning disks with two replica shards and force-merging disabled.

C

Export telemetry older than 90 days to offline cold tape storage and re-index the JSON files into Elasticsearch upon audit request.

D

Transition telemetry older than 90 days to the Cold or Frozen tier using searchable snapshots mounted directly against cloud object storage.

Sections you finish are checked off in the contents.