2.3 Security Data Storage, Data Streams & ILM
Key Takeaways
Elastic Data Streams provide an append-only abstraction over multiple time-series backing indices, automatically managing write indexing and time-based partitioning through the required @timestamp field.
Data streams follow the standardized naming taxonomy of type-dataset-namespace (such as logs-endpoint.events.process-production), enabling granular Role-Based Access Control, dataset-specific retention, and modular template inheritance.
Composable index templates assemble modular component templates containing index settings (shards, replicas, ILM policies) and ECS mappings, matching incoming streams via index patterns.
Index Lifecycle Management (ILM) governs automated phase transitions through Hot, Warm, Cold, Frozen, and Delete phases based on primary shard size and age, with min_age calculated strictly from index rollover time.
Searchable snapshots on the Cold tier (fully mounted) and Frozen tier (partially mounted) reduce storage costs by backing indices with snapshots in object storage while keeping them searchable with KQL and ES|QL.
In an enterprise SIEM, security telemetry arrives as a continuous, high-volume, timestamped sequence of events. Because security audit records, network flows, and endpoint process executions represent historical facts, they are fundamentally immutable once recorded. Managing millions of time-series events using traditional, static Elasticsearch indices required manual index rotation, custom alias management, and complex reindexing scripts. Modern Elastic Stack architectures replace this operational complexity with Elastic Data Streams and Index Lifecycle Management (ILM).
1. Elastic Data Streams Architecture
A Data Stream is a named, logical abstraction that allows an analyst or ingest client to treat time-series data as a single continuous resource, while Elasticsearch automatically partitions the underlying data across multiple hidden physical indices known as backing indices.
Logical Data Stream Interface
(logs-endpoint.events.process-production)
|
+------------------------+------------------------+
| | |
v (Read Queries) v (Read Queries) v (Read & Write)
+------------------+ +------------------+ +------------------+
| Backing Index | | Backing Index | | Backing Index |
| .ds-...-000001 | | .ds-...-000002 | | .ds-...-000003 |
| (Read-Only) | | (Read-Only) | | (Write Index) |
+------------------+ +------------------+ +------------------+
Core Properties of Data Streams
- Append-Only Operations: Data streams are purpose-built for time-series logging. They strictly prohibit direct document updates (
_update) and document replacement with static IDs (PUT <stream>/_doc/<id>). Ingested events must be submitted using thecreateoperation (op_type=create), guaranteeing the forensic immutability of security audit logs. Deletions can only be executed by targeting specific backing indices directly or via_delete_by_query. - Backing Indices Naming Convention: Underlying physical indices are hidden indices named according to the deterministic pattern:
.ds-<data-stream-name>-<yyyy.MM.dd>-<generation>Thegenerationis a 6-digit zero-padded integer (e.g.,000001,000002) that increments sequentially upon every rollover. - Single Active Write Index: Exactly one backing index serves as the active write index (the backing index with the highest generation number). All incoming write requests directed to the data stream name are routed exclusively to this index. Read requests query across all backing indices associated with the stream.
- Mandatory
@timestampField: Every document indexed into a data stream must contain an@timestampfield mapped to thedateordate_nanosfield type.
2. Data Stream Naming Taxonomy: <type>-<dataset>-<namespace>
Elastic enforces a standardized, three-part naming convention for data streams across all integrations and security data sources:
The Three Components
type: High-level generic data category. In SIEM, this is almost universallylogsfor events, ormetricsfor performance telemetry.dataset: Identifies the specific source system, application, or event format. Examples includewindows.sysmon_operational,suricata.eve,aws.cloudtrail,cisco.asa, andendpoint.events.process.namespace: An organization-defined partition representing an environment, business division, cloud tenant, or regulatory classification (such asproduction,corp,dmz,pci_dss, ordev).
Example Stream Name
logs-endpoint.events.process-production
type:logsdataset:endpoint.events.processnamespace:production
SecOps Benefits of the Taxonomy
- Granular Role-Based Access Control (RBAC): SOC administrators can grant tier-1 analysts read privileges to
logs-*-productionwhile restricting sensitive audit streams likelogs-aws.cloudtrail-complianceto senior forensic investigators. - Dataset-Specific Retention: High-volume, short-retention network telemetry (e.g.,
logs-network.traffic-*) can be assigned an ILM policy that purges data after 30 days, while authentication audit records (e.g.,logs-system.auth-*) are retained for 365 days. - Modular Template Inheritance: Index templates match wildcards across types, datasets, or namespaces, ensuring uniform mappings across disparate security tools.
3. Composable Index Templates and Component Templates
In modern Elasticsearch, indices and data streams are configured using Composable Index Templates (_index_template), which assemble reusable building blocks called Component Templates (_component_template).
+--------------------------------------------------------------+
| Component Templates (_component_template) |
| +-------------------------+ +-------------------------+ |
| | Settings Template | | Mappings Template | |
| | (shards, replicas, ILM) | | (ECS field mappings) | |
| +-------------------------+ +-------------------------+ |
+------------------------------+-------------------------------+
|
v (Inherited into)
+--------------------------------------------------------------+
| Composable Index Template (_index_template) |
| - Index Pattern: logs-windows.*-production |
| - Data Stream: {} (Enables Data Stream Creation) |
| - Priority: 200 |
+------------------------------+-------------------------------+
|
v (Automatically Configures)
+--------------------------------------------------------------+
| Data Stream |
| logs-windows.sysmon_operational-production |
+--------------------------------------------------------------+
Component Templates
Component templates modularize configuration settings across the cluster:
- Settings Component Templates: Define cluster-level behavior such as
number_of_shards: 2,number_of_replicas: 1,refresh_interval: "5s", and the assigned ILM policy name (index.lifecycle.name: "security-logs-ilm-policy"). - Mappings Component Templates: Define the Elastic Common Schema (ECS) field definitions, explicit field types (
keyword,ip,date), and dynamic templates for nested fields.
Composable Index Templates
A composable index template binds component templates together:
index_patterns: Specifies wildcard patterns matching incoming index names (e.g.,["logs-windows.*-*"]).data_stream: {}: An empty JSON object block. Its presence explicitly instructs Elasticsearch to create a Data Stream rather than a standard index when a matching write request is received.composed_of: An ordered array listing the names of component templates to apply.priority: An integer defining template precedence. If a data stream matches multiple templates (e.g.,logs-*-*with priority 100 andlogs-windows.*-*with priority 200), the template with the highest priority wins.
4. Index Lifecycle Management (ILM) Deep-Dive
Index Lifecycle Management (ILM) automates the progression of backing indices through up to five distinct lifecycle phases to balance high-speed search performance against infrastructure storage expenditure.
| ILM Phase | Minimum Age (min_age) | Primary Configured Actions | Target Storage Tier | SecOps Purpose & Operational Impact |
|---|---|---|---|---|
| Hot | 0ms (Immediate upon creation) | rollover (Triggered by size/age) | data_hot (NVMe/SSD) | Active write index, continuous ingestion, real-time alert rule evaluation |
| Warm | Configured delay from rollover (e.g., 7d) | readonly, forcemerge (1 segment), shrink | data_warm (Fast SSD / HDD) | Read-only index, segment optimization, active forensic investigation and triage |
| Cold | Configured delay from rollover (e.g., 30d) | searchable_snapshot (fully mounted) | data_cold (full local copy of snapshot data) | Rarely queried telemetry; no replicas needed because the snapshot provides resilience |
| Frozen | Configured delay from rollover (e.g., 90d) | searchable_snapshot (partially mounted) | data_frozen (object store + shared cache) | Compliance retention, annual threat hunting baselining, minimal local disk |
| Delete | Configured delay from rollover (e.g., 365d) | delete | None (Purged) | Automated purge of data exceeding statutory or regulatory retention limits |
Phase 1: Hot Phase and Rollover Triggers
The Hot phase handles active write operations. The active backing index remains in the Hot phase until at least one configured rollover condition is met:
max_primary_shard_size: The recommended primary rollover trigger for security data. Elastic's default ILM policies for integrations roll over at 50 GB per primary shard (or 30 days), which keeps shards within the recommended 10 GB to 50 GB range.max_age: Maximum elapsed time before forcing a rollover (e.g.,30d). Ensures that low-volume datasets roll over regularly and transition through the lifecycle rather than remaining in the Hot tier indefinitely.max_docs: Triggers rollover after a fixed number of documents have been indexed.max_size: Evaluates total index size across both primary and replica shards.
When rollover executes:
- A new backing index is created with an incremented generation number (e.g.,
.ds-...-000002). - The new backing index becomes the active write index.
- The previous backing index becomes read-only and begins its lifecycle progression toward subsequent phases.
Phase 2: Warm Phase and Segment Optimization
In the Warm phase, telemetry is read-only and queried for secondary investigations and baseline calculations. Two critical actions execute here:
forcemerge(max_num_segments: 1): Merges all smaller Lucene segments into a single segment file. This permanently purges deleted document tombstones, optimizes BKD trees and inverted indices, and substantially reduces JVM heap overhead.shrink: Consolidates primary shard count (e.g., shrinking a 4-shard index down to 1 shard), further reducing cluster-wide shard overhead.
Phase 3 & 4: Cold and Frozen Tiers with Searchable Snapshots
Searchable Snapshots eliminate the requirement to retain full primary and replica shard copies on expensive local block storage:
- Fully Mounted Indices (Cold Tier): The index is restored from the snapshot repository onto local
data_colddisks as a complete copy. Searches run against local data, and no replica shards are needed because the snapshot itself protects the data. - Partially Mounted Indices (Frozen Tier): Only a shared local cache is kept on
data_frozennodes. When a search needs data that is not cached, the node fetches it from the object store, so searches are slower but local storage needs are very small.
Phase 5: Delete Phase
Once an index's age exceeds the configured retention window (e.g., min_age: 365d), ILM executes the delete action, permanently removing the backing index and releasing cloud storage allocations.
5. The Critical min_age Calculation Rule (Exam Trap)
A frequent source of confusion on the Elastic SIEM exam is how the min_age parameter is calculated.
Important
The Fundamental ILM Timing Rule: For backing indices managed by Data Streams and rolled over via ILM, min_age is calculated relative to the time of rollover, NOT the index creation time or the timestamps of the documents inside the index.
Step-by-Step Scenario Analysis
- Day 0 (October 1st): A new backing index
.ds-logs-network-default-2026.10.01-000001is created in the Hot phase. - Day 15 (October 15th): The index reaches 50 GB and satisfies
max_primary_shard_size: 50GB. The index rolls over on October 15th. - Warm Phase Configuration: The assigned ILM policy defines the Warm phase with
min_age: 10d. - Transition Date: The backing index does not enter the Warm phase on October 10th (10 days from creation). It enters the Warm phase on October 25th (October 15th rollover time + 10 days
min_age).
If an index does not roll over (e.g., a static standalone index), min_age is calculated from the index creation time. But for all Data Streams governed by rollover, rollover time is the reference timestamp.
A security data stream is governed by an Index Lifecycle Management (ILM) policy with rollover configured at max_primary_shard_size: 50GB and a Warm phase transition configured with min_age: 7d. A new backing index begins ingesting events on September 1st. Due to moderate ingestion rates, the backing index reaches 50 GB and rolls over on September 14th. On what date will the backing index transition into the Warm phase?
September 8th, exactly 7 days after the backing index was initially created.
September 14th, immediately upon satisfying the 50 GB rollover criteria.
September 21st, exactly 7 days after the backing index rolled over.
October 1st, exactly 30 days after the start of the initial ingestion calendar month.
A security analyst attempts to update an existing document within an active Elastic Data Stream using the POST logs-endpoint.events.process-default/_update/doc-1001 API. What is the expected behavior and architectural reason?
The request is rejected with an error because Data Streams are append-only structures that prohibit direct document updates and static ID mutations.
The request succeeds and overwrites the document across all backing indices simultaneously.
The request automatically creates a new replica shard in the Frozen tier to store the document delta.
The request succeeds only if the target document resides within the active write backing index.
An ILM policy governing security audit logs includes a forcemerge action configured with max_num_segments: 1 within its Warm phase definition. What is the primary operational benefit of executing this action on rolled-over backing indices?
It transforms the backing index from an inverted index structure into a columnar BKD tree for real-time write acceleration.
It automatically executes synthetic detection alert tests against historical logs to validate correlation rules.
It consolidates multiple smaller Lucene segments into a single segment, expelling deleted document tombstones and optimizing search performance.
It immediately converts the backing index into a searchable snapshot without requiring an external snapshot repository.
Sections you finish are checked off in the contents.