9.2 Custom Query & Threshold Detection Rules

Key Takeaways

  • Custom Query rules evaluate KQL or Lucene query syntax against defined index patterns on a scheduled recurring interval, generating one alert per matching document or grouped alert batch.

  • The effective search window for scheduled rules is calculated as the rule execution interval plus additional look-back time (W = interval + look-back), creating an overlapping buffer that prevents missed alerts due to network latency and ingestion pipeline delays.

  • Threshold rules identify volume-based anomalies by aggregating matching events within the look-back window and triggering an alert only when the event count reaches or exceeds a defined threshold value.

  • Group by fields (up to three, such as source.ip or user.name) make threshold rules count per entity rather than globally, and the optional Count setting adds a unique-value (cardinality) condition.

  • Inappropriate look-back intervals or missing grouping fields lead to severe alert degradation—either generating redundant duplicate alerts on overlapping windows or obscuring distributed low-and-slow brute force attacks.

Last updated: September 2026

Core Detection Rule Paradigms in Elastic Security

While prebuilt rules provide comprehensive coverage against common commodity malware and recognized threat actor tactics, every enterprise presents a unique attack surface. Proprietary internal applications, bespoke network topologies, specialized administrative scripts, and organization-specific compliance mandates require SIEM analysts to engineer custom detection logic. Elastic Security provides several specialized rule types to meet these demands, with Custom Query Rules and Threshold Rules forming the fundamental workhorses of scheduled telemetry analysis.

Understanding the mathematical relationships between rule execution schedules, ingestion latency, search windows, and aggregation grouping is critical for authoring rules that reliably surface genuine security threats while eliminating false positives.


Custom Query Rules: Architecture and Formulation

Custom Query rules are the most widely deployed rule type in Elastic Security. They evaluate a search query against one or more index patterns on a recurring schedule. When documents matching the query criteria are returned, the Detection Engine generates an alert for each hit (or bundles them into an alert group).

+-------------------------------------------------------------+
|                      Custom Query Rule                      |
|  Query: process.name : "certutil.exe" and                   |
|         process.args : ("-urlcache" or "-split")            |
|  Indices: logs-endpoint.events.*, winlogbeat-*              |
+-------------------------------------------------------------+
                               |
                               v
               Elasticsearch Evaluates Query DSL
                               |
         +---------------------+---------------------+
         |                                           |
[ Document 1 Matches ]                      [ Document 2 Matches ]
         |                                           |
         v                                           v
[ Alert Signal Created ]                    [ Alert Signal Created ]
(Entity: host-01, PID: 3412)                (Entity: host-09, PID: 8812)

1. Query Languages: KQL vs. Lucene

When authoring Custom Query rules, analysts can write logic in either Kibana Query Language (KQL) or standard Lucene syntax:

  • Kibana Query Language (KQL):
    • Default and recommended language for most query rules.
    • Provides intuitive syntax, automated field completion, and seamless handling of nested ECS fields.
    • Example: event.category: "process" and process.name: "powershell.exe" and process.args: ("-enc" or "-encodedcommand")
    • Safe execution: KQL treats missing fields gracefully without throwing runtime exceptions.
  • Lucene Query Syntax:
    • Required when detection logic demands advanced search features unsupported by KQL, such as regular expressions, fuzzy matching, or field proximity searches.
    • Regex matching syntax: process.command_line: /.*-e(nc|ncodedcommand).*/
    • Range queries: destination.port: [1024 TO 65535]

2. Scoping Index Patterns and Data Streams

Custom Query rules must explicitly declare the index patterns or data streams they inspect. A common performance mistake is scoping rules against overly broad wildcards (e.g., * or logs-*).

  • Best practice: Scope queries strictly to the relevant ECS data streams, such as logs-endpoint.events.process-* for process execution, logs-network.* for firewall traffic, or logs-system.security-* for Windows event logs.
  • Scoping to targeted data streams reduces shard-level execution overhead and dramatically speeds up query evaluation.

Ingestion Delay Mitigation and the Search Window Formula

A critical challenge in real-time threat detection is accounting for ingestion delay (also known as telemetry latency). Understanding how the Detection Engine handles temporal windows prevents critical alerts from being missed.

The Latency Challenge

Consider a security event occurring on an endpoint at exactly 12:00:00. That timestamp is permanently stamped on the event as @timestamp. However, the event must traverse multiple architectural stages:

  1. The endpoint agent buffers and flushes the log.
  2. The log transmits across the WAN to an ingestion queue (e.g., Apache Kafka or Logstash).
  3. Ingest pipelines parse, grok, and enrich the event.
  4. Elasticsearch indexes the document into a data stream shard.

If this pipeline takes 90 seconds, the document is indexed into Elasticsearch at 12:01:30. If a detection rule executes on a strict 1-minute interval searching only from [now - 1m TO now], the query executing at 12:01:00 will look from 12:00:00 to 12:01:00. At 12:01:00, the event has not yet arrived in Elasticsearch! When the next run executes at 12:02:00, it queries 12:01:00 to 12:02:00. The event timestamped at 12:00:00 falls outside both search windows and is never evaluated, resulting in a missed detection.

The Look-Back Buffer Solution

To solve ingestion delay, the Detection Engine incorporates an Additional Look-Back Time parameter. The rule searches a rolling window that extends further back in time than the run interval.

The effective temporal search window (WW) evaluated by the rule is calculated as:

W=interval+look-backW = \text{interval} + \text{look-back}

Timeline:
| [==================== Effective Search Window (W) ====================] |
|                                                                         |
+---------------------------------------+---------------------------------+
|        Additional Look-Back Time      |        Rule Run Interval        |
|               (e.g., 2m)              |               (e.g., 5m)        |
+---------------------------------------+---------------------------------+
^                                                                         ^
now - (5m + 2m) = now - 7m                                               now

For example, if a rule runs every 5 minutes (interval = 5m) with an additional look-back of 2 minutes (look-back = 2m), each execution searches a 7-minute historical window (W=5m+2m=7mW = 5\text{m} + 2\text{m} = 7\text{m}):

  • Run at 12:05: Evaluates documents with @timestamp between 11:58:00 and 12:05:00.
  • Run at 12:10: Evaluates documents with @timestamp between 12:03:00 and 12:10:00.

Notice that the 2-minute window between 12:03:00 and 12:05:00 is searched twice. To prevent duplicate alerts for documents identified in multiple runs, the Detection Engine applies an automated Signal Deduplication Engine based on internal document IDs and rule signatures. If an alert has already been generated for a specific underlying document, subsequent overlapping runs ignore it. Elastic recommends at least 1 minute of additional look-back, and a rule that runs every 5 minutes with 1 minute of look-back searches the last 6 minutes on each run.


Threshold Detection Rules: Architecture and Mechanics

Many attack techniques cannot be identified by examining a single isolated event. A single failed SSH login is a routine operational occurrence; 50 failed SSH logins within 3 minutes targeting a single server is an active brute force attack. Threshold rules address this challenge by aggregating events over a specified time window and alerting only when the document count reaches or exceeds a designated numerical threshold.

+--------------------------------------------------------------------------+
| Threshold Rule: SSH Brute Force Detection                                |
| Query: event.category: "authentication" and event.outcome: "failure"     |
| Group By: source.ip                                                      |
| Threshold: count() >= 10                                                 |
| Window: 5 minutes                                                        |
+--------------------------------------------------------------------------+
                                    |
                                    v
       [ Aggregation by source.ip over [now - (5m + 1m) TO now] ]
                                    |
        +---------------------------+---------------------------+
        |                                                       |
  source.ip: 10.0.0.15                                    source.ip: 198.51.100.82
  Count: 2 failures                                       Count: 42 failures
  Threshold >= 10? [ NO ]                                 Threshold >= 10? [ YES ]
        |                                                       |
  (No Alert Fired)                                        [ ALERT SIGNAL FIRED ]
                                                          Entity: 198.51.100.82
                                                          Alert Count: 42 events

1. The Group By Field: Entity Cardinality and Partitioning

A fundamental requirement of threshold rule engineering is defining the Group by field(s). The group-by field determines how the aggregation buckets are segmented:

  • Without Group By (Global Threshold): The query calculates a single global count across all matching documents in the enterprise. If the threshold is set to 20 failed logins, and 20 different users across 5,000 corporate workstations each mistype their password once, the rule fires. This produces massive false positives.
  • With Single Group By (Entity Partitioning): Grouping by source.ip partitions the count per unique originating IP address. Grouping by user.name partitions the count per target account. An alert fires only when an individual entity crosses the threshold.
  • Composite / Multi-Field Group By: Grouping by multiple fields (up to three, e.g., source.ip AND user.name) identifies specific attacker-target pairings. This detects when a specific host attempts multiple logins against a specific user account.
  • Count (Cardinality) Condition: A threshold rule can also require a minimum number of unique values of another field. For example, Group by source.ip, Threshold ≥ 20 and Count user.name ≥ 10 alerts only when one IP produced at least 20 matching events against at least 10 different accounts, the typical shape of password spraying.
  • Synthetic Alerts: A threshold alert summarizes many events, so it contains only the Group by fields and counts rather than every field of the source documents. Sending the alert to Timeline lists all the matching events, including ones that did not reach the threshold.

2. Threshold Alert Suppression and Noise Control

When an ongoing attack generates hundreds or thousands of matching events within a short timeframe, an unmanaged threshold rule could generate an alert on every single scheduled execution. To manage analyst fatigue, Elastic Security allows threshold rules to configure Alert Suppression (a Platinum feature, in technical preview for threshold rules in 8.15):

  • When enabled, once an alert is generated for an entity (e.g., source.ip: 198.51.100.82), further alerts for that same entity are suppressed for a designated duration (e.g., 1 hour), even if subsequent executions continue to detect counts exceeding the threshold.

Practical SecOps Threat Detection Scenarios

Scenario A: RDP / Active Directory Brute Force Attack

  • Objective: Detect adversaries attempting rapid credential guessing against Windows Domain Controllers.
  • Rule Type: Threshold Rule
  • Index Pattern: winlogbeat-*, logs-system.security-*
  • Query: event.category: "authentication" and event.action: "logon-failed" and winlog.logon.type: "RemoteInteractive"
  • Group By: source.ip, user.name
  • Threshold: Count ≥10\ge 10
  • Schedule: Run every 5 minutes, Additional look-back: 2 minutes.

Scenario B: Reconnaissance & Port Scanning

  • Objective: Surface external network scanners probing perimeter firewalls.
  • Rule Type: Threshold Rule
  • Index Pattern: logs-firewall.*, logs-suricata.*
  • Query: event.category: "network" and network.transport: "tcp" and event.type: "denied"
  • Group By: source.ip
  • Threshold: Count ≥100\ge 100
  • Schedule: Run every 2 minutes, Additional look-back: 1 minute.

Scenario C: Ransomware Mass File Modification Canary

  • Objective: Detect rapid file tampering or mass renaming indicative of active ransomware encryption routines.
  • Rule Type: Threshold Rule
  • Index Pattern: logs-endpoint.events.file-*
  • Query: event.category: "file" and event.action: ("rename" or "deletion") and file.extension: ("locked" or "crypted" or "enc")
  • Group By: host.id, process.entity_id
  • Threshold: Count ≥50\ge 50
  • Schedule: Run every 1 minute, Additional look-back: 1 minute.

The "Low-and-Slow" Evasion Challenge

A known limitation of threshold rules is vulnerability to "low-and-slow" attacks. If an adversary knows an organization alerts on ≥10\ge 10 failed authentications within 5 minutes, the attacker can pace attempts to occur once every 2 minutes (2.5 attempts per 5 minutes). The threshold is never crossed. SIEM analysts counter low-and-slow tactics by deploying complementary Machine Learning anomaly detection rules or expanding threshold windows to hours or days with lower proportionate thresholds.


Custom Query Rules vs. Threshold Rules Comparison

The following table contrasts the technical, architectural, and operational characteristics of Custom Query rules versus Threshold rules:

SpecificationCustom Query RuleThreshold Rule
Primary PurposeDetect specific, high-fidelity single events or signaturesDetect volume-based anomalous aggregations over time
Query StructureKQL or Lucene search expressionKQL or Lucene search expression + Aggregation filter
Execution EngineStandard Elasticsearch Query DSL (hits search)Elasticsearch Terms / Date Histogram Aggregations
Alert GenerationOne alert per matching document (or 1 grouped alert)One alert per grouped entity bucket exceeding count (NN)
Group By FieldOptional (used for alert suppression)Optional, but usually essential for entity-level thresholds (up to 3 fields)
Search Window (WW)W=interval+look-backW = \text{interval} + \text{look-back}W=interval+look-backW = \text{interval} + \text{look-back}
Typical Threat Use CasesLOLBin execution, credential dumping, malware hashesBrute force logins, port scans, DDoS, ransomware file churn
Primary False Positive RiskBenign administrative scripts matching broad queriesHigh baseline activity crossing global un-partitioned thresholds
Primary Evasion RiskAdversary alters command-line parameters or filenamesAdversary slows attack rate below threshold window ("low-and-slow")
Loading diagram...
Scheduled Detection Rule Execution Window and Ingestion Buffer Calculation
Test Your Knowledge

A detection engineer configures a scheduled Custom Query rule with a run interval of 10 minutes (interval = 10m). Perimeter firewall logs experience an average ingestion delay of 3 minutes before becoming searchable in Elasticsearch. To ensure that late-arriving events are never missed, how should the rule's look-back time be configured, and what is the resulting effective search window (W)?

A

Configure an additional look-back time of at least 3 minutes (e.g. 3m to 5m), resulting in an effective search window of 13 to 15 minutes (W = interval + look-back).

B

Set the additional look-back time to 0 minutes, because Elasticsearch automatically recalculates timestamps upon index refresh.

C

Reduce the run interval to 3 minutes and set the look-back time to -3 minutes to eliminate overlapping shard scans.

D

Configure the look-back time to exactly 10 minutes and disable the deduplication engine so that events generate duplicate alerts across cycles.

Test Your Knowledge

A SOC deploys a threshold detection rule to identify Active Directory password spraying, configured with the query 'event.category: authentication and event.outcome: failure', a threshold count of 20, and a 5-minute schedule. Shortly after deployment, the rule triggers false positive alerts every morning during shift changes, even though no single user account or workstation experienced more than one or two failed logins. What architectural flaw in the threshold rule configuration caused this issue?

A

The query language was set to Lucene instead of Kibana Query Language (KQL).

B

The rule index pattern included domain controller security event streams.

C

The effective look-back window exceeded the 50 GB index lifecycle management rollover threshold.

D

The rule lacked a Group by field (such as 'source.ip' or 'user.name'), so it calculated a single global count across the entire organization.

Test Your Knowledge

A security analyst needs to detect an external credential stuffing attack against an enterprise web application. The attacker rotates through hundreds of distinct usernames, attempting exactly one login per user from a single malicious IP address (203.0.113.50), accumulating 60 failed login events within 4 minutes. Which threshold rule grouping configuration will successfully trigger an alert on this behavior while minimizing benign noise?

A

Group by 'user.name' with a threshold of 60 events.

B

Group by 'source.ip' with a threshold of 25 events.

C

Group by both 'user.name' and 'destination.port' with a threshold of 1 event.

D

Group by 'host.name' with an execution interval of 24 hours.

Sections you finish are checked off in the contents.