10.2 New Terms, ES|QL & Machine Learning Detection Rules

Key Takeaways

  • New Terms detection rules track novel categorical values in streaming telemetry by comparing terms in the active search window against an established historical baseline look-back window.

  • To mitigate false-positive alert floods during cold-start deployment, New Terms rules require adequate historical index depth and careful configuration of baseline look-back durations.

  • ES|QL detection rules execute vectorized, piped queries within the detection engine, enabling in-line mathematical calculations, dynamic ratio thresholding, and multi-stage statistical filtering.

  • Machine Learning (ML) rules integrate unsupervised anomaly detection jobs with the detection engine, mapping continuous anomaly scores (0 to 100) to discrete alert severity ratings.

  • Selecting among New Terms, ES|QL, and ML rules depends on the threat scenario, computational overhead, baseline requirements, and whether the detection logic is statistical, ratio-driven, or novelty-based.

Last updated: September 2026

While traditional Custom Query and Threshold rules provide robust defense against known, explicit attack patterns, modern adversaries frequently execute attacks that leave no static signature. Attackers utilize "living-off-the-land" binaries (LOLBins), compromise legitimate administrative user accounts, establish slow-and-low command-and-control channels, or dynamically generate algorithmic domain names (DGA). Detecting these sophisticated tactics requires advanced analytical engines capable of establishing behavioral baselines, executing mathematical aggregations, and running unsupervised statistical models.

Elastic Security provides three advanced rule architectures designed for behavioral and statistical detection: New Terms Rules, ES|QL Detection Rules, and Machine Learning (ML) Rules. Mastering the implementation mechanics, operational trade-offs, and appropriate use cases for each rule type is essential for enterprise detection engineering.


New Terms Detection Rules: Architecture & Baselining

Security analysts frequently need to answer the question: "Has this entity ever performed this action before in our environment?" Examples include:

  • Has this user account ever authenticated to this domain controller before?
  • Has this internal server ever established an outbound connection to this external country or autonomous system (ASN)?
  • Has this executable binary ever run on this critical database host?

Answering these questions using standard query rules would require maintaining complex external state databases or writing unmaintainable Lucene queries containing thousands of inverted values. New Terms Rules solve this by tracking the historical frequency of terms directly within Elasticsearch.

+---------------------------------------------------------------------------------------------------+
|                             NEW TERMS RULE BASELINING ARCHITECTURE                                |
+---------------------------------------------------------------------------------------------------+
|                                                                                                   |
|   HISTORICAL BASELINE LOOK-BACK WINDOW                           CURRENT SEARCH WINDOW            |
|   [ now-30d  =========================================>  now-5m ] [ now-5m ========> now ]        |
|                                                                 |                                 |
|   Observed Historical Terms for 'process.name' on srv-app01:    | Newly Observed Term:            |
|   - "java.exe"                                                  | - "powershell.exe"              |
|   - "cmd.exe"                                                   |                                 |
|   - "conhost.exe"                                               |                                 |
|   - "w3wp.exe"                                                  |                                 |
|                                                                 |                                 |
|                   |                                                             |                 |
|                   +------------------------------+------------------------------+                 |
|                                                  |                                                |
|                                                  v                                                |
|                                  +-------------------------------+                                |
|                                  |   TERM COMPARISON ENGINE      |                                |
|                                  | Is "powershell.exe" in the    |                                |
|                                  | historical 30-day baseline?   |                                |
|                                  +-------------------------------+                                |
|                                                  |                                                |
|                                           NO (First-Seen!)                                        |
|                                                  |                                                |
|                                                  v                                                |
|                                  [ GENERATE SECURITY ALERT ]                                      |
|                                  Severity: Medium / High                                          |
|                                  New Term: process.name = "powershell.exe"                        |
|                                                                                                   |
+---------------------------------------------------------------------------------------------------+

Operational Mechanics

When configuring a New Terms rule, the detection engineer defines:

  1. Target Field: The primary ECS keyword or IP field to monitor for novel values (e.g., user.name, destination.ip, process.executable).
  2. History Look-Back Window: The historical baseline duration against which current activity is evaluated (e.g., 30d or 14d).
  3. Search Window / Frequency: The current evaluation timeframe (e.g., every 5m or 1h).
  4. Multiple Fields (up to three): Selecting several fields detects a combination of values never seen together before. For example, user.name plus host.name detects a user logging into a specific host for the first time, even if that user routinely logs into other workstations. The history window must be larger than the rule interval plus additional look-back time.

The Cold-Start Problem & Baseline Building

A major operational pitfall when deploying New Terms rules is the cold-start phenomenon. If an engineer creates a New Terms rule with a 30-day look-back window in a cluster that has only retained 2 days of logs, or if a rule targets an extremely volatile field (e.g., ephemeral source ports), the engine will classify thousands of routine operations as "first-seen," generating an overwhelming alert storm.

Best Practices for New Terms Baselining:

  • Verify Telemetry Depth: Ensure the underlying data streams have continuously collected logs for the full duration of the configured look-back window before enabling the rule in production.
  • Scope Target Fields: Apply strict source filters to eliminate volatile, high-cardinality background noise. For example, when detecting first-seen outbound network destinations, filter for specific critical servers (host.name: "db-prod-*") and exclude internal subnets and enterprise CDN ranges.
  • Staging / Mute Period: Deploy newly authored New Terms rules in an inactive or notification-suppressed staging mode for 7 to 14 days to observe term churn before routing alerts to Tier 1 operational queues.

ES|QL Detection Rules: Vectorized Mathematical & Multi-Stage Analysis

With the introduction of the Elasticsearch Query Language (ES|QL), Elastic Security gained a transformative detection paradigm. Traditional query languages (Lucene and KQL) operate on a filter-and-retrieve model, while Event Query Language (EQL) specializes in chronological event sequencing. However, neither easily supports inline mathematical calculations, statistical aggregations across groups, or multi-stage piped transformations.

ES|QL Detection Rules execute piped, vectorized queries directly inside the detection engine, processing telemetry in memory using column-oriented data blocks. This unlocks advanced threat detection scenarios that previously required external custom scripts or complex SOAR workflows.

Key Capabilities of ES|QL in Detection Engineering

  • In-Line Mathematical Calculations (EVAL): Compute dynamic metrics such as ratios, percentage changes, duration differences, or cryptographic entropy values directly within the query.
  • Multi-Stage Statistical Aggregations (STATS ... BY): Aggregate counts, distinct values, sums, averages, and percentiles across multiple categorical dimensions simultaneously.
  • Inline Data Enrichment (ENRICH): Cross-reference streaming telemetry against pre-compiled Elasticsearch enrich policies (e.g., asset vulnerability scores, user risk tiers, threat intelligence IP lists) mid-pipeline.
  • Dynamic Conditional Logic (CASE / COALESCE): Implement branching logic to evaluate complex operational scenarios in a single rule.

SecOps Scenario: Detecting Password Spraying via Failure-to-Success Ratios

Consider detecting a distributed password spraying campaign. An adversary targets hundreds of user accounts across an enterprise, but ensures that each individual IP address only attempts 3 failed logins to evade standard threshold rules. However, from the perspective of the authentication service, the ratio of failed logins to successful logins spikes abnormally for specific target domains.

An ES|QL detection rule detects this by calculating dynamic failure percentages:

FROM logs-auth.*
| EVAL failed = CASE(event.outcome == "failure", 1, 0),
       succeeded = CASE(event.outcome == "success", 1, 0)
| STATS 
    total_attempts = COUNT(*),
    failed_attempts = SUM(failed),
    successful_attempts = SUM(succeeded)
  BY source.ip, user.domain
| EVAL failure_ratio = (failed_attempts * 100.0) / total_attempts
| WHERE total_attempts >= 20 AND failure_ratio >= 85.0
| SORT failure_ratio DESC
| LIMIT 50

In this single pipeline, ES|QL:

  1. Searches authentication telemetry within the rule's schedule window (interval plus additional look-back), which the detection engine applies automatically.
  2. Uses EVAL with CASE to flag failures and successes, then STATS ... BY to total them per source.ip and user.domain. ES|QL 8.15 has no COUNT_IF function, so conditional counts are built this way.
  3. Calculates the failure_ratio percentage.
  4. Keeps sources with at least 20 attempts and a failure ratio of 85% or more.
  5. Creates one alert per remaining row. The lower of LIMIT and the rule's Max alerts per run setting (default 100) caps the number of alerts.

Aggregating vs. Non-Aggregating ES|QL Rules

  • Aggregating queries (using STATS ... BY) create alerts that contain only the columns the query returns. Because events in the look-back overlap can be counted in two runs, they may create duplicate alerts.
  • Non-aggregating queries keep the source event fields. To let the engine deduplicate them, add METADATA _id, _index, _version after FROM (for example FROM logs-* METADATA _id, _index, _version) and do not DROP those fields.

Machine Learning (ML) Detection Rules: Unsupervised Anomaly Detection

Certain adversary techniques exhibit no static signature and cannot be easily modeled with fixed mathematical thresholds. For example:

  • How many DNS requests per minute constitute an anomaly for an enterprise domain with fluctuating business hours?
  • Which parent-child process relationships are genuinely abnormal across an environment of 50,000 heterogeneous servers?
  • How does a security analyst detect compromised credentials being used at 3:00 AM by an employee who has never previously logged in outside of standard European business hours?

Machine Learning Detection Rules integrate Elastic Security with unsupervised anomaly detection jobs running on dedicated Elasticsearch ML nodes. Rather than evaluating hardcoded rules, these rules continuously monitor the outputs of statistical models that learn the baseline behavior of entities over days and weeks.

+---------------------------------------------------------------------------------------------------+
|                         MACHINE LEARNING DETECTION RULE ARCHITECTURE                              |
+---------------------------------------------------------------------------------------------------+
|                                                                                                   |
|   [ RAW TELEMETRY ] ===> [ ML ANOMALY DETECTION ENGINE ] ===> [ STATISTICAL MODEL BASELINE ]       |
|   - logs-endpoint.*       - Evaluates metric distributions     - Diurnal / weekly cyclic rhythms  |
|   - logs-network.*        - Detects rare discrete events       - Dynamic probability bands        |
|                                          |                                                        |
|                                          v                                                        |
|                           [ ANOMALY EVALUATION & SCORING ]                                        |
|                           - Anomaly Score: 0 to 100                                               |
|                           - Computes probability of occurrence                                    |
|                                          |                                                        |
|                                          v                                                        |
|                           [ ML DETECTION RULE EVALUATION ]                                        |
|                           - Threshold: Anomaly Score >= 75                                        |
|                           - Severity Mapping: Critical                                            |
|                                          |                                                        |
|                                          v                                                        |
|                        [ GENERATE SECURITY ALERT IN SIEM ]                                        |
|                        - .alerts-security.alerts-*                                                |
|                                                                                                   |
+---------------------------------------------------------------------------------------------------+

Core Prebuilt ML Security Jobs

Elastic Security provides dozens of prebuilt ML jobs mapped to specific MITRE ATT&CK tactics:

  • Rare Process Execution (rare function): Continuously builds a statistical model of executable names and file paths across endpoint hosts. When a process runs that is statistically rare across both the specific host and the broader enterprise peer group, the model records an anomaly.
  • DGA (Domain Generation Algorithm) Detection: Elastic's DGA detection package pairs a supervised model, which scores DNS query names for DGA-like character patterns in an ingest pipeline, with anomaly jobs and rules that flag hosts showing unusually high DGA activity.
  • Network Spike and Exfiltration Detection: Models normal network byte volume distributions between internal subnets and external CIDR blocks, adjusting dynamically for weekends and holidays. A sudden 50 GB transfer during non-business hours triggers an immediate anomaly.
  • Unusual System User Activity: Identifies service accounts executing interactive shells or standard end-users executing administrative utilities.

Anomaly Score (0–100) and Severity Mapping

Elastic Machine Learning assigns every detected anomaly a normalized Anomaly Score ranging from 0 to 100. The score represents the statistical probability and impact of the deviation:

Anomaly Score RangeML Anomaly RatingExample Rule Severity You Might ChooseOperational Significance
75 to 100CriticalCriticalExtreme statistical outlier. Event has a near-zero probability of occurring under normal operational baselines (e.g., mass data exfiltration or active ransomware encryption).
50 to 74Major / HighHighHighly unusual behavior that strongly deviates from established historical and peer-group patterns (e.g., rare binary execution in system directories).
25 to 49Minor / MediumMediumModerate statistical deviation. Warranted for review when correlated with other indicators or alerts on the same host.
0 to 24Warning / LowLowMinor fluctuations or low-confidence deviations. Typically retained in ML indices for trend modeling but suppressed from operational triage queues.

In the configuration of an ML Detection Rule, the analyst specifies an Anomaly Score Threshold (e.g., score >= 75). The rule periodically queries the machine learning results index (.ml-anomalies-*) and generates a security alert only when an anomaly meeting or exceeding that threshold is recorded.


Architectural Comparison: New Terms vs. ES|QL vs. Machine Learning Rules

The following table contrasts the three advanced rule types across core engineering criteria:

Feature / AttributeNew Terms Detection RulesES|QL Detection RulesMachine Learning (ML) Rules
Core Detection ParadigmNovelty tracking (first-seen terms against history).Piped vectorized transformation and math.Unsupervised statistical modeling and baselining.
Baseline RequirementHistorical index retention spanning look-back window.None (evaluates live telemetry within query window).7 to 30 days of active training data for model convergence.
Mathematical CapabilitiesNone (discrete term existence check).Extensive: Ratios, percentages, EVAL, stats, enrich.Automated probability distribution and cyclical curve fitting.
Computational OverheadModerate; inverted index terms queries.Low to Moderate; highly optimized vectorized engine.Dedicated ML node CPU/RAM allocation (model_memory_limit).
Cold-Start SensitivityHigh: Generates false positives if baseline is shallow.None: Rules execute deterministically on current data.Moderate: Requires model training phase before activation.
Primary Threat ScenariosFirst-seen user on DC, novel external IP connection.Password spraying ratios, exfiltration threshold math.DGA domain generation, rare process execution, traffic spikes.
License / Node RolesBasic tier; runs on normal data nodes.Basic tier; runs on normal data nodes.Requires Platinum or higher and at least one node with the ml role.
Loading diagram...
Advanced Detection Rule Architecture: New Terms, ES|QL & ML Anomaly Detection
Test Your Knowledge

A detection engineer is deploying a New Terms detection rule to identify when an administrative user account logs into a sensitive production server for the first time. The engineer sets the target field to 'user.name' and configures a history look-back window of 30 days. Immediately upon activating the rule, the SOC is flooded with 450 alerts for routine system administrators. What is the most likely root cause of this alert storm?

A

New Terms rules cannot track user identities and must only be bound to IP address fields.

B

The underlying cluster indices only contained 3 days of historical telemetry, causing the engine to classify all active administrators as newly observed terms.

C

The rule was configured without an unsupervised Machine Learning job running on dedicated ML nodes.

D

The engineer failed to define an ES|QL STATS aggregation pipeline to calculate term variance.

Test Your Knowledge

An enterprise SOC utilizes Elastic Machine Learning to detect algorithmically generated domains (DGA) used by malware for command-and-control beaconing. The ML model outputs continuous anomaly scores between 0 and 100 into the '.ml-anomalies-*' index. The SOC lead wants an Elastic Security detection rule that generates alerts only for high-confidence anomalies representing severe statistical deviations from normal DNS traffic. Which Anomaly Score threshold and severity mapping should the engineer configure?

A

Anomaly score threshold >= 10 with severity mapped to Informational.

B

Anomaly score threshold >= 30 with severity mapped to Low.

C

Anomaly score threshold >= 50 with severity mapped to Warning.

D

Anomaly score threshold >= 75 with severity mapped to Critical (or High).

Test Your Knowledge

A detection engineering team needs to author a detection rule that identifies distributed brute-force attacks against cloud identity services. The rule must calculate the percentage ratio of failed authentication attempts to total attempts for each source IP address over a rolling 15-minute window, alerting only when total attempts exceed 50 and the failure ratio exceeds 80%. Which Elastic Security rule type provides native support for this in-line mathematical calculation?

A

An ES|QL Detection Rule that uses EVAL with CASE to flag failures, STATS ... BY source.ip to total them, and a post-aggregation WHERE on the calculated ratio.

B

A standard Custom Query rule using Lucene wildcards and boolean NOT operators.

C

An Indicator Match rule mapped against AlienVault OTX pulses.

D

A legacy Kibana Canvas workpad scheduled via cron.

Sections you finish are checked off in the contents.