4.2 Kibana Query Language (KQL) & Lucene for SecOps

Key Takeaways

  • KQL is the default query language in Kibana, providing field and value autocomplete, simple boolean operators, and no regex or fuzzy syntax (which avoids some very costly query patterns).

  • KQL distinguishes between exact matching on keyword fields and analyzed tokenized matching on text fields, requiring analysts to understand ECS multi-field conventions.

  • Unlike Lucene, KQL boolean operators (and, or, not) are case-insensitive and bind with explicit precedence rules where 'and' takes priority over 'or'.

  • Multi-value grouping in KQL allows concise querying across multiple parameters (e.g., source.ip: (10.0.0.1 or 10.0.0.2)) without repeatedly referencing the field name.

  • On keyword fields, unanchored leading wildcards (string) cannot seek into the term dictionary and must scan every unique term; ECS maps process.command_line as a wildcard field that is built for such patterns.

Last updated: September 2026

Efficient threat hunting and alert triage require precision query construction. When scouring terabytes of streaming security telemetry, an imprecise query returns overwhelming noise, whereas an inefficient query exhausts Elasticsearch cluster compute resources. Kibana provides two primary query syntaxes in Discover and Elastic Security: Kibana Query Language (KQL) and traditional Lucene Query Syntax.

While Lucene remains accessible for legacy queries and advanced regular expressions, KQL is the default language across modern Elastic Security workflows. Mastering KQL syntax mechanics, schema awareness, and operator precedence is an indispensable operational skill for any SIEM analyst.


KQL Syntax Architecture for Security Analysts

KQL was developed to overcome the usability pitfalls of raw Lucene. It understands the underlying Elasticsearch data types defined in the active Data View, offering automatic syntax validation and contextual autocomplete for ECS field names and observed values.

Free Text vs. Field-Specific Queries

  • Free Text Search: Entering raw strings without field delimiters (e.g., mimikatz or unauthorized) instructs Elasticsearch to search across all analyzed text fields or the consolidated multi-field string representations. While helpful for broad, exploratory triage, free text searching is computationally expensive and frequently yields false positives across benign log segments.
  • Field-Specific Queries: Prepending an ECS field name followed by a colon (e.g., user.name: "admin") restricts evaluation to the specified field index, delivering sub-second response times and deterministic accuracy.

Keyword, Text and Wildcard Field Semantics

ECS uses three string field types, and KQL matches each one differently:

Field typeHow it is indexedECS examplesHow a KQL value matches
keywordOne exact, case-sensitive term per valueprocess.name, process.executable, user.name, host.nameThe whole value exactly, or a wildcard pattern such as process.name: power*
text / match_only_textAnalyzed into lower-case tokensmessage, process.command_line.textAny token; quote a phrase to require the words in order
wildcardOptimized for pattern matching on long, varied valuesprocess.command_line, url.fullThe whole value or a wildcard pattern such as *Invoke-Mimikatz*

Examples:

  • process.executable: "C:\\Windows\\System32\\cmd.exe" must match the stored keyword value exactly, including case.
  • message: "failed password" matches documents whose message contains the phrase failed password. Without quotes, message: failed password matches both words in any order.
  • process.command_line: *-EncodedCommand* works well because ECS maps process.command_line as a wildcard field. For token-style search, use its .text multi-field: process.command_line.text: encodedcommand.
  • Non-ECS fields created by dynamic mapping usually become text with a .keyword sub-field (my_app.msg and my_app.msg.keyword). Use the .keyword version for exact matches, sorting and aggregations.

Wildcard Matching Mechanics

KQL wildcard syntax supports only * (zero or more characters). A wildcard only works when the value is not in quotes; inside quotes, * is a literal character.

  • Trailing Wildcards (Prefix Matching): process.name: power* matches powershell.exe, powerpnt.exe, and powertoy.exe. Prefix matching is efficient because Elasticsearch can seek straight to the prefix in the sorted term dictionary.
  • Middle Wildcards: file.path: /home/*/authorized_keys matches paths across any user home directory.
  • Leading Wildcards: process.name: *powershell.exe matches any process ending in powershell.exe. On a keyword field, an unanchored leading wildcard cannot seek into the term dictionary, so Elasticsearch must test every unique term in every shard. The query:allowLeadingWildcards advanced setting controls whether KQL accepts them at all. Fields mapped as wildcard, such as process.command_line, are built for these patterns and handle them much more efficiently.

Boolean Logic and Operator Precedence

KQL supports three core logical operators: and, or, and not.

  1. Case-Insensitivity: Unlike Lucene, which mandates strict uppercase AND, OR, and NOT, KQL operators are case-insensitive (and, AND, And are identical).
  2. Precedence Hierarchy: In KQL, and takes precedence over or. Consider the following query:
event.category: network and destination.port: 443 or destination.port: 80

Because and binds more tightly than or, Elasticsearch interprets this as:

(event.category: network and destination.port: 443) or (destination.port: 80)

This unintended query returns all network events on port 443, plus every event in the entire database where destination port is 80 (including file, process, and authentication records). To enforce the desired logic, explicit parenthetical grouping is mandatory:

event.category: network and (destination.port: 443 or destination.port: 80)
  1. Negation with not: The not operator negates the condition immediately following it:
process.name: "cmd.exe" and not user.name: ("SYSTEM" or "LOCAL SERVICE")

Range Queries and Comparison Operators

KQL uses intuitive mathematical comparison operators (>=, <=, >, <) for numeric, date, and IP address ranges:

  • Numeric port boundaries: destination.port >= 1024 and destination.port <= 65535
  • Network packet thresholds: network.bytes > 10485760 (traffic exceeding 10 MB)
  • Relative date math: @timestamp >= now-1h and @timestamp < now

Multi-Value Field Grouping

KQL provides an abbreviated grouping syntax that eliminates repetitive field declarations when matching multiple values against the same field:

// Standard verbose syntax:
source.ip: 10.0.0.1 or source.ip: 10.0.0.2 or source.ip: 10.0.0.3

// Streamlined KQL multi-value syntax:
source.ip: (10.0.0.1 or 10.0.0.2 or 10.0.0.3)

When applied to ECS array fields (such as related.ip or threat.technique.id), KQL matches if any value in the array satisfies the condition.


KQL vs. Lucene Syntax Comparison

While KQL is the default, analysts can toggle the query bar in Kibana to use classic Lucene syntax. Lucene provides low-level control but introduces strict syntax requirements and hazards.

+--------------------------------------------------------------------------------+
| FEATURE / BEHAVIOR          | KIBANA QUERY LANGUAGE (KQL) | LUCENE CLASSIC     |
+-----------------------------+-----------------------------+--------------------+
| Boolean Casing              | and, or, not (Any case)     | AND, OR, NOT (UPPER)|
| Range Queries               | port >= 1024 and port <= 5k | port:[1024 TO 5000]|
| Exclusive Range             | port > 1024 and port < 5000 | port:{1024 TO 5000}|
| Field Existence             | user.name: *                | _exists_:user.name |
| Regular Expressions         | Not natively supported      | /.*powershell.*/   |
| Fuzzy String Search         | Not supported               | administrator~1    |
| Proximity Searching         | Not supported               | "failed login"~3   |
| Type-Ahead Autocomplete     | Fully schema-aware          | Limited / None     |
+--------------------------------------------------------------------------------+

Side-by-Side Syntax Table for Common SecOps Query Patterns

The following table illustrates identical operational objectives formulated in KQL versus Lucene:

Operational Hunting ObjectiveKQL SyntaxLucene Syntax
Failed Logins by Userevent.category: authentication and event.outcome: failure and user.name: "admin"event.category:authentication AND event.outcome:failure AND user.name:"admin"
Restricted Port Rangedestination.port >= 1024 and destination.port <= 49151destination.port:[1024 TO 49151]
Multiple Suspicious Binariesprocess.name: ("powershell.exe" or "cmd.exe" or "wscript.exe")process.name:("powershell.exe" OR "cmd.exe" OR "wscript.exe")
Excluding Service Accountsnot user.name: ("SYSTEM" or "LOCAL SERVICE")NOT user.name:("SYSTEM" OR "LOCAL SERVICE")
Field Presence Checkhttp.request.referrer: *_exists_:http.request.referrer
Subnet Traffic Matchsource.ip: 192.168.1.0/24source.ip:"192.168.1.0/24"
Prefix Wildcard Processprocess.name: certutil*process.name:certutil*

Why KQL is the Standard for Elastic Security

  1. Fewer Expensive Constructs: Lucene syntax lets an analyst write regular expressions and fuzzy searches (e.g., /.*.*/ or admin~2) that can be very costly on large clusters. KQL has no regex or fuzzy syntax, so these patterns cannot be written in the KQL bar at all.
  2. Autocomplete and Syntax Checking: KQL suggests field names and values from the active data view, and the query bar flags malformed syntax before the search runs.
  3. Escape Sequence Simplicity: Lucene treats special characters as reserved symbols requiring backslash escaping. KQL requires minimal escaping, substantially reducing syntax errors during high-stress triage.

SOC Query Performance Pitfalls & Optimization Guidelines

  1. Avoid Unanchored Leading Wildcards on Keyword Fields: A query like file.name: *mimikatz* forces Elasticsearch to test every unique term of that keyword field in every shard. ECS maps process.command_line as a wildcard field, which is designed for *term* patterns, but it still pays to narrow the search with structured ECS filters first:
    // Optimized: Restrict to process execution events on Windows endpoints first
    event.category: process and host.os.family: windows and process.name: powershell.exe and process.command_line: *Invoke-Mimikatz*
    
  2. Use Structured Keyword Fields Over Raw Full-Text: Querying event.action: "logon-failed" against a keyword field is faster and far more precise than querying message: "failed logon" against an analyzed text field.
  3. Quote Literal Strings with Special Characters: When querying command lines, file paths, or registry keys containing colons or slashes, enclose values in double quotes (process.executable: "/tmp/.cache/loader") or escape the special characters \():<>"* with a backslash. Single quotes are not string delimiters in KQL.
Test Your Knowledge

A Tier 2 security analyst is conducting an investigation in Kibana Discover to identify potential egress communication. The analyst needs to find all network events where destination port is within the registered port range (1024 through 49151 inclusive), but wishes to explicitly exclude standard web proxy traffic on port 8080. Which KQL query correctly implements this logic?

A

destination.port:[1024 TO 49151] NOT destination.port:8080

B

destination.port >= 1024 and destination.port <= 49151 and not destination.port: 8080

C

destination.port >= 1024 or destination.port <= 49151 and destination.port != 8080

D

event.category: network and (destination.port: 1024-49151) without destination.port: 8080

Test Your Knowledge

During a threat hunting exercise across hundreds of millions of endpoint event records, an analyst observes that querying file.name: vssadmin (a keyword field) results in significant cluster search latency and shard CPU spikes. What is the root architectural cause of this performance degradation?

A

Kibana compiles KQL queries into Painless scripts that must execute inside the Java Virtual Machine for every record

B

Elasticsearch disables caching for all queries that target fields mapped under the process ECS namespace

C

The query bar enforces a strict client-side timeout that retransmits the query repeatedly across all primary nodes

D

Leading wildcards prevent Elasticsearch from seeking into the keyword field's sorted term dictionary, so it must test every unique term across shard segments

Test Your Knowledge

An analyst migrating legacy search saved objects to modern Kibana Discover wants to understand the behavioral differences between Lucene and KQL. Which statement accurately highlights an operational rule governing KQL?

A

KQL requires all logical operators to be capitalized (AND, OR, NOT) to be recognized as boolean expressions

B

KQL supports case-insensitive boolean operators (and, or, not), enforces higher precedence for 'and' over 'or', and utilizes comparative operators like '>=' for range filtering

C

KQL relies on bracket syntax like '[1024 TO 65535]' for numeric range queries and uses 'exists' for field presence

D

KQL supports advanced Lucene fuzzy syntax ('term~2') and unbounded regex execution directly within the query bar

Sections you finish are checked off in the contents.