3.2 Ingest Pipelines, Processors & Normalization Workflows

Key Takeaways

  • Ingest pipelines execute pre-indexing transformations directly on Elasticsearch ingest nodes, converting raw log strings into structured ECS documents.

  • The dissect processor provides high-throughput delimiter splitting for structured logs, whereas grok uses regular expression patterns to parse complex, variable log formats.

  • Conditional processor execution via Painless 'if' expressions ensures targeted execution and prevents runtime exceptions on missing fields.

  • Production SecOps pipelines require defensive design, utilizing 'ignore_missing: true' and 'on_failure' handlers to avoid dropping unparsed documents.

  • The _simulate API enables safe, in-memory validation of ingest pipelines against sample log payloads prior to production deployment.

Last updated: September 2026

Ingest Pipeline Architecture in the Elastic Stack

In modern SecOps architectures, log parsing and schema normalization occur directly within Elasticsearch using Ingest Pipelines. An ingest pipeline is a configured sequence of processors running on Elasticsearch ingest nodes that intercept, parse, enrich, and transform documents before they are written to a data stream or index.

An Ingest Pipeline definition contains three core elements:

  1. description: Operational context explaining the pipeline's purpose.
  2. processors: An ordered array of individual transformation steps executed sequentially on each incoming document.
  3. on_failure: A fallback array of processors executed if an unhandled exception occurs in the primary processor sequence.
PUT _ingest/pipeline/logs-cisco_asa-normalization
{
  "description": "Parses and normalizes Cisco ASA firewall syslog to ECS",
  "processors": [
    {
      "grok": {
        "field": "message",
        "patterns": ["%ASA-%{INT:log.syslog.severity.code}-%{INT:cisco.asa.message_id}: %{GREEDYDATA:cisco.asa.message}"]
      }
    },
    {
      "set": {
        "field": "event.category",
        "value": "network"
      }
    }
  ],
  "on_failure": [
    {
      "append": {
        "field": "tags",
        "value": ["_cisco_asa_parse_failure"]
      }
    }
  ]
}

Conditional Execution with Painless (if Clauses)

Processors can execute conditionally using the if parameter, which evaluates an inline Painless script. The document context is accessible via the ctx map:

  • Field existence checks: "if": "ctx.host?.name != null"
  • Dataset matching: "if": "ctx.event?.dataset == 'cisco.asa' && ctx.cisco?.asa?.message_id == '106023'"
  • Regex pattern evaluation: "if": "ctx.message =~ /^(?i)failed password/"

Using if conditions prevents unnecessary processing overhead, eliminates runtime exceptions caused by missing keys, and enables modular, multi-source pipeline routing.

Pipeline Validation via the _simulate API

Before attaching a pipeline to a production data stream, engineers validate its behavior using the _simulate API. This runs the pipeline in memory against test payloads without writing documents to disk:

POST _ingest/pipeline/_simulate
{
  "pipeline": {
    "processors": [
      {
        "dissect": {
          "field": "message",
          "pattern": "%{syslog_timestamp} %{host.name} %{process.name}[%{process.pid}]: %{log.message}"
        }
      }
    ]
  },
  "docs": [
    {
      "_source": {
        "message": "Sep 30 10:14:22 srv-dc01 sshd[4192]: Failed password for invalid user admin from 192.168.10.50 port 54822"
      }
    }
  ]
}

The API response returns the resulting transformed document or highlights the exact processor that failed, including the underlying error message and stack trace.


Core Processors in SecOps Normalization

Security pipelines leverage specialized processors to convert raw strings into validated ECS fields.

Parsing: dissect vs. grok

Parsing unstructured log lines is the most resource-intensive phase of ingestion. Analysts must evaluate the trade-offs between dissect and grok:

Characteristicdissect Processorgrok Processor
MechanismDelimiter-based string splitting (no regex)Regular expression pattern matching with regex libraries
Resource UtilizationExtremely low CPU and memory footprintModerate to high CPU overhead due to regex backtracking
Format FlexibilityStrict; requires predictable, fixed delimiter structuresHigh; handles variable tokens, optional fields, and alternate formats
Type CoercionExtracts all matched tokens as stringsSupports inline data typing (e.g., %{NUMBER:source.port:long})
Primary Use CasesStructured syslog, CEF, W3C web access logs, CSVComplex multi-format logs, Linux auditd, arbitrary app logs
  • dissect Syntax: Matches field names enclosed in %{} separated by static characters:
    {
      "dissect": {
        "field": "message",
        "pattern": "%{timestamp} %{host.name} %{process.name}: %{event.original}"
      }
    }
    
  • grok Syntax: Utilizes built-in ECS pattern libraries:
    {
      "grok": {
        "field": "message",
        "patterns": [
          "%{IP:source.ip}:%{POSINT:source.port:long} -> %{IP:destination.ip}:%{POSINT:destination.port:long}",
          "Connection from %{IP:source.ip}"
        ]
      }
    }
    

Field Transformation Processors

  • set: Hardcodes ECS fields or copies values between attributes. Supports dynamic Mustache templating:
    {
      "set": {
        "field": "event.kind",
        "value": "event"
      }
    }
    
  • rename: Maps proprietary vendor keys to standard ECS naming. Always configure "ignore_missing": true to handle optional vendor fields:
    {
      "rename": {
        "field": "src_ip",
        "target_field": "source.ip",
        "ignore_missing": true
      }
    }
    
  • convert: Coerces data types (e.g., converting port strings to integers or byte counters to long):
    {
      "convert": {
        "field": "destination.port",
        "type": "long",
        "ignore_missing": true
      }
    }
    
  • remove: Strips intermediate staging fields, redundant raw payloads, or sensitive authentication tokens:
    {
      "remove": {
        "field": ["temp_token", "unparsed_header"],
        "ignore_missing": true
      }
    }
    

Enrichment Processors: date, geoip, and user_agent

  • date: Normalizes vendor timestamps into the standardized ISO-8601 @timestamp field. If omitted, @timestamp reflects indexing time rather than event occurrence, distorting timeline analysis:
    {
      "date": {
        "field": "vendor_time",
        "target_field": "@timestamp",
        "formats": ["yyyy-MM-dd HH:mm:ss", "UNIX_MS", "ISO8601"],
        "timezone": "UTC"
      }
    }
    
  • geoip: Enriches IP addresses with geographic coordinates, country, city, and Autonomous System (ASN) metadata from GeoLite2 databases that Elasticsearch downloads and keeps updated automatically:
    {
      "geoip": {
        "field": "source.ip",
        "target_field": "source.geo",
        "ignore_missing": true
      }
    }
    
  • user_agent: Parses raw HTTP User-Agent strings, extracting browser engine, operating system, and device classifications into user_agent.*.

Custom Attributes: labels.* vs. tags

SecOps workflows frequently require custom metadata during ingestion:

  • labels: A key-value dictionary of string pairs (e.g., labels.environment: "production", labels.compliance_tier: "pci-dss"). Used for structured filtering and organizational grouping.
  • tags: A flat array of keyword strings used as operational markers (e.g., tags: ["forwarded", "vpn_gateway", "_grokparsefailure"]).

Fault Tolerance and Pipeline Error Handling

In high-throughput SOC ingestion, malformed log lines are inevitable. An unhandled exception stops the pipeline, and the document is not indexed. Dropping security logs creates critical visibility blind spots.

Defensive pipeline architecture incorporates two layers of error management:

  1. Processor-Level on_failure: Catches errors from a single processor so the rest of the pipeline can continue (the example tags any geoip lookup error rather than failing the document):
    {
      "geoip": {
        "field": "source.ip",
        "target_field": "source.geo",
        "on_failure": [
          {
            "append": {
              "field": "tags",
              "value": ["_geoip_lookup_failure"]
            }
          }
        ]
      }
    }
    
  2. Pipeline-Level on_failure: Captures any unhandled processor failure, extracts the error message, tags the document, and guarantees indexing so the raw message (event.original) remains retrievable for forensic investigation:
    "on_failure": [
      {
        "set": {
          "field": "error.message",
          "value": "{{ _ingest.on_failure_message }}"
        }
      },
      {
        "append": {
          "field": "tags",
          "value": ["_pipeline_failure"]
        }
      }
    ]
    
Loading diagram...
Ingest Pipeline Transformation & Error Handling Workflow
Test Your Knowledge

A security engineer needs to normalize hundreds of thousands of structured web proxy access logs per second that follow a strict, space-delimited format. Why is the dissect processor preferred over the grok processor for this workload?

A

dissect performs deterministic delimiter-based splitting without regex engine overhead, delivering substantially higher throughput and lower CPU utilization.

B

dissect automatically enriches IP addresses with GeoIP location data during string extraction.

C

dissect automatically coerces all extracted numeric strings into long integers.

D

grok cannot extract timestamp fields into standard ISO-8601 formats.

Test Your Knowledge

An ingest pipeline fails intermittently when processing legacy application logs because some records omit the client_ip field, causing downstream processors to throw exceptions. How can the engineer prevent these documents from being dropped?

A

Enable cluster-wide lenient ingestion in the elasticsearch.yml configuration.

B

Insert a Painless script that generates random IP addresses when fields are missing.

C

Add 'ignore_missing: true' to the affected processors and define an on_failure block that appends an error tag.

D

Replace the rename processor with a convert processor set to type 'ip'.

Test Your Knowledge

Which Elasticsearch API should a security analyst use to test an ingest pipeline's grok patterns and Painless conditionals against real log samples in memory without writing any data to persistent storage?

A

POST _cluster/nodes/hot_threads

B

POST _ingest/pipeline/_simulate

C

GET _security/role_mapping/_verify

D

POST _scripts/painless/_execute

Sections you finish are checked off in the contents.