11.2 Transform Processor In Action
Key Takeaways
The transform processor (available in the OpenTelemetry Collector Contrib distribution) is the primary engine executing OTTL statements across traces, metrics, and logs pipelines.
Transform processor configuration organizes statements by signal and context blocks (trace_statements, metric_statements, log_statements), executing each list sequentially.
OTTL editors such as set, replace_pattern, replace_all_patterns, delete_key, keep_keys, and merge_maps change telemetry, while converters such as ParseJSON, ExtractPatterns, Concat, IsMatch, and Int return values for arguments and conditions.
The ParseJSON() function parses stringified JSON payloads into structured maps, enabling extraction of nested properties into queryable OpenTelemetry attributes.
The error_mode setting controls behavior on statement errors: 'ignore' (the default: log and continue), 'silent' (continue without logging), and 'propagate' (return the error, which drops the payload).
11.2 Transform Processor In Action
Quick Answer: The
transformprocessor (part of the OpenTelemetry Collector Contrib distribution) is the execution engine that applies OTTL transformation statements to telemetry streams. Configured inside theprocessors:pipeline section, it organizes transformations into signal-specific lists (trace_statements,metric_statements,log_statements). In the basic style each list holds plain statements with context-qualified paths (for exampleset(span.attributes["x"], "y")), and the processor infers the context; the advanced style shown below groups statements under an explicitcontext(such asspan,datapoint, orlog) with optional conditions and a per-grouperror_mode. Key functions include the editorsreplace_pattern(PII redaction),set(attribute injection and normalization), anddelete_key/keep_keys(sanitization), and the converterParseJSON(turning stringified JSON bodies into maps).
In enterprise environments, the Collector acts as the central data gateway. Before forwarding data to multiple backend destinations (such as Datadog, Prometheus, Elasticsearch, or AWS CloudWatch), telemetry must often be transformed to meet strict compliance mandates, strip sensitive data, align with organizational naming conventions, or populate missing fields. The transform processor provides the computational power to perform these operations in-stream without needing external code or sidecars.
Architecture and Configuration of the Transform Processor
The transform processor processes batches of telemetry records as they flow through the pipeline. It parses each OTTL statement at startup, compiles it into internal Go abstract syntax trees, and executes the statements sequentially on each record.
Core YAML Configuration Structure
processors:
transform:
error_mode: ignore # ignore | silent | propagate
trace_statements:
- context: span
statements:
- set(attributes["processed.by"], "otel-collector")
metric_statements:
- context: datapoint
statements:
- set(attributes["cluster"], "us-east-k8s") where metric.name == "system.cpu.load"
log_statements:
- context: log
statements:
- set(severity_number, 17) where IsMatch(body, "(?i)error|exception")
The error_mode Directive
When evaluating statements over millions of unpredictable user payloads, occasional runtime errors occur (e.g., regex parsing failures, invalid type conversions, or accessing non-existent map elements). The error_mode setting dictates how the processor handles these conditions:
ignore(the default whenerror_modeis not set, and the recommended mode): Logs the error and continues with the next statement. Telemetry is preserved.silent: Suppresses error logging completely and proceeds to the next statement. Used in high-volume production where expected missing fields would otherwise flood collector logs.propagate: Returns the error up the pipeline, which drops the payload from the Collector. Use it when strict compliance requires that data which failed a redaction step is never exported.
Essential OTTL Function Catalog
OTTL functions come in two kinds. Editors (lowercase names) change telemetry and begin every statement. Converters (UpperCamelCase names) return a value that you pass to an editor or test in a where condition.
1. Editors: Changing Telemetry
set(target, value): Writesvaluetotarget, creating the attribute if it does not exist and overwriting it if it does.replace_pattern(target, regex, replacement): Replaces every match ofregexinside the string fieldtarget. The replacement can reference capture groups ($1); inside a Collector configuration file each$must be written as$$so that environment-variable substitution leaves it alone.replace_all_patterns(target, mode, regex, replacement): Applies the same replacement to every value (mode"value") or every key (mode"key") of a map such asattributes.delete_key(target, key): Deleteskeyfrom the maptarget(e.g.,delete_key(attributes, "authorization.token")); a missing key is a no-op.keep_keys(target, keys): Drops every key in the map except those listed, implementing strict allow-listing.merge_maps(target, source, strategy): Merges a map (often the result of a converter) intotargetusinginsert,update, orupsert.
2. Converters: Strings
Concat(values[], delimiter): Joins values with a delimiter (e.g.,Concat([attributes["env"], attributes["service"]], "-")).Split(target, delimiter): Splits a string into a list.Substring(target, start, length): Returns part of a string (byte offsets).ToLowerCase(target)/ToUpperCase(target): Change case.
3. Converters: Parsing and Extraction
ParseJSON(target): Parses a JSON string into a map; index it directly (ParseJSON(body)["user_id"]) or merge it withmerge_maps.ExtractPatterns(target, pattern): Returns a map of the named capture groups inpattern; combine it withmerge_maps(attributes, ExtractPatterns(body, "..."), "upsert").
4. Converters: Types and Predicates
Int(val)/Double(val)/String(val)/Bool(val): Convert values between types.IsMatch(target, regex): Returnstruewhen the string matches the regex; used inwhereconditions.HasPrefix(target, prefix)/HasSuffix(target, suffix): Fast string checks that return Boolean values.
Real-World OTTL Use Cases
Use Case 1: Masking PII & Sensitive Headers (PCI-DSS & GDPR)
Observability pipelines must prevent personally identifiable information (PII) and secret credentials from leaking into telemetry backends.
processors:
transform/pii_redaction:
error_mode: ignore
trace_statements:
- context: span
statements:
# 1. Mask credit card numbers (13-16 digits) with [REDACTED_PCI]
- replace_pattern(attributes["http.request.body"], "\\b(?:\\d[ -]*?){13,16}\\b", "[REDACTED_PCI]") where attributes["http.request.body"] != nil
# 2. Scrub Bearer authentication tokens from HTTP authorization headers
- replace_pattern(attributes["http.request.header.authorization"], "Bearer [A-Za-z0-9\\-\\._~\\+/]+=*", "Bearer [REDACTED_TOKEN]") where attributes["http.request.header.authorization"] != nil
# 3. Completely delete user social security numbers from span attributes
- delete_key(attributes, "user.ssn")
Use Case 2: Normalizing Legacy Attributes to Modern Semantic Conventions
When microservices transition across OpenTelemetry SDK versions, attribute names often change (for instance, older HTTP conventions used http.status_code, whereas newer semantic conventions mandate http.response.status_code).
processors:
transform/normalize_semantics:
error_mode: ignore
trace_statements:
- context: span
statements:
# Migrate http.status_code to http.response.status_code if modern key is absent
- set(attributes["http.response.status_code"], attributes["http.status_code"]) where attributes["http.response.status_code"] == nil and attributes["http.status_code"] != nil
# Delete legacy key once migrated
- delete_key(attributes, "http.status_code") where attributes["http.response.status_code"] != nil
# Ensure response status code is an Integer
- set(attributes["http.response.status_code"], Int(attributes["http.response.status_code"]) ) where attributes["http.response.status_code"] != nil
Use Case 3: Parsing Structured Fields from Unstructured Log Bodies
Legacy application frameworks frequently output JSON-formatted logs as raw text strings directly into the body field. Storing the body as a raw string makes searching and dashboarding difficult. OTTL can unpack the JSON string into top-level attributes:
processors:
transform/json_logs:
error_mode: ignore
log_statements:
- context: log
statements:
# Extract user_id and ip_address from JSON log body
# ($$ is how a regex end anchor is written inside a Collector config file)
- set(attributes["user.id"], ParseJSON(body)["user_id"]) where IsMatch(body, "^\\{.*\\}$$")
- set(attributes["client.ip"], ParseJSON(body)["client_ip"]) where IsMatch(body, "^\\{.*\\}$$")
# Extract transaction amount and cast to Float
- set(attributes["transaction.amount"], Double(ParseJSON(body)["amount"])) where IsMatch(body, "^\\{.*\\}$$")
Use Case 4: Dynamic Log Severity Assignment
Many legacy applications output all log entries with severity_number: 0 (Unspecified) or default to INFO (severity 9) even when reporting fatal exceptions. OTTL can inspect the message content and assign the appropriate OpenTelemetry severity number:
processors:
transform/log_severity:
error_mode: ignore
log_statements:
- context: log
statements:
# If body contains FATAL or CRITICAL, elevate severity to FATAL (21)
- set(severity_number, 21) where IsMatch(body, "(?i)fatal|critical")
- set(severity_text, "FATAL") where severity_number == 21
# If body contains ERROR or EXCEPTION, elevate severity to ERROR (17)
- set(severity_number, 17) where IsMatch(body, "(?i)error|exception") and severity_number < 17
- set(severity_text, "ERROR") where severity_number == 17
Production Pipeline Integration: End-to-End Collector Config
The following configuration demonstrates how the transform processor coordinates with the batch processor inside the collector pipeline:
receivers:
otlp:
protocols:
grpc: { endpoint: "0.0.0.0:4317" }
http: { endpoint: "0.0.0.0:4318" }
processors:
batch:
send_batch_size: 1024
timeout: 1s
transform/cleanup:
error_mode: ignore
trace_statements:
- context: span
statements:
- replace_pattern(attributes["db.query.text"], "(?i)password=\\S+", "password=[REDACTED]")
log_statements:
- context: log
statements:
- set(attributes["log.extracted_user"], ParseJSON(body)["user"]) where IsMatch(body, "^\\{.*\\}$$")
exporters:
otlp:
endpoint: "telemetry-backend:4317"
tls: { insecure: true }
service:
pipelines:
traces:
receivers: [otlp]
processors: [transform/cleanup, batch]
exporters: [otlp]
logs:
receivers: [otlp]
processors: [transform/cleanup, batch]
exporters: [otlp]
A security audit reveals that client credit card numbers are inadvertently being logged in distributed trace span attributes under the key 'payment.card_number'. To comply with PCI-DSS regulations, the security team mandates that any 16-digit credit card number must be replaced with the string '[REDACTED]' before telemetry is exported to any backend. Which transform processor configuration accomplishes this without dropping the span?
A trace_statements block in the span context executing: replace_pattern(attributes["payment.card_number"], "\\b\\d{16}\\b", "[REDACTED]")
A trace_statements block in the resource context executing: delete_key(attributes, "payment.card_number")
A metric_statements block in the datapoint context executing: set(attributes["payment.card_number"], "[REDACTED]")
A log_statements block in the log context executing: ParseJSON(attributes["payment.card_number"]) where attributes["payment.card_number"] != nil
An enterprise web service outputs log records where the log record body is a raw stringified JSON document: '{"user_id": "usr_4920", "client_ip": "10.0.4.12", "status": "failed"}'. The operations team wants to parse this JSON string and promote 'user_id' and 'client_ip' into dedicated top-level OpenTelemetry log attributes for rapid indexing in Elasticsearch. Which OTTL function should be used inside the log context statements?
ExtractPatterns(body, "user_id=(.*)")
ParseJSON(body)
Split(body, ",")
Concat([attributes["user_id"], body], "")
A site reliability engineer configures an OTTL statement to cast an attribute string value to an integer using Int(attributes["http.response.status_code"]). However, during peak production traffic, several upstream services emit spans where 'http.response.status_code' contains non-numeric strings like 'N/A'. With error_mode set to propagate, the transform processor halts and drops entire batches of spans. Which modification resolves this issue while maintaining pipeline stability?
Change the context from span to spanevent to prevent errors from escalating to the root tracer provider
Replace the transform processor with an SDK BatchSpanProcessor in every upstream application runtime
Change error_mode to ignore (or silent) and guard the statement with a condition like: where IsMatch(attributes["http.response.status_code"], "^[0-9]+$")
Wrap the attribute key in single quotes instead of double quotes to force automatic type coercion in OTTL
Sections you finish are checked off in the contents.