8.3 Stream Processing: Windowing, Watermarks, and Triggers
Key Takeaways
Stream processing distinguishes between Event Time (the moment a data event occurred at its origin) and Processing Time (the moment the event reaches the processing pipeline runner).
Apache Beam provides four primary windowing strategies: Fixed (Tumbling) windows for non-overlapping intervals, Sliding (Hopping) windows for overlapping rolling metrics, Session windows for variable intervals based on inactivity gaps, and Global window for all-time aggregation.
The Watermark is Dataflow's monotonically advancing metric that estimates progress in Event Time, representing the system's confidence that all data up to that event timestamp has been observed.
Late-arriving data occurs when events have timestamps older than the current watermark; Dataflow manages this via Allowed Lateness, customizable triggers (early, on-time, late), and accumulation modes (accumulating vs. discarding).
The Dead-Letter Queue (DLQ) design pattern isolates unparseable, malformed, or poisoned messages by writing them to a separate sink (such as a Pub/Sub topic or Cloud Storage bucket) to prevent catastrophic pipeline stalls.
Stream Processing: Windowing, Watermarks, and Triggers
Core Focus: Processing unbounded streaming datasets requires robust temporal reasoning. In distributed systems, data rarely arrives in the order it occurred. Mastering the distinction between Event Time and Processing Time, configuring windowing strategies, tracking event completeness via Watermarks, managing late data with Triggers and Allowed Lateness, and safeguarding pipelines with Dead-Letter Queues (DLQs) are essential skills for Google Cloud data practitioners.
In batch processing, the dataset is finite and bounded. The processing engine reads all records, groups them by key, and calculates aggregations with a known start and end. In streaming processing, data is continuous, infinite, and unbounded. A streaming system cannot wait for "all data to arrive" before calculating a sum or average, because the data stream never ends.
To make sense of unbounded streams, processing systems must partition data into finite temporal chunks. This process is called Windowing.
Temporal Foundations: Event Time vs. Processing Time
Understanding the two domains of time is the foundational prerequisite for stream processing:
+-----------------------------------------------------------------------------------------+
| Event Time vs. Processing Time Skew |
+-----------------------------------------------------------------------------------------+
| [Event Occurs] [Record Processed] |
| Timestamp: 12:01:15 Timestamp: 12:05:45 |
| Source: Mobile App / IoT Sensor Runner: Cloud Dataflow |
| | ^ |
| +--------------> Network Delay / Offline Buffering --------------->+ |
| |
| Event Time = 12:01:15 Processing Time = 12:05:45|
| |
| Time Skew = Processing Time - Event Time = 4 minutes 30 seconds |
+-----------------------------------------------------------------------------------------+
- Event Time: The time at which the event actually occurred on the client device, sensor, or application. This timestamp is generated by the event producer and embedded directly in the data payload (e.g.,
"timestamp": "2026-10-08T14:32:01.450Z"). - Processing Time: The time at which the event is observed and processed by the execution engine (Cloud Dataflow worker). This timestamp is determined by the system clock of the processing infrastructure.
Why Processing Time Fails for Business Logic
In an ideal world with zero network latency, Event Time and Processing Time would be identical. However, in distributed real-world systems, network congestion, cellular dead zones, device offline buffering, and service restarts introduce variable delays.
If a mobile player scores 500 points in an online game at 12:01 PM (Event Time), enters a subway tunnel with no reception, and re-establishes connectivity at 12:45 PM (Processing Time), calculating hourly leaderboard bonuses using Processing Time would assign the player's score to the wrong hour. Correct stream analytics must group and aggregate data according to Event Time.
Windowing Strategies
Windowing divides an unbounded PCollection into finite logical slices based on the timestamps of individual elements. Apache Beam provides four primary windowing strategies:
+-----------------------------------------------------------------------------------------+
| Beam Windowing Strategies |
+-----------------------------------------------------------------------------------------+
| 1. Fixed (Tumbling) Windows: Non-overlapping, uniform intervals |
| [--- Window 1 (0-5m) ---][--- Window 2 (5-10m) ---][--- Window 3 (10-15m) ---] |
| |
| 2. Sliding (Hopping) Windows: Overlapping, periodic intervals |
| [------ Window A (0-10m) ------] |
| [------ Window B (2-12m) ------] |
| [------ Window C (4-14m) ------] |
| |
| 3. Session Windows: Dynamic per-key intervals bounded by inactivity gaps |
| User 1: [--- Session A (30m) ---] [--- Session B (15m) ---] |
| User 2: [----------- Session C (60m) -----------] |
| |
| 4. Global Window: Single window spanning all time across the entire stream |
| [----------------------------- Entire Stream Timeline ----------------------------] |
+-----------------------------------------------------------------------------------------+
1. Fixed (Tumbling) Windows
Fixed windows partition the timeline into discrete, non-overlapping, contiguous time chunks of uniform duration (e.g., 5-minute fixed windows, 1-hour fixed windows).
- Behavior: Every element belongs to exactly one window. For example, in a 5-minute fixed window, events with timestamps between
12:00:00and12:04:59.999fall into Window 1; events at12:05:00fall into Window 2. - Use Case: Periodic summary reporting, such as calculating total page views per hour or total sales revenue every 15 minutes.
from apache_beam import window
fixed_windowed_pcoll = (
raw_events
| "ApplyFixedWindow" >> beam.WindowInto(window.FixedWindows(5 * 60)) # 5 minutes
| "SumPerUser" >> beam.CombinePerKey(sum)
)
2. Sliding (Hopping) Windows
Sliding windows partition the timeline into uniform, overlapping time intervals. They are defined by two parameters: Window Duration (length of the window) and Slide Period (frequency at which a new window begins).
- Behavior: Because the slide period is shorter than the window duration, individual elements fall into multiple concurrent windows. For example, a window duration of 10 minutes with a slide period of 2 minutes creates windows spanning
[12:00-12:10),[12:02-12:12),[12:04-12:14). An event occurring at12:03is duplicated across both the first and second windows. - Use Case: Calculating continuous moving averages, rolling 24-hour stock volatility, or real-time traffic congestion rates.
sliding_windowed_pcoll = (
sensor_telemetry
| "ApplySlidingWindow" >> beam.WindowInto(
window.SlidingWindows(size=10 * 60, period=2 * 60) # 10m duration, 2m slide
)
)
3. Session Windows
Session windows define boundaries dynamically on a per-key basis. Instead of using fixed clock intervals, a session window is bounded by periods of user inactivity, known as the Gap Duration.
- Behavior: When an event arrives for a specific key (e.g.,
user_id), it is placed into an active session. If subsequent events for that same key arrive within the gap duration, the session window expands to encompass them. If no events arrive for longer than the gap duration, the session window closes. Sessions are independent for each individual key. - Use Case: Web analytics tracking user visits, mobile gaming engagement analysis, and fraud detection identifying clusters of suspicious banking transactions.
session_windowed_pcoll = (
clickstream_events
| "ApplySessionWindow" >> beam.WindowInto(
window.Sessions(gap_size=30 * 60) # 30-minute inactivity gap
)
)
4. Global Window
The Global Window is the default windowing strategy in Apache Beam. It encompasses all data across all time in a single window.
- Behavior: For bounded batch datasets, the Global Window works naturally, emitting the final aggregate when the entire dataset is read. For unbounded streams, the global window's default trigger would fire only at the end of the window, which never comes. Beam therefore rejects a
GroupByKeyorCombineon an unboundedPCollectionin the global window unless you set a non-default trigger or switch to another windowing strategy.
Watermarks: Measuring Event-Time Completeness
In an event-time streaming system, how does Cloud Dataflow know when all data for a specific window (e.g., [12:00 - 12:05)) has arrived so it can close the window and emit the final total?
This is solved by the Watermark. The watermark is a system metric that measures progress in Event Time.
- Definition: A watermark of timestamp represents the system's assertion that: "We believe we have observed all events with an event timestamp ."
- Monotonic Progression: Watermarks advance monotonically forward in time. As Dataflow observes elements flowing through ingestion connectors (such as Cloud Pub/Sub), it computes an estimated watermark based on message publish times, pending backlogs, and network latency.
- Window Closure: When the watermark advances past the end boundary of a window (e.g., the watermark reaches
12:05:01for the[12:00 - 12:05)window), Dataflow considers the window complete and triggers downstream aggregations.
Timeline (Event Time) ------------------------------------------------------------------------>
[12:00 ------------- Window 1 ------------- 12:05) | [12:05 ------------- Window 2 ------------- 12:10)
^
Watermark (Tw)
When Tw reaches 12:05:00, Window 1 fires its on-time pane.
Handling Late Data: Allowed Lateness and Triggers
In real-world networks, assumptions fail. Even with advanced watermark heuristics, an event generated at 12:03:00 might arrive after the watermark has already advanced to 12:06:00. This record is classified as Late-Arriving Data.
By default, Apache Beam discards any late-arriving data that appears after the watermark has passed the end of its window. To prevent data loss, practitioners configure Allowed Lateness and Triggers.
1. Allowed Lateness
Allowed Lateness specifies how long Dataflow should preserve window state in memory or storage after the watermark has passed the window end boundary.
- If a window spans
[12:00 - 12:05)and the allowed lateness is set to2 hours, Dataflow will retain the window's intermediate state until the watermark reaches14:05:00. - Any record arriving with an event timestamp in that window during those 2 hours will be incorporated into an updated calculation.
- Once the watermark passes
Window End + Allowed Lateness, the window state is permanently purged. Any subsequent records for that window are permanently dropped.
2. Triggers: Controlling When Panes Fire
A Trigger determines the exact conditions under which a window emits an intermediate or final result (known as a pane). Beam supports three operational trigger categories:
- On-Time Trigger: Fires automatically when the watermark passes the end of the window.
- Early (Speculative) Triggers: Fires before the watermark reaches the end of the window (e.g., emitting a speculative count every 1 minute or every 10,000 records). Useful for low-latency dashboards that need real-time approximations before all data arrives.
- Late Triggers: Fires after the watermark has passed the window end, whenever a late record arrives within the allowed lateness window.
3. Accumulation Modes
When multiple panes are emitted for the same window (due to speculative or late triggers), the practitioner specifies an Accumulation Mode:
- Accumulating Mode: Each pane includes the full cumulative total of all records seen so far in that window. (e.g., Pane 1 emits
100; late record of5arrives; Pane 2 emits105). - Discarding Mode: Each pane includes only the delta change since the previous firing. (e.g., Pane 1 emits
100; late record of5arrives; Pane 2 emits5). This is standard when writing incremental adjustments to downstream databases.
from apache_beam.transforms.trigger import (
AfterWatermark, AfterProcessingTime, Repeatedly, AccumulationMode
)
windowed_with_triggers = (
raw_stream
| "WindowWithLateness" >> beam.WindowInto(
window.FixedWindows(10 * 60), # 10-minute window
trigger=AfterWatermark(
early=AfterProcessingTime(60), # Speculative pane every 60s
late=AfterProcessingTime(30) # Late pane 30s after late event
),
allowed_lateness=3600, # Keep state for 1 hour after watermark
accumulation_mode=AccumulationMode.ACCUMULATING
)
)
The Dead-Letter Queue (DLQ) Pattern
In continuous stream processing, pipeline resilience is paramount. If a single malformed, unparseable, or schema-violating payload (a "poison pill" message) causes an unhandled exception inside a DoFn, the worker task will crash.
Because Cloud Dataflow automatically retries failed tasks to ensure fault tolerance, a corrupted record will trigger an infinite retry loop. This halts watermark progression, creates catastrophic backpressure, consumes worker CPU, and prevents all subsequent valid messages from being processed.
+-----------------------------------------------------------------------------------------+
| Dead-Letter Queue (DLQ) Architecture |
+-----------------------------------------------------------------------------------------+
| [Cloud Pub/Sub Stream] |
| | |
| v |
| [ParDo Validation / Parsing] |
| / \ |
| Valid Record / \ Corrupted / Unparseable |
| v v |
| [Main PCollection] [Tagged Side Output: DLQ] |
| | | |
| v v |
| [Sink: BigQuery] [Sink: Cloud Storage / Pub/Sub Topic] |
| | |
| v |
| [Alerting & Remediation] |
+-----------------------------------------------------------------------------------------+
Implementation Best Practices:
- Wrap in Try/Except Blocks: Never permit raw parsing logic (such as
json.loads(element)or string-to-integer casting) to execute unprotected inside aDoFn. - Route via Tagged Outputs: Catch parsing exceptions and emit the raw, unmodified record along with error metadata (stack trace, error message, ingestion timestamp) to a tagged side output.
- Write to a Durable Quarantine Sink: Direct the dead-letter
PCollectionto a dedicated Cloud Storage bucket (e.g., partitioned by date) or a dedicated Cloud Pub/Sub dead-letter topic. - Alert and Replay: Configure Cloud Monitoring alerts on the DLQ output rate. After engineering fixes the upstream schema bug or parser logic, the quarantined records can be re-injected into the pipeline for reprocessing without data loss.
A mobile gaming analytics platform captures in-game player events. Players frequently enter subway tunnels or lose cellular connectivity, causing events to buffer on client devices and arrive at the processing pipeline minutes or hours later. The pipeline must compute total in-game score bonuses grouped by each player's active gameplay bursts, closing a player's session when no activity occurs for 15 consecutive minutes. Which windowing strategy is required?
Session windows with a 15-minute gap duration
Sliding (Hopping) windows with a 15-minute duration and a 1-minute slide
Fixed (Tumbling) windows with a 15-minute duration
Global window with a Processing Time trigger set to 15 minutes
A streaming pipeline on Cloud Dataflow processes real-time financial trades using 5-minute fixed windows based on Event Time. Due to intermittent network latency between trading desks, some trade events arrive after the watermark has advanced past the corresponding window's end boundary. The engineering team wants to process these late trades and update downstream reporting dashboards rather than discarding them immediately. What combination of Beam features must be configured?
Switch the pipeline from Event Time to Processing Time evaluation
Configure a dead-letter queue on Cloud Storage to store dropped records
Extend the fixed window duration to 1 hour and disable the watermark
Set an Allowed Lateness period and configure a trigger to fire on late data
A continuous Cloud Dataflow streaming pipeline consuming from Cloud Pub/Sub encounters an unhandled JSON deserialization exception whenever a mobile client sends a corrupted payload. The unhandled exception causes worker tasks to fail repeatedly, triggering infinite retries, backpressure, and pipeline processing delays. What architectural pattern should the data engineer implement to prevent malformed records from stalling the pipeline?
Increase the Cloud Pub/Sub subscription acknowledgement deadline to 600 seconds
Increase Compute Engine worker VM memory so workers can absorb the malformed payloads without crashing
Catch the exception inside the ParDo and send the bad payloads to a dead-letter side output
Configure an early speculative trigger on the Global Window to discard unparseable events
Sections you finish are checked off in the contents.