9.4 Monitoring Pipelines: Dataflow Job UI, Cloud Logging & Cloud Monitoring
Key Takeaways
The Dataflow job UI's job graph, execution details, job metrics, autoscaling charts, and pipeline logs show where a job is slow or failing.
Rising data freshness, system latency, and backlog together mean input is outpacing processing; check for a single bottleneck step before adding workers.
Cloud Logging supports log queries, log-based metrics, and sinks that route logs to BigQuery, Cloud Storage, or Pub/Sub.
INFORMATION_SCHEMA.JOBS answers BigQuery usage questions such as which queries consumed the most slot time or bytes billed.
Cloud Monitoring alerting policies pair a metric condition, such as oldest_unacked_message_age above 900 seconds for 5 minutes, with notification channels.
9.4 Monitoring Pipelines: Dataflow Job UI, Cloud Logging & Cloud Monitoring
Core Focus: The exam guide asks you to monitor Dataflow pipeline progress using the Dataflow job UI and to review and analyze logs in Cloud Logging and Cloud Monitoring. A pipeline that runs without observation fails silently; this section shows where to look and what each signal means.
The Dataflow Monitoring Interface
Open Dataflow > Jobs in the console to see running, completed, and failed jobs (with jobs from the last 30 days), then open a job:
| View | What it shows | Use it to |
|---|---|---|
| Job graph | The pipeline's steps and stages, with a job summary, job logs, and per-step details | See which step is slow or failing and read its throughput |
| Execution details | Stages over time, data freshness for streaming jobs, worker progress for batch jobs | Find the stage that is holding back the pipeline |
| Job metrics | Charts such as throughput, CPU utilization, data freshness, system latency, backlog, and I/O | Watch trends and spot spikes |
| Autoscaling | Current and target workers and the reason for scaling decisions | Tell whether the job is short of workers or capped by max_num_workers or quota |
| Estimated cost | Cost from resource usage so far | Catch an expensive runaway job |
| Recommendations | Suggestions for performance, cost, and errors | Fix common misconfigurations |
| Pipeline logs | Job logs from the service and worker logs from your code | Read the stack trace behind a failure |
| Data sampling | Sampled elements at each step | Check what malformed input actually looks like |
Reading Streaming Health
- Data freshness is how far behind real time the pipeline's output is, based on the oldest data still being processed.
- System latency is how long elements wait to be processed.
- Backlog (for example, unacknowledged Pub/Sub messages) shows input piling up.
Rising freshness, latency, and backlog together mean input is arriving faster than the pipeline can process it. Before adding workers, check whether one step is the bottleneck: a slow external API or a sink at its quota creates backpressure on upstream steps, and more workers will not help. Dataflow also offers turnkey alerts for streaming jobs that monitor backlog and resource use.
Cloud Logging: Finding What Went Wrong
All the data services write logs to Cloud Logging: Dataflow job and worker logs, Composer (Airflow) task logs, Dataform invocation logs, and Cloud Audit Logs for BigQuery jobs.
resource.type="dataflow_step"
resource.labels.job_id="2026-10-08_12_34_56-1234567890"
severity>=ERROR
resource.type="bigquery_project"
protoPayload.methodName="google.cloud.bigquery.v2.JobService.InsertJob"
severity=ERROR
| Feature | Purpose |
|---|---|
| Logs Explorer | Search logs with the logging query language by resource, severity, job ID, or text |
| Log-based metrics | Count matching log entries (for example, "dead-letter writes") to chart and alert on |
| Log sinks | Route logs to BigQuery (analysis with SQL), Cloud Storage (cheap long-term retention), or Pub/Sub (other tools) |
| Log buckets and retention | Keep logs longer than the default for audits |
For BigQuery itself, the INFORMATION_SCHEMA.JOBS views are often faster than logs for usage questions:
-- The 10 queries that consumed the most slot time in the last 7 days
SELECT user_email, job_id, total_slot_ms, total_bytes_billed, query
FROM `region-us`.INFORMATION_SCHEMA.JOBS
WHERE creation_time >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
AND job_type = 'QUERY'
ORDER BY total_slot_ms DESC
LIMIT 10;
Cloud Monitoring: Dashboards and Alerts
Cloud Monitoring collects metrics from every service and turns them into dashboards and alerting policies.
| Metric | What it tells you |
|---|---|
pubsub.googleapis.com/subscription/num_undelivered_messages | How many messages are waiting in a subscription |
pubsub.googleapis.com/subscription/oldest_unacked_message_age | How old the oldest waiting message is, in seconds (a direct freshness SLA signal) |
dataflow.googleapis.com/job/system_lag | The longest time an item has waited to be processed in a streaming job |
dataflow.googleapis.com/job/data_watermark_age | How far the pipeline's data watermark is behind real time |
bigquery.googleapis.com/slots/allocated | Slots currently allocated to a project or reservation |
An alerting policy combines a condition (for example, oldest_unacked_message_age above 900 seconds for 5 minutes) with notification channels such as email, SMS, Slack, PagerDuty, Pub/Sub, or webhooks. Conditions can be built in the console or written in PromQL. Uptime checks and log-based alerts cover services and log events that have no metric.
Troubleshooting Playbook
| Symptom | Where to look | Likely fix |
|---|---|---|
| A batch Dataflow job failed | Job graph, then worker logs for the failing step | Fix the code or bad input; add a dead-letter output |
| Streaming freshness keeps growing | Execution details and job metrics; check one step's throughput | Remove the bottleneck (batch API calls, raise sink quota) or raise max_num_workers |
| Pub/Sub backlog age climbing | Monitoring chart for oldest_unacked_message_age | Scale or fix the subscriber; check for poison messages |
| A scheduled query stopped updating a table | Scheduled queries run history; Cloud Logging for the transfer run | Fix permissions (service account), quota, or SQL errors |
| BigQuery costs jumped | INFORMATION_SCHEMA.JOBS by user and bytes billed | Add partition filters, maximum_bytes_billed, or materialized views |
Common Exam Traps
- Scaling workers to fix backpressure: If one sink or API limits throughput, more workers only pile up more waiting work. Fix the bottleneck.
- Confusing system lag with per-element processing time: Lag measures how long the oldest item has been waiting, not how long your code takes per element.
- Keeping logs only in the default bucket: Default retention is short for Data Access logs (30 days); route logs to BigQuery or Cloud Storage with a sink when audits need longer.
- Alerting on a single spike: Use a duration window (for example, 5 minutes) so alerts fire on sustained problems, not momentary blips.
An on-call data practitioner is investigating a severe degradation alert on a mission-critical Cloud Dataflow streaming pipeline consuming transaction messages from Cloud Pub/Sub. The Cloud Monitoring dashboard indicates that Pub/Sub oldest unacknowledged message age is climbing rapidly, and the Dataflow monitoring console shows high System Lag with red visual indicators on a stage performing external REST API enrichments. What does this diagnostic data reveal, and how should it be addressed?
The Apache Beam SDK has encountered a fatal memory leak, so the streaming pipeline must be converted to a batch pipeline.
The external REST API is throttling the enrichment step, causing backpressure that inflates system lag; cache, batch, or decouple those calls.
Dataflow worker VMs have run out of disk space, so workers must be switched from SSD persistent disks to standard persistent disks.
The Pub/Sub subscription has reached its message storage quota, so the topic should be deleted and recreated to clear the backlog.
A batch Dataflow job failed overnight. Where should a practitioner look first in the Dataflow monitoring interface to find the exception raised by the pipeline code?
The autoscaling charts showing worker counts over time
The estimated cost panel for the failed job
The project-level monitoring dashboard for all Dataflow jobs
The failing step in the job graph, then its worker logs
An on-call team must be paged when any message in a Pub/Sub subscription has waited longer than 15 minutes, sustained for 5 minutes. What should they configure?
A longer acknowledgement deadline on the subscription so messages wait less often
A BigQuery scheduled query that runs every minute and emails the on-call engineer
An alerting policy on oldest_unacked_message_age above 900 seconds for 5 minutes
A log sink that exports all Pub/Sub logs for the subscription to Cloud Storage
Sections you finish are checked off in the contents.