9.4 Monitoring Pipelines: Dataflow Job UI, Cloud Logging & Cloud Monitoring

Key Takeaways

  • The Dataflow job UI's job graph, execution details, job metrics, autoscaling charts, and pipeline logs show where a job is slow or failing.

  • Rising data freshness, system latency, and backlog together mean input is outpacing processing; check for a single bottleneck step before adding workers.

  • Cloud Logging supports log queries, log-based metrics, and sinks that route logs to BigQuery, Cloud Storage, or Pub/Sub.

  • INFORMATION_SCHEMA.JOBS answers BigQuery usage questions such as which queries consumed the most slot time or bytes billed.

  • Cloud Monitoring alerting policies pair a metric condition, such as oldest_unacked_message_age above 900 seconds for 5 minutes, with notification channels.

Last updated: October 2026

9.4 Monitoring Pipelines: Dataflow Job UI, Cloud Logging & Cloud Monitoring

Core Focus: The exam guide asks you to monitor Dataflow pipeline progress using the Dataflow job UI and to review and analyze logs in Cloud Logging and Cloud Monitoring. A pipeline that runs without observation fails silently; this section shows where to look and what each signal means.


The Dataflow Monitoring Interface

Open Dataflow > Jobs in the console to see running, completed, and failed jobs (with jobs from the last 30 days), then open a job:

ViewWhat it showsUse it to
Job graphThe pipeline's steps and stages, with a job summary, job logs, and per-step detailsSee which step is slow or failing and read its throughput
Execution detailsStages over time, data freshness for streaming jobs, worker progress for batch jobsFind the stage that is holding back the pipeline
Job metricsCharts such as throughput, CPU utilization, data freshness, system latency, backlog, and I/OWatch trends and spot spikes
AutoscalingCurrent and target workers and the reason for scaling decisionsTell whether the job is short of workers or capped by max_num_workers or quota
Estimated costCost from resource usage so farCatch an expensive runaway job
RecommendationsSuggestions for performance, cost, and errorsFix common misconfigurations
Pipeline logsJob logs from the service and worker logs from your codeRead the stack trace behind a failure
Data samplingSampled elements at each stepCheck what malformed input actually looks like

Reading Streaming Health

  • Data freshness is how far behind real time the pipeline's output is, based on the oldest data still being processed.
  • System latency is how long elements wait to be processed.
  • Backlog (for example, unacknowledged Pub/Sub messages) shows input piling up.

Rising freshness, latency, and backlog together mean input is arriving faster than the pipeline can process it. Before adding workers, check whether one step is the bottleneck: a slow external API or a sink at its quota creates backpressure on upstream steps, and more workers will not help. Dataflow also offers turnkey alerts for streaming jobs that monitor backlog and resource use.


Cloud Logging: Finding What Went Wrong

All the data services write logs to Cloud Logging: Dataflow job and worker logs, Composer (Airflow) task logs, Dataform invocation logs, and Cloud Audit Logs for BigQuery jobs.

resource.type="dataflow_step"
resource.labels.job_id="2026-10-08_12_34_56-1234567890"
severity>=ERROR
resource.type="bigquery_project"
protoPayload.methodName="google.cloud.bigquery.v2.JobService.InsertJob"
severity=ERROR
FeaturePurpose
Logs ExplorerSearch logs with the logging query language by resource, severity, job ID, or text
Log-based metricsCount matching log entries (for example, "dead-letter writes") to chart and alert on
Log sinksRoute logs to BigQuery (analysis with SQL), Cloud Storage (cheap long-term retention), or Pub/Sub (other tools)
Log buckets and retentionKeep logs longer than the default for audits

For BigQuery itself, the INFORMATION_SCHEMA.JOBS views are often faster than logs for usage questions:

-- The 10 queries that consumed the most slot time in the last 7 days
SELECT user_email, job_id, total_slot_ms, total_bytes_billed, query
FROM `region-us`.INFORMATION_SCHEMA.JOBS
WHERE creation_time >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
  AND job_type = 'QUERY'
ORDER BY total_slot_ms DESC
LIMIT 10;

Cloud Monitoring: Dashboards and Alerts

Cloud Monitoring collects metrics from every service and turns them into dashboards and alerting policies.

MetricWhat it tells you
pubsub.googleapis.com/subscription/num_undelivered_messagesHow many messages are waiting in a subscription
pubsub.googleapis.com/subscription/oldest_unacked_message_ageHow old the oldest waiting message is, in seconds (a direct freshness SLA signal)
dataflow.googleapis.com/job/system_lagThe longest time an item has waited to be processed in a streaming job
dataflow.googleapis.com/job/data_watermark_ageHow far the pipeline's data watermark is behind real time
bigquery.googleapis.com/slots/allocatedSlots currently allocated to a project or reservation

An alerting policy combines a condition (for example, oldest_unacked_message_age above 900 seconds for 5 minutes) with notification channels such as email, SMS, Slack, PagerDuty, Pub/Sub, or webhooks. Conditions can be built in the console or written in PromQL. Uptime checks and log-based alerts cover services and log events that have no metric.


Troubleshooting Playbook

SymptomWhere to lookLikely fix
A batch Dataflow job failedJob graph, then worker logs for the failing stepFix the code or bad input; add a dead-letter output
Streaming freshness keeps growingExecution details and job metrics; check one step's throughputRemove the bottleneck (batch API calls, raise sink quota) or raise max_num_workers
Pub/Sub backlog age climbingMonitoring chart for oldest_unacked_message_ageScale or fix the subscriber; check for poison messages
A scheduled query stopped updating a tableScheduled queries run history; Cloud Logging for the transfer runFix permissions (service account), quota, or SQL errors
BigQuery costs jumpedINFORMATION_SCHEMA.JOBS by user and bytes billedAdd partition filters, maximum_bytes_billed, or materialized views

Common Exam Traps

  • Scaling workers to fix backpressure: If one sink or API limits throughput, more workers only pile up more waiting work. Fix the bottleneck.
  • Confusing system lag with per-element processing time: Lag measures how long the oldest item has been waiting, not how long your code takes per element.
  • Keeping logs only in the default bucket: Default retention is short for Data Access logs (30 days); route logs to BigQuery or Cloud Storage with a sink when audits need longer.
  • Alerting on a single spike: Use a duration window (for example, 5 minutes) so alerts fire on sustained problems, not momentary blips.
Test Your Knowledge

An on-call data practitioner is investigating a severe degradation alert on a mission-critical Cloud Dataflow streaming pipeline consuming transaction messages from Cloud Pub/Sub. The Cloud Monitoring dashboard indicates that Pub/Sub oldest unacknowledged message age is climbing rapidly, and the Dataflow monitoring console shows high System Lag with red visual indicators on a stage performing external REST API enrichments. What does this diagnostic data reveal, and how should it be addressed?

A

The Apache Beam SDK has encountered a fatal memory leak, so the streaming pipeline must be converted to a batch pipeline.

B

The external REST API is throttling the enrichment step, causing backpressure that inflates system lag; cache, batch, or decouple those calls.

C

Dataflow worker VMs have run out of disk space, so workers must be switched from SSD persistent disks to standard persistent disks.

D

The Pub/Sub subscription has reached its message storage quota, so the topic should be deleted and recreated to clear the backlog.

Test Your Knowledge

A batch Dataflow job failed overnight. Where should a practitioner look first in the Dataflow monitoring interface to find the exception raised by the pipeline code?

A

The autoscaling charts showing worker counts over time

B

The estimated cost panel for the failed job

C

The project-level monitoring dashboard for all Dataflow jobs

D

The failing step in the job graph, then its worker logs

Test Your Knowledge

An on-call team must be paged when any message in a Pub/Sub subscription has waited longer than 15 minutes, sustained for 5 minutes. What should they configure?

A

A longer acknowledgement deadline on the subscription so messages wait less often

B

A BigQuery scheduled query that runs every minute and emails the on-call engineer

C

An alerting policy on oldest_unacked_message_age above 900 seconds for 5 minutes

D

A log sink that exports all Pub/Sub logs for the subscription to Cloud Storage

Sections you finish are checked off in the contents.