9.3 Event-Driven Ingestion: Pub/Sub to BigQuery & Eventarc Triggers
Key Takeaways
A BigQuery subscription writes Pub/Sub messages straight into a BigQuery table through the Storage Write API, with no Dataflow job or subscriber code.
BigQuery subscriptions deliver at least once and can map fields with the topic schema or the table schema and write message metadata such as message_id and publish_time.
Use a Dataflow streaming pipeline (for example the Pub/Sub to BigQuery template) when messages need parsing, validation, enrichment, or windowed aggregation before loading.
Eventarc routes CloudEvents, such as google.cloud.storage.object.v1.finalized, to Cloud Run, Cloud Run functions, Workflows, GKE, or internal HTTP endpoints.
Dead-letter topics (5 delivery attempts by default, configurable from 5 to 100) isolate poison messages so the rest of the stream keeps flowing.
9.3 Event-Driven Ingestion: Pub/Sub to BigQuery & Eventarc Triggers
Core Focus: The exam guide asks you to identify use cases for event-driven ingestion from Pub/Sub to BigQuery and to use Eventarc triggers in event-driven pipelines. Event-driven designs start work when data arrives instead of on a clock, which cuts latency and avoids paying for jobs that find nothing to do.
Traditional pipelines run on fixed schedules such as a nightly cron job. That adds latency, wastes runs when no new data exists, and breaks when upstream producers are late. Event-driven pipelines react to a message or a file the moment it arrives.
Event-Driven vs. Time-Based Scheduling
Time-Based (Cron / Polling):
[ Clock: 02:00 UTC ] ---> (Triggers Pipeline) ---> [ Reads Cloud Storage Bucket ]
- Problem 1: If data lands at 02:05 UTC, it waits almost 24 hours.
- Problem 2: If data is late, the job processes empty or stale input.
Event-Driven (Pub/Sub / Eventarc):
[ File Lands in GCS ] ---> [ Eventarc: object finalized ] ---> (Triggers Pipeline)
- Processing starts seconds after the upload completes.
- Nothing runs (and nothing is billed) while waiting.
| Dimension | Time-Based Scheduling | Event-Driven Architecture |
|---|---|---|
| Trigger | Wall-clock time (cron) | A state change: a message published, a file finalized |
| Freshness | Bounded by the schedule interval | Seconds after the event |
| Idle cost | Jobs run even with no new data | Little or none while idle |
| Late producers | Need guessed buffer windows | Work starts when data actually arrives |
| Typical tools | Cloud Scheduler, BigQuery scheduled queries | Pub/Sub, Eventarc, Cloud Run, Cloud Run functions, Workflows |
Pub/Sub Basics
Pub/Sub is Google Cloud's managed, asynchronous messaging service. Publishers send messages to a topic; each subscription on the topic receives its own copy of every message.
| Subscription type | How messages are delivered | Typical consumer |
|---|---|---|
| Pull | The consumer requests messages and acknowledges them | Dataflow streaming jobs, custom workers |
| Push | Pub/Sub sends each message as an HTTPS POST | Cloud Run services, Cloud Run functions |
| BigQuery (export) | Pub/Sub writes messages straight into a BigQuery table | No consumer code at all |
| Cloud Storage (export) | Pub/Sub writes batches of messages to files in a bucket | Raw archives, data lake landing zones |
A dead-letter topic catches messages that fail delivery more than a configured number of attempts (5 by default, configurable from 5 to 100), so one bad message does not block the rest.
Three Ways to Get Pub/Sub Data into BigQuery
1. BigQuery Subscription (No Pipeline)
A BigQuery subscription writes each message to an existing BigQuery table as it arrives, using the Storage Write API behind the scenes. There is no Dataflow job or subscriber client to run.
- Schema options: use the topic schema (Avro or Protocol Buffers) mapped to table columns; use the table schema to map JSON message fields to matching columns; or, with neither option, write the raw message to a
datacolumn. - Metadata: optionally write the subscription name, message ID, publish time, and attributes to extra columns for auditing.
- Failures: messages that cannot be written are negatively acknowledged and retried, and move to the dead-letter topic if one is configured.
- Delivery: at least once, so a downstream view or
MERGEshould deduplicate onmessage_idor a business key if duplicates matter. - Light changes: a Single Message Transform (a small JavaScript function) can adjust a message before it is written.
- Permissions: the Pub/Sub service agent (or a service account you specify) needs permission to write to the table, such as BigQuery Data Editor on it.
gcloud pubsub subscriptions create orders-to-bq \
--topic=orders \
--bigquery-table=my-project:sales.orders_raw \
--use-table-schema \
--write-metadata \
--dead-letter-topic=orders-dlq
2. Dataflow Streaming Pipeline
Use Dataflow (for example the Google-provided Pub/Sub to BigQuery template, or your own Apache Beam pipeline) when messages need real transformation: parsing, validation, enrichment with lookups, windowed aggregation, or routing bad records to a dead-letter output. The Pub/Sub-to-BigQuery template enforces exactly-once processing by default; a BigQuery subscription is at-least-once.
3. Custom Subscriber
A Cloud Run service on a push subscription can call the Storage Write API itself. Choose this only when you need custom logic that neither option above provides.
| Requirement | Best choice |
|---|---|
| Land JSON events in a table with no transformation and no code | BigQuery subscription |
| Parse, validate, enrich, or aggregate (windows) before loading | Dataflow streaming pipeline |
| Keep a raw copy of every message as files | Cloud Storage subscription |
| Load batches of files on a schedule, not events | Load jobs or BigQuery Data Transfer Service (Section 3.1) |
Eventarc: Routing Events to Pipelines
Eventarc routes events in the open CloudEvents format from Google Cloud services, Pub/Sub topics, and other sources to a destination.
- Direct events: For example,
google.cloud.storage.object.v1.finalized, emitted when an upload to a bucket completes. Filters narrow it to one bucket. - Cloud Audit Logs events: React to API calls recorded in audit logs (for example, a BigQuery job completing or a table being created) without changing the producer.
- Pub/Sub messages: Any message published to a topic can trigger a destination.
- Destinations: Cloud Run services, Cloud Run functions (formerly Cloud Functions), Workflows, GKE services, and internal HTTP endpoints in a VPC network.
Triggering Data Services from Eventarc
The exam guide mentions Eventarc with Dataform, Dataflow, Cloud Functions, Cloud Run, and Cloud Composer. Eventarc sends events to Cloud Run, Cloud Run functions, or Workflows, and that target then starts the data service through its API:
| Goal | Eventarc destination | What the destination does |
|---|---|---|
| Run a Dataflow template for each uploaded file | Cloud Run function or Workflows | Calls the Dataflow API to launch a Flex Template with the file path |
| Refresh Dataform models when new data lands | Workflows | Compiles the repository and creates a Dataform workflow invocation |
| Start an Airflow DAG in Cloud Composer | Cloud Run function | Calls the Airflow REST API to trigger a DAG run |
| Validate and load a file into BigQuery | Cloud Run service | Checks the file, then starts a BigQuery load job |
Failure Handling in Event-Driven Pipelines
- Idempotency: Events can be delivered more than once. Use
MERGEon a business key, deterministic job IDs, or deduplicating views so a retry does not double-count. - Dead-letter handling: Route messages that keep failing to a dead-letter topic or bucket instead of retrying forever, and alert on its size.
- Backoff: Retry calls to rate-limited APIs (HTTP 429 or 503) with exponential backoff and jitter.
Common Exam Traps
- Building a Dataflow job just to copy messages into BigQuery: If no transformation is needed, a BigQuery subscription removes the pipeline entirely.
- Assuming exactly-once from a BigQuery subscription: It is at-least-once; deduplicate downstream when it matters.
- Polling a bucket: A function that lists a bucket every minute adds latency and API cost; an Eventarc trigger on object finalization starts work immediately.
- Silently dropping bad messages: Use a dead-letter topic so failures are kept, counted, and replayable.
A financial services organization requires that CSV files uploaded by external vendors to a Cloud Storage bucket be processed by an automated ingestion pipeline within seconds of arrival. The pipeline must scale to zero when no files are uploaded and avoid periodic polling loops. Which Google Cloud architectural pattern satisfies these requirements?
A BigQuery scheduled query running every minute to inspect Cloud Storage external table metadata for new files.
An Eventarc trigger on Cloud Storage object-finalized events that invokes a serverless Cloud Run service.
A cron job running on a persistent Compute Engine VM that executes gsutil ls every 10 seconds.
A Cloud Composer DAG running a continuous GCSObjectExistenceSensor in mode='poke' with a 5-second interval.
Mobile apps publish JSON purchase events to a Pub/Sub topic. Analysts need the events in a BigQuery table within seconds; no transformation is required, and the team does not want to operate a pipeline. What should they create?
A Dataflow streaming job running a custom Apache Beam pipeline that parses each event
A Cloud Composer DAG with a Pub/Sub sensor that inserts rows into the table
A BigQuery subscription on the topic using the table schema
A nightly BigQuery load job that reads exported event files from Cloud Storage
A high-velocity IoT telemetry ingestion pipeline consumes semi-structured JSON events from Cloud Pub/Sub. Occasionally, devices transmit corrupted or malformed payloads that cause parsing exceptions in the processing application. To maintain continuous pipeline availability, ensure auditability of bad data, and alert engineers when corrupted messages surge, how should the pipeline be architected?
Configure the parsing application to silently drop invalid payloads and run an hourly scheduled query that searches for missing sequence numbers.
Configure the Pub/Sub subscription with the maximum acknowledgment deadline and disable automatic message retries for the subscription.
Instruct the IoT devices to publish corrupted payloads directly to Cloud Logging instead of to the Pub/Sub topic.
A Pub/Sub dead-letter topic with a delivery-attempt limit, bad messages archived to Cloud Storage, and an alert on dead-letter volume.
Sections you finish are checked off in the contents.