10.4 Secure Data Sharing with Analytics Hub (BigQuery Sharing)

Key Takeaways

  • Analytics Hub provides a serverless, zero-copy data sharing platform built natively on BigQuery, eliminating data duplication and brittle extract pipelines.

  • Data Exchanges serve as secure management containers for publishing shared listings across internal business units, commercial partners, or public marketplaces.

  • Subscribing to a listing creates a read-only linked dataset in the subscriber's project that references the publisher's live data in place with real-time consistency.

  • Cost decoupling ensures publishers pay only for data storage, while subscribers pay for their own query execution compute (slots or bytes scanned).

  • Analytics Hub Data Clean Rooms allow multiple organizations to perform privacy-preserving collaborative analysis on combined first-party data without exposing raw records or PII.

Last updated: October 2026

Secure Data Sharing with Analytics Hub

Core Focus: Modern organizations must share data across internal departments, corporate subsidiaries, external suppliers, and commercial partners. Traditional sharing mechanisms—such as exporting CSV files to SFTP servers or replicating tables across cloud buckets—introduce severe data staleness, ballooning storage costs, and governance vulnerabilities. Google Cloud Analytics Hub delivers a modern, zero-copy data sharing architecture natively inside BigQuery. For the Associate Data Practitioner exam, you must understand Analytics Hub core entities (Exchanges, Listings, Linked Datasets), cost and security governance models, and collaborative Data Clean Rooms.

As data ecosystems expand, the demand for timely, governed data exchange has outgrown legacy file-transfer architectures. In a modern data mesh or inter-organizational partnership, data providers must retain centralized governance over their datasets while allowing subscribers to query live, pristine data on demand without copying terabytes across network boundaries.


The Shortcomings of Legacy Data Sharing Architectures

Historically, when an organization needed to share analytical datasets with external partners or subsidiaries, data engineering teams relied on batch extraction and transfer pipelines:

+-------------------------------------------------------------------------+
|                    Legacy Data Sharing Anti-Pattern                     |
|                                                                         |
|   [ Publisher BigQuery ]                                                |
|             |                                                           |
|             v                                                           |
|     (Nightly bq extract)                                                |
|             v                                                           |
|   [ Cloud Storage Bucket ] ---> [ SFTP Server / External Bucket ]       |
|                                                |                        |
|                                                v                        |
|                                    (Subscriber Ingestion ETL)           |
|                                                v                        |
|                                    [ Subscriber BigQuery ]              |
+-------------------------------------------------------------------------+

This legacy architecture creates severe operational and compliance problems:

  1. Data Staleness: By the time a nightly batch export finishes extracting, transferring, and loading into the subscriber's warehouse, the data is already hours out of date.
  2. Runaway Storage Duplication: Replicating a 100 TB analytical table across 50 partner organizations creates 5 petabytes of redundant, paid storage across the cloud ecosystem.
  3. Fragile Pipeline Maintenance: Custom ETL scripts, network tunnels, SFTP credentials, and transfer schedules frequently fail, requiring ongoing engineering intervention and monitoring.
  4. Total Loss of Governance and Auditability: Once raw data files are transferred to an external bucket or server, the publisher loses all control. The publisher cannot revoke access, cannot audit who queries the records, and cannot enforce data deletion requests mandated by privacy laws such as GDPR.

Analytics Hub Architecture: Zero-Copy Data Sharing

Analytics Hub, now branded BigQuery sharing (IAM roles still use the analyticshub prefix, and the exam guide says Analytics Hub), is a fully managed, serverless data exchange platform built directly into BigQuery. It leverages Google Cloud's decoupled compute (Dremel) and storage (Colossus) architecture to enable in-place, zero-copy data sharing across projects, organizations, and geographic regions.

+-----------------------------------------------------------------------------------------+
|                              Analytics Hub Architecture                                 |
|                                                                                         |
|   [ Data Publisher Project ]                                                            |
|   Primary BigQuery Dataset (Storage Billed to Publisher)                                |
|             |                                                                           |
|             v                                                                           |
|   [ Analytics Hub Data Exchange ]                                                       |
|   └── Listing: "Daily Retail Indicators" (Metadata, Documentation, Terms)               |
|             |                                                                           |
|             | (Authorized Subscription)                                                 |
|             v                                                                           |
|   [ Subscriber Project A ]                  [ Subscriber Project B ]                    |
|   Linked Dataset (Read-Only)                Linked Dataset (Read-Only)                  |
|   (Query Compute Billed to Sub A)           (Query Compute Billed to Sub B)             |
+-----------------------------------------------------------------------------------------+

Core Architectural Entities

Analytics Hub structures data sharing around three fundamental constructs:

  1. Data Exchanges:

    • A secure, centralized container within Analytics Hub where listings are published and discovered.
    • Private Exchanges: Scoped internally to an enterprise organization, enabling seamless data mesh sharing across departments (e.g., Marketing, Supply Chain, Finance) without project-level IAM sprawling.
    • Cross-Organization Exchanges: Private exchanges configured to allow specific external Google Cloud organizations or business partners to discover and subscribe to curated listings.
    • Public Exchanges: Public catalogs hosted on Google Cloud Marketplace containing open data sets (e.g., Google Public Datasets, NOAA weather data, public health figures).
  2. Listings:

    • A curated package of shared data published within an exchange.
    • A listing references an underlying BigQuery dataset or BigLake table. It bundles descriptive metadata, schema documentation, categories, data provider contact details, and usage terms.
    • Publishers control who can view and subscribe to each listing using IAM permissions (roles/analyticshub.viewer, roles/analyticshub.subscriber).
  3. Subscriptions & Linked Datasets:

    • When an authorized user browses a Data Exchange and subscribes to a listing, Analytics Hub provisions a Linked Dataset in the subscriber's designated Google Cloud project.
    • Mechanics of a Linked Dataset: A linked dataset is a read-only symbolic reference pointing directly to the publisher's live storage blocks in BigQuery.
    • Zero-Copy Guarantee: No physical bytes are copied, serialized, or transferred across projects. Storage remains strictly inside the publisher's original dataset.
    • Real-Time Data Freshness: As soon as the publisher inserts, appends, or updates records in the primary table, subscribers querying the linked dataset see the updated records instantly with zero pipeline latency.

Security, Governance, and Cost Separation

Analytics Hub provides clean architectural separation between security governance and analytical compute billing.

1. Centralized Publisher Governance

  • Revocation on Demand: If a commercial contract ends or an unauthorized access pattern is identified, the publisher can immediately revoke the subscriber's listing access. The subscriber's linked dataset instantly becomes inactive, without requiring manual data purging.
  • Curated Exposure: Publishers decide exactly which dataset, tables, or authorized views go into a listing, and can turn on data egress restrictions so subscribers cannot copy or export the shared data.
  • Complete Auditability: Publishers inspect Cloud Monitoring metrics and Cloud Audit Logs to track sharing metrics across all listings: total active subscriptions, query counts, and aggregated bytes scanned by subscribers.

2. Cross-Project Security Isolation

Subscribers do not receive direct IAM permissions inside the publisher's project. They require only standard BigQuery querying permissions (roles/bigquery.dataViewer on their local linked dataset and roles/bigquery.jobUser on their local project). This eliminates the security risk of granting external partners IAM credentials within internal production projects.

3. Clear Cost Decoupling

Cost CategoryEntity Responsible for PaymentTechnical Mechanism
Storage CostsData PublisherThe publisher stores the physical data blocks in their primary BigQuery dataset and pays standard BigQuery active or long-term storage fees.
Query Compute CostsData SubscriberWhen a subscriber queries a linked dataset, the query job executes within the subscriber's project. The subscriber pays for the compute slots or on-demand scanned bytes consumed by the query.
Data Egress CostsNone (Within Region)Data sharing within the same regional location incurs zero network egress charges because queries access storage in place via Google's internal Jupiter network.

Exam Rule: In Analytics Hub, subscribers always pay for their own query compute, while publishers pay for data storage. Queries across linked datasets must reside in the same region or multi-region (e.g., sharing from US multi-region to US multi-region). Cross-region linked datasets are not supported without explicit replication.


Analytics Hub Data Clean Rooms

A Data Clean Room in Analytics Hub is a privacy-enhancing collaborative environment that enables multiple organizations to join and analyze their respective first-party datasets without sharing raw underlying records or PII.

+-------------------------------------------------------------------------+
|                    Analytics Hub Data Clean Room                        |
|                                                                         |
|   [ Retail Enterprise Data ]             [ Media Streaming Data ]       |
|   - Customer Purchase Records            - Ad Impression Events         |
|   - Loyalty ID, Purchase Total           - Loyalty ID, Stream Watch Time|
|                 |                                      |                |
|                 +------------------+-------------------+                |
|                                    |                                    |
|                                    v                                    |
|                     [ Clean Room Analysis Rules ]                       |
|                     - Only Aggregate Queries Allowed                    |
|                     - Threshold: Minimum Cohort Size > 50               |
|                     - Raw PII Projections Blocked                       |
|                                    |                                    |
|                                    v                                    |
|                     [ Governed Campaign Insights ]                      |
|                     "Campaign conversion lift: 18.4%"                   |
+-------------------------------------------------------------------------+

How Clean Rooms Operate

  1. Multi-Party Collaboration: Consider a major retailer and an online advertising network. The retailer wants to measure how many users purchased products after seeing an advertisement on the streaming network.
  2. Privacy Risks: Neither organization is legally permitted to hand over raw customer identities or full transaction logs to the other.
  3. Clean Room Environment: Both parties contribute their datasets to an Analytics Hub Data Clean Room. Within the clean room, analysts execute collaborative queries that join the datasets on common identifiers (such as hashed email addresses).
  4. Analysis Rules & Differential Privacy: Clean rooms enforce strict analysis rules:
    • Aggregate-Only Queries: Queries must include GROUP BY aggregations. Selecting individual rows (SELECT *) is forbidden.
    • Aggregation Thresholds: Results are suppressed if any grouping cohort contains fewer than a defined threshold of individuals (e.g., threshold = 50), preventing adversaries from isolating individual customer records through narrow filtering.
    • No Raw Column Projections: Participating organizations can only view aggregate statistical output (e.g., campaign reach, conversion rates, average order values), ensuring raw underlying PII never leaves the publisher's boundary.

Common Exam Traps & Real-World Scenarios

Trap 1: Recommending Cross-Project Bucket Copying for Data Sharing

  • The Scenario: A business intelligence vendor needs to share nightly product catalogs with 200 corporate retail clients.
  • The Trap: Designing an automated pipeline using Cloud Storage Transfer Service or Cloud Data Fusion to copy Parquet files into each client's storage bucket.
  • The Reality: This is an inefficient anti-pattern resulting in 200x storage replication, high egress/transfer costs, and substantial pipeline fragility. Analytics Hub is the designated Google Cloud solution, enabling all 200 clients to subscribe to a single shared listing via zero-copy linked datasets.

Trap 2: Believing Publishers Pay for Subscriber Query Execution

  • The Scenario: An executive fears publishing a dataset on Analytics Hub because thousands of external subscribers might execute expensive queries that inflate the publisher's cloud bill.
  • The Trap: Assuming query compute is billed to the dataset owner.
  • The Reality: Analytics Hub enforces complete cost decoupling. Query compute jobs run entirely within the subscriber's project and are billed to the subscriber's Cloud Billing account. The publisher pays only for their own underlying storage.

Trap 3: Cross-Region Linked Dataset Querying

  • The Scenario: A publisher hosts a primary BigQuery dataset in the europe-west1 (Belgium) region and publishes a listing. A subscriber attempts to link the dataset into a project configured in us-central1 (Iowa) to run joins with local tables.
  • The Trap: Expecting linked datasets to support cross-region queries natively.
  • The Reality: BigQuery requires compute slots and storage to reside in the same physical region or multi-region location to execute queries. The subscriber's linked dataset must be provisioned in the same region as the publisher's underlying dataset (europe-west1). Cross-region analytics require explicit table replication via BigQuery cross-region dataset replication.
Test Your Knowledge

A telecommunications company wants to share a 50 TB network performance analytics dataset with dozens of independent hardware maintenance vendors. The requirements specify that vendors must query the live dataset with real-time freshness, no data can be copied or duplicated across projects, and the telecommunications company must not incur any compute charges when vendors run analytical queries. Which architecture should the company deploy?

A

Deploy Cloud Data Fusion pipelines to replicate the tables nightly into each vendor's individual BigQuery project.

B

Create a private Data Exchange in Analytics Hub, publish the dataset as a listing, and instruct vendors to subscribe by creating linked datasets in their own projects.

C

Export the dataset to Cloud Storage as Parquet files and grant the vendors roles/storage.objectViewer on the bucket.

D

Grant dataset-level roles/bigquery.dataViewer directly to each vendor's service account within the telecommunications company's production project.

Test Your Knowledge

When an enterprise subscribes to a shared listing published on Google Cloud Analytics Hub, how does the resulting linked dataset function within the subscriber's BigQuery project?

A

It is a read-only pointer to the publisher's live data in place, so subscribers see updates without any copy being made.

B

It provisions a BigQuery materialized view that must be refreshed manually on a Cloud Scheduler timetable.

C

It configures a BigQuery external table backed by CSV files stored in the publisher's Cloud Storage bucket.

D

It triggers a background BigQuery Data Transfer Service job that copies the source tables into standard BigQuery physical storage.

Test Your Knowledge

A national grocery chain and a consumer packaged goods (CPG) manufacturer want to combine their first-party customer loyalty and media advertising datasets to evaluate campaign conversion performance. Both companies have strict privacy policies forbidding the disclosure of raw customer identities, individual transactions, or PII to external partners. Which Analytics Hub feature enables this collaborative analysis while enforcing privacy guarantees?

A

Public data exchanges combined with Cloud Storage signed URLs for each partner

B

BigQuery data clean rooms whose analysis rules enforce aggregation thresholds

C

Uniform bucket-level access combined with IAM Conditions on the shared bucket

D

Authorized datasets shared with each partner's users through primitive Viewer roles

Sections you finish are checked off in the contents.