10.1 IAM Principles & Role Hierarchies for Data Services
Key Takeaways
The Principle of Least Privilege requires granting identities only the minimum necessary permissions required to execute their specific analytical or pipeline functions.
Google Cloud resource hierarchy enforces downward inheritance (Organization -> Folder -> Project -> Resource); permissions granted at a parent level cannot be revoked or restricted at a child level.
Basic (primitive) roles like Viewer and Editor are dangerous anti-patterns in production data systems; fine-grained predefined roles or custom roles must be used instead.
Executing BigQuery queries requires two distinct IAM permissions: compute job authorization at the project level (roles/bigquery.jobUser) and data read authorization on the target dataset (roles/bigquery.dataViewer).
Automated pipelines must run under dedicated user-managed service accounts using Workload Identity Federation or runtime impersonation rather than long-lived, exportable JSON service account keys.
IAM Principles & Role Hierarchies for Data Services
Core Focus: Modern data platforms centralize vast volumes of sensitive organizational data, making robust access control and governance foundational to enterprise security. Google Cloud Identity and Access Management (IAM) defines who (identity) can do what (permission) on which resource. For the Google Cloud Associate Data Practitioner examination, you must master resource hierarchy inheritance, the operational distinctions between basic, predefined, and custom roles, resource-level dataset and bucket bindings, and the critical role pairings required to run queries and manage automated data pipelines.
Securing cloud data architectures requires shifting away from legacy perimeter security models toward identity-centric governance. In an analytical environment like Google Cloud, raw data lakes, staging zones, and dimensional data warehouses coexist within shared projects. Applying the Principle of Least Privilege ensures that data engineers, business intelligence analysts, compliance auditors, and automated pipeline workloads access only the exact data resources required for their responsibilities—minimizing the blast radius of compromised credentials or accidental data modifications.
The Principle of Least Privilege in Data Governance
The Principle of Least Privilege (PoLP) dictates that any identity—whether a human analyst, an automated service account, or a third-party application—must be granted only the minimum set of permissions necessary to perform its intended task, and only for the duration required.
In cloud data engineering, violating least privilege introduces severe operational and security risks:
- Data Exfiltration and Privacy Violations: Granting broad project-level read permissions allows unauthorized users to inspect sensitive datasets containing Personally Identifiable Information (PII) or confidential financial records.
- Runaway Analytical Compute Costs: Granting unconstrained query execution privileges without quota restrictions or job boundaries allows inexperienced users to trigger multi-terabyte unpartitioned queries that rapidly consume cloud budgets.
- Accidental Schema and Data Destruction: Granting administrative or edit privileges to reporting users creates the risk of accidental table drops, schema truncation, or erroneous data overrides.
- Compliance Non-Compliance: Regulations such as GDPR, HIPAA, and PCI-DSS require demonstrable separation of duties, ensuring that personnel who manage infrastructure cannot arbitrarily view or export protected customer records.
Google Cloud Resource Hierarchy and Permission Inheritance
Google Cloud organizes all resources into a strict four-tier hierarchy. Understanding how permissions flow down this tree is essential for designing data access boundaries.
+-------------------------------------------------------------------------+
| Google Cloud Organization |
| (example.com) |
| | |
| v |
| Folders (e.g., Data Platform) |
| | |
| v |
| Projects (e.g., analytics-prod) |
| | |
| +---------------------------+---------------------------+ |
| | | |
| v v |
| BigQuery Datasets Cloud Storage |
| (marketing_dw, finance_dw) Buckets & Objects |
+-------------------------------------------------------------------------+
The Permissive Union Rule
IAM policy inheritance operates on a strictly additive (permissive union) model:
- Downward Inheritance: Permissions granted at an upper level of the hierarchy (such as an Organization or Project) are automatically inherited by all child resources underneath.
- No Child Restrictions: You cannot revoke or narrow an inherited permission at a lower level. For example, if a user is granted
roles/bigquery.dataViewerat the project level, they can read every BigQuery dataset within that project. You cannot remove their access from a single sensitive dataset (e.g.,payroll_dw) within that project using standard IAM grants. - Effective Policy: The effective permissions on any individual resource are the mathematical union of all IAM bindings granted at that resource, its parent project, its parent folders, and the root organization.
IAM Deny Policies
To address the limitation of additive inheritance, Google Cloud provides IAM Deny Policies. Deny policies allow security administrators to define explicit restrictions at the organization, folder, or project level that override allow rules. For example, a deny policy can explicitly block all non-security personnel from calling bigquery.tables.getData on datasets containing restricted customer data, regardless of what project-level allow roles they possess.
Types of IAM Roles in Data Workloads
Google Cloud classifies IAM roles into three distinct tiers: Basic (Primitive), Predefined, and Custom roles.
+---------------------------------------------------------------------------------------+
| IAM Role Hierarchy |
| |
| [ Basic Roles ] [ Predefined Roles ] [ Custom Roles ] |
| - Owner - roles/bigquery.dataViewer - Tailored sets of |
| - Editor - roles/bigquery.jobUser specific |
| - Viewer - roles/storage.objectViewer permissions |
| (Coarse-grained; dangerous (Fine-grained; service- (Minimizes excess |
| anti-pattern in production) specific; recommended) privilege) |
+---------------------------------------------------------------------------------------+
1. Basic (Primitive) Roles: The Production Anti-Pattern
Basic roles represent legacy Google Cloud permissions: Owner, Editor, and Viewer.
roles/viewer: Grants read-only access to almost all Google Cloud resources across the entire project.roles/editor: Grants modify and delete permissions across almost all resources in the project, including creating VMs, dropping databases, and modifying storage buckets.roles/owner: Grants full control, including billing management and the ability to modify IAM permissions.
Exam Warning: Basic roles are a major anti-pattern for production data systems. Granting
roles/editorto a data engineer grants them permission to delete firewall rules, deploy unmonitored compute instances, and alter IAM policies outside their domain. Grantingroles/viewerto an analyst exposes every dataset, secret, and storage bucket in the project. Never use primitive roles in production.
2. Predefined Roles: Fine-Grained Service Specialization
Predefined roles are authored and maintained by Google Cloud. They bundle granular permissions tailored to specific job responsibilities within individual services.
Key BigQuery Predefined Roles
| Predefined Role Name | Level Typically Granted | Core Permissions & Purpose |
|---|---|---|
roles/bigquery.dataViewer | Dataset or Table | Read data and metadata from tables and views (bigquery.tables.getData, bigquery.tables.list). Cannot run query jobs on its own. |
roles/bigquery.dataEditor | Dataset or Table | Create, append, update, and truncate tables and views within the dataset. Cannot delete the dataset itself. |
roles/bigquery.dataOwner | Dataset | Full control over the dataset, including creating tables, updating schema, altering dataset ACLs, and deleting the dataset. |
roles/bigquery.jobUser | Project | Run BigQuery jobs (queries, table exports, table copies, load jobs). Allows allocating query compute resources. |
roles/bigquery.user | Project | Run query jobs AND create new BigQuery datasets within the project. |
roles/bigquery.admin | Project | Full administrative control over all BigQuery resources, datasets, jobs, reservations, and slot allocations. |
Key Cloud Storage Predefined Roles
| Predefined Role Name | Level Typically Granted | Core Permissions & Purpose |
|---|---|---|
roles/storage.objectViewer | Bucket or Project | Read object data and list bucket contents (storage.objects.get, storage.objects.list). |
roles/storage.objectCreator | Bucket or Project | Write new objects to the bucket. Cannot read, list, overwrite, or delete existing objects (ideal for write-only dropzones). |
roles/storage.objectUser | Bucket or Project | Read, write, and delete objects within a bucket. Ideal for automated staging and transformation pipelines. |
roles/storage.objectAdmin | Bucket or Project | Full control over objects, including viewing and modifying object-level ACLs. |
roles/storage.admin | Bucket or Project | Full administrative control over buckets and objects, including bucket creation, deletion, and lifecycle rules. |
3. Custom Roles
When predefined roles bundle more permissions than an organization's compliance policy permits, administrators author Custom Roles. Custom roles assemble an exact list of granular permissions (e.g., combining bigquery.tables.get and bigquery.tables.list without granting bigquery.tables.getData to allow schema inspection without data access).
Custom Role Trade-Offs:
- Advantage: Absolute enforcement of least privilege.
- Disadvantage: High maintenance overhead. Google Cloud regularly updates services with new API capabilities; predefined roles incorporate these automatically, whereas custom roles require manual updates.
- Scope: Can be defined at the Project or Organization level, but cannot be defined at the Folder level.
Resource-Level IAM Bindings for Sensitive Data Isolation
To maintain strong security boundaries, data architectures must avoid granting data access roles at the project level whenever possible. Instead, organizations leverage resource-level IAM bindings.
Dataset-Level Isolation in BigQuery
Consider an enterprise analytics project containing three datasets:
marketing_public_analytics(general web analytics)sales_pipeline_reporting(internal business metrics)payroll_executive_records(highly sensitive executive compensation)
If the data team granted roles/bigquery.dataViewer at the project level, every user in the project could read payroll_executive_records.
By scoping permissions to the dataset level:
- General business analysts receive
roles/bigquery.dataViewerbound strictly tomarketing_public_analyticsandsales_pipeline_reporting. - Only HR and executive leadership identities receive
roles/bigquery.dataViewerbound directly topayroll_executive_records. - To execute queries, all analysts receive
roles/bigquery.jobUserat the project level.
Project: corp-analytics
│
├── IAM Binding (Project Level): User group "analysts" has roles/bigquery.jobUser
│
├── Dataset: marketing_public_analytics
│ └── IAM Binding (Dataset Level): "analysts" has roles/bigquery.dataViewer
│
└── Dataset: payroll_executive_records
└── IAM Binding (Dataset Level): "hr-execs" has roles/bigquery.dataViewer
Service Accounts & Credential Delegation
A Service Account is a special Google Cloud identity used by non-human workloads—such as Cloud Dataflow jobs, Cloud Composer Airflow workers, Cloud Functions, and Dataproc clusters—to authenticate and interact with Google Cloud APIs.
Default vs. User-Managed Service Accounts
When Google Cloud APIs are enabled, the platform often generates default service accounts (e.g., the Compute Engine default service account <project-number>-compute@developer.gserviceaccount.com). Historically, default service accounts were automatically granted the primitive roles/editor role on the project (newer organizations block that automatic grant by default).
Best Practice: Never run production data pipelines under default service accounts. Always provision User-Managed Service Accounts named for their specific workload (e.g.,
sa-dataflow-ingestion@analytics-prod.iam.gserviceaccount.com) endowed exclusively with the minimal predefined roles required for the pipeline.
Eliminating Service Account Keys: Workload Identity & Impersonation
Service account keys are downloadable JSON files containing long-lived cryptographic private keys. They represent one of the greatest security vulnerabilities in cloud data engineering:
- Keys are frequently leaked into public Git repositories.
- Keys lack automatic rotation mechanisms.
- Keys bypass corporate multi-factor authentication (MFA) and single sign-on (SSO).
Modern Google Cloud architectures eliminate downloadable keys using two core mechanisms:
- Workload Identity Federation: Enables workloads running outside Google Cloud (such as GitHub Actions CI/CD pipelines, Kubernetes clusters on AWS, or on-premises servers) to authenticate using short-lived OpenID Connect (OIDC) or SAML tokens exchanged directly for Google Cloud access tokens.
- Service Account Impersonation: Human developers and CI/CD pipelines do not possess service account keys. Instead, their authenticated human identity is granted
roles/iam.serviceAccountTokenCreatoron the target service account. When they need to execute an automated job, their local CLI or client library calls the Google Cloud IAM credentials API to generate a temporary, short-lived (1-hour) OAuth access token.
Critical Exam Traps & Role Combination Requirements
The "Two-Role BigQuery Dilemma"
One of the most frequently tested IAM concepts on the Associate Data Practitioner exam is the requirement of two distinct roles to execute a BigQuery query.
When a user queries a table via the BigQuery console, CLI (bq query), or client library:
- Compute / Execution Permission: BigQuery must spin up an execution job within a project to allocate computing slots and process the SQL query. This requires the
bigquery.jobs.createpermission, provided byroles/bigquery.jobUser(orroles/bigquery.user) granted at the project level. - Data / Storage Permission: BigQuery must read the underlying storage blocks belonging to the queried table or view. This requires the
bigquery.tables.getDatapermission, provided byroles/bigquery.dataViewergranted at the dataset, table, or project level.
+-------------------------------------------------------------------------+
| The Two-Role BigQuery Requirement |
| |
| User Identity |
| | |
| +---> [ Project Level ] ==> roles/bigquery.jobUser |
| | (Permission: bigquery.jobs.create) |
| | |
| +---> [ Dataset Level ] ==> roles/bigquery.dataViewer |
| (Permission: bigquery.tables.getData)|
+-------------------------------------------------------------------------+
If a user has roles/bigquery.dataViewer on the dataset but no project role, their query fails with an Access Denied error stating that they lack the bigquery.jobs.create permission in the project.
If a user has roles/bigquery.jobUser on the project but no dataset role, their query fails with an Access Denied error saying they don't have permission to query the table (or that it may not exist).
A data analyst needs to run exploratory SQL queries against a specific reporting dataset named 'retail_marts' in the 'analytics-prod' project. The security team mandates strict adherence to the Principle of Least Privilege: the analyst must not be able to read any other datasets in the project, create new datasets, or modify existing tables. Which IAM configuration satisfies these requirements?
Grant roles/bigquery.user at the 'analytics-prod' project level, and grant roles/bigquery.dataEditor on the 'retail_marts' dataset.
Grant roles/bigquery.jobUser at the 'analytics-prod' project level, and grant roles/bigquery.dataViewer scoped directly to the 'retail_marts' dataset.
Grant roles/bigquery.dataViewer at the 'analytics-prod' project level, and grant roles/bigquery.jobUser on the 'retail_marts' dataset.
Grant roles/viewer at the 'analytics-prod' project level, and remove read permissions from all other datasets.
A cloud data engineering team is configuring an automated Cloud Dataflow batch pipeline that reads JSON files from a Cloud Storage bucket and writes transformed records into a BigQuery staging table. In accordance with Google Cloud enterprise security guidelines, how should authentication and authorization for the Dataflow workers be configured?
A dedicated worker service account with roles/dataflow.worker, storage.objectViewer on the bucket, and bigquery.dataEditor on the dataset.
Download a JSON key for the Compute Engine default service account, encrypt it with Cloud KMS, and pass the key path to the Dataflow pipeline as an execution argument.
Create a custom IAM role containing all BigQuery and Cloud Storage permissions, and attach an API key to the Dataflow pipeline runner.
Execute the Dataflow pipeline under the Compute Engine default service account, which holds the primitive roles/editor role at the project level.
A newly hired data scientist was assigned the primitive roles/viewer role at the Google Cloud project level. The data governance officer notices that the data scientist is able to view raw employee payroll tables located inside a restricted dataset within that same project, despite the dataset ACL containing no explicit grant for that user. Why does the data scientist have access to the payroll records?
The roles/viewer role includes administrative policy-editing permissions that allow users to self-authorize access to any dataset.
The BigQuery service automatically makes all datasets readable to all authenticated corporate Google accounts unless an external firewall rule blocks them.
BigQuery datasets do not support independent IAM bindings and can only inherit policies defined at the folder level of the resource hierarchy.
Project-level grants are inherited by every dataset in the project and cannot be revoked by dataset-level allow policies (only deny policies restrict them).
Sections you finish are checked off in the contents.