11.2 High Availability & Disaster Recovery Architectures

Key Takeaways

  • Disaster recovery architectures are bounded by Recovery Point Objective (RPO), the maximum acceptable data loss in time, and Recovery Time Objective (RTO), the maximum acceptable downtime to restore service.

  • Cloud Storage provides Regional (99.9% availability), Dual-Region (99.95% availability), and Multi-Region (99.95% availability) configurations; Dual-Region with Turbo Replication guarantees 100% replication within 15 minutes to back a strict 15-minute RPO.

  • Cloud SQL Regional High Availability provides synchronous block-level persistent disk replication between primary and standby instances across two zones, enabling automatic failover in about a minute with zero data loss (RPO = 0).

  • Cross-region disaster recovery for Cloud SQL leverages asynchronous read replicas that can be manually promoted to primary during a regional outage, trading non-zero replication lag for geographical resilience.

  • Automated Cloud SQL backups and transaction log archiving enable Point-in-Time Recovery (PITR) to restore a database to a specific second within the transaction log retention window (1 to 7 days on Enterprise edition, 1 to 35 days on Enterprise Plus).

Last updated: October 2026

High Availability & Disaster Recovery Architectures

Core Focus: Enterprise data architectures must remain resilient against localized hardware failures, datacenter network interruptions, and catastrophic regional disasters. Designing resilient cloud architectures requires balancing business continuity metrics—specifically Recovery Point Objective (RPO) and Recovery Time Objective (RTO)—against infrastructure cost and complexity. This section examines high availability (HA) and disaster recovery (DR) patterns across Google Cloud Storage and Cloud SQL.

Business continuity planning begins by establishing clear Service Level Objectives (SLOs) anchored to business impact. An architecture that guarantees instant recovery across continents is economically wasteful if applied to transient development logs, while an under-engineered backup strategy can shutter an enterprise if applied to a core financial transaction ledger.


Foundational DR Concepts: RPO, RTO, and Continuity Strategies

Disaster recovery designs are governed by two fundamental operational parameters:

Recovery Point Objective (RPO) vs. Recovery Time Objective (RTO)

  • Recovery Point Objective (RPO): The maximum acceptable amount of data loss measured in time preceding a disruptive incident. RPO defines the age of the data that must be recovered from backup or replication systems for normal operations to resume. An RPO of zero means zero data loss can be tolerated; an RPO of 24 hours means the organization can tolerate losing a full day of transactional records.
  • Recovery Time Objective (RTO): The maximum acceptable duration of system downtime measured in clock time. RTO defines how quickly the system must be restored and operational following a failure. An RTO of 5 minutes demands rapid, automated failover; an RTO of 24 hours allows for manual intervention, server provisioning, and sequential backup restorations.
Disaster Timeline:
[Last Safe State / Sync] <------- RPO -------> [Disaster Strikes] <------- RTO -------> [Service Restored]
                        (Lost Data Gap)                          (Downtime Gap)

Architectural DR Spectrum

Business continuity designs represent a spectrum trading cost for RPO/RTO minimization:

  1. Cold Standby (Backup & Restore): Data is backed up periodically to remote object storage. Compute resources are not provisioned in the recovery region until disaster strikes. Lowest infrastructure cost, but highest RTO (hours to days) and RPO (hours to days depending on backup frequency).
  2. Warm Standby (Pilot Light / Scaled Replica): A minimal replica of the environment runs continuously in a secondary region. Databases replicate asynchronously. When the primary fails, the secondary database is promoted and compute auto-scales to absorb traffic. Moderate cost, low RPO (seconds to minutes), moderate RTO (minutes).
  3. Active-Active Multi-Region: Fully provisioned infrastructure actively serves read and write traffic across multiple geographically separated regions simultaneously. Near-zero RPO and near-zero RTO, but highest cost and operational complexity.

Cloud Storage Redundancy Architectures

Google Cloud Storage provides 99.999999999% (eleven 9s) annual durability across all storage bucket location types. However, availability and disaster resilience vary depending on geographic location placement.

1. Regional Buckets

  • Architecture: Objects are distributed redundantly across multiple independent availability zones within a single geographic region (e.g., us-central1).
  • Availability SLA: 99.9% for Standard storage.
  • Cost & Performance: Lowest storage cost; optimal throughput and zero cross-region network egress charges when accessed by compute resources (Compute Engine, Dataproc, Dataflow, BigQuery) located in that same region.
  • Resilience Boundary: Resilient against individual datacenter or zone outages. Vulnerable to catastrophic regional disasters (e.g., severe natural disaster taking down an entire metropolitan area).

2. Dual-Region Buckets

  • Architecture: Data is stored actively across two specific, geographically separated regions (e.g., us-central1 and us-east1).
  • Availability SLA: 99.95% for Standard storage.
  • Failover: Fully automated, active-active read and write availability. If one region suffers a total outage, client requests route seamlessly to the healthy region with zero downtime.
  • Turbo Replication: By default, cross-region replication is asynchronous (Google targets replicating 99.9% of objects within 1 hour). For mission-critical workloads, organizations can enable Turbo Replication. Turbo Replication provides a legally binding Service Level Agreement guaranteeing that 100% of newly written or updated objects replicate across regions within 15 minutes. This enables architectures with a strict 15-minute RPO.

3. Multi-Region Buckets

  • Architecture: Data is redundantly distributed across at least two geographic regions within a broad continental boundary (e.g., us, eu, asia).
  • Availability SLA: 99.95% for Standard storage.
  • Primary Use Case: Content distribution networks (CDNs), globally distributed analytics, and highest resilience against nationwide regional disruptions. Storage cost is higher than Regional, and replication occurs asynchronously across the continent.

Bucket Redundancy Decision Matrix

Location TypeTypical Availability SLACross-Zone ResilienceCross-Region ResilienceReplication MechanismTarget Workload
Regional99.9%Yes (multi-zone)NoSynchronous within regionIn-region processing (Dataproc, BigQuery, Dataflow), lowest cost.
Dual-Region99.95%Yes (multi-zone)Yes (pair of 2 regions)Asynchronous (or Turbo: 15-min SLA)Mission-critical enterprise data lakes, strict 15-minute RPO compliance.
Multi-Region99.95%Yes (multi-zone)Yes (entire continent)AsynchronousGlobal content delivery, public data dissemination, broad DR resilience.

Cloud SQL High Availability (HA) Architecture

Relational databases (Cloud SQL for MySQL, PostgreSQL, and SQL Server) serve as the transactional backbone of enterprise operations. To prevent downtime from hardware or zonal failures, Google Cloud provides Cloud SQL Regional High Availability.

Regional High Availability Mechanics

When High Availability is enabled, Cloud SQL provisions two instances within the same region:

  1. Primary Instance: Located in the primary zone (e.g., us-central1-a). Actively accepts read and write traffic from client applications.
  2. Standby Instance: Located in a secondary zone (e.g., us-central1-b). Does not serve client traffic during normal operations.
  3. Synchronous Block-Level Replication: The primary and standby instances do not rely on database-level replication streams for HA. Instead, they utilize synchronous persistent disk replication at the block storage layer. Every write transaction executed on the primary instance is synchronously committed to both the primary zone disk and the secondary zone disk before the transaction commits.
  4. Automated Failover: A health check monitor continuously polls the primary instance. If the primary instance or its underlying zone fails, Cloud SQL triggers automated failover:
    • The standby instance mounts the replicated disk.
    • The standby instance initializes database recovery.
    • The primary IP address automatically transitions to the standby instance.
    • Failover Duration: Failover typically takes about a minute; Cloud SQL Enterprise Plus edition is designed for much shorter interruptions.
    • RPO and RTO: Because storage replication is synchronous, RPO = 0 (zero data loss). The RTO is roughly a minute.
Cloud SQL Regional HA Architecture:
[Zone A] Primary Instance  <==== Synchronous Storage Replication ====> [Zone B] Standby Instance
        |                                                                      |
    [Serves Writes/Reads]                                               [Idle / Health Check]
        |                                                                      |
        +---------------------- Shared Regional IP ----------------------------+

Cloud SQL Disaster Recovery: Cross-Region Read Replicas

While Regional HA protects against zonal outages, it cannot protect against a catastrophic regional disruption that affects all zones within a region. Cross-region disaster recovery requires a multi-region strategy.

Cross-Region Read Replicas

Cloud SQL allows administrators to create Cross-Region Read Replicas:

  • Asynchronous Replication: The primary instance in Region A continuously streams transaction logs (binary logs in MySQL, WAL in PostgreSQL) to the replica in Region B over Google's global fiber backbone.
  • Latency & RPO: Because replication across geographic regions is asynchronous to prevent dragging down primary transactional write performance, there is a small replication lag. During a regional disaster, the RPO equals the replication lag at the time of failure (typically seconds).
  • Disaster Recovery Promotion: If Region A experiences a total disaster, Cloud SQL does not automatically fail over across regions. Cross-region failover is a deliberate decision that you initiate (Cloud SQL Enterprise Plus adds advanced disaster recovery with a designated DR replica you can switch over to, but you still trigger it):
    1. The administrator or orchestration script calls gcloud sql instances promote-replica.
    2. The read replica detaches from the replication stream and is promoted to a standalone, read-write primary database instance.
    3. Applications are re-routed to the new primary instance endpoint in Region B.
    4. Failover Duration: Promotion typically completes within minutes, so the RTO is minutes rather than seconds.

Cloud SQL Backups and Point-in-Time Recovery (PITR)

Backups and replication serve different purposes. High availability and replicas protect against infrastructure hardware failures, but they replicate human errors instantly: if a developer accidentally drops a production table (DROP TABLE customers), that drop statement replicates synchronously to the HA standby and asynchronously to read replicas within milliseconds.

To protect against corruption and human error, Cloud SQL utilizes Automated Backups and Point-in-Time Recovery (PITR).

Automated Backups

  • Scheduled daily during a configurable backup window.
  • Incremental storage: only modified blocks are written, minimizing storage usage.
  • Stored redundantly across multiple regions for disaster resilience.
  • Retention is configurable up to 365 days.

Point-in-Time Recovery (PITR)

  • Mechanism: PITR leverages continuous transaction logging (binary logs for MySQL, Write-Ahead Logs for PostgreSQL) alongside automated daily backups.
  • Capabilities: Allows administrators to restore the database to its state at a specific second within the transaction log retention window: 1 to 7 days on Cloud SQL Enterprise edition (default 7) and 1 to 35 days on Enterprise Plus (default 14).
  • Recovery Workflow: Cloud SQL creates a brand-new database instance, restores the closest preceding base backup, and plays forward the transaction logs up to the exact target microsecond preceding the corrupting command.

Disaster Recovery Decision Matrix

The following matrix provides a clear operational comparison for architecting data reliability on Google Cloud:

Strategy / ComponentScopeTypical RPOTypical RTOAutomated Failover?Primary Failure Guard
Cloud SQL Regional HAZonal (intra-region)0 (Zero data loss)About 1 minuteYes (automatic)Zonal hardware, VM, or datacenter failure.
Cloud SQL Cross-Region ReplicaRegional (inter-region)Seconds (replication lag)MinutesNo (manual promotion)Complete regional disaster.
Cloud SQL PITRDatabase instanceNear-zero (exact second)10 - 30+ minutesNo (provisions new instance)Accidental table drop, bad batch update, corruption.
Cloud Storage RegionalZonal (intra-region)0 (within region)0Yes (intra-zone)Single datacenter outage within region.
Cloud Storage Dual-RegionRegional (pair)15 min (with Turbo)0Yes (active-active)Full regional outage; strict 15-minute RPO.
Cloud Storage Multi-RegionContinentalAsynchronous (~1 hr)0Yes (active-active)Multi-region disaster, global distribution.

Exam Traps and Best Practices

Exam Tip: Distinguish between Zonal HA and Regional DR. Cloud SQL High Availability is strictly a zonal redundancy mechanism within a single region. If an exam question specifies surviving a complete regional outage, Cloud SQL HA alone is insufficient—you must implement Cross-Region Read Replicas.

Trap 1: Expecting Automatic Failover for Cross-Region Read Replicas

  • The Trap: Believing that if a primary Cloud SQL region goes down, Cloud SQL will automatically promote a cross-region read replica.
  • The Reality: Automatic failover applies exclusively to Regional HA standby instances across zones in the same region. Cross-region replica promotion is strictly manual to prevent split-brain scenarios and accidental cross-region application rerouting.

Trap 2: Believing HA Standby Protects Against Accidental Data Deletion

  • The Trap: Relying on Cloud SQL Regional HA to recover from a rogue TRUNCATE TABLE command.
  • The Reality: HA uses synchronous storage replication. The TRUNCATE command writes to both disks synchronously. The standby disk is corrupted at the exact same instant as the primary. Restoring from human error requires Point-in-Time Recovery (PITR).

Trap 3: Conflating Dual-Region Default Replication with Turbo Replication

  • The Trap: Assuming standard Dual-Region Cloud Storage guarantees a 15-minute RPO out of the box.
  • The Reality: Standard Dual-Region replication asynchronously replicates 99.9% of data within 1 hour. To achieve a contractual 15-minute RPO backed by a financial SLA, administrators must explicitly enable Turbo Replication.
Test Your Knowledge

A financial enterprise requires a transactional Cloud SQL for PostgreSQL database that can withstand an entire zone outage with zero data loss (RPO = 0) and fail over automatically within about a minute. Which configuration satisfies these requirements?

A

Deploy a primary Cloud SQL instance paired with an asynchronous cross-region read replica.

B

Deploy a single-zone Cloud SQL instance with automated daily backups and Point-in-Time Recovery (PITR) enabled.

C

Deploy two Cloud SQL instances in the same zone behind a Cloud Load Balancer with round-robin traffic routing.

D

Configure Cloud SQL Regional High Availability (HA) with a standby instance located in a separate zone within the same region.

Test Your Knowledge

A healthcare company stores critical patient medical imaging files in Google Cloud Storage. Compliance mandates require that all data must survive a total regional facility disaster with a certified Recovery Point Objective (RPO) of no more than 15 minutes. How should the bucket be architected?

A

Create a Dual-Region Cloud Storage bucket across two geographical regions with Turbo Replication enabled.

B

Create a Regional Cloud Storage bucket and schedule an hourly Storage Transfer Service job to copy files to a backup region.

C

Create a Regional Cloud Storage bucket with Object Versioning and a 15-minute Object Lifecycle Management deletion rule.

D

Create a Multi-Region Cloud Storage bucket using the Archive storage class.

Test Your Knowledge

A database administrator accidentally runs a batch SQL script that updates every customer record to an invalid status at 14:23:15 UTC on a production Cloud SQL MySQL instance configured with Regional High Availability. How should the administrator restore the database to its pristine state immediately before the incident?

A

Trigger a manual failover to the standby HA instance in the secondary zone.

B

Perform a Point-in-Time Recovery (PITR) to a new instance, specifying a timestamp of 14:23:14 UTC.

C

Revert the instance to the previous midnight automated backup snapshot and discard today's transaction logs.

D

Promote the cross-region read replica to primary.

Sections you finish are checked off in the contents.