4.2 vSAN Fault Domains & Stretched Clusters

Key Takeaways

  • vSAN Fault Domains group ESXi hosts by physical failure boundaries (racks, enclosures, or power feeds), requiring a minimum of 2n + 1 fault domains to support Failures to Tolerate (FTT) = n with RAID-1.

  • A vSAN Stretched Cluster provides synchronous zero-RPO active/active replication across two data sites (Preferred and Secondary), coordinated by a third independent Witness Host.

  • Network latency across the Inter-Site Link (ISL) between the Preferred and Secondary data sites must not exceed 5 ms Round Trip Time (RTT), with at least 10 Gbps dedicated bandwidth.

  • The witness host sits at a third site on a routed network, with up to 200 ms RTT for stretched clusters of up to 10 hosts per site (100 ms for larger ones); it stores only witness components.

  • Split-brain avoidance relies on quorum voting: an object remains accessible only if more than 50% of its component votes are online; during an ISL partition, the Preferred Site and Witness maintain quorum.

Last updated: September 2026

4.2 vSAN Fault Domains & Stretched Clusters

Enterprise storage architectures must survive physical infrastructure failures that extend beyond individual drive or host crashes. In high-density data centers, the loss of an entire server rack due to a top-of-rack (ToR) switch failure, a power distribution unit (PDU) trip, or a cooling anomaly can simultaneously take down multiple hosts. Furthermore, mission-critical applications frequently demand continuous availability across entire geographic availability zones or metro campuses.

VMware vSAN provides two primary mechanisms to address these catastrophic failure scopes: vSAN Fault Domains for local rack-level protection, and vSAN Stretched Clusters for metro-distance multi-site zero-data-loss availability.


Physical Rack Resiliency: vSAN Fault Domains

By default, vSAN treats each individual ESXi host as an independent failure domain. When a virtual machine storage policy mandates redundant mirrors (e.g., RAID-1 FTT=1), vSAN places replica component 1 on Host A, replica component 2 on Host B, and the witness component on Host C.

However, if Host A, Host B, and Host C all reside within the same physical server rack, a failure of that rack's power or ToR switch will take all three components offline simultaneously, causing complete data unavailability. vSAN Fault Domains solve this problem by allowing administrators to map ESXi hosts to physical boundaries such as racks, server enclosures, blade chassis, or distinct electrical circuits.

+-----------------------------------------------------------------------------+
|                        vSAN Fault Domain Architecture                       |
+-----------------------------------------------------------------------------+
|                                                                             |
|   +-------------------+  +-------------------+  +-------------------+       |
|   |  Fault Domain 1   |  |  Fault Domain 2   |  |  Fault Domain 3   |       |
|   |     (Rack 01)     |  |     (Rack 02)     |  |     (Rack 03)     |       |
|   | +---------------+ |  | +---------------+ |  | +---------------+ |       |
|   | |  ESXi Host 1  | |  | |  ESXi Host 3  | |  | |  ESXi Host 5  | |       |
|   | +---------------+ |  | +---------------+ |  | +---------------+ |       |
|   | |  ESXi Host 2  | |  | |  ESXi Host 4  | |  | |  ESXi Host 6  | |       |
|   | +---------------+ |  | +---------------+ |  | +---------------+ |       |
|   |         |         |  |         |         |  |         |         |       |
|   |  [Data Comp 1]    |  |  [Data Comp 2]    |  |  [Witness Comp]   |       |
|   +-------------------+  +-------------------+  +-------------------+       |
|             |                      |                      |                 |
|             +----------------------+----------------------+                 |
|                                    |                                        |
|            [ VMDK Object Protected Across 3 Physical Racks ]                |
+-----------------------------------------------------------------------------+

Fault Domain Configuration Rules

  1. Minimum Domain Count Rules:
    • To support FTT=1 with RAID-1 (Mirroring): Minimum 3 fault domains are required (Data Component 1 in Domain 1, Data Component 2 in Domain 2, Witness in Domain 3).
    • To support FTT=2 with RAID-1 (Mirroring): Minimum 5 fault domains are required (2n + 1).
    • To support FTT=3 with RAID-1 (Mirroring): Minimum 7 fault domains are required.
    • To support RAID-5 (3+1 Erasure Coding, OSA): Minimum 4 fault domains are required (3 data fragments + 1 parity fragment, each in a different fault domain). vSAN ESA instead uses RAID-5 as 2+1 (3 fault domains) or 4+1 (adaptive, used with 6 or more).
    • To support RAID-6 (4+2 Erasure Coding): Minimum 6 fault domains are required (4 data fragments + 2 parity fragments across 6 separate domains).
  2. Symmetrical Sizing Mandate:
    • Administrators must configure fault domains with approximately equal storage capacity and compute resources. If Rack 1 contains 100 TB of capacity while Rack 2 and Rack 3 contain only 20 TB each, vSAN cannot balance components evenly across the three domains. The cluster will encounter premature capacity exhaustion in the smaller domains, causing storage policy compliance errors.
  3. Unassigned Hosts:
    • Any host not explicitly placed into a custom fault domain remains in its own default individual domain. Mixing explicitly named multi-host fault domains with unassigned single-host domains often leads to unexpected component placement failures.

Stretched Cluster Architecture & Topology

A vSAN Stretched Cluster extends a single vSAN cluster across two physically separated sites at metro distance, limited in practice by the inter-site latency requirement. This provides active/active compute, synchronous storage replication with a Recovery Point Objective of zero (RPO = 0), and automated failover via vSphere High Availability (HA) with a near-zero Recovery Time Objective (RTO).

A stretched cluster consists of three distinct logical entities:

  1. Preferred Data Site:
    • Hosts active workload virtual machines and primary data components under normal operations. Designating one site as Preferred provides a deterministic tie-breaker during network split events.
  2. Secondary Data Site:
    • Houses mirrored data components synchronously replicated from the Preferred site. Virtual machines can also run on the Secondary site during normal operations when configured with site affinity rules.
  3. Witness Host:
    • A dedicated virtual or physical ESXi instance deployed at an independent third location (such as a cloud provider or third data center) completely outside the failure domain of Site A and Site B.
    • Critical Role: The Witness Host never runs virtual machine compute workloads and never stores user data blocks. It hosts only small witness components (metadata, no VM data) and holds the tie-breaking votes that maintain quorum.

Network Connectivity & Latency Prerequisites

Stretched clusters impose rigorous network requirements due to the synchronous nature of storage writes across the inter-site fabric.

Data Site-to-Data Site (Inter-Site Link - ISL)

  • Round Trip Time (RTT) Latency: Must be 5 ms RTT or less, because every write is mirrored synchronously to both sites.
  • Bandwidth: 10 Gbps or more is recommended for all-flash clusters, sized from peak write bandwidth. VMware's sizing formula multiplies the write bandwidth by a data multiplier (1.4) and a resynchronization multiplier (1.25): required bandwidth = write bandwidth x 1.4 x 1.25.
  • Topology: The vSAN network between the data sites can be stretched Layer 2 (recommended) or routed Layer 3, with low jitter and minimal packet loss.

Data Sites-to-Witness Host Connection

  • Round Trip Time (RTT) Latency: Up to 200 ms RTT for stretched clusters with up to 10 hosts per site; 100 ms RTT for larger clusters.
  • Bandwidth: Sized from the number of components, because witness traffic is only metadata. VMware's rule of thumb is about 2 Mbps for every 1,000 components.
  • Traffic Isolation: Witness network traffic is separated from inter-site data traffic by binding a dedicated VMkernel interface (vmk) tagged with the vSAN Witness traffic type, typically routed over standard corporate WAN or VPN links.

Quorum, Voting, and Split-Brain Avoidance

To prevent the catastrophic scenario known as split-brain—where both sites believe the other has failed, mount the same virtual disks independently, and write conflicting data—vSAN relies on a strict quorum voting algorithm.

The 50% Quorum Rule

Every vSAN object is decomposed into components, and each component is assigned a number of integer votes by the CMMDS directory service. An object remains accessible to ESXi hosts if and only if strictly greater than 50% of the total component votes are reachable.

In a standard 2-site stretched cluster configured with FTT=1 Mirroring:

  • Preferred Site data components possess equal voting weight.
  • Secondary Site data components possess equal voting weight.
  • The Witness Host possesses the balance of votes, acting as the definitive tie-breaker.
+-----------------------------------------------------------------------------+
|                   Stretched Cluster Quorum Voting Model                     |
+-----------------------------------------------------------------------------+
|                                                                             |
|      [ Preferred Site A ]                             [ Secondary Site B ]  |
|     Data Mirror Components                           Data Mirror Components |
|          (40% Votes)                                      (40% Votes)       |
|               \                                                /            |
|                \                                              /             |
|                 \                                            /              |
|                  +------------------+-----------------------+               |
|                                     |                                       |
|                             [ Witness Host ]                                |
|                        Witness Metadata Components                          |
|                                (20% Votes)                                  |
+-----------------------------------------------------------------------------+

Operational Failure Scenarios & Behavioral Outcomes

  1. Inter-Site Link (ISL) Failure (Partition Event):

    • Condition: The communication link between Site A and Site B fails completely, but both Site A and Site B can still communicate independently with the Witness Host.
    • Quorum Evaluation:
      • Site A sees its own components (40%) + Witness components (20%) = 60% total votes (> 50% quorum achieved).
      • Site B sees its own components (40%), but cannot reach Site A. When Site B queries the Witness, the Witness recognizes that Site A is the designated Preferred Site. The Witness grants its votes exclusively to Site A.
    • Outcome: Site A maintains full quorum and continues processing I/O uninterrupted. Storage on Site B goes offline. Virtual machines running on Site B lose access to storage; vSphere HA detects host unreachability and restarts those VMs on hosts in Site A.
  2. Complete Preferred Site Outage (Site A Failure):

    • Condition: Site A suffers a catastrophic power grid collapse. Both data links and witness connections to Site A drop.
    • Quorum Evaluation: Site B sees its local components (40%) + Witness components (20%) = 60% total votes (> 50% quorum achieved).
    • Outcome: Site B successfully forms cluster quorum with the Witness Host. vSphere HA initiates automated failover, restarting all production workloads on Site B ESXi hosts.
  3. Witness Host Failure:

    • Condition: The virtual machine hosting the Witness appliance crashes or its WAN connection drops. Both Data Site A and Data Site B remain fully healthy and connected across the low-latency ISL.
    • Quorum Evaluation: Site A (40%) + Site B (40%) = 80% total votes (> 50% quorum achieved).
    • Outcome: The cluster stays fully operational with no workload downtime, but it has lost its tie-breaker. If the ISL also fails while the witness is down, each site holds only half the votes, so objects lose quorum at both sites until connectivity returns.

2-Node vSAN Clusters

A common variant of the stretched cluster architecture is the vSAN 2-Node Cluster, designed specifically for Remote Office / Branch Office (ROBO) and retail edge environments.

  • Direct-Connect Architecture: In a 2-Node deployment, the two ESXi physical data nodes are directly connected back-to-back using 10GbE or 25GbE copper/optical direct-attach cables for vSAN data and vMotion traffic. This eliminates the expense of costly 10GbE physical switches at remote locations.
  • Shared Witness Appliance: A single physical or virtual ESXi host in a central data center can act as the Witness Host for multiple independent 2-Node ROBO clusters, significantly lowering per-site licensing and hardware costs.

Network Latency and Bandwidth Specifications

Connection PathMax Latency (RTT)Min BandwidthTraffic TypeProtocol / Routing
Data Site to Data Site (ISL)<= 5 ms RTT10 Gbps (All-Flash); 1 GbE (Hybrid)vSAN Data, vMotion, VM TrafficStretched L2 or Low-Latency Routed L3
Data Site to Witness (up to 10 hosts per site)<= 200 ms RTT~2 Mbps per 1,000 componentsvSAN Witness Heartbeats / MetadataRouted Layer 3 (Standard WAN / VPN)
Data Site to Witness (more than 10 hosts per site)<= 100 ms RTT~2 Mbps per 1,000 componentsvSAN Witness Heartbeats / MetadataRouted Layer 3 (High-Bandwidth WAN)
2-Node Direct Connect Data Link< 1 ms RTT10 Gbps / 25 GbpsvSAN Data & vMotionDirect Crossover Cable (Point-to-Point)
Loading diagram...
vSAN Stretched Cluster Architecture and Quorum Network Topologies
Test Your Knowledge

In a production vSAN Stretched Cluster spanning Site A (Preferred) and Site B (Secondary), a complete fiber cut severs the Inter-Site Link (ISL). Both Site A and Site B maintain healthy, unhindered network connectivity to the third-site Witness Host. What is the immediate operational outcome of this partition event?

A

Both Site A and Site B immediately suspend all virtual machines to prevent split-brain data corruption until the ISL recovers.

B

Site A maintains quorum by pairing with the Witness Host and continues processing I/O, while Site B storage goes offline and vSphere HA restarts Site B workloads on Site A.

C

Site B assumes the active master role because it has fewer active I/O transactions, forcing Site A to become an isolated standby site.

D

The cluster splits into two independent, autonomous vSAN clusters, allowing both sites to independently accept writes to local VMDK mirrors.

Test Your Knowledge

A network infrastructure team is provisioning dark fiber and Layer 3 routing for an enterprise vSAN Stretched Cluster between two primary data centers running all-flash workloads. What are the maximum allowable round-trip time (RTT) network latency and minimum recommended bandwidth required across the Inter-Site Link (ISL)?

A

Maximum 20 ms RTT latency and minimum 1 Gbps bandwidth.

B

Maximum 100 ms RTT latency and minimum 10 Gbps bandwidth.

C

Maximum 1 ms RTT latency and minimum 40 Gbps bandwidth.

D

Maximum 5 ms RTT latency and minimum 10 Gbps bandwidth.

Test Your Knowledge

A system administrator is designing a high-availability vSAN cluster across physical server racks to protect virtual machine objects against simultaneous rack power outages. If the storage policy is configured with Primary Level of Failures to Tolerate (PFTT) = 1 using RAID-1 (Mirroring), what is the minimum number of vSAN Fault Domains that must be configured?

A

3 Fault Domains

B

2 Fault Domains

C

4 Fault Domains

D

5 Fault Domains

Sections you finish are checked off in the contents.