2.1 vSphere High Availability (HA) Architecture
Key Takeaways
The Fault Domain Manager (FDM) agent runs on every ESXi host in an HA cluster; the host connected to the most datastores is elected master, and a tie goes to the host with the lexically highest Managed Object ID (MOID).
HA hosts exchange management-network heartbeats every second on port 8182, and by default HA picks two heartbeat datastores per host (up to five) to tell host isolation apart from host failure.
Host isolation responses determine VM behavior when network connectivity fails: Disabled (leave VMs running), Shut down and restart VMs (graceful guest OS shutdown with fallback to hard stop), and Power off and restart VMs (immediate termination).
Admission control policies guarantee failover capacity through cluster resource percentage calculation, slot policies based on maximum CPU/memory reservations, or dedicated failover hosts.
VM Component Protection (VMCP) reacts to storage failures: Permanent Device Loss (PDL) triggers an immediate restart, while All Paths Down (APD) waits for the 140-second Misc.APDTimeout plus a VMCP delay of 3 minutes by default.
vSphere High Availability (HA) Architecture
VMware vSphere High Availability (HA) provides foundational compute resiliency by pooling virtual machines and the ESXi hosts they reside upon into a managed failure domain. When a physical host experiences an outage, hardware fault, or management network isolation, vSphere HA automatically detects the failure condition and restarts affected virtual machines on surviving ESXi hosts with available compute capacity.
Fault Domain Manager (FDM) & Master Host Election
The operational core of vSphere HA is the Fault Domain Manager (FDM) agent (fdm daemon), which installs automatically on every ESXi host when HA is enabled on a cluster. The FDM agent communicates directly with the local host daemon (hostd) through internal vSphere Infrastructure APIs and interfaces with vCenter Server to receive cluster state configurations and protected virtual machine inventories.
Within an HA cluster, hosts operate in one of two distinct roles:
- Master Host: A single ESXi host responsible for monitoring subordinate hosts, tracking the power state and placement of all protected virtual machines, coordinating restart sequences following a failure, and reporting cluster health and alarms to vCenter Server.
- Subordinate (Slave) Hosts: All remaining hosts in the cluster. Subordinates monitor locally running virtual machines, report runtime state changes to the master host, monitor the master host's heartbeats, and participate in master election if the master fails.
The Master Election Algorithm
A master election initiates automatically whenever vSphere HA is enabled, when the current master host fails or is placed into maintenance mode, or when a network partition separates cluster members. The election algorithm evaluates candidates using specific hierarchical criteria:
- Datastore Connectivity: The host with access to the greatest number of connected datastores wins the election. Maximizing storage visibility ensures the elected master can communicate with and monitor as many virtual machine state files as possible.
- MOID Lexical Tie-Break: If several hosts see the same number of datastores, the host with the lexically highest Managed Object ID (MOID) wins. The comparison is text, character by character, so
host-98beatshost-102because the character 9 sorts above 1, even though 102 is the larger number.
+--------------------------------------------------------------------------+
| HA Master Election Algorithm |
+--------------------------------------------------------------------------+
| 1. Trigger: HA enabled, master failure, maintenance, or network split |
| 2. Datastore Count: Host with access to highest datastore count wins |
| 3. Tie-Break: Lexically highest host MOID wins ('host-98' > 'host-102') |
| 4. Role Assignment: Winner becomes Master; remaining become Subordinates|
+--------------------------------------------------------------------------+
Heartbeat Mechanisms & Failure Detection
vSphere HA relies on two independent, redundant heartbeat channels to assess host and network health: Management Network Heartbeats and Datastore Heartbeats.
Management Network Heartbeats
The master and its subordinate hosts exchange network heartbeats every second over port 8182. When the master stops receiving a subordinate's heartbeats, it runs further checks before declaring the host failed:
- The master attempts to ping the subordinate's management IP address via ICMP.
- The master inspects the subordinate host's datastore heartbeat files.
Datastore Heartbeats
Datastore heartbeats serve as an out-of-band communication mechanism to differentiate between an isolated/partitioned host and a completely deceased host:
- By default vSphere HA selects two heartbeat datastores per host (
das.heartbeatDsPerHost, maximum 5). With fewer than two shared datastores available, HA still runs but raises a configuration warning. - Datastores can be selected automatically by vCenter or manually designated by the administrator to prioritize specific shared storage arrays.
- Each host creates and locks a heartbeat file within the hidden
.vSphere-HAdirectory located at the root of the designated datastores. ESXi hosts periodically write a timestamp to their heartbeat file.
Host Failure States
By cross-referencing management network heartbeats, ICMP pings, and datastore heartbeat timestamps, the master host categorizes host states into three distinct conditions:
| Failure State | Network Heartbeat (UDP 8182) | ICMP Ping to Management IP | Datastore Heartbeat Active? | Master Host Action |
|---|---|---|---|---|
| Host Failed | Stopped | Fails | No (Timestamp stagnant) | Declares host dead; immediately initiates VM restarts on surviving hosts |
| Host Network Partitioned | Stopped between partitions | May succeed locally | Yes (Writing to storage) | Sub-cluster forms; isolated partition elects secondary master; VMs remain running |
| Host Network Isolated | Stopped to all peers | Fails to default gateway | Yes (Writing to storage) | Host executes its configured Isolation Response; master waits for lock release |
Host Network Isolation Mechanics
A host enters the Host Network Isolated state when it loses management network connectivity with all other cluster hosts AND fails to ping its configured isolation address (by default, the management network default gateway; custom addresses can be added with das.isolationaddress0 through das.isolationaddress9). While cut off from the management network, the host still accesses its storage fabric and writes to its datastore heartbeat files.
Host Isolation Responses
When an ESXi host detects that it is isolated, it executes its configured Host Isolation Response. Selecting the correct response depends on workload characteristics, clustering requirements, and storage locking protocols:
1. Disabled
- Mechanics: The isolated host takes no action on its running virtual machines. VMs continue running uninterrupted.
- Operational Context: Recommended when network transient drops are common and VMs do not rely on management network reachability. However, because the isolated host retains the VMFS file locks (
.lck), the master host cannot restart the VMs elsewhere, leaving workloads inaccessible if their tenant network is also severed.
2. Shut Down and Restart VMs
- Mechanics: The isolated host initiates a graceful guest operating system shutdown via VMware Tools. It waits for the duration configured by
das.isolationShutdownTimeout(default: 300 seconds). If the guest OS fails to shut down within this window, the host issues a hard power-off command. Once the host powers off the VM and releases the storage lock, the master host restarts the VM on a reachable cluster host. - Operational Context: Ideal for transactional applications, messaging brokers, and database servers where sudden termination risks file system or database corruption.
3. Power Off and Restart VMs
- Mechanics: The isolated host immediately hard powers off all running virtual machines. Storage locks are released instantly, allowing the master host to initiate restarts on surviving hosts with minimal delay.
- Operational Context: Recommended for stateless workloads, clustered applications with native replication, or enterprise environments where minimizing the Recovery Time Objective (RTO) is prioritized over graceful OS termination.
Admission Control Policies
Admission Control is a cluster-level gatekeeping mechanism that guarantees sufficient compute capacity is reserved to accommodate host failures. If a virtual machine power-on or vMotion migration would violate the configured failover constraints, vCenter blocks the action.
| Admission Control Policy | Calculation Method | Advantages | Disadvantages & Operational Risks |
|---|---|---|---|
| Cluster Resource Percentage | Administrator specifies a dedicated percentage of total cluster CPU and memory reserved for failover (e.g., 25% in a 4-host cluster). | Dynamically adjusts as hosts are added or removed; accommodates variable VM sizing evenly. | Large outlier VMs with high reservations can exhaust unreserved capacity unexpectedly. |
| Slot Policy (Fixed or Calculated) | A "slot" represents the compute capacity required for one VM. The slot size is defined by the largest CPU reservation and largest memory reservation among all running VMs. | Guarantees restart capacity for any VM regardless of cluster distribution. | Slot Fragmentation: An outlier VM with an 8 GHz CPU reservation and 64 GB RAM forces all slots to that size, artificially lowering total available slots. |
| Dedicated Failover Hosts | One or more ESXi hosts are designated strictly as standby failover targets and run zero workloads during normal operation. | Completely predictable failover placement; zero resource contention on surviving nodes. | Inefficient; expensive compute and memory resources sit idle 99% of the operational lifecycle. |
Performance Degradation VMs Tolerate: This setting (default 100%) makes vSphere HA raise a warning when a failover would leave VMs with less performance than you are willing to accept. At 0%, HA warns whenever failover capacity cannot preserve the current performance. It is a warning, not an extra admission-control block.
VM Component Protection (VMCP)
Traditional vSphere HA only monitored ESXi host availability and guest OS heartbeats. VM Component Protection (VMCP) extends HA detection to storage interconnect failures, specifically addressing Permanent Device Loss (PDL) and All Paths Down (APD) conditions on block (FC, iSCSI, FCoE) and file (NFS) datastores.
Permanent Device Loss (PDL)
A PDL event occurs when a storage array explicitly notifies the ESXi kernel that a storage device is permanently unavailable and unrecoverable. The array returns specific SCSI Sense Codes (such as SCSI 0x5 / ASC 0x25 / ASCQ 0x00: Logical Unit Not Supported).
- Response: The PDL response can be Disabled, Issue events, or Power off and restart VMs. With restart selected, VMCP immediately powers off the affected VM and restarts it on a host that still has healthy access to the datastore.
All Paths Down (APD)
An APD event represents a condition where all storage communication paths become unreachable (for example, severed fiber cables or failed fabric switches), but no SCSI sense codes are returned. The hypervisor cannot determine whether the outage is momentary or permanent.
- Default APD Timeout: ESXi starts an APD timer set by the host advanced option
Misc.APDTimeout, 140 seconds by default. During this window ESXi keeps retrying the I/O. - Response: If the APD condition persists past the timeout, VMCP waits for its response recovery delay (3 minutes by default) and then applies the configured policy:
- Disabled or Issue events: log the condition without recovering the VM.
- Power off and restart VMs, Conservative restart policy: act only if HA can confirm another host can restart the VM.
- Power off and restart VMs, Aggressive restart policy: act even if HA cannot confirm that capacity or datastore access exists elsewhere. If the paths return before the delay ends, the VM is left running.
+--------------------------------------------------------------------------+
| VMCP Storage Failure Handling |
+--------------------------------------------------------------------------+
| Storage Outage Occurs |
| ├── SCSI Sense Code Returned (PDL) ─────────────────────────────────┐ |
| │ └── Immediate Action: Terminate VM & Restart on Surviving Host │ |
| └── No SCSI Code Returned (APD) │ |
| └── 140-Second APD Timer Runs │ |
| ├── Path Recovers: I/O Resumes Without Downtime │ |
| └── Timer + 3-min VMCP delay expire: VM restarted elsewhere ─┘ |
+--------------------------------------------------------------------------+
In a 4-node vSphere HA cluster, an election is triggered after the master host encounters an unexpected kernel panic. Host A and Host B are connected to 6 shared datastores, Host C is connected to 5 datastores, and Host D is connected to 4 datastores. Host A has Managed Object ID 'host-115' and Host B has Managed Object ID 'host-89'. Which host is elected as the new HA master?
Host C, because it possesses an odd number of datastores which prevents quorum deadlocks
Host A, because its numeric Managed Object ID suffix (115) is numerically greater than Host B's suffix (89)
Host B, because it ties for the highest datastore count and 'host-89' is lexically higher than 'host-115'
Host D, because lowest datastore connectivity forces secondary election fallbacks
An administrator manages an HA cluster hosting a latency-critical transactional database. If an ESXi host experiences a network isolation event, the database must be recovered on a healthy host as rapidly as possible without waiting for OS shutdown sequences. Which Host Isolation Response should be selected?
Power off and restart VMs
Shut down and restart VMs
Disabled
Restart Guest OS
A storage controller encounters a hardware failure that severs all Fibre Channel paths to an ESXi host. The storage array returns no SCSI sense codes to the hypervisor. How does VM Component Protection (VMCP) classify this condition, and what is the default timer duration before automated recovery occurs?
Permanent Device Loss (PDL), triggering an immediate VM termination without delay
Storage Quorum Fault, requiring manual administrative intervention within 60 seconds
Transient LUN Drop, with automated failover occurring after a fixed 30-second pause
All Paths Down (APD), initiating a default 140-second timeout before executing recovery actions
Sections you finish are checked off in the contents.