2.5 Fault Tolerance, Proactive HA & Advanced Cluster Options
Key Takeaways
vSphere Fault Tolerance keeps a live secondary VM on another host in lockstep with the primary, so a host failure causes no downtime and no loss of in-memory state.
A Fault Tolerance VM can have up to 8 vCPUs with vSphere Enterprise Plus or vSphere Foundation, and up to 2 vCPUs with vSphere Standard or Enterprise.
By default a host can run 4 fault-tolerant VMs (das.maxftvmsperhost) and 8 FT vCPUs in total (das.maxftvcpusperhost), and VMware recommends a dedicated 10 Gbit FT logging network.
Fault Tolerance does not support VM snapshots, Storage vMotion, linked clones, vVols datastores, I/O filters, or VBS-enabled VMs.
Proactive HA uses vendor hardware health providers and DRS to move VMs off degraded hosts, using Quarantine mode, Mixed mode, or Maintenance mode.
2.5 Fault Tolerance, Proactive HA & Advanced Cluster Options
vSphere HA (Section 2.1) restarts VMs after a failure, so the guest reboots and in-memory state is lost. Fault Tolerance (FT) and Proactive HA fill the gaps: FT avoids the restart entirely, and Proactive HA acts before a failing host goes down.
How Fault Tolerance Works
When you turn on FT for a VM, vSphere creates a Secondary VM on a different host and keeps it continuously synchronized with the Primary VM over a dedicated FT logging network. If the primary's host fails, the secondary immediately takes over with the VM's entire running state preserved: no guest reboot, no lost connections, no lost in-memory data. vSphere HA then creates a new secondary on another host to restore protection.
| vSphere HA | vSphere Fault Tolerance | |
|---|---|---|
| Protection model | Restart the VM on another host | Instant failover to a running secondary |
| Downtime | Guest reboot (minutes) | None |
| In-memory state | Lost | Preserved |
| Resource cost | Spare capacity for restarts | Double compute, memory, and storage for each protected VM |
| Scope | Whole cluster | Selected VMs |
Fault Tolerance Use Cases (1.4.5)
VMware's documentation lists these use cases:
- Applications that must always be available, especially those with long-lasting client connections that users want to keep through a hardware failure.
- Custom applications that have no other way of doing clustering.
- Cases where custom clustering solutions would be too complicated to configure and maintain.
- On-demand protection during critical periods, such as quarter-end reporting, with FT turned off afterward.
FT protects against host failure. It does not protect against guest OS crashes or application bugs, because the secondary mirrors the same software state.
Requirements and Limits
| Area | Requirement |
|---|---|
| Cluster | vSphere HA enabled; at least two hosts with access to the VM's storage |
| CPU | vMotion-compatible CPUs with hardware MMU virtualization (Intel EPT or AMD RVI); Intel Sandy Bridge or later, AMD Bulldozer or later |
| Networking | A VMkernel adapter enabled for Fault Tolerance logging, ideally a dedicated, low-latency 10 Gbit network, plus vMotion networking |
| vCPUs per FT VM | Up to 8 with vSphere Enterprise Plus or vSphere Foundation (VVF); up to 2 with vSphere Standard or Enterprise |
| FT VMs per host | 4 by default (das.maxftvmsperhost, set to 0 to disable the check) |
| FT vCPUs per host | 8 by default (das.maxftvcpusperhost) |
| Storage | The secondary keeps its own copy of the virtual disks, which can be on a different datastore. Thin and thick disks are supported |
Features Not Supported with FT
- Snapshots must be removed or committed before FT can be turned on.
- Storage vMotion cannot be used on an FT VM.
- Linked clones cannot use FT, and you cannot create a linked clone from an FT VM.
- vVols datastores, storage-based policy management (except on vSAN), and I/O filters.
- VBS-enabled VMs.
- VM encryption became supported with FT in vSphere 7.0 Update 2. FT is also not supported on NSX-created port groups.
Turning FT On
Right-click the VM, choose Fault Tolerance > Turn On Fault Tolerance, select a datastore for the secondary's files and a host for the secondary, and finish. From the same menu you can Test Failover, Test Restart Secondary, Suspend, Migrate Secondary, or Turn Off FT.
Proactive HA (4.6)
Proactive HA integrates vCenter with a hardware vendor's health provider plug-in (from the server OEM's management integration). When the provider reports a degraded component, such as a failed fan, a lost power supply, or correctable memory errors, DRS moves VMs off the host before it fails.
| Setting | Options |
|---|---|
| Automation level | Manual (recommendations) or Automated (DRS acts) |
| Quarantine mode | For all failures: DRS avoids placing VMs on the host and evacuates it when that won't hurt performance or break rules |
| Mixed mode | Quarantine for moderate failures, maintenance mode for severe failures |
| Maintenance mode | For all failures: the host is fully evacuated |
Proactive HA requires DRS and a supported vendor provider. It complements HA; it does not replace it.
HA and DRS Advanced Options You Should Recognize
| Option | Purpose |
|---|---|
| VM restart priority (Lowest, Low, Medium, High, Highest) | Order in which HA restarts VMs |
| Start next priority VMs when | Resources allocated, powered on, guest heartbeats detected, or app heartbeats detected, plus an optional delay |
| VM Monitoring | Restart VMs whose VMware Tools heartbeats and I/O stop (VM monitoring or VM and application monitoring) |
das.isolationaddress0-9 | Extra isolation addresses to ping |
das.usedefaultisolationaddress | Whether to also ping the default gateway |
das.heartbeatdsperhost | Heartbeat datastores per host (default 2, maximum 5) |
das.isolationshutdowntimeout | Wait for guest shutdown in the "Shut down and restart" isolation response (default 300 seconds) |
| Admission control policy | Cluster resource percentage (default), slot policy, or dedicated failover hosts (Section 2.1) |
| DRS VM distribution and CPU over-commitment options | Balance VM counts per host and cap the vCPU-to-pCPU ratio |
| Scalable shares (vSphere 7 and later) | Make resource pool shares scale with the number of VMs in each pool |
Exam Traps
- FT is not a backup or DR tool. Both VMs run the same software, so a guest crash or corruption affects both. For site loss, use SRM or a stretched cluster.
- Licensing sets the vCPU ceiling. A 4-vCPU VM can use FT on Enterprise Plus or VVF, but not on Standard or Enterprise, which cap FT at 2 vCPUs.
- FT doubles resources. Each protected VM consumes its compute, memory, and storage twice, plus FT logging bandwidth. Plan capacity before protecting many VMs.
- Proactive HA needs a vendor provider and DRS. Without a hardware health provider registered in vCenter, there is nothing to react to.
An organization runs a legacy application that has no clustering support and holds long-lived client connections that must survive a host failure without interruption. Which vSphere feature should protect this VM?
vSphere HA with Highest restart priority
vSphere Fault Tolerance
Proactive HA in Quarantine mode
vSphere Replication with a 5-minute RPO
A cluster uses Proactive HA. The administrator wants hosts with moderate hardware degradation to be avoided for new VM placement, but hosts with severe degradation to be evacuated completely. Which remediation setting achieves this?
Quarantine mode for all failures
Maintenance mode for all failures
Mixed mode
Manual automation level with VM Monitoring
An administrator tries to turn on Fault Tolerance for a VM, but the operation is blocked. The VM runs on a VMFS datastore, has 4 vCPUs, and the cluster is licensed with vSphere Enterprise Plus. What is the most likely blocker?
The VM has existing snapshots
The VM uses a thin-provisioned virtual disk
The VM has more than 2 vCPUs
The datastore is VMFS-6 rather than VMFS-5
Sections you finish are checked off in the contents.