2.5 Fault Tolerance, Proactive HA & Advanced Cluster Options

Key Takeaways

  • vSphere Fault Tolerance keeps a live secondary VM on another host in lockstep with the primary, so a host failure causes no downtime and no loss of in-memory state.

  • A Fault Tolerance VM can have up to 8 vCPUs with vSphere Enterprise Plus or vSphere Foundation, and up to 2 vCPUs with vSphere Standard or Enterprise.

  • By default a host can run 4 fault-tolerant VMs (das.maxftvmsperhost) and 8 FT vCPUs in total (das.maxftvcpusperhost), and VMware recommends a dedicated 10 Gbit FT logging network.

  • Fault Tolerance does not support VM snapshots, Storage vMotion, linked clones, vVols datastores, I/O filters, or VBS-enabled VMs.

  • Proactive HA uses vendor hardware health providers and DRS to move VMs off degraded hosts, using Quarantine mode, Mixed mode, or Maintenance mode.

Last updated: September 2026

2.5 Fault Tolerance, Proactive HA & Advanced Cluster Options

vSphere HA (Section 2.1) restarts VMs after a failure, so the guest reboots and in-memory state is lost. Fault Tolerance (FT) and Proactive HA fill the gaps: FT avoids the restart entirely, and Proactive HA acts before a failing host goes down.

How Fault Tolerance Works

When you turn on FT for a VM, vSphere creates a Secondary VM on a different host and keeps it continuously synchronized with the Primary VM over a dedicated FT logging network. If the primary's host fails, the secondary immediately takes over with the VM's entire running state preserved: no guest reboot, no lost connections, no lost in-memory data. vSphere HA then creates a new secondary on another host to restore protection.

vSphere HAvSphere Fault Tolerance
Protection modelRestart the VM on another hostInstant failover to a running secondary
DowntimeGuest reboot (minutes)None
In-memory stateLostPreserved
Resource costSpare capacity for restartsDouble compute, memory, and storage for each protected VM
ScopeWhole clusterSelected VMs

Fault Tolerance Use Cases (1.4.5)

VMware's documentation lists these use cases:

  • Applications that must always be available, especially those with long-lasting client connections that users want to keep through a hardware failure.
  • Custom applications that have no other way of doing clustering.
  • Cases where custom clustering solutions would be too complicated to configure and maintain.
  • On-demand protection during critical periods, such as quarter-end reporting, with FT turned off afterward.

FT protects against host failure. It does not protect against guest OS crashes or application bugs, because the secondary mirrors the same software state.

Requirements and Limits

AreaRequirement
ClustervSphere HA enabled; at least two hosts with access to the VM's storage
CPUvMotion-compatible CPUs with hardware MMU virtualization (Intel EPT or AMD RVI); Intel Sandy Bridge or later, AMD Bulldozer or later
NetworkingA VMkernel adapter enabled for Fault Tolerance logging, ideally a dedicated, low-latency 10 Gbit network, plus vMotion networking
vCPUs per FT VMUp to 8 with vSphere Enterprise Plus or vSphere Foundation (VVF); up to 2 with vSphere Standard or Enterprise
FT VMs per host4 by default (das.maxftvmsperhost, set to 0 to disable the check)
FT vCPUs per host8 by default (das.maxftvcpusperhost)
StorageThe secondary keeps its own copy of the virtual disks, which can be on a different datastore. Thin and thick disks are supported

Features Not Supported with FT

  • Snapshots must be removed or committed before FT can be turned on.
  • Storage vMotion cannot be used on an FT VM.
  • Linked clones cannot use FT, and you cannot create a linked clone from an FT VM.
  • vVols datastores, storage-based policy management (except on vSAN), and I/O filters.
  • VBS-enabled VMs.
  • VM encryption became supported with FT in vSphere 7.0 Update 2. FT is also not supported on NSX-created port groups.

Turning FT On

Right-click the VM, choose Fault Tolerance > Turn On Fault Tolerance, select a datastore for the secondary's files and a host for the secondary, and finish. From the same menu you can Test Failover, Test Restart Secondary, Suspend, Migrate Secondary, or Turn Off FT.

Proactive HA (4.6)

Proactive HA integrates vCenter with a hardware vendor's health provider plug-in (from the server OEM's management integration). When the provider reports a degraded component, such as a failed fan, a lost power supply, or correctable memory errors, DRS moves VMs off the host before it fails.

SettingOptions
Automation levelManual (recommendations) or Automated (DRS acts)
Quarantine modeFor all failures: DRS avoids placing VMs on the host and evacuates it when that won't hurt performance or break rules
Mixed modeQuarantine for moderate failures, maintenance mode for severe failures
Maintenance modeFor all failures: the host is fully evacuated

Proactive HA requires DRS and a supported vendor provider. It complements HA; it does not replace it.

HA and DRS Advanced Options You Should Recognize

OptionPurpose
VM restart priority (Lowest, Low, Medium, High, Highest)Order in which HA restarts VMs
Start next priority VMs whenResources allocated, powered on, guest heartbeats detected, or app heartbeats detected, plus an optional delay
VM MonitoringRestart VMs whose VMware Tools heartbeats and I/O stop (VM monitoring or VM and application monitoring)
das.isolationaddress0-9Extra isolation addresses to ping
das.usedefaultisolationaddressWhether to also ping the default gateway
das.heartbeatdsperhostHeartbeat datastores per host (default 2, maximum 5)
das.isolationshutdowntimeoutWait for guest shutdown in the "Shut down and restart" isolation response (default 300 seconds)
Admission control policyCluster resource percentage (default), slot policy, or dedicated failover hosts (Section 2.1)
DRS VM distribution and CPU over-commitment optionsBalance VM counts per host and cap the vCPU-to-pCPU ratio
Scalable shares (vSphere 7 and later)Make resource pool shares scale with the number of VMs in each pool

Exam Traps

  • FT is not a backup or DR tool. Both VMs run the same software, so a guest crash or corruption affects both. For site loss, use SRM or a stretched cluster.
  • Licensing sets the vCPU ceiling. A 4-vCPU VM can use FT on Enterprise Plus or VVF, but not on Standard or Enterprise, which cap FT at 2 vCPUs.
  • FT doubles resources. Each protected VM consumes its compute, memory, and storage twice, plus FT logging bandwidth. Plan capacity before protecting many VMs.
  • Proactive HA needs a vendor provider and DRS. Without a hardware health provider registered in vCenter, there is nothing to react to.
Test Your Knowledge

An organization runs a legacy application that has no clustering support and holds long-lived client connections that must survive a host failure without interruption. Which vSphere feature should protect this VM?

A

vSphere HA with Highest restart priority

B

vSphere Fault Tolerance

C

Proactive HA in Quarantine mode

D

vSphere Replication with a 5-minute RPO

Test Your Knowledge

A cluster uses Proactive HA. The administrator wants hosts with moderate hardware degradation to be avoided for new VM placement, but hosts with severe degradation to be evacuated completely. Which remediation setting achieves this?

A

Quarantine mode for all failures

B

Maintenance mode for all failures

C

Mixed mode

D

Manual automation level with VM Monitoring

Test Your Knowledge

An administrator tries to turn on Fault Tolerance for a VM, but the operation is blocked. The VM runs on a VMFS datastore, has 4 vCPUs, and the cluster is licensed with vSphere Enterprise Plus. What is the most likely blocker?

A

The VM has existing snapshots

B

The VM uses a thin-provisioned virtual disk

C

The VM has more than 2 vCPUs

D

The datastore is VMFS-6 rather than VMFS-5

Sections you finish are checked off in the contents.