4.4 vSAN Maintenance, Health Checks & Monitoring
Key Takeaways
Skyline Health for vSAN continuously audits cluster state against the online VMware Hardware Compatibility List (HCL), verifying storage controller firmware, NVMe driver compliance, and network packet drop thresholds.
Entering Maintenance Mode provides three distinct evacuation modes: Ensure Accessibility (default, guarantees read/write access with temporary vulnerability), Full Data Migration (preserves full redundancy), and No Data Migration (evacuates zero components).
Permanent hardware failures immediately mark components as Degraded, triggering an instant rebuild, whereas host reboots or transient disconnections mark components as Absent, triggering a default 60-minute clomd delay timer.
The 60-minute delay is the vSAN Object Repair Timer, set under the cluster's vSAN Advanced Options (the host-level advanced option is VSAN.ClomRepairDelay), and can be lengthened for long maintenance windows.
When VM I/O and resync traffic contend, vSAN Adaptive Resync guarantees resync about 20% of the bandwidth and leaves the rest to VM I/O; with no contention, resync can use all available bandwidth.
4.4 vSAN Maintenance, Health Checks & Monitoring
Day-two operations in a software-defined storage environment require rigorous health diagnostics, proactive lifecycle management, and strict operational discipline during maintenance procedures. Because vSAN aggregates local server hardware into a distributed cluster, common administrative tasks—such as updating server firmware, applying ESXi security patches, or replacing failed solid-state drives—directly impact distributed storage availability and network bandwidth.
Understanding the diagnostic capabilities of Skyline Health for vSAN, the exact behavioral differences between Maintenance Mode data evacuation options, and the operational lifecycle of Degraded vs. Absent components is critical for real-world cluster management and the VCP-DCV examination.
Skyline Health for vSAN & HCL Compliance
Skyline Health for vSAN (formerly vSAN Health Check) is an intelligent, automated diagnostic framework integrated natively into vCenter Server and the ESXi VMkernel. It continuously monitors hundreds of hardware and software parameters, comparing cluster telemetry against best practices and the official VMware Compatibility Guide (VCG/HCL).
+-----------------------------------------------------------------------------+
| Skyline Health for vSAN Engine |
+-----------------------------------------------------------------------------+
| |
| +---------------------------------------------------------------------+ |
| | Online VMware Cloud Service (vSAN HCL Database & Telemetry Rules) | |
| +----------------------------------+----------------------------------+ |
| | (Periodic Database Updates) |
| +----------------------------------v----------------------------------+ |
| | vCenter Server Skyline Health Daemon | |
| +----------------------------------+----------------------------------+ |
| | |
| +------------------------+------------------------+ |
| | | | |
| +---------v-----------+ +---------v-----------+ +---------v-----------+ |
| | Hardware / HCL Check| | Network Health Check| | Physical Disk Health| |
| | Controller Firmware | | MTU / Large Ping | | SMART Attributes | |
| | Driver Compliance | | Dropped Packets | | Wear-Out Indicators | |
| +---------------------+ +---------------------+ +---------------------+ |
+-----------------------------------------------------------------------------+
Core Health Check Categories
- Hardware Compatibility (HCL):
- Queries PCI vendor IDs, device IDs, subsystem IDs, driver versions, and firmware revisions for all storage controllers, NVMe drives, and SAS/SATA host bus adapters.
- Compares active hardware states against the downloaded HCL database. If an administrator applies an uncertified ESXi patch containing a newer storage driver that has not been certified for the specific controller firmware, Skyline Health generates an immediate Red Warning.
- Network Health:
- Ping / Large Packet Ping Test: Transmits standard (1500-byte) and jumbo (9000-byte) ICMP frames across all vSAN VMkernel adapters with the Don't Fragment (DF) bit set, detecting MTU mismatches and packet drops.
- Unicast Agent Verification: Audits CMMDS unicast neighbor tables to ensure every ESXi host can communicate directly with the designated cluster master and backup master.
- Physical Disks & SMART Telemetry:
- Monitors physical drive health, SCSI error codes, write endurance wear-out percentages on flash drives, and latency anomalies. If a flash drive exceeds its rated write endurance threshold, Skyline Health flags the drive for proactive replacement before data loss occurs.
- vSAN Cluster & Capacity:
- Tracks cluster-wide disk space consumption, alerting administrators when utilization crosses the 80% threshold.
- Audits the status of the Operations Reserve and Host Rebuild Reserve, verifying that sufficient capacity remains to absorb an unexpected host failure.
Host Maintenance Mode Evacuation Options
When an administrator places an ESXi host into Maintenance Mode (e.g., to upgrade firmware, replace motherboard memory, or apply hypervisor patches), vSAN cannot simply ignore the storage components residing on that host's local drives. The administrator is prompted to select one of three distinct Data Evacuation Modes:
+-----------------------------------------------------------------------------+
| Maintenance Mode Data Evacuation Modes |
+-----------------------------------------------------------------------------+
| |
| [ Ensure Accessibility (Default) ] |
| - Migrates ONLY components necessary to keep VMs active. |
| - Leaves redundant mirrors on maintenance host. |
| - Temporary vulnerability window (FTT=0 for affected objects). |
| - Fast, minimal network traffic. Ideal for reboots & short patches. |
| |
| [ Full Data Migration ] |
| - Evacuates ALL components to other hosts in the cluster. |
| - Preserves 100% redundancy and policy compliance at all times. |
| - Generates massive network I/O; requires significant free capacity. |
| - Ideal for permanent host decommissioning or extensive hardware service. |
| |
| [ No Data Migration ] |
| - Evacuates ZERO components. Leaves all data in place. |
| - Objects with mirrors elsewhere stay online; non-redundant objects drop. |
| - Zero evacuation time. Used during full cluster shutdown. |
+-----------------------------------------------------------------------------+
1. Ensure Accessibility (Default Mode)
- Operational Mechanics: vSAN analyzes all storage objects residing across the cluster. It migrates only those components that are strictly required to ensure that every virtual machine remains accessible (i.e., retains quorum and read/write capability). If an object has a healthy mirror on another host, vSAN does not evacuate the local copy.
- Redundancy State: Virtual machines remain fully online and accessible, but their redundancy level is temporarily compromised. During the maintenance window, the object operates with reduced tolerance (effectively FTT = 0 for FTT = 1 objects). If a secondary drive or host failure occurs while the first host is in maintenance mode, data loss or VM downtime will occur.
- Recommended Use Case: Fast hypervisor reboots, rolling host updates via vSphere Lifecycle Manager (vLCM), and short-duration maintenance windows.
2. Full Data Migration
- Operational Mechanics: vSAN systematically evacuates every single component residing on all disk groups of the target host, copying them to healthy storage devices on remaining cluster nodes before allowing the host to enter maintenance mode.
- Redundancy State: Full redundancy and 100% storage policy compliance are maintained continuously throughout the maintenance period.
- Capacity & Time Prerequisites: Requires that the remaining hosts in the cluster possess sufficient unreserved free capacity to absorb the evacuated data. Evacuation can take multiple hours on dense hosts and saturates the vSAN network.
- Recommended Use Case: Permanently removing or decommissioning an ESXi host from a cluster, or conducting extensive hardware maintenance exceeding several hours.
3. No Data Migration
- Operational Mechanics: vSAN evacuates zero storage components from the host. The host enters maintenance mode immediately without moving any data blocks.
- Redundancy State: Any virtual machine object that has components on this host without healthy mirrors on other hosts will lose quorum and become immediately Inaccessible (powering off or crashing). Objects with compliant mirrors elsewhere remain online with reduced redundancy.
- Recommended Use Case: Coordinated, cluster-wide cold maintenance shutdowns where all virtual machines are powered off and all ESXi hosts are rebooted simultaneously.
Component Failure States: Degraded vs. Absent
When a physical storage device or ESXi host stops responding, vSAN differentiates between permanent hardware destruction and temporary disconnections by classifying components into two distinct states:
1. Degraded State (Permanent Failure)
- Root Cause: Permanent, irrecoverable hardware failure. Examples include persistent SCSI read/write sense errors, bad media sectors, unrecoverable drive controller detachment, or catastrophic flash NAND degradation.
- vSAN Response: The VMkernel recognizes that the component cannot recover. It immediately marks the component as Degraded and instructs CLOM to begin an instant rebuild on remaining healthy hosts/disks in the cluster, provided sufficient free capacity exists.
2. Absent State (Transient Failure)
- Root Cause: Temporary loss of connectivity or planned reboot. Examples include an ESXi host rebooting, a physical network cable being temporarily unplugged, or an administrator placing a host into maintenance mode with Ensure Accessibility.
- vSAN Response: Because the physical storage device may return unharmed within minutes, initiating an immediate cluster-wide rebuild would flood the network with unnecessary terabytes of replication traffic. Instead, vSAN marks the component as Absent and starts the clomd rebuild delay timer.
The clomd Rebuild Delay Timer
- Default Value: 60 minutes.
- Operational Flow:
- When a host drops offline unexpectedly, components enter the Absent state, and the 60-minute countdown begins.
- If the host boots back up and rejoins the cluster within 60 minutes, vSAN terminates the timer. Rather than rebuilding the entire object, it performs a fast delta resynchronization, updating only the specific blocks that were modified while the host was offline.
- If the host remains absent when the timer reaches 0 (60 minutes expire), vSAN assumes the node is permanently lost. It transitions the components into an active rebuild state, allocating fresh storage on surviving hosts to restore full policy compliance.
- Advanced Configuration: Change the Object Repair Timer under Cluster > Configure > vSAN > Services > Advanced Options, which is backed by the ESXi advanced option
VSAN.ClomRepairDelay(in minutes). Setting it to120, for example, gives a 2-hour buffer during long maintenance. Setting it too low causes unnecessary resyncs.
Object Resynchronization & Traffic Shaping
Whenever components are evacuated, rebuilds are initiated, or storage policies are changed (e.g., converting a VMDK from RAID-1 to RAID-5), vSAN initiates Resync Traffic across the network.
Impact on Guest I/O
Uncontrolled resynchronization can saturate host HBAs and 10GbE/25GbE network links, causing severe latency spikes (noisy neighbor effect) on production guest virtual machines.
Adaptive Resync
Modern vSAN implementations feature Adaptive Resync algorithms that dynamically govern bandwidth allocation:
- No contention: Resync traffic can use all available bandwidth to restore compliance quickly.
- Contention: When VM I/O and resync compete, Adaptive Resync keeps roughly 20% of the bandwidth for resync so rebuilds keep progressing, and leaves about 80% for VM I/O.
- Monitoring: Track resync progress under Cluster > Monitor > vSAN > Resyncing Objects, where you can also trigger an immediate repair instead of waiting for the repair timer.
Maintenance Mode Evacuation Comparison
| Attribute | Ensure Accessibility | Full Data Migration | No Data Migration |
|---|---|---|---|
| Default Selection | Yes | No | No |
| Data Evacuated | Only components required for VM quorum | 100% of all components | 0% of components |
| Redundancy Maintained | No (Operates at FTT=0 risk) | Yes (100% redundancy preserved) | No (Non-redundant VMs go offline) |
| Free Capacity Required | Minimal | High (Must hold host's entire payload) | None |
| Time Required | Very Fast (Minutes) | Slow (Hours depending on data volume) | Instantaneous |
| Primary Exam Use Case | Routine ESXi patching & fast reboots | Decommissioning host or extended repairs | Coordinated cluster-wide power down |
A system administrator is executing automated rolling patch upgrades across a 6-node all-flash vSAN cluster using vSphere Lifecycle Manager (vLCM). Each host requires an operating system reboot following installation. What is the recommended and default Maintenance Mode data evacuation option for this maintenance workflow?
Full Data Migration
No Data Migration
Evacuate Powered Off Virtual Machines Only
Ensure Accessibility
An ESXi host in a healthy 8-node vSAN cluster unexpectedly suffers a total power supply failure and abruptly shuts down. How does vSAN immediately classify the storage components residing on the powered-off host, and what automated action follows by default?
Components are marked as Degraded, and vSAN begins rebuilding them immediately across the surviving hosts.
Components are marked as Absent, and vSAN initiates a 60-minute clomd delay countdown before triggering an active rebuild.
Components are marked as Orphaned, and vSAN prompts the administrator to manually authorize data evacuation.
Components are marked as Inaccessible, and all virtual machines residing on the cluster are suspended by vCenter.
A physical flash capacity SSD in an all-flash vSAN disk group generates persistent unrecoverable SCSI read/write sense errors and stops acknowledging I/O commands. How does the vSAN VMkernel respond to this specific hardware failure?
vSAN marks the component as Absent and waits 60 minutes for the disk to reset before initiating a rebuild.
vSAN places the parent ESXi host into Maintenance Mode using Ensure Accessibility.
vSAN immediately marks the component as Degraded and initiates an active rebuild on available healthy drives without waiting for a delay timer.
vSAN shuts down all virtual machines with VMDKs touching the failed disk group to prevent data corruption.
Sections you finish are checked off in the contents.