8.4 Troubleshooting Common Infrastructure Scenarios

Key Takeaways

  • All Paths Down (APD) represents a transient storage fabric interruption where paths become unavailable without explicit SCSI sense codes, triggering an APD timeout (default 140 seconds) before VMCP executes configured response policies.

  • Permanent Device Loss (PDL) indicates an unrecoverable storage failure or unmapped LUN accompanied by SCSI sense code 0x5 0x25 0x00, enabling vSphere HA / VMCP to immediately terminate affected VMs and restart them on healthy cluster hosts without waiting for a timeout.

  • In vSphere HA, a host isolation occurs when an ESXi host cannot communicate with cluster peers over the management network and fails to ping its default gateway or custom isolation address (das.isolationaddress), using datastore heartbeats to verify that the host remains alive before taking isolation actions.

  • An ESXi host showing 'Not Responding' in vCenter while virtual machines remain operational is typically resolved safely by restarting management daemons via /sbin/services.sh restart or DCUI without impacting running VMs.

  • Virtual machine disk locks preventing power-on or snapshot consolidation can be tracked to the locking host's MAC address using vmkfstools -D /vmfs/volumes/.../<vm>-flat.vmdk, and consolidation is resolved using the Consolidate workflow after removing stale snapshot references.

Last updated: September 2026

8.4 Troubleshooting Common Infrastructure Scenarios

Enterprise virtualization environments inevitably encounter critical hardware, fabric, and configuration failures. The VCP-DCV exam rigorously tests an administrator's ability to diagnose failure symptoms, isolate root causes, and execute non-disruptive remediation workflows. This section explores four core production troubleshooting domains: storage outages (APD vs. PDL), vSphere HA partition and isolation events, unresponsive host management agents, and locked virtual machine disks during snapshot consolidation.


Storage Outage Troubleshooting: All Paths Down (APD) vs. Permanent Device Loss (PDL)

When an ESXi host loses connectivity to a storage LUN, the VMkernel classifies the outage into one of two fundamentally distinct conditions: All Paths Down (APD) or Permanent Device Loss (PDL). Understanding this architectural divergence is essential for configuring vSphere HA VM Component Protection (VMCP).

+-----------------------------------------------------------------------------------------+
| Storage Outage Classification: APD vs. PDL                                              |
|                                                                                         |
| [ Physical LUN / Datastore Unreachable ]                                                |
|                    |                                                                    |
|                    +---> Does the array return SCSI Sense 0x5 0x25 0x00?                |
|                                  |                                                      |
|               +------------------+------------------+                                   |
|               | NO (Silent link failure)            | YES (Authoritative failure)       |
|               v                                     v                                   |
|   [ All Paths Down (APD) ]              [ Permanent Device Loss (PDL) ]                 |
|   • Transient fabric/switch drop        • Unmapped LUN or dead array controller         |
|   • No SCSI sense code returned         • SCSI sense code 0x5 0x25 0x00 returned        |
|   • ESXi retries I/O continuously       • ESXi immediately halts all I/O retries        |
|   • APD Timeout begins (default 140s)   • VMCP terminates VMs immediately               |
|   • VMCP triggers: Conservative or      • vSphere HA restarts VMs on healthy            |
|     Aggressive restart after timeout      hosts with datastore access                   |
+-----------------------------------------------------------------------------------------+

All Paths Down (APD)

  • Definition & Cause: APD occurs when all physical paths to a storage device become unavailable, but the storage array cannot be contacted to provide an authoritative status. This commonly stems from severed fiber cables, failed SAN fabric switches, or misconfigured network zoning.
  • Hypervisor Behavior: Because the hypervisor cannot determine whether the outage is momentary (e.g., a 30-second switch reboot) or permanent, ESXi keeps retrying the I/O, and guest file systems stall while they wait.
  • APD Timeout: ESXi initializes an internal timer called the APD Timeout (configured via advanced parameter Misc.APDTimeout, set to 140 seconds by default). During this 140-second window, no destructive actions are taken. If paths recover before 140 seconds elapse, I/O resumes seamlessly without VM termination.
  • VMCP APD Policy Response: If the APD condition persists past the 140-second timeout and the VMCP response recovery delay (3 minutes by default), VMCP executes the configured policy:
    • Conservative: VMCP powers off the affected virtual machines only if vSphere HA verifies that another host in the cluster has active access to the datastore and sufficient compute capacity to restart them.
    • Aggressive: VMCP immediately terminates the virtual machines even if vSphere HA cannot guarantee that another host can restart them.

Permanent Device Loss (PDL)

  • Definition & Cause: PDL occurs when a storage device becomes permanently unavailable, and the storage array explicitly and authoritatively informs the ESXi host via SCSI sense code 0x5 0x25 0x00 (LOGICAL UNIT NOT SUPPORTED) or another PDL code such as 0x4 0x4c 0x0 (LOGICAL UNIT FAILED SELF-CONFIGURATION). This happens when a storage administrator unmaps or deletes a LUN from an array storage group while virtual machines are running on it, or when an entire storage controller fails catastrophically.
  • Hypervisor Behavior: Because the array explicitly confirmed that the device is gone, ESXi immediately terminates all I/O retries. No timeout counter is required.
  • VMCP PDL Policy Response: VMCP immediately terminates the running virtual machines and triggers vSphere HA to restart them on healthy hosts that retain access to an alternate datastore or replicated copy.

APD vs. PDL Comparison Matrix

Technical AttributeAll Paths Down (APD)Permanent Device Loss (PDL)
Underlying CauseTransient fabric/switch link lossLUN unmapped, deleted, or destroyed on array
Array ResponseNone (silent timeout)SCSI Sense Code 0x5 0x25 0x00
I/O HandlingVMkernel retries I/O continuouslyVMkernel drops I/O immediately
Default Timeout140 seconds (Misc.APDTimeout)0 seconds (Immediate response)
VMCP Action TimingExecutes only after APD timeout expiresExecutes immediately upon detection
VM Restart StrategyConservative or Aggressive policyImmediate Power Off and Restart on healthy host
Loading diagram...
Storage Outage VMCP Decision Workflow

vSphere HA Host Isolation vs. Network Partitioning

vSphere High Availability (HA) relies on continuous management network communication between the Master (Primary) Host and Subordinate (Secondary) Hosts via the Fault Domain Manager (fdm) service. When network disruptions occur, vSphere HA distinguishes between an isolated host and a partitioned cluster.

Host Isolation Mechanics

A host enters an Isolated State when it satisfies three concurrent criteria:

  1. It ceases receiving heartbeats from the HA Master host over the management network.
  2. It fails to elect itself as Master because it cannot discover any other cluster hosts over the network.
  3. It attempts to ping its configured Isolation Address (by default, the management network default gateway) and the ping fails.
  • Isolation Addresses (das.isolationaddress): By default, ESXi pings the default gateway configured on vmk0. If corporate firewalls block ICMP echo requests to the gateway, administrators must specify reliable network addresses using the advanced cluster settings das.isolationaddress0 through das.isolationaddress9.
  • Datastore Heartbeating: Crucially, before executing an isolation response, the isolated host and the master host check Datastore Heartbeats (which use heartbeat files in the .vSphere-HA directory on shared datastores). If the master host sees that the isolated host is still updating its datastore heartbeat file, the master knows the isolated host is still alive and running VMs—preventing duplicate split-brain restarts.
  • Isolation Responses: Configured at the cluster level:
    • Disabled (leave VMs powered on): Recommended for environments where false isolations occur or where workloads must continue processing even without management network connectivity.
    • Power Off and Restart VMs: Hard power off; VMs are terminated immediately and restarted by HA on surviving hosts.
    • Shutdown and Restart VMs: Gracefully shuts down guest OS via VMware Tools (waiting up to das.isolationshutdowntimeout, default 300 seconds); if graceful shutdown fails, it hard powers off and restarts.

Network Partitioning

A Network Partition occurs when a cluster is split into two or more isolated network segments, but hosts within each segment can still communicate with one another. Each partition elects its own Master host. Because datastore heartbeating informs each master which VMs are running on which hosts, hosts do not redundantly restart virtual machines across partitions, preventing split-brain corruption.

Host Disconnects & Virtual Machine Disk Lock Remediation

Troubleshooting "Not Responding" ESXi Hosts

When an ESXi host shows "Not Responding" or "Disconnected" in vCenter Server, administrators must follow a disciplined diagnostic methodology before performing hard reboots:

+-----------------------------------------------------------------------------------+
| ESXi Host "Not Responding" Diagnostic Workflow                                    |
|                                                                                   |
| 1. Network Connectivity Check:                                                    |
|    Ping ESXi management IP (ICMP) and test TLS handshake: `nc -zv <host-ip> 443` |
|    - If ping fails -> Physical network cable, switch port, or VLAN trunk failure.  |
|    - If ping succeeds -> Host management agents hung.                             |
|                                                                                   |
| 2. Direct Management Console Access:                                              |
|    Connect via IPMI / iLO / iDRAC virtual console or physical monitor.            |
|    - If Purple Screen of Death (PSOD) -> Capture crash dump, review kernel panic. |
|    - If DCUI is visible -> Press F2 to authenticate.                              |
|                                                                                   |
| 3. Management Agent Restart (Non-Disruptive):                                     |
|    - Via DCUI: Troubleshooting Options -> Restart Management Agents               |
|    - Via SSH: `/sbin/services.sh restart`                                        |
|      or individual restarts: `/etc/init.d/hostd restart && /etc/init.d/vpxa restart`|
+-----------------------------------------------------------------------------------+

Exam Trap: Restarting management agents (services.sh restart, hostd, or vpxa) does not reboot the host or impact running virtual machines. Virtual machine execution threads run independently inside the VMkernel (vmx worlds). Do NOT power-cycle physical hosts when a simple management agent restart resolves the vCenter disconnect.

Locked Virtual Machine Disks (.vmdk) & Snapshot Consolidation

During snapshot creation, backup operations (via vStorage APIs for Data Protection / VADP), or replication, virtual disks can become locked, resulting in errors such as "An error occurred while consolidating disks" or "Cannot open the disk /vmfs/volumes/.../<vm>.vmdk: Failed to lock the file."

  1. Underlying Locking Mechanism: VMFS uses distributed on-disk locking mechanisms (SCSI-2 reservations or Atomic Test and Set / ATS) to ensure that only one ESXi host accesses a base -flat.vmdk or delta disk at any time.
  2. Identifying the Host Holding the Lock (vmkfstools -D): When a .vmdk file is locked, an administrator can execute vmkfstools -D against the flat disk file from the ESXi Shell:
    vmkfstools -D /vmfs/volumes/Datastore01/AppVM/AppVM-flat.vmdk
    
    In the lock details it prints (also written to /var/log/vmkernel.log), find the owner field, for example: mode 1, owner 4ff58b73-11e24b56-7f9a-0026b9a1e1f5 The last 12 hex digits (0026b9a1e1f5) are the MAC address of the ESXi host holding the lock. An owner of all zeros means no host holds it. The administrator can then log into that specific host to investigate why it has retained the lock (e.g., a hung backup proxy appliance holding a hot-added disk).
  3. Snapshot Consolidation Procedure:
    • In the vSphere Client, a warning banner states: "Virtual machine disks consolidation needed."
    • Step 1: Verify that external backup appliances (VADP proxies) have completely detached the VM's virtual disks from their own virtual hardware.
    • Step 2: Right-click the virtual machine in the vSphere Client -> select Snapshots -> select Consolidate.
    • Step 3: Confirm the consolidation task completes successfully, merging redundant delta disks into the base virtual disk.
Test Your Knowledge

A storage array administrator accidentally unmaps a production LUN from a storage group assigned to an ESXi 8.0 cluster. The ESXi hosts receive SCSI sense code 0x5 0x25 0x00 from the array. If VM Component Protection (VMCP) is enabled with default settings, what action does vSphere HA take?

A

VMCP immediately terminates the affected virtual machines and restarts them on surviving cluster hosts that retain access to an alternate datastore

B

VMCP initializes a 140-second APD countdown timer and keeps the virtual machines running while retrying I/O

C

VMCP gracefully shuts down the guest operating systems via VMware Tools and leaves them powered off indefinitely

D

VMCP pauses the virtual machine execution threads and triggers an automated Storage vMotion to migrate them to another datastore

Test Your Knowledge

An administrator attempts to power on a virtual machine but receives the error: 'Cannot open the disk AppServer.vmdk: Failed to lock the file.' Which command-line diagnostic tool should the administrator execute on the flat VMDK file to identify the MAC address of the ESXi host currently holding the file lock?

A

esxcli storage core device list -d AppServer-flat.vmdk

B

esxtop -v -d 5 -n 1

C

vmkfstools -D /vmfs/volumes/Datastore01/AppServer/AppServer-flat.vmdk

D

vSphere-land-lock-analyzer --inspect AppServer.vmdk

Test Your Knowledge

A network switch configuration error severs management network connectivity between an ESXi host and the rest of its vSphere HA cluster. The host cannot communicate with the HA Master and cannot ping its default gateway. What mechanism does vSphere HA employ to verify whether the isolated host has crashed or is still running virtual machines before executing an isolation response?

A

vCenter Server Agent (vpxa) polls the host via UDP port 902 heartbeats

B

Datastore Heartbeating verifies host liveness via heartbeat files on shared datastores

C

The ESXi VMkernel transmits ARP broadcast probes across the VMotion network

D

The Direct Console User Interface (DCUI) initiates an automated reboot sequence

Sections you finish are checked off in the contents.