8.2 Command-Line Diagnostics & esxtop
Key Takeaways
esxtop is the interactive real-time diagnostic tool on ESXi; its screens include CPU (c), memory (m), disk adapter (d), disk device (u), disk VM (v), network (n), and power (p).
Customize esxtop with field selection (f), field order (o), and the refresh interval (s, default 5 seconds), then press W to save the layout to ~/.esxtop50rc.
Running esxtop in batch mode (esxtop -b -d <delay> -n <iterations> > file.csv) allows administrators to collect non-interactive diagnostic telemetry during intermittent performance anomalies for analysis in Windows Performance Monitor (Perfmon).
Key CPU indicators in esxtop are %RDY (ready time), %CSTP (co-stop, a sign of oversized SMP VMs), %SYS (VMkernel work on the world's behalf), and %MLMTD (time held back by a configured CPU limit).
Storage queuing issues in esxtop are identified when KAVG/cmd exceeds 2 ms or QAVG/cmd spikes above 0 ms, indicating that queue depths at the LUN (disk.sched.maxQLen) or HBA adapter are saturated.
8.2 Command-Line Diagnostics & esxtop
While vCenter Server performance charts provide excellent macro-level views and historical trending, troubleshooting acute, high-impact performance incidents requires direct, high-frequency command-line telemetry. The definitive real-time diagnostic utility in VMware vSphere is esxtop. Executing directly within the ESXi Shell or via SSH (and remotely via resxtop from the vSphere Command-Line Interface), esxtop exposes the raw, un-averaged scheduling state of every active world, process, adapter, and LUN on the hypervisor.
esxtop Architecture & Interactive Navigation
In ESXi scheduling terminology, every executing entity is designated as a World. A virtual machine consists of multiple worlds (such as the main vmx process, vCPU execution threads, and asynchronous I/O threads) grouped together under a Cartel or Group. esxtop provides dedicated interactive screens tailored to specific hardware and virtualization subsystems.
Primary Screen Modes
Administrators switch between subsystem views by pressing specific keyboard shortcuts:
| Shortcut | Display Screen Mode | Scope & Primary Diagnostic Target |
|---|---|---|
c | CPU Mode (Default) | vCPU scheduling, %RDY, %CSTP, %SYS, %WAIT, %MLMTD, and physical core utilization. |
m | Memory Mode | Host physical RAM allocation, NUMA node balance, ballooning (MCTLSZ), swap (SWCUR), and compression cache. |
d | Storage Adapter Mode | Physical Host Bus Adapters (e.g., vmhba0), throughput, queue depth, commands/sec, aborts. |
u | Storage Device Mode | Physical LUNs, naa-formatted identifiers, queue stats (QAVG), array response times (DAVG). |
v | Storage Virtual Machine Mode | Per-VM virtual disk performance, guest-perceived latency (GAVG), and VMDK IOPS. |
n | Network Mode | Physical uplinks (vmnic), virtual switches, VM ports, packet rates, and dropped packet percentages (%DRPRX, %DRPTX). |
p | Power Management Mode | Physical processor C-states, P-states, and dynamic voltage/frequency stepping. |
Interactive Navigation & Display Customization
esxtop allows extensive manipulation of displayed data fields and refresh behavior:
- Change Refresh Interval (
s): By default,esxtoprefreshes every 5 seconds. Pressingsprompts for a new refresh delay in seconds. While 2 seconds is useful for high-resolution analysis, setting the delay below 2 seconds on a heavily loaded host imposes measurable CPU overhead. - Field Selection (
f): Pressingfopens the field selection screen, displaying all available counters for the active mode. Pressing the corresponding letter toggles a counter column on or off. Upper-case letters denote enabled fields; lower-case letters denote hidden fields. - Field Order (
o): Pressingolets you change the order of the displayed columns. - Sorting: In CPU mode, press
Uto sort by%USED,Rto sort by%RDY, andNto sort by group ID. In memory mode,Msorts by memory size. - Save Custom Configuration (
W): PressingWwrites the current display configuration, active fields, sort order, and refresh rate to the configuration file~/.esxtop50rcin the user's home directory (/root/.esxtop50rc). Whenesxtopis launched in the future, it automatically loads this custom profile. - Toggle VM Only View (
V): In CPU and memory modes, pressing uppercaseVfilters the output to show only virtual machine worlds, hiding internal ESXi VMkernel system daemons.
Batch Mode Data Collection for Post-Incident Analysis
When troubleshooting intermittent or overnight performance anomalies, administrators cannot monitor the interactive esxtop screen in real time. For these situations, esxtop supports a non-interactive Batch Mode that outputs raw CSV telemetry to a file.
Batch Execution Syntax
esxtop -b -d 10 -n 360 | gzip -9c > /vmfs/volumes/datastore1/esxtop_capture.csv.gz
Batch Parameter Breakdown
-b: Enables batch mode operation. In this mode,esxtopdoes not initialize the interactive terminal UI; it streams comma-separated values to standard output.-d 10: Sets the sampling interval to 10 seconds (default is 5 seconds).-n 360: Specifies the total number of iterations before terminating. At 10-second intervals, 360 iterations captures exactly 1 hour of continuous diagnostic telemetry (10s × 360 = 3600 seconds).gzip -9c: Pipes the massive text stream directly into gzip compression. A raw batch capture for a dense host can generate hundreds of megabytes of text; compression reduces file size by over 90%.- Storage Destination: Always direct the output file to a persistent VMFS or NFS datastore (
/vmfs/volumes/...). Never output batch files to the ramdisk/tmpor/var/logon the host, as this can fill root memory partitions and crash ESXi management agents.
Analyzing Batch Captures with Windows Perfmon
Once collected, the .csv file can be decompressed and loaded into Windows Performance Monitor (perfmon.exe):
- Launch
perfmon.exeon a Windows workstation. - Right-click the graph area and select Properties -> Source tab.
- Select Log files, browse to the uncompressed
esxtop_capture.csvfile, and click Apply. - Navigate to the Data tab to add specific ESXi counters (e.g.,
Physical CPU(0)\% Util Time,Group Cpu(VM-Name)\% Ready Time) for graphical analysis.
Critical esxtop Counters Cheat Sheet
Mastering esxtop requires memorizing specific counter acronyms and their operational significance across CPU, memory, storage, and networking.
CPU Counters (c mode)
%USED: The percentage of physical CPU core cycles consumed by the world. Reflects actual compute workload.%SYS: Percentage of time the VMkernel spend executing system services on behalf of the world. High%SYS(> 10-15%) indicates excessive I/O trap handling, driver interrupt thrashing, or high virtualization overhead.%WAIT: Total percentage of time the world was in a wait state. Important Exam Distinction:%WAITincludes both intentional idle time (waiting for user input or guest timer ticks) and wait time for hypervisor resources. A high%WAITon an idle VM is completely normal.%VMWAIT: A critical subset of%WAIT. Measures the percentage of time the world spent waiting specifically for VMkernel activities to complete—such as waiting for storage I/O completions, memory decompression, or memory swap reads. Normal values are< 1-2%. A high%VMWAITsignals a storage or memory bottleneck, not a CPU problem.%RDY: Percentage of time the world was ready to execute but queued for a physical CPU. Warning:> 5%; Critical:> 10%.%CSTP: Co-stop percentage. Indicates SMP synchronization delay. Critical:> 3%.%MLMTD: Percentage of time the world was ready to run but was held back by a configured CPU Limit. If%MLMTD > 0%, the VM is being throttled by policy, not by physical contention.
Memory Counters (m mode)
MCTLSZ: How much guest memory (in MB) the balloon driver has currently reclaimed.MCTLTGT: The balloon size (in MB) the VMkernel wants to reach.SWCUR: Current amount of memory (in MB) swapped out to the virtual machine's.vswpfile on the datastore. Any non-zero value that is actively increasing indicates severe host RAM starvation.CACHESZ/CACHEUSD: Size of the VM's compression cache and how much of it is in use (MB).ZIP/sandUNZIP/s: The rate of memory pages being compressed and decompressed per second.
Storage Counters (d, u, and v modes)
DAVG/cmd: Average device latency per command (in milliseconds). Measures SAN fabric and storage array response time. Target:< 15-20 msfor HDD,< 2-3 msfor SSD/NVMe.KAVG/cmd: Average VMkernel latency per command (in milliseconds). Measures queuing time inside ESXi. Target:< 1-2 ms. Elevated values indicate HBA queue depth or LUN queue depth exhaustion.GAVG/cmd: Total guest-perceived latency (DAVG + KAVG). Target:< 15-20 ms.QAVG/cmd: Average time a command spent in the device queue before transmission. Any sustained value> 0 msindicates queue saturation.ABRTS/s: Aborted commands per second. Target is0.00. Values> 0indicate SAN fabric disconnects, zoning errors, or unresponsive storage arrays.ACTV: Number of active commands currently issued to the device.
Network Counters (n mode)
%DRPRX: Percentage of incoming packets dropped by the virtual switch port. Indicates guest OS vNIC buffer exhaustion.%DRPTX: Percentage of outgoing packets dropped by the virtual switch or uplink port. Indicates physical switch port flow control or uplink congestion.PKTRX/sandPKTTX/s: Packets received and transmitted per second.
Comprehensive esxtop Diagnostic Reference Table
| Mode | Key Counter | Healthy Baseline | Warning / Action Threshold | Primary Root Cause |
|---|---|---|---|---|
CPU (c) | %RDY | < 5% | > 5% - 10% | Host CPU overcommitment; run queue saturation |
CPU (c) | %CSTP | 0% - 1% | > 3% | Too many vCPUs allocated to VM; co-scheduling skew |
CPU (c) | %MLMTD | 0% | > 0% | Configured CPU Limit in VM resource settings |
CPU (c) | %VMWAIT | < 1% | > 2% | VM waiting on disk I/O, paging, or swap completion |
Memory (m) | SWCUR | 0 MB | > 0 MB | Physical RAM exhaustion; hypervisor paging to .vswp |
Memory (m) | MCTLSZ | 0 MB | Escalating MB | Host memory ballooning active; reclaim in progress |
Storage (u) | KAVG/cmd | < 1 ms | > 2 ms | Adapter or LUN queue depth saturation (disk.sched.maxQLen) |
Storage (u) | DAVG/cmd | < 10 ms | > 20 ms | Storage array overload, SAN fabric congestion |
Storage (u) | ABRTS/s | 0.00 | > 0.00 | Storage path drops, LUN reset, timeout reached |
Network (n) | %DRPRX | 0.0% | > 0.5% | In-guest ring buffer overrun; vNIC driver mismatch |
While investigating a virtual machine that is executing slowly despite low host utilization, an administrator inspects esxtop in CPU mode ('c'). The administrator observes: %USED = 48.0%, %RDY = 22.5%, %CSTP = 0.1%, and %MLMTD = 22.4%. What is causing the high ready time on this virtual machine?
The physical ESXi host compute cores are severely oversubscribed by competing workloads
The virtual machine is throttled because its processing has reached an administrative CPU Limit
The virtual machine has too many vCPUs allocated, triggering severe co-scheduling penalties
The guest operating system is thrashing memory pages to its internal swap file
An administrator customizes the esxtop CPU display screen by adding specific field columns, adjusting column sort order, and setting the refresh delay to 3 seconds. Which interactive keystroke must the administrator press to permanently save these display preferences across future CLI sessions?
Press 'S' to save the current configuration to the system registry
Press 'c' followed by 'Ctrl+S' to commit changes to the VMkernel boot bank
Press 'f' to open the configuration menu and select 'Write to disk'
Press 'W' to save the configuration file to ~/.esxtop50rc
An engineer monitors a LUN in esxtop storage device mode ('u') and notes the following metrics: DAVG/cmd = 1.8 ms, KAVG/cmd = 14.2 ms, and QAVG/cmd = 12.1 ms. What is the root cause of this storage latency?
The ESXi storage device or adapter queue depth is exhausted, causing commands to queue in the VMkernel
The SAN storage array disk spindles are experiencing heavy seek contention
The Fibre Channel fabric switch ports are generating frame CRC errors
The guest operating system is generating non-aligned 4KB I/O requests
Sections you finish are checked off in the contents.