8.2 Command-Line Diagnostics & esxtop

Key Takeaways

  • esxtop is the interactive real-time diagnostic tool on ESXi; its screens include CPU (c), memory (m), disk adapter (d), disk device (u), disk VM (v), network (n), and power (p).

  • Customize esxtop with field selection (f), field order (o), and the refresh interval (s, default 5 seconds), then press W to save the layout to ~/.esxtop50rc.

  • Running esxtop in batch mode (esxtop -b -d <delay> -n <iterations> > file.csv) allows administrators to collect non-interactive diagnostic telemetry during intermittent performance anomalies for analysis in Windows Performance Monitor (Perfmon).

  • Key CPU indicators in esxtop are %RDY (ready time), %CSTP (co-stop, a sign of oversized SMP VMs), %SYS (VMkernel work on the world's behalf), and %MLMTD (time held back by a configured CPU limit).

  • Storage queuing issues in esxtop are identified when KAVG/cmd exceeds 2 ms or QAVG/cmd spikes above 0 ms, indicating that queue depths at the LUN (disk.sched.maxQLen) or HBA adapter are saturated.

Last updated: September 2026

8.2 Command-Line Diagnostics & esxtop

While vCenter Server performance charts provide excellent macro-level views and historical trending, troubleshooting acute, high-impact performance incidents requires direct, high-frequency command-line telemetry. The definitive real-time diagnostic utility in VMware vSphere is esxtop. Executing directly within the ESXi Shell or via SSH (and remotely via resxtop from the vSphere Command-Line Interface), esxtop exposes the raw, un-averaged scheduling state of every active world, process, adapter, and LUN on the hypervisor.


esxtop Architecture & Interactive Navigation

In ESXi scheduling terminology, every executing entity is designated as a World. A virtual machine consists of multiple worlds (such as the main vmx process, vCPU execution threads, and asynchronous I/O threads) grouped together under a Cartel or Group. esxtop provides dedicated interactive screens tailored to specific hardware and virtualization subsystems.

Primary Screen Modes

Administrators switch between subsystem views by pressing specific keyboard shortcuts:

ShortcutDisplay Screen ModeScope & Primary Diagnostic Target
cCPU Mode (Default)vCPU scheduling, %RDY, %CSTP, %SYS, %WAIT, %MLMTD, and physical core utilization.
mMemory ModeHost physical RAM allocation, NUMA node balance, ballooning (MCTLSZ), swap (SWCUR), and compression cache.
dStorage Adapter ModePhysical Host Bus Adapters (e.g., vmhba0), throughput, queue depth, commands/sec, aborts.
uStorage Device ModePhysical LUNs, naa-formatted identifiers, queue stats (QAVG), array response times (DAVG).
vStorage Virtual Machine ModePer-VM virtual disk performance, guest-perceived latency (GAVG), and VMDK IOPS.
nNetwork ModePhysical uplinks (vmnic), virtual switches, VM ports, packet rates, and dropped packet percentages (%DRPRX, %DRPTX).
pPower Management ModePhysical processor C-states, P-states, and dynamic voltage/frequency stepping.

Interactive Navigation & Display Customization

esxtop allows extensive manipulation of displayed data fields and refresh behavior:

  • Change Refresh Interval (s): By default, esxtop refreshes every 5 seconds. Pressing s prompts for a new refresh delay in seconds. While 2 seconds is useful for high-resolution analysis, setting the delay below 2 seconds on a heavily loaded host imposes measurable CPU overhead.
  • Field Selection (f): Pressing f opens the field selection screen, displaying all available counters for the active mode. Pressing the corresponding letter toggles a counter column on or off. Upper-case letters denote enabled fields; lower-case letters denote hidden fields.
  • Field Order (o): Pressing o lets you change the order of the displayed columns.
  • Sorting: In CPU mode, press U to sort by %USED, R to sort by %RDY, and N to sort by group ID. In memory mode, M sorts by memory size.
  • Save Custom Configuration (W): Pressing W writes the current display configuration, active fields, sort order, and refresh rate to the configuration file ~/.esxtop50rc in the user's home directory (/root/.esxtop50rc). When esxtop is launched in the future, it automatically loads this custom profile.
  • Toggle VM Only View (V): In CPU and memory modes, pressing uppercase V filters the output to show only virtual machine worlds, hiding internal ESXi VMkernel system daemons.

Batch Mode Data Collection for Post-Incident Analysis

When troubleshooting intermittent or overnight performance anomalies, administrators cannot monitor the interactive esxtop screen in real time. For these situations, esxtop supports a non-interactive Batch Mode that outputs raw CSV telemetry to a file.

Batch Execution Syntax

esxtop -b -d 10 -n 360 | gzip -9c > /vmfs/volumes/datastore1/esxtop_capture.csv.gz

Batch Parameter Breakdown

  • -b: Enables batch mode operation. In this mode, esxtop does not initialize the interactive terminal UI; it streams comma-separated values to standard output.
  • -d 10: Sets the sampling interval to 10 seconds (default is 5 seconds).
  • -n 360: Specifies the total number of iterations before terminating. At 10-second intervals, 360 iterations captures exactly 1 hour of continuous diagnostic telemetry (10s × 360 = 3600 seconds).
  • gzip -9c: Pipes the massive text stream directly into gzip compression. A raw batch capture for a dense host can generate hundreds of megabytes of text; compression reduces file size by over 90%.
  • Storage Destination: Always direct the output file to a persistent VMFS or NFS datastore (/vmfs/volumes/...). Never output batch files to the ramdisk /tmp or /var/log on the host, as this can fill root memory partitions and crash ESXi management agents.

Analyzing Batch Captures with Windows Perfmon

Once collected, the .csv file can be decompressed and loaded into Windows Performance Monitor (perfmon.exe):

  1. Launch perfmon.exe on a Windows workstation.
  2. Right-click the graph area and select Properties -> Source tab.
  3. Select Log files, browse to the uncompressed esxtop_capture.csv file, and click Apply.
  4. Navigate to the Data tab to add specific ESXi counters (e.g., Physical CPU(0)\% Util Time, Group Cpu(VM-Name)\% Ready Time) for graphical analysis.
Loading diagram...
esxtop CPU Performance Troubleshooting Decision Flowchart

Critical esxtop Counters Cheat Sheet

Mastering esxtop requires memorizing specific counter acronyms and their operational significance across CPU, memory, storage, and networking.

CPU Counters (c mode)

  • %USED: The percentage of physical CPU core cycles consumed by the world. Reflects actual compute workload.
  • %SYS: Percentage of time the VMkernel spend executing system services on behalf of the world. High %SYS (> 10-15%) indicates excessive I/O trap handling, driver interrupt thrashing, or high virtualization overhead.
  • %WAIT: Total percentage of time the world was in a wait state. Important Exam Distinction: %WAIT includes both intentional idle time (waiting for user input or guest timer ticks) and wait time for hypervisor resources. A high %WAIT on an idle VM is completely normal.
  • %VMWAIT: A critical subset of %WAIT. Measures the percentage of time the world spent waiting specifically for VMkernel activities to complete—such as waiting for storage I/O completions, memory decompression, or memory swap reads. Normal values are < 1-2%. A high %VMWAIT signals a storage or memory bottleneck, not a CPU problem.
  • %RDY: Percentage of time the world was ready to execute but queued for a physical CPU. Warning: > 5%; Critical: > 10%.
  • %CSTP: Co-stop percentage. Indicates SMP synchronization delay. Critical: > 3%.
  • %MLMTD: Percentage of time the world was ready to run but was held back by a configured CPU Limit. If %MLMTD > 0%, the VM is being throttled by policy, not by physical contention.

Memory Counters (m mode)

  • MCTLSZ: How much guest memory (in MB) the balloon driver has currently reclaimed.
  • MCTLTGT: The balloon size (in MB) the VMkernel wants to reach.
  • SWCUR: Current amount of memory (in MB) swapped out to the virtual machine's .vswp file on the datastore. Any non-zero value that is actively increasing indicates severe host RAM starvation.
  • CACHESZ / CACHEUSD: Size of the VM's compression cache and how much of it is in use (MB).
  • ZIP/s and UNZIP/s: The rate of memory pages being compressed and decompressed per second.

Storage Counters (d, u, and v modes)

  • DAVG/cmd: Average device latency per command (in milliseconds). Measures SAN fabric and storage array response time. Target: < 15-20 ms for HDD, < 2-3 ms for SSD/NVMe.
  • KAVG/cmd: Average VMkernel latency per command (in milliseconds). Measures queuing time inside ESXi. Target: < 1-2 ms. Elevated values indicate HBA queue depth or LUN queue depth exhaustion.
  • GAVG/cmd: Total guest-perceived latency (DAVG + KAVG). Target: < 15-20 ms.
  • QAVG/cmd: Average time a command spent in the device queue before transmission. Any sustained value > 0 ms indicates queue saturation.
  • ABRTS/s: Aborted commands per second. Target is 0.00. Values > 0 indicate SAN fabric disconnects, zoning errors, or unresponsive storage arrays.
  • ACTV: Number of active commands currently issued to the device.

Network Counters (n mode)

  • %DRPRX: Percentage of incoming packets dropped by the virtual switch port. Indicates guest OS vNIC buffer exhaustion.
  • %DRPTX: Percentage of outgoing packets dropped by the virtual switch or uplink port. Indicates physical switch port flow control or uplink congestion.
  • PKTRX/s and PKTTX/s: Packets received and transmitted per second.

Comprehensive esxtop Diagnostic Reference Table

ModeKey CounterHealthy BaselineWarning / Action ThresholdPrimary Root Cause
CPU (c)%RDY< 5%> 5% - 10%Host CPU overcommitment; run queue saturation
CPU (c)%CSTP0% - 1%> 3%Too many vCPUs allocated to VM; co-scheduling skew
CPU (c)%MLMTD0%> 0%Configured CPU Limit in VM resource settings
CPU (c)%VMWAIT< 1%> 2%VM waiting on disk I/O, paging, or swap completion
Memory (m)SWCUR0 MB> 0 MBPhysical RAM exhaustion; hypervisor paging to .vswp
Memory (m)MCTLSZ0 MBEscalating MBHost memory ballooning active; reclaim in progress
Storage (u)KAVG/cmd< 1 ms> 2 msAdapter or LUN queue depth saturation (disk.sched.maxQLen)
Storage (u)DAVG/cmd< 10 ms> 20 msStorage array overload, SAN fabric congestion
Storage (u)ABRTS/s0.00> 0.00Storage path drops, LUN reset, timeout reached
Network (n)%DRPRX0.0%> 0.5%In-guest ring buffer overrun; vNIC driver mismatch
Test Your Knowledge

While investigating a virtual machine that is executing slowly despite low host utilization, an administrator inspects esxtop in CPU mode ('c'). The administrator observes: %USED = 48.0%, %RDY = 22.5%, %CSTP = 0.1%, and %MLMTD = 22.4%. What is causing the high ready time on this virtual machine?

A

The physical ESXi host compute cores are severely oversubscribed by competing workloads

B

The virtual machine is throttled because its processing has reached an administrative CPU Limit

C

The virtual machine has too many vCPUs allocated, triggering severe co-scheduling penalties

D

The guest operating system is thrashing memory pages to its internal swap file

Test Your Knowledge

An administrator customizes the esxtop CPU display screen by adding specific field columns, adjusting column sort order, and setting the refresh delay to 3 seconds. Which interactive keystroke must the administrator press to permanently save these display preferences across future CLI sessions?

A

Press 'S' to save the current configuration to the system registry

B

Press 'c' followed by 'Ctrl+S' to commit changes to the VMkernel boot bank

C

Press 'f' to open the configuration menu and select 'Write to disk'

D

Press 'W' to save the configuration file to ~/.esxtop50rc

Test Your Knowledge

An engineer monitors a LUN in esxtop storage device mode ('u') and notes the following metrics: DAVG/cmd = 1.8 ms, KAVG/cmd = 14.2 ms, and QAVG/cmd = 12.1 ms. What is the root cause of this storage latency?

A

The ESXi storage device or adapter queue depth is exhausted, causing commands to queue in the VMkernel

B

The SAN storage array disk spindles are experiencing heavy seek contention

C

The Fibre Channel fabric switch ports are generating frame CRC errors

D

The guest operating system is generating non-aligned 4KB I/O requests

Sections you finish are checked off in the contents.