8.1 Performance Monitoring & Metrics

Key Takeaways

  • vCenter Server collects real-time performance statistics at 20-second intervals and retains them in volatile memory for 1 hour before rolling them up into historical database intervals.

  • Historical performance statistics rollups span four intervals: Past Day (5-minute interval, rollup 1), Past Week (30-minute interval, rollup 2), Past Month (2-hour interval, rollup 3), and Past Year (1-day interval, rollup 4).

  • CPU contention shows up as %RDY (ready time); common rules of thumb flag more than 5% per vCPU for investigation and more than 10% as severe, and %CSTP above about 3% suggests the VM has more vCPUs than it can use efficiently.

  • ESXi sets its memory states relative to the host's minFree value: High (400%), Clear (100%), Soft (64%), Hard (32%), and Low (16%). Reclamation escalates from large-page breaking for TPS to ballooning, then compression and swapping, then swapping plus blocking.

  • Storage I/O latency decomposes into Device Latency (DAVG/cmd, target < 15-20 ms) and Kernel Queuing Latency (KAVG/cmd, target < 1-2 ms), where GAVG/cmd represents total guest-perceived latency (GAVG = DAVG + KAVG).

Last updated: September 2026

8.1 Performance Monitoring & Metrics

Performance optimization and bottleneck identification are critical responsibilities for a vSphere administrator. VMware vCenter Server and VMware ESXi provide a comprehensive performance measurement infrastructure that captures compute, memory, storage, and networking metrics across virtual machines, resource pools, ESXi hosts, and clusters. To accurately diagnose performance degradation and prepare for the VCP-DCV exam, administrators must understand how performance data is sampled and aggregated, the architectural thresholds governing subsystem bottlenecks, and the precise escalation sequence of hypervisor resource reclamation.


vCenter Server Performance Charting Architecture

vCenter Server utilizes a tiered statistical collection architecture that balances real-time diagnostic visibility against database storage consumption. Performance data collection operates across two operational tiers: real-time collection and historical rollups.

+---------------------------------------------------------------------------------------+
| vCenter Server Performance Statistics Architecture                                   |
|                                                                                       |
| [ ESXi Host ] --(20-sec raw samples)--> [ Real-Time Memory Cache (vCenter / Host) ]  |
|                                                      |                                |
|                                            (1-hour retention window)                  |
|                                                      v                                |
|                                          [ Rollup 1: Past Day ]                       |
|                                          5-minute interval / 1-day retention          |
|                                                      v                                |
|                                          [ Rollup 2: Past Week ]                      |
|                                          30-minute interval / 1-week retention        |
|                                                      v                                |
|                                          [ Rollup 3: Past Month ]                     |
|                                          2-hour interval / 1-month retention          |
|                                                      v                                |
|                                          [ Rollup 4: Past Year ]                      |
|                                          1-day interval / 1-year retention            |
+---------------------------------------------------------------------------------------+

Real-Time vs. Historical Statistics Collection

  • Real-Time Statistics: ESXi collects performance metrics directly from the VMkernel every 20 seconds. These raw 20-second samples are stored in volatile RAM on the ESXi host and inside the vCenter Server Appliance memory cache. Real-time statistics are held for exactly 1 hour (providing 180 data points). Because they reside in volatile memory, real-time metrics do not incur database write I/O. However, once 60 minutes elapse, real-time samples must either be rolled up into historical statistics or discarded.
  • Historical Statistics (Database Rollups): To preserve performance history over extended periods without exhausting database capacity, vCenter Server executes automated statistical rollup jobs at predefined intervals. During each rollup pass, vCenter aggregates smaller intervals into larger summary blocks using four statistical calculations: Average, Minimum, Maximum, and Summation.

Historical Rollup Intervals and Retention Matrix

Rollup LevelTime Period DisplayedSample Interval DurationNumber of Samples AggregatedDatabase Retention Period
Real-timePast 1 Hour20 secondsRaw counter sample1 hour (in memory)
Rollup 1Past Day5 minutes15 raw samples (20s × 15 = 5m)1 day (24 hours)
Rollup 2Past Week30 minutes6 Rollup 1 samples (5m × 6 = 30m)1 week (7 days)
Rollup 3Past Month2 hours4 Rollup 2 samples (30m × 4 = 2h)1 month (30 days)
Rollup 4Past Year1 day (24 hours)12 Rollup 3 samples (2h × 12 = 24h)1 year (365 days)

Statistics Collection Levels

vCenter Server defines four hierarchical Statistics Collection Levels (Level 1 through Level 4). Higher levels collect more granular counter types but significantly expand the vCenter PostgreSQL database footprint:

  • Level 1 (Default): Captures basic performance metrics including aggregated CPU, memory, network, and disk usage, plus system uptime. It includes average statistics only and does not track per-device metrics.
  • Level 2: All Level 1 metrics plus additional CPU, disk, memory, and other counters, but not minimum and maximum rollup values.
  • Level 3: Level 1 and 2 metrics plus device-level counters (individual disks, NICs, vCPUs), still without minimum and maximum rollups.
  • Level 4: Every counter vCenter supports, including minimum and maximum rollup values. It is the most detailed level and the most expensive in database space.

Compute & Memory Diagnostic Metrics

Virtual machines share physical compute cores and host RAM under the arbitration of the VMkernel scheduler. Understanding specific counters is crucial to isolating whether a VM is suffering from CPU starvation, co-scheduling latency, or memory reclamation pressure.

Core CPU Performance Metrics

  1. CPU Usage (usagemhz and usage %): Measures the actual compute cycles processed by the virtual machine or host. High CPU usage is not necessarily indicative of a problem if the workload requires processing power and ready time remains low.
  2. CPU Ready Time (%RDY): Represents the percentage of time that a virtual machine was ready to execute instructions on a physical CPU core but was forced to wait in the VMkernel run queue because physical processing cores were fully occupied by other workloads. The thresholds below are widely used rules of thumb, not VMware limits.
    • Normal Baseline: < 5% per vCPU.
    • Warning Threshold: 5% - 10% per vCPU (indicates emerging compute contention).
    • Critical Threshold: > 10% per vCPU (severe performance degradation; users experience sluggish application response).
  3. CPU Co-Stop (%CSTP): Tracks the percentage of time a multi-vCPU (SMP) virtual machine was stopped and paused by the VMkernel co-scheduler while waiting for sibling vCPUs to achieve synchronization.
    • Underlying Mechanism: ESXi requires that the multiple vCPUs of an SMP virtual machine advance their execution clocks in near lock-step. When one vCPU executes faster than another, the VMkernel deliberately halts the faster vCPU to allow the lagging vCPU to catch up.
    • Threshold: %CSTP > 3% indicates excessive co-scheduling penalty.
    • Exam Trap & Remediation: When a VM experiences high %CSTP, administrators often intuitively attempt to allocate more vCPUs to solve the sluggishness. This exacerbates the problem. The definitive fix for high %CSTP is to reduce the number of vCPUs assigned to the virtual machine, allowing the VMkernel scheduler to place the remaining vCPUs much more rapidly across available physical cores.
  4. CPU Max Limited (%MLMTD): The percentage of time a VM was ready to run but was throttled by the VMkernel because it reached its configured CPU Limit (in MHz). If %MLMTD > 0%, the workload is being artificially restricted by administrative configuration, not physical resource exhaustion.

Core Memory Performance Metrics

  • Consumed Memory: The total amount of physical host memory allocated to the virtual machine. Because guest operating systems touch and zero memory pages during boot and never return them unless instructed, consumed memory frequently approaches 100% of configured vRAM and does not reflect real-time memory need.
  • Active Memory: The hypervisor's statistical estimate of the memory pages actively read or written by the guest OS within the recent sampling window. This is the true metric for sizing virtual machine memory allocations.
  • Granted Memory: The amount of physical memory mapped to the virtual machine, including shared pages.
  • Ballooned Memory (vmmemctl): Memory reclaimed from the guest OS by inflating the balloon driver inside the guest. The guest OS decides which idle pages to allocate to the balloon driver, preserving guest paging efficiency.
  • Swapped Memory (.vswp): Memory forced out of host RAM and written directly to the virtual machine's .vswp file on the datastore. Because storage latency is orders of magnitude slower than physical RAM, host swapping causes catastrophic workload degradation.

ESXi Memory Reclamation Escalation Sequence

In vSphere 6.0 and later, the VMkernel defines host memory states as percentages of a host-specific minFree value, and each state unlocks a more expensive technique (older documents quote the ESX-era 6%/4%/2%/1% free-memory thresholds, which no longer apply):

  1. High (400% of minFree) and Clear (100% of minFree): Normal operation. Transparent Page Sharing (TPS) works within each VM (inter-VM sharing only if salting is relaxed), and between High and Clear, ESXi starts breaking large pages so TPS can share them.
  2. Soft (64% of minFree): The VMkernel inflates the balloon driver (vmmemctl) in guests, which makes each guest give up memory it chooses, paging to its own swap if needed.
  3. Hard (32% of minFree): Ballooning cannot keep up, so ESXi compresses pages into the compression cache and swaps unreserved memory to the .vswp file.
  4. Low (16% of minFree): ESXi keeps swapping and also blocks VMs that exceed their memory targets until free memory recovers.
Loading diagram...
ESXi Memory Reclamation Escalation States

Storage, Network, and Alarm Management Architecture

Storage bottlenecks are among the most frequent causes of enterprise application failure. To troubleshoot storage performance, administrators must break down storage latency into distinct hypervisor and physical infrastructure layers.

Storage Latency Decomposition

Total storage latency experienced by a virtual machine guest operating system is represented as Guest Average Latency (GAVG/cmd). The VMkernel decomposes this value into two discrete measurements:

GAVG/cmd=KAVG/cmd+DAVG/cmd\text{GAVG/cmd} = \text{KAVG/cmd} + \text{DAVG/cmd}
+-----------------------------------------------------------------------------------+
| Storage Latency Decomposition Architecture                                        |
|                                                                                   |
| [ Virtual Machine Guest OS ]                                                      |
|             |                                                                     |
|             |  GAVG/cmd (Total Guest Latency, target < 15-20 ms)                  |
|             v                                                                     |
| [ ESXi VMkernel Storage Stack ]                                                   |
|   - LUN Queues / Adapter Queues                                                   |
|             |  KAVG/cmd (Kernel Queuing Latency, target < 1-2 ms)                 |
|             v                                                                     |
| [ Storage Fabric & SAN / Storage Controller ]                                     |
|   - FC Switches / IP Network / Physical Disks / SSDs                              |
|             |  DAVG/cmd (Physical Device Latency, target < 15-20 ms)              |
|             v                                                                     |
| [ Physical Storage Array ]                                                        |
+-----------------------------------------------------------------------------------+
  • Kernel Latency (KAVG/cmd): Measures the average time a SCSI command spends waiting inside the ESXi VMkernel storage queue before being passed to the physical Host Bus Adapter (HBA). Under normal operating conditions, KAVG/cmd must remain below 1 to 2 milliseconds. If KAVG/cmd exceeds 2 ms, the ESXi host is queuing requests due to saturated adapter queue depth (disk.sched.maxQLen) or storage path throttle limits.
  • Device Latency (DAVG/cmd): Measures the time elapsed between the HBA transmitting the SCSI command onto the physical SAN fabric and receiving a completion status from the storage array. It represents the combined latency of the physical switches, SAN controllers, cache modules, and physical disks. For enterprise spinning disk arrays, DAVG/cmd should remain under 15 to 20 ms; for all-flash and NVMe-oF arrays, DAVG/cmd should remain under 1 to 3 ms.
  • Queue Latency (QAVG/cmd): Average time spent per command queued inside the storage driver or device queue. Spikes indicate queue starvation.
  • Command Aborts (ABRTS/s): The number of SCSI commands aborted per second. Any sustained value greater than 0 indicates physical SAN timeouts, fabric zoning flaps, or failing storage controllers.

Network Performance Metrics

  • Network Dropped Packets (%DRPRX and %DRPTX): Percentage of received or transmitted packets dropped by the virtual switch or physical network interface card (pNIC). Under healthy conditions, dropped packets must be 0%.
    • High %DRPRX on a virtual port indicates that the guest OS or virtual NIC driver cannot process incoming packets fast enough, overflowing the receive ring buffer.
    • High %DRPTX on an uplink port indicates physical network switch port congestion or link speed/duplex mismatch.
  • Data Rate (MbRX/s and MbTX/s): Measures Megabits received or transmitted per second. Used to detect bandwidth saturation on teamed uplinks.

Performance Metrics Threshold Reference Table

CounterSubsystemHealthy TargetWarning ThresholdCritical Action Threshold
%RDYCPU< 5% per vCPU5% - 10%> 10% (Reduce host consolidation ratio, migrate VMs)
%CSTPCPU0% - 1%1% - 3%> 3% (Right-size VM: reduce assigned vCPU count)
%MLMTDCPU0%> 0%> 0% (Remove or increase artificial CPU Limit)
MCTLSZMemory0 MB> 0 MBEscalating balloon size (Add physical RAM to host)
SWCURMemory0 MB> 0 MBActive hypervisor swapping (Host RAM exhausted)
KAVG/cmdStorage< 1 ms1 - 2 ms> 2 ms (Adjust LUN queue depth, multipathing)
DAVG/cmdStorage< 10 ms15 - 20 ms> 25 ms (SAN array/fabric bottleneck)
%DRPRX / %DRPTXNetwork0%> 0.5%> 1% (Check pNIC buffer, vNIC type, port team)

Alarms and Notification Architecture in vCenter

vCenter Server includes a built-in alarm subsystem that monitors inventory objects against performance thresholds and system events:

  • Alarm Triggers: Alarms can be triggered by specific conditions or state (e.g., CPU usage exceeds 90% for 5 minutes) or by discrete events (e.g., host connection lost, power supply redundancy degraded).
  • Tolerance Intervals and Duration: Metric-based alarms include a time condition (e.g., "Trigger alert if metric exceeds threshold for 300 seconds") to prevent false alarms from momentary usage spikes.
  • Automated Actions: Alarms can execute automated remediation and alert actions including sending Email notifications (SMTP), generating SNMP traps to enterprise monitoring systems, or executing a custom script/PowerCLI workflow on the vCenter Server.

Using Performance Charts (Objective 5.10)

Select any object (vCenter, data center, cluster, host, resource pool, VM, or datastore) and open Monitor > Performance:

ViewUse It For
OverviewPredefined charts for the object (CPU, memory, disk, network) over a chosen period; quick triage
AdvancedCustom charts. Chart Options let you pick the metric group, time span (real-time or a historical interval), chart type (line, stacked, or pie where applicable), the objects, and the counters. Save the settings as a named chart option

Practical tips:

  • Real-time charts show the last hour at 20-second resolution; use them for live issues. Historical views use the rollup intervals above.
  • Counters that exist only at higher statistics levels (for example per-device metrics) stay empty until you raise the level for that interval.
  • Charts can be exported (for example to CSV, PNG, or JPEG) to attach to support cases or change reviews.
  • Compare CPU Ready and Co-stop per VM, Balloon and Swapped memory, and disk latency counters alongside usage, because high usage alone is not a problem.
Test Your Knowledge

A database administrator reports that a critical 8-vCPU virtual machine is experiencing severe latency. Performance charting in the vSphere Client reveals an average CPU Ready Time (%RDY) of 1.2%, but CPU Co-Stop (%CSTP) is consistently hovering at 8.5%. What is the most effective administrative action to resolve this performance problem?

A

Increase the virtual machine's CPU reservation to 100% of configured clock cycles

B

Allocate 8 additional vCPUs to the virtual machine to double its execution parallelism

C

Reduce the number of vCPUs assigned to the virtual machine from 8 to 4 or 2

D

Change the virtual machine's CPU share value from Normal to High

Test Your Knowledge

An ESXi 8.0 host experiences severe memory pressure during a sudden compute workload spike. Which memory reclamation mechanism does ESXi rely on when host free memory reaches the Soft state (64% of the host's minFree value)?

A

Memory Ballooning via the vmmemctl guest driver

B

Hypervisor Swapping of unreserved memory directly to the .vswp disk file

C

Memory Compression inside the host compression cache

D

Inter-VM Transparent Page Sharing across isolated salting domains

Test Your Knowledge

A systems engineer is analyzing an alert reporting high storage latency on a virtual machine disk. In the performance charts, GAVG/cmd is 48 ms, DAVG/cmd is 46 ms, and KAVG/cmd is 2 ms. What does this telemetry indicate about the root cause of the storage bottleneck?

A

The ESXi host storage adapter queue depth is saturated, causing excessive hypervisor queuing

B

The virtual machine's virtual SCSI controller driver is corrupt and requires a VMware Tools upgrade

C

The physical network switch port connected to the ESXi management vmnic is dropping packets

D

The physical storage array, SAN fabric, or disk spindles are experiencing latency

Sections you finish are checked off in the contents.