2.4 Virtual Machine Advanced Compute Configuration
Key Takeaways
Virtual hardware version 20 (vSphere 8.0) supports up to 768 vCPUs and 24 TB of memory per VM, limits first reached with hardware version 18 in vSphere 7.0 U1, and adds features such as Device Groups and DirectPath I/O with Device Virtualization Extensions (DVX).
vCPU topology design (Sockets vs Cores per Socket) directly influences virtual NUMA (vNUMA) node formation; crossing physical NUMA boundaries incurs memory access latency penalties.
Latency Sensitivity 'High' requires a full CPU reservation and a 100% memory reservation and gives each vCPU exclusive access to a physical core; vSphere 8's 'High with Hyperthreading' gives each vCPU exclusive access to one hyperthread instead.
ESXi executes memory reclamation through an established four-stage hierarchy: Transparent Page Sharing (TPS), Ballooning (vmmemctl), Memory Compression, and Hypervisor Swapping (.vswp).
The size of a virtual machine's swap file (.vswp) equals Configured RAM minus Reserved RAM; full 100% memory reservations create a 0-byte swap file, saving datastore capacity.
Virtual Machine Advanced Compute Configuration
Optimizing virtual machine compute performance requires deep architectural alignment between guest operating system workloads and the physical hardware topology of ESXi hosts. Configuring virtual hardware versions, Non-Uniform Memory Access (NUMA) topologies, latency sensitivity parameters, and hypervisor memory reclamation mechanisms ensures mission-critical workloads operate with peak throughput and deterministic latency.
Virtual Hardware 20 & Scalability Limits
With each major release of VMware vSphere, the virtual hardware specification (virtual machine compatibility level) advances, exposing new physical hardware primitives and expanding resource scalability limits. Virtual Hardware version 20 (vHW 20) arrived with vSphere 8.0. The large jump in VM size came earlier, with hardware version 18 in vSphere 7.0 Update 1:
| Capability | Hardware 17 (vSphere 7.0) | Hardware 18 / 19 (vSphere 7.0 U1 / U2) | Hardware 20 (vSphere 8.0) |
|---|---|---|---|
| Max vCPUs per VM | 256 | 768 | 768 |
| Max Memory per VM | 6 TB | 24 TB | 24 TB |
| Notable additions | Virtual watchdog timer, precision clock, vSGX | Larger VMs (HW 18); UEFI HTTP boot needs HW 19 or later | Device Groups, DirectPath I/O with Device Virtualization Extensions (DVX), vTPM provisioning policy, latency sensitivity with hyper-threading |
Upgrading virtual hardware exposes new instruction sets and hardware capabilities to the guest OS. However, virtual hardware upgrades must be managed deliberately because newer hardware versions prevent virtual machines from running on older ESXi hosts lacking support for that version.
NUMA, vNUMA, and vCPU Topology Design
Modern enterprise multi-socket servers employ Non-Uniform Memory Access (NUMA) architecture. In a NUMA system, physical processors and physical memory are divided into discrete NUMA nodes:
- Local Memory Access: When a CPU core accesses memory attached to its own socket, latency is lowest.
- Remote Memory Access: When a core fetches memory attached to another socket, the request crosses the interconnect (such as Intel UPI or AMD Infinity Fabric), and latency rises noticeably.
+--------------------------------------------------------------------------+
| Physical NUMA Node Topology |
+--------------------------------------------------------------------------+
| Physical Socket 0 (NUMA Node 0) Physical Socket 1 (NUMA Node 1) |
| +-----------------------------+ +-----------------------------+ |
| | 16 CPU Cores | | 16 CPU Cores | |
| | 128 GB Local Memory | | 128 GB Local Memory | |
| +-----------------------------+ +-----------------------------+ |
| ▲ ▲ |
| └────── Intel UPI / AMD Fabric ───────┘ |
| (Remote Access Latency Penalty) |
+--------------------------------------------------------------------------+
Virtual NUMA (vNUMA)
When a virtual machine is sized with more than 8 vCPUs (or configured to exceed the core count of a single physical NUMA node), the ESXi kernel automatically creates a Virtual NUMA (vNUMA) topology inside the guest OS. This informs the guest operating system's kernel scheduler (such as the Linux or Windows NUMA-aware kernel) how to place memory allocations and process threads across virtual NUMA nodes that correspond to physical NUMA nodes.
Sockets vs. Cores per Socket Best Practice
In early vSphere versions, changing the "Cores per Socket" setting directly altered how vSphere structured vNUMA nodes. Since vSphere 6.5 and later, vSphere decouples vCPU topology from vNUMA formation:
- By default, ESXi automatically calculates the optimal vNUMA topology based on total vCPU count and physical host NUMA architecture.
- Administrators should configure virtual machines with 1 socket and multiple cores, or multiple sockets with 1 core per socket, primarily to conform to guest OS licensing constraints (e.g., operating systems licensed on a per-socket basis).
- Wide VM Risk: If a VM is allocated 24 vCPUs on a physical host with 16 cores per NUMA node, the VM becomes a "Wide VM" spanning two NUMA nodes. Half of its memory accesses will traverse remote interconnects unless the workload inside the guest is strictly NUMA-aware.
Latency Sensitivity: Normal vs. High
For most standard enterprise workloads, ESXi defaults to Latency Sensitivity: Normal, utilizing dynamic co-scheduling, thread preemption, and shared physical core balancing. However, for extreme jitter-sensitive workloads (such as high-frequency trading platforms, real-time audio/telecom processing, and NFV data planes), vSphere provides Latency Sensitivity: High.
When Latency Sensitivity is configured to High:
- Exclusive Physical Core Allocation: High latency sensitivity requires a full CPU reservation as well as the memory reservation. ESXi then gives each vCPU exclusive access to a physical core, and the partner hyper-thread stays idle so neighbors cannot interfere. vSphere 8's High with Hyperthreading option (virtual hyperthreading) instead gives each vCPU exclusive access to one hyper-thread, placing consecutive vCPU pairs on the two hyper-threads of a core.
- 100% Memory Reservation: The virtual machine is automatically required to have a 100% memory reservation. Memory pages are pinned in physical RAM; zero ballooning or hypervisor swapping can ever occur.
- Reduced Virtualization Overhead: ESXi tunes interrupt delivery and idle handling for the VM to cut jitter and give guest threads lower, more predictable latency.
- Operational Constraints: Latency-sensitive VMs require that physical host CPU capacity is not overcommitted. If a physical core is dedicated exclusively to a vCPU, DRS cannot freely place other workloads on that physical core.
Memory Allocation & The Four-Stage Reclamation Ladder
When ESXi hosts experience progressive memory pressure, the VMkernel reclaims memory through a four-stage hierarchy. Each stage is tied to a host memory state. In vSphere 6.0 and later, those states are defined as percentages of a host-specific value called minFree: High (400% of minFree), Clear (100%), Soft (64%), Hard (32%), and Low (16%).
+--------------------------------------------------------------------------+
| ESXi Memory Reclamation Hierarchy |
+--------------------------------------------------------------------------+
| Stage 1: Transparent Page Sharing (TPS) [Lowest Performance Overhead] |
| └── De-duplicates identical memory pages via Copy-on-Write |
| |
| Stage 2: Memory Ballooning (vmmemctl) [Guest-Cooperative Reclamation] |
| └── Inflates driver in guest; guest OS pages to internal swap |
| |
| Stage 3: Memory Compression Cache [Host Memory In-Flight Compression] |
| └── Pages compressible to 2 KB or less go to compression cache |
| |
| Stage 4: Hypervisor Swapping (.vswp) [Highest Overhead / Last Resort] |
| └── Forces unreserved memory pages to disk storage |
+--------------------------------------------------------------------------+
Stage 1: Transparent Page Sharing (TPS)
- Operation: ESXi scans memory pages for identical contents, consolidates duplicate pages into a single physical memory address, and marks the page as Copy-on-Write (COW). If a VM writes to the shared page, a private copy is created instantly.
- Security Salting: Since 2014, vSphere salts TPS by default for security reasons, which mitigates side-channel attacks, so sharing happens within a VM only. Inter-VM sharing requires changing the host option
Mem.ShareForceSaltingor giving VMs the samesched.mem.pshare.saltvalue. - Large Pages: When free memory drops below the High state toward Clear, ESXi starts breaking large pages into small pages so TPS can collapse them.
Stage 2: Memory Ballooning (vmmemctl)
- Operation: As host free memory approaches the Soft state (64% of minFree), the VMkernel signals the
vmmemctlballoon driver that VMware Tools installs in the guest. - The balloon driver inflates, requesting memory from the guest OS. The guest OS views the balloon as an active process consuming memory and relies on its own internal virtual memory manager to page idle memory to the guest OS swap file.
- The balloon driver passes the freed physical addresses back to the VMkernel, reclaiming host physical RAM with zero hypervisor-level disk I/O.
Stage 3: Memory Compression
- Operation: When free memory reaches the Hard state (32% of minFree) and ballooning cannot keep pace, ESXi compresses and swaps pages. It tries compression first for pages that are about to be swapped.
- If a 4 KB page compresses to 2 KB or less (50% or better), it goes into the VM's in-memory compression cache, which is capped by default at 10% of the VM's configured memory (
Mem.MemZipMaxPct). Reading a compressed page back from RAM is far faster than reading it from swap on disk.
Stage 4: Hypervisor Swapping (.vswp)
- Operation: Hypervisor swapping starts with compression in the Hard state. In the Low state (16% of minFree), ESXi keeps swapping and also blocks VMs that exceed their memory targets until free memory recovers.
- The VMkernel directly writes guest memory pages to the virtual machine's
.vswpfile on the datastore. Because disk I/O latency is orders of magnitude slower than physical RAM, application throughput plummets catastrophically.
Swap File Sizing & Placement Mechanics
The physical size of a virtual machine's swap file (.vswp) is calculated using a strict mathematical formula:
- If a VM is configured with 32 GB RAM and has a 0 GB reservation, the host creates a 32 GB
.vswpfile on the datastore when the VM powers on. - If a VM is configured with 32 GB RAM and has a 32 GB reservation (100% reservation), the
.vswpfile size is 0 bytes (or a small metadata header).
Administrators can configure the swap file location:
- Same directory as the VM (Default): Keeps swap alongside VMDKs on shared storage. Enables seamless vMotion migrations without transferring swap files across hosts.
- Dedicated host datastore: Places swap files on local high-speed host SSDs/NVMe drives. This saves expensive SAN/vSAN storage capacity, but requires transferring swap files across the network during vMotion migrations.
CPU Affinity and Other Per-VM Compute Settings (Objective 7.3)
Objective 7.3 lists several per-VM settings you manage in Edit Settings. Per-VM EVC and latency sensitivity are covered above. The other one you must recognize is CPU affinity.
Scheduling Affinity
In Edit Settings > Virtual Hardware > CPU > Scheduling Affinity, you enter the physical logical CPUs a VM may run on, for example 0-3 or 4,6.
| Consideration | Effect |
|---|---|
| Placement | The VM's vCPUs are restricted to the listed logical CPUs |
| Migration | After vMotion the affinity may no longer apply, because the destination host can have a different number of processors |
| Performance | Limits the scheduler's load balancing, and the NUMA scheduler may be unable to manage a VM pinned with affinity |
| Reservations | CPU admission control ignores affinity, so a pinned VM might not receive its full reservation |
| Isolation | Does not reserve those CPUs; other VMs can still use them unless you pin them elsewhere too |
VMware advises avoiding manual CPU affinity. Treat it as a last resort for special licensing or troubleshooting cases. Reservations, shares, latency sensitivity, and vNUMA settings usually solve the same problem without these side effects.
Other Settings Worth Knowing
| Setting | Where | Purpose |
|---|---|---|
| CPU and memory hot add | Edit Settings > CPU / Memory | Add resources while the VM runs (hot-adding CPUs disables vNUMA for that VM) |
| Cores per socket | Edit Settings > CPU | Present a socket and core layout, mainly for guest licensing |
| Reservation / limit / shares | Edit Settings > CPU / Memory | Guarantee, cap, or prioritize resources |
| VM compatibility (hardware version) | Actions > Compatibility | Upgrade the virtual hardware (the VM must be powered off; plan a guest reboot) |
Exam Trap: CPU affinity does not dedicate cores to a VM, and it can undermine reservations, NUMA placement, and load balancing. If the requirement is dedicated cores for a low-latency workload, the answer is Latency Sensitivity High with full CPU and memory reservations, not CPU affinity.
A virtual machine is configured with 64 GB of vRAM and has an explicit memory reservation of 48 GB. Upon powering on the virtual machine, how much datastore disk space will ESXi allocate for this virtual machine's .vswp swap file?
16 GB
64 GB
48 GB
0 GB
When an ESXi host experiences escalating memory pressure, what is the exact operational sequence in which memory reclamation mechanisms are executed?
Hypervisor Swapping, followed by Memory Compression, Ballooning, and Transparent Page Sharing
Memory Compression, followed by Hypervisor Swapping, Transparent Page Sharing, and Ballooning
Transparent Page Sharing, followed by Ballooning (vmmemctl), Memory Compression, and Hypervisor Swapping (.vswp)
Ballooning, followed immediately by Hypervisor Swapping, bypassing Memory Compression if SSD storage is detected
An architect is configuring a financial trading virtual machine that needs the lowest and most consistent CPU latency possible, with minimal hypervisor jitter. What configuration setting should be applied?
Enable Memory Ballooning with a 0% reservation to allow dynamic memory expansion
Set Latency Sensitivity to 'High', which allocates exclusive physical cores and mandates 100% memory reservations
Configure a 'Should run on' DRS soft affinity rule with Priority 5 aggressive migration thresholds
Deploy Virtual Hardware version 17 with unconstrained multi-socket vNUMA interleaving
Sections you finish are checked off in the contents.