4.1 vSAN Core Architecture & Disk Groups
Key Takeaways
vSAN is a software-defined, hyperconverged storage tier embedded natively into the ESXi VMkernel, eliminating the need for virtual storage appliances (VSAs).
The Original Storage Architecture (OSA) organizes physical storage into disk groups with exactly 1 flash cache drive and 1 to 7 capacity drives, supporting up to 5 disk groups (35 capacity drives) per host.
Hybrid OSA splits the cache device into 70% read cache and 30% write buffer; all-flash OSA uses the cache device only as a write buffer, and vSAN 8 raised the logical write-buffer limit per disk group from 600 GB to 1.6 TB.
The Express Storage Architecture (ESA) in vSAN 8 eliminates disk groups entirely, employing a single-tier NVMe storage pool managed by the high-performance vSAN Log-Structured File System (vSAN LFS).
Physical networking requirements mandate 1 GbE dedicated for hybrid OSA, 10 GbE shared or dedicated for all-flash OSA, and 25 GbE minimum for Express Storage Architecture ReadyNodes.
4.1 vSAN Core Architecture & Disk Groups
VMware vSAN is a distributed, software-defined storage tier embedded natively into the ESXi hypervisor kernel. Rather than relying on traditional external storage area networks (SAN) or network-attached storage (NAS) arrays—or routing I/O through intermediary Virtual Storage Appliances (VSAs)—vSAN pools server-attached flash devices and magnetic disks across an ESXi cluster into a single shared, distributed object store.
Because vSAN is integrated directly into the VMkernel I/O path, it delivers lower CPU overhead, deterministic sub-millisecond latencies, and linear scalability. With the release of vSAN 8, VMware introduced two parallel operational architectures:
- Original Storage Architecture (OSA): The proven two-tier disk group architecture supporting both hybrid and all-flash media configurations.
- Express Storage Architecture (ESA): A single-tier, next-generation architecture optimized exclusively for high-performance NVMe flash media and modern multi-core server platforms.
The Native VMkernel Storage Stack
Traditional virtual storage appliances run as virtual machines that intercept SCSI commands, translating them through virtual network drivers. In contrast, vSAN operates directly inside the VMkernel using five core distributed modules:
+-----------------------------------------------------------------------------+
| vSAN Native VMkernel Architecture |
+-----------------------------------------------------------------------------+
| |
| [ Virtual Machine (VMDK) ] |
| | |
| +--------------------------------v------------------------------------+ |
| | VMkernel Storage Stack | |
| | | |
| | +-------------------------------------------------------------+ | |
| | | CLOM (Cluster Level Object Manager) | | |
| | | Validates policies, calculates layouts & initiates rebuilds | | |
| | +-------------------------------------------------------------+ | |
| | | | |
| | +----------------------------v--------------------------------+ | |
| | | DOM (Distributed Object Manager) | | |
| | | Coordinates distributed I/O, mirrors, parity & split-brain | | |
| | +-------------------------------------------------------------+ | |
| | | | | |
| | | (Local I/O) | (Network I/O) | |
| | +-------v---------------------+ +-------v--------------+ | |
| | | LSOM (Local Storage Obj Mgr)| | RDT (Reliable | | |
| | | Manages local SSDs/HDDs | | Datagram Transport) | | |
| | +-----------------------------+ +----------------------+ | |
| | | | |
| | +--------------------------------------------------v--------------+ | |
| | | CMMDS (Cluster Monitoring, Membership, and Directory Services) | | |
| | | Maintains cluster membership, heartbeat & object metadata state | | |
| | +-----------------------------------------------------------------+ | |
| +---------------------------------------------------------------------+ |
| | |
| [ Physical Storage Devices (NVMe / SAS / SATA) ] |
+-----------------------------------------------------------------------------+
- CLOM (Cluster Level Object Manager): Validates storage policy compliance before object creation, calculates optimal component placement layouts across cluster nodes, and coordinates data evacuation and rebuild workflows during maintenance or failures.
- DOM (Distributed Object Manager): Acts as the traffic coordinator for distributed storage objects. Each object has a DOM Client (on the host executing the VM) and a DOM Owner (the authoritative node managing the object's components, mirrors, and witnesses). It ensures consistent read/write routing across hosts.
- LSOM (Local Storage Object Manager): Interfaces directly with local physical storage devices within the host, handling block allocation, error reporting, drive-level health monitoring, and local caching.
- CMMDS (Cluster Monitoring, Membership, and Directory Services): The distributed directory and heartbeat service. It discovers cluster nodes, monitors cluster membership, tracks physical drive states, and maintains an active catalog of all object and component locations.
- RDT (Reliable Datagram Transport): A high-performance, kernel-level transport protocol operating over TCP/IP that handles inter-host communication for replication, resynchronization, and witness voting.
Original Storage Architecture (OSA) Mechanics
The Original Storage Architecture employs a two-tier, hierarchical storage design built around Disk Groups. A disk group is a logical container on an ESXi host that binds a dedicated cache device to multiple capacity devices.
Disk Group Rules & Maximums
- Cache Tier Requirement: Exactly 1 flash device per disk group. You cannot configure multiple cache drives within a single disk group.
- Capacity Tier Requirement: Between 1 and 7 capacity devices per disk group.
- Disk Groups per Host: An ESXi host can host a maximum of 5 disk groups.
- Maximum Drives per Host: 5 cache drives + 35 capacity drives = 40 physical storage drives maximum per host.
- Cluster Storage Minimum: At least one host in the cluster must contribute storage. However, production clusters require all hosts to contribute symmetrical disk groups to prevent capacity imbalances and uneven resync bottlenecks.
Hybrid vs. All-Flash OSA Configurations
OSA supports two distinct operational modes that dictate caching algorithms and hardware requirements:
| Architectural Attribute | Hybrid OSA | All-Flash OSA |
|---|---|---|
| Cache Device Media | Flash (SAS, SATA, or NVMe SSD) | High-Endurance Flash (NVMe or SAS SSD) |
| Capacity Device Media | Magnetic Spinning Disks (SAS/NL-SAS/SATA) | Flash (SATA, SAS, or NVMe SSD) |
| Cache Sizing Rule | 10% Rule: Cache must equal >= 10% of anticipated consumed capacity | Sized for write endurance and burst write buffering |
| Read Cache Allocation | 70% of Cache Drive: Dynamically caches frequently accessed read blocks | 0%: Reads are serviced directly from flash capacity drives |
| Write Buffer Allocation | 30% of Cache Drive: Absorbs incoming writes, commits to flash, and destages to spinning disks | 100% of Cache Drive: Absorbs incoming writes, commits to flash, and destages to flash capacity |
| Maximum Usable Write Buffer | Limited by 30% of physical cache device size | 1.6 TB logical limit per disk group in vSAN 8 (600 GB in vSAN 7 and earlier) |
| Deduplication & Compression | Not Supported | Supported (Configured per cluster) |
| Erasure Coding (RAID-5/6) | Not Supported (RAID-1 Mirroring only) | Supported (Configured via VM Storage Policy) |
Important
The All-Flash Write Buffer Limit: In an all-flash OSA configuration, vSAN dedicates the cache device to the write buffer. vSAN 7 and earlier capped the logical write buffer at 600 GB per disk group. vSAN 8 raises the limit to 1.6 TB, so a 1.6 TB cache device can be used in full as write buffer on vSAN 8, while vSAN 7 would address only 600 GB of it. Across five disk groups, a vSAN 8 OSA host can therefore have up to 8 TB of write buffer.
Express Storage Architecture (ESA) in vSAN 8
vSAN 8 introduced the Express Storage Architecture (ESA), designed from the ground up to exploit fast NVMe drives, modern high-core-count processors, and 25GbE/100GbE networks.
Single-Tier Storage Pool (No Disk Groups)
In ESA, the legacy concept of disk groups is completely eradicated:
- Uniform Drive Role: There is no dedicated cache tier. All NVMe drives contribute simultaneously to both capacity and performance.
- Storage Pool Construction: Drives are claimed into a single, cluster-wide storage pool per host. An administrator simply selects the certified NVMe drives, and vSAN adds them to the unified pool.
- Failure Isolation: In OSA, the failure of a single cache drive takes the entire disk group (up to 7 capacity drives) offline. In ESA, because each drive operates independently within the storage pool, a drive failure only affects the components residing on that specific physical device, dramatically shrinking the failure domain.
vSAN Log-Structured File System (vSAN LFS)
ESA replaces the legacy object engine with the vSAN Log-Structured File System (vSAN LFS):
- Write Path Optimization: Incoming write I/O is written sequentially to a fast, log-structured write log before being packaged and written in full stripes to persistent capacity. This eliminates small random write amplification.
- Upper-Layer Data Services: In ESA, Compression and Encryption occur at the highest layer of the storage stack before data is replicated across the network. If an incoming block compresses from 4 KB to 2 KB, only 2 KB of data is transmitted across the vSAN VMkernel network to secondary hosts, cutting network fabric utilization in half.
- Native High-Performance Snapshots: Traditional vSphere VM snapshots rely on redo logs that add read overhead and require consolidation when deleted. ESA builds a native snapshot engine into the vSAN LFS metadata layer, so snapshots have much less performance impact and deleting them avoids redo-log consolidation.
Sizing, ReadyNodes, and Reserve Capacities
Designing a stable vSAN cluster requires accounting for compute, memory, and capacity overheads alongside operational growth buffers.
vSAN ReadyNodes
VMware validates hardware through vSAN ReadyNode profiles—pre-certified server configurations from major OEM vendors tailored to specific workload profiles:
- ReadyNode Profiles: Categorized by performance and capacity tiers (e.g., ESA-2 through ESA-8, or OSA AF-4 through AF-8). ReadyNodes prescribe validated CPU socket counts, memory configurations, certified storage controller HBA models, drive endurance tiers, and network adapter chipsets.
- Self-Built / Build-Your-Own (BYO): Permitted in OSA, but every individual component (controller, firmware, driver, SSD model) must be strictly certified on the VMware vSAN Compatibility Guide (VCG/HCL). For ESA, deployments must strictly comply with official ESA ReadyNode specifications.
vSAN Reserve Capacity (Operations & Host Rebuild Reserve)
Historically, administrators were advised to maintain a manual "20% to 30% slack space" rule to accommodate component rebuilds, snapshot growth, and rebalancing. Modern vSphere releases formally integrate this through automated Reserve Capacity settings:
- Operations Reserve (Slack Space):
- Capacity reserved for internal vSAN background operations, including policy reconfigurations, object rebalancing, and dynamic snapshot consolidation.
- vSAN sizes it automatically based on cluster capacity.
- Host Rebuild Reserve:
- Capacity reserved to ensure that if a single host fails completely, the remaining hosts have sufficient free disk space to immediately rebuild all affected components to 100% policy compliance.
- vSAN dynamically calculates this based on
1 / N(where N is the number of hosts in the cluster). In a 4-host cluster, the Host Rebuild Reserve consumes 25% of cluster capacity; in an 8-host cluster, it consumes only 12.5%.
CPU and RAM Overhead
- CPU Overhead: Varies with the data services enabled (encryption, compression, deduplication).
- Memory and CPU Minimums:
- OSA: hosts need at least 32 GB of memory to support the maximum configuration of 5 disk groups with 7 capacity devices each.
- ESA: hosts need at least 128 GB of memory and 16 CPU cores.
Networking Infrastructure Prerequisites
Storage performance in a hyperconverged architecture is inextricably tied to physical network quality. vSAN does not utilize traditional SCSI HBAs; the physical Ethernet switch fabric is the storage backplane.
Bandwidth Mandates
- Hybrid OSA: 1 GbE dedicated physical uplinks minimum. (10 GbE highly recommended).
- All-Flash OSA: 10 GbE dedicated or shared physical uplinks minimum. When the links are shared with vMotion or VM traffic, use Network I/O Control (NIOC) shares or reservations on a vSphere Distributed Switch to protect vSAN traffic.
- Express Storage Architecture (ESA): 25 GbE physical uplinks minimum are required for production environments. 100 GbE is recommended for dense NVMe configurations.
Unicast vs. Legacy Multicast
- Prior to vSAN 6.6, cluster heartbeats and directory services mandated Layer 2 / Layer 3 Multicast (IGMP Snooping and PIM routing).
- vSAN 6.6 and all modern versions operate exclusively via Unicast. Multicast is completely eliminated from the architecture. Cluster membership is distributed via CMMDS using direct unicast communication between the designated cluster master, backup master, and agent nodes.
MTU & Jumbo Frames
- vSAN fully supports Jumbo Frames (MTU 9000), which can reduce CPU interrupt overhead on high-throughput workloads.
- Crucial Rule: Jumbo frames must be configured consistently end-to-end across the entire path: VMkernel adapters (
vmk), vSphere Distributed Switch, physical switch ports, and any intermediate IP routing gateways. A mismatch where an ESXi host transmits MTU 9000 packets to a physical switch configured for MTU 1500 will result in silent packet drops, leading to immediate cluster partitioning and host disconnections.
Architectural Comparison: OSA vs. ESA
| Feature / Metric | Original Storage Architecture (OSA) | Express Storage Architecture (ESA) |
|---|---|---|
| Storage Topology | Two-Tier: Disk Groups (1 Cache + 1-7 Capacity) | Single-Tier: Unified NVMe Storage Pool |
| Media Types | Hybrid (SSD + HDD) or All-Flash (SSD + SSD) | All-NVMe Flash (TLC / QLC) only |
| Max Disk Groups | 5 Disk Groups per host | None (Single Storage Pool) |
| Drive Failure Impact | Cache drive failure takes entire disk group offline | Single drive failure affects only local components |
| Max Write Buffer | 1.6 TB logical limit per disk group (vSAN 8) | No dedicated cache tier; log-structured writes |
| Snapshot Architecture | Redo-log (delta disk) snapshots | Native vSAN LFS snapshots |
| Compression & Encryption | Applied at LSOM layer (after network transmission) | Applied at top layer (before network replication) |
| Minimum Network | 1 GbE (Hybrid) / 10 GbE (All-Flash) | 25 GbE (Production ReadyNodes) |
| Minimum Host Count | 2 hosts (ROBO/Direct-Connect) / 3 hosts (Standard) | 3 hosts minimum for standard deployment |
A storage architect is designing an all-flash vSAN cluster using the Original Storage Architecture (OSA). Each ESXi host contains high-performance storage controllers and multiple flash drives. What is the maximum number of capacity drives that a single ESXi host can contribute to the vSAN datastore across all configured disk groups?
28 capacity drives
32 capacity drives
35 capacity drives
40 capacity drives
An administrator deploys an all-flash vSAN 8 cluster using the Original Storage Architecture (OSA) and equips each disk group with a 1.6 TB enterprise NVMe SSD as the dedicated cache tier. How does vSAN 8 use this cache device?
It uses the device as the disk group's write buffer, up to vSAN 8's 1.6 TB logical write-buffer limit
It allocates 70% of the drive (1,120 GB) for read caching and 30% (480 GB) for write buffering
It uses only 600 GB as write buffer and leaves the remaining 1 TB unaddressed
It reserves 600 GB for the write buffer and assigns the remaining 1,000 GB as a secondary read cache
Which architectural characteristic accurately describes the vSAN Express Storage Architecture (ESA) in vSAN 8 compared to the legacy Original Storage Architecture (OSA)?
ESA requires a dedicated 1.6 TB NVMe drive per host exclusively for write caching before destaging to QLC storage pools.
ESA processes data compression and data-at-rest encryption at the local storage driver layer after network replication.
ESA utilizes multi-hop IGMP snooping multicast trees to replicate storage metadata across cluster nodes.
ESA eliminates disk groups in favor of a single-tier NVMe storage pool and builds native snapshots into its log-structured file system.
Sections you finish are checked off in the contents.