3.2 VMFS Datastore Architecture & Management
Key Takeaways
VMFS-6 keeps a 1 MB file block size, replaces VMFS-5's 8 KB sub-blocks with 64 KB small file blocks and 512 MB large file blocks, and adds automatic background space reclamation (UNMAP).
VMFS cluster distributed locking uses Atomic Test and Set (ATS) hardware acceleration to replace legacy SCSI-2 full-LUN reservations with sector-level locking.
Extending a datastore by expanding an existing LUN partition preserves data integrity, whereas spanning multiple LUN extents creates a multi-disk failure dependency.
Thick Provision Eager Zeroed zeroes every block at creation, so first writes carry no zeroing penalty. It is required for shared disks that use the multi-writer flag and for WSFC clustered VMDKs, but not for vSphere Fault Tolerance.
3.2 VMFS Datastore Architecture & Management
The Virtual Machine File System (VMFS) is a high-performance, clustered filesystem designed specifically for virtual machine storage. VMFS allows multiple autonomous ESXi hosts concurrent read and write access to the same shared block storage pool (Fibre Channel, FCoE, or iSCSI LUNs). By managing concurrent access directly within the filesystem metadata, VMFS enables core vSphere enterprise capabilities including vSphere vMotion, High Availability (HA), and Distributed Resource Scheduler (DRS).
Understanding VMFS internal architecture, locking mechanisms, capacity expansion operations, and disk provisioning formats is essential for maintaining enterprise performance and availability.
VMFS-5 vs. VMFS-6 Architectural Evolution
vSphere 6.5 introduced VMFS-6, representing a major architectural redesign of the underlying filesystem structure compared to legacy VMFS-5.
1. Unified Block Size and Sub-Block Allocation
- VMFS-5 Block Architecture: VMFS-5 unified the filesystem block size to 1 MB for regular files, supporting virtual disk files up to 62 TB in size. However, to store small files (such as
.vmxconfiguration files, logs, and small descriptor files) without wasting full 1 MB blocks, VMFS-5 used 8 KB sub-blocks. As a datastore hosted thousands of small files, managing millions of discrete 8 KB sub-blocks generated heavy metadata fragmentation and locking overhead. - VMFS-6 Block Architecture: VMFS-6 eliminates 8 KB sub-blocks. Instead, it utilizes two distinct, highly optimized internal block types:
- Small File Blocks (SFB): Sized at 64 KB, drastically reducing the total number of metadata pointers required for small files.
- Large File Blocks (LFB): 512 MB regions that VMFS-6 uses to allocate large thick-provisioned files efficiently. The file block size itself stays 1 MB.
2. Storage Media & 4K Native (4Kn) Drive Support
Traditional enterprise hard drives and arrays operated with physical sector sizes of 512 bytes (512n). Modern high-density spinning drives and NVMe storage devices utilize physical sector sizes of 4096 bytes (4K Native).
- VMFS-5: Supports 512n drives and 512-byte emulation (512e) drives. VMFS-5 cannot run on native 4Kn storage devices.
- VMFS-6: Fully supports 512e and native 4Kn drives (under supported storage controllers and NVMe over Fabrics). VMFS-6 aligns data structures natively to 4K boundaries, preventing the severe read-modify-write performance penalties that occur when 512-byte filesystem writes misalign with 4K physical sectors.
3. Space Reclamation (UNMAP)
When files or virtual disks are deleted or migrated off a datastore, the underlying blocks become unallocated at the filesystem layer. However, thin-provisioned storage arrays cannot detect filesystem deletions unless an explicit SCSI UNMAP command (SCSI opcode 0x42) informs the array controller that those logical block addresses (LBAs) are no longer needed.
- VMFS-5 Manual UNMAP: VMFS-5 did not support automatic background space reclamation. Administrators were forced to run manual command-line operations during off-peak hours using
esxcli storage vmfs unmap -u <Datastore-UUID>or scheduled scripts. - VMFS-6 Automatic Background UNMAP: VMFS-6 integrates an internal, continuous, asynchronous space reclamation engine. When a VMDK is deleted or moved, VMFS-6 tracks the freed blocks and quietly issues SCSI UNMAP requests to the array in the background. The reclamation rate can be configured in vCenter or via CLI:
- Priority method, Low (default): Sends UNMAP at 25-50 MB/s so guest I/O is not affected (settable in the vSphere Client or with esxcli).
- Priority method, Medium or High: Roughly two or three times the Low rate (50-100 MB/s, or over 100 MB/s), set with esxcli only.
- Fixed method: You specify the reclamation rate in MB/s. Set it by editing the space reclamation settings of an existing datastore in the vSphere Client.
- None: Disables automatic reclamation.
4. Upgrade Path: Migration Requirement
There is no in-place, non-destructive upgrade from VMFS-5 to VMFS-6. Because VMFS-6 introduces fundamental changes to the disk layout, metadata block structures, and allocation bitmaps, an existing VMFS-5 volume cannot be converted. Administrators must create a new VMFS-6 datastore on newly provisioned storage LUNs and migrate running virtual machines using Storage vMotion.
Cluster Distributed Locking: SCSI-2 vs. Atomic Test and Set (ATS)
Because VMFS is a shared cluster filesystem, multiple ESXi hosts can issue read and write requests to the same LUN simultaneously. However, filesystem metadata operations—such as creating a new file, powering on a virtual machine (acquiring a file lock), expanding a VMDK, or creating a snapshot—require synchronization across all hosts to prevent metadata corruption.
+-----------------------------------------------------------------------------+
| VMFS Distributed Locking Mechanisms |
+-----------------------------------------------------------------------------+
| |
| [ Legacy SCSI-2 LUN Reservation ] (Entire LUN Locked) |
| |
| Host A -----------> [ SCSI RESERVE ] ----------> [ Entire LUN 1 (Locked) ]
| Host B -----------> [ I/O Blocked ] (Conflict) | |
| Host C -----------> [ I/O Blocked ] (Conflict) | |
| v |
| Host A -----------> [ SCSI RELEASE ] -----------> [ LUN Unlocked ] |
| |
| ----------------------------------------------------------------------- |
| |
| [ VAAI ATS Hardware Assisted Locking ] (Discrete Sector Locked) |
| |
| Host A ----> [ COMPARE AND WRITE ] ----> [ Sector 0x4A (Locked) ] |
| Host B ----> [ Reads / Writes I/O ] ---> [ Sector 0x9B (Free) ] |
| Host C ----> [ Reads / Writes I/O ] ---> [ Sector 0x1F (Free) ] |
| |
| Result: No LUN-wide locking conflicts. All hosts proceed simultaneously. |
+-----------------------------------------------------------------------------+
Legacy SCSI-2 Reservations
In early versions of ESXi and VMFS, metadata operations relied on standard SCSI-2 RESERVE (opcode 0x16) and RELEASE (opcode 0x17) commands.
- Mechanism: When Host A needed to update datastore metadata, it issued a SCSI-2
RESERVEcommand. The storage array controller locked the entire LUN exclusively for Host A. - The Bottleneck: While Host A held the lock, any read or write I/O issued by Host B or Host C to any virtual machine on that same LUN was rejected with a SCSI Reservation Conflict status (
0x18). The other hosts were forced to back off and retry, leading to major I/O latency spikes and limiting cluster scalability to a small number of hosts per datastore.
Atomic Test and Set (ATS)
Modern vSphere environments rely on Hardware-Assisted Locking, also known as Atomic Test and Set (ATS), an enterprise feature of vSphere APIs for Array Integration (VAAI) using the SCSI COMPARE AND WRITE command (opcode 0x89).
- Mechanism: Instead of locking the entire LUN, the ESXi host instructs the array controller to atomically inspect and update only a specific 512-byte metadata sector on disk.
- Execution: The host transmits both the expected current data and the new metadata. The storage array verifies whether the sector matches the expected data; if it matches, the array writes the new data as one atomic operation. If another host updated the sector in the interim, the operation fails and is immediately re-tried.
- Benefit: Regular guest VM I/O across the rest of the LUN continues uninterrupted. VMFS-5 and VMFS-6 datastores created on ATS-capable storage use ATS-only locking, so many hosts can share one datastore without SCSI reservation conflicts.
Datastore Expansion: LUN Expansion vs. Multi-Extent Spanning
When a VMFS datastore approaches capacity, administrators can expand the filesystem using one of two methods:
1. Extending an Existing LUN (Recommended Best Practice)
- Workflow:
- The SAN administrator increases the physical capacity of the existing LUN on the storage array.
- The vSphere administrator performs a storage adapter rescan in vCenter Server.
- The administrator selects the datastore, chooses Increase Datastore Capacity, and selects the contiguous unpartitioned space on that existing device.
- ESXi expands the partition table and grows the filesystem without interrupting running virtual machines.
- Advantages: The datastore remains mapped 1:1 to a single physical LUN. There is no added architectural complexity, and performance, array-based replication, and snapshot operations remain predictable.
2. Adding Extents (Multi-Extent Spanning)
- Workflow: The administrator selects an existing VMFS datastore and adds a completely separate, additional LUN as an extent. The VMFS volume spans across multiple distinct LUNs (up to a maximum of 32 extents, aggregate volume limit 64 TB).
- Severe Operational Risks:
- Extent Failure Risk: If the head (first) extent goes offline, the entire datastore becomes inaccessible. If another extent goes offline, the datastore is marked degraded: any VM with blocks on that extent loses access, and the resulting failed I/O can make
hostdunresponsive. Recovery may require restoring affected VMs from backup, which is why Broadcom recommends single-extent datastores. - Asymmetric Performance: If extents reside on arrays or storage tiers with different spindle speeds, RAID types, or controller cache configurations, VM performance becomes erratic and unpredictable.
- Incompatibility: Spanned datastores complicate or break array-level features such as synchronous replication, storage-level snapshots, and storage tiering.
- Extent Failure Risk: If the head (first) extent goes offline, the entire datastore becomes inaccessible. If another extent goes offline, the datastore is marked degraded: any VM with blocks on that extent loses access, and the resulting failed I/O can make
Caution
Exam Trap & Architecture Best Practice: Never use multi-extent spanned datastores in production environments unless temporarily consolidating or evacuating data. Always expand capacity by increasing the size of the underlying LUN on the storage array.
Virtual Disk Provisioning Formats
When creating a virtual machine disk (VMDK), vSphere offers three disk provisioning formats. The choice impacts provisioning speed, runtime write latency, storage overhead, and enterprise feature support.
+-----------------------------------------------------------------------------+
| Virtual Disk Provisioning Formats Overview |
+-----------------------------------------------------------------------------+
| |
| 1. Thin Provisioning |
| Creation: Instant (0 GB physical allocation) |
| First Write: Allocates block + zeroes block + writes data |
| Disk Size on Datastore grows dynamically as guest writes |
| |
| 2. Thick Provision Lazy Zeroed |
| Creation: Fast (Full disk space allocated on datastore) |
| First Write: Zeroes block + writes data |
| Blocks reserved upfront, but wiped only on initial write |
| |
| 3. Thick Provision Eager Zeroed |
| Creation: Slowest (Full disk space allocated AND all blocks zeroed) |
| First Write: Writes data directly with ZERO zeroing penalty |
| * Required for multi-writer shared disks (e.g., Oracle RAC) |
+-----------------------------------------------------------------------------+
1. Thin Provision
- Allocation: Consumes zero physical blocks at creation time (except for minimal metadata pointers). Disk space is allocated dynamically from the datastore as the guest operating system writes new data.
- Performance: Incurs a minor latency penalty on the first write to any unwritten block, because the VMkernel must allocate the block, zero it out (for security), and update filesystem metadata before committing the guest data.
- Management Risk: Allows storage overcommitment. If multiple VMs consume more capacity than the physical datastore possesses, the datastore will reach 100% capacity. When a thin-provisioned VM attempts to write to a full datastore, ESXi pauses the virtual machine and displays an alert in vCenter until additional storage capacity is added.
2. Thick Provision Lazy Zeroed
- Allocation: The entire requested capacity is reserved and allocated on the datastore at creation time. No other VM can consume this physical space.
- Behavior: The allocated blocks on the physical storage media contain old, unzeroed residual data from prior operations. When the guest OS writes to a block for the first time, ESXi writes zeros across the block immediately before writing the guest payload.
- Performance: Fast creation time; immune to out-of-space datastore pauses. Suffers a slight latency penalty on initial block writes, but subsequent writes to initialized blocks execute at native speed.
3. Thick Provision Eager Zeroed
- Allocation: The full capacity is allocated and reserved on the datastore upfront, and every single block is overwritten with zeros during disk creation.
- Behavior: Because all blocks are already sanitized, guest operating systems write directly to disk without any allocation or zeroing latency penalty from the very first I/O.
- Where It Is Required:
- Multi-Writer Flag Configurations: Shared disks used by Oracle Real Application Clusters (RAC) with the multi-writer flag, and clustered VMDKs used by Microsoft Windows Server Failover Clustering (WSFC), must be Eager Zeroed Thick.
- Not required for Fault Tolerance: Legacy single-vCPU FT once required eager-zeroed disks, but current multi-processor FT supports thin and thick disks.
VMFS-5 vs. VMFS-6 Feature Matrix
| Capability | VMFS-5 | VMFS-6 |
|---|---|---|
| Maximum Datastore Size | 64 TB | 64 TB |
| Maximum Virtual Disk Size | 62 TB | 62 TB |
| Base Block Size | 1 MB | 1 MB |
| Sub-Block Allocation | 8 KB sub-blocks | 64 KB Small File Blocks (SFB) |
| Large File Allocation | 1 MB file blocks | 512 MB Large File Blocks (1 MB file block size) |
| Drive Sector Format Support | 512n and 512e | 512e and 4K Native (4Kn) |
| Space Reclamation (UNMAP) | Manual CLI (esxcli storage vmfs unmap) | Automatic, background asynchronous |
| Reclamation Rate | Manual execution | Low priority (25-50 MB/s, default), Medium, High, or a fixed MB/s rate |
| Default Distributed Locking | ATS-only | ATS-only |
| In-Place Upgrade Path | From VMFS-3 to VMFS-5 supported | Not supported (requires Storage vMotion) |
Creating and Modifying Datastores (Objective 7.4.1)
Create a Datastore
Right-click a host, cluster, or data center and choose Storage > New Datastore:
- VMFS: name it, select a device (LUN) visible to the host, choose VMFS 6, set the partition layout (usually all available space), and review the block size (1 MB) and space reclamation settings. Rescan the other hosts so they mount it.
- NFS: choose NFS 3 or NFS 4.1, then enter the name, export path (folder), and server address(es). For 4.1 you can add multiple server addresses and Kerberos. Select the hosts to mount it on, with the same version everywhere.
- vVols: select a storage container that the array exposes through its VASA provider.
Modify a Datastore
| Task | Notes |
|---|---|
| Increase capacity | Grow the LUN on the array, rescan, then Increase Datastore Capacity into the same device (preferred) or add an extent |
| Rename / browse | Rename freely; browse files to upload ISOs or find orphaned VMDKs |
| Space reclamation | Edit the priority or fixed-rate settings on VMFS 6 |
| Unmount | Before unmounting, the datastore must hold no running VMs, not be in a datastore cluster managed by Storage DRS, have SIOC disabled, and not be used for vSphere HA heartbeating |
| Delete | Destroys the VMFS volume and its contents; unmount from every host first as a safety step |
| Maintenance mode | For datastores in a datastore cluster; Storage DRS evacuates them |
An infrastructure team manages a vSphere 7.0 cluster connected to an older VMFS-5 datastore. The storage team requests that automatic background space reclamation (UNMAP) be enabled on the datastore. How should the administrator accomplish this objective?
Enable the Automatic Space Reclamation checkbox in the VMFS-5 datastore properties within vCenter Server.
Provision a new VMFS-6 datastore on newly allocated storage and use Storage vMotion to migrate the virtual machines.
Run esxcli storage vmfs unmap set --automatic=true from the ESXi console of the cluster master.
Execute an in-place filesystem upgrade from VMFS-5 to VMFS-6 using the vSphere Client.
A database administrator is configuring an enterprise clustered database across two virtual machines running on separate ESXi hosts using the multi-writer flag. Which virtual disk provisioning format is mandatory for these shared disks?
Thin Provision with automatic space reclamation enabled.
Thick Provision Lazy Zeroed with an attached Raw Device Mapping.
Thick Provision Lazy Zeroed formatted with 64 KB allocation unit size.
Thick Provision Eager Zeroed.
A production VMFS-6 datastore is running low on disk space. An administrator considers spanning the datastore across a second LUN by adding an extent. What is the primary operational risk associated with this multi-extent configuration?
Losing the head extent makes the whole datastore inaccessible, and losing any other extent makes every VM with data on it inaccessible
The datastore will automatically revert from VMFS-6 to VMFS-5 distributed locking semantics
Virtual machines on the secondary extent cannot participate in vSphere vMotion or vSphere HA
Automatic background space reclamation (UNMAP) will permanently stop operating across all datastores in vCenter
Sections you finish are checked off in the contents.