1.3 vSphere Distributed Services Engine (DSE) & DPUs

Key Takeaways

  • vSphere Distributed Services Engine (DSE), introduced in vSphere 8.0, offloads network packet processing and NSX security services from server CPUs to PCIe Data Processing Units (DPUs / SmartNICs).

  • DSE executes an ARM-compiled instance of the ESXi VMkernel directly on the DPU hardware, operating synchronously alongside the primary host x86 hypervisor.

  • Offloading infrastructure functions to DPUs returns x86 CPU cycles that software switching, overlay encapsulation, and packet inspection would otherwise consume.

  • To offload a VM port group, use a vSphere Distributed Switch 8.0.0 or later whose Network Offloads compatibility is set to the DPU vendor (AMD Pensando or NVIDIA BlueField); Network I/O Control is disabled on that switch.

  • DPU lifecycle management is fully integrated into vSphere Lifecycle Manager (vLCM), updating host ESXi images and DPU firmware in a unified remediation workflow.

Last updated: September 2026

1.3 vSphere Distributed Services Engine (DSE) & DPUs

As enterprise network bandwidth escalated from 10 GbE to 25 GbE, 40 GbE, and 100 GbE, standard server processors experienced an increasingly severe "infrastructure tax." Modern software-defined networking, packet encapsulation (Geneve/VXLAN), distributed firewall inspection, and storage virtualization consume substantial host CPU cycles simply routing packets and enforcing security policies. In vSphere 8.0, VMware addressed this bottleneck by introducing the vSphere Distributed Services Engine (DSE), enabling hypervisor offloads to Data Processing Units (DPUs), commonly known as SmartNICs.


The I/O Tax & The Hardware DPU Paradigm Shift

In a standard vSphere architecture without DSE, whenever a virtual machine transmits or receives network traffic, the host's x86-64 CPU cores must execute:

  • Virtual distributed switch (vDS) forwarding and frame replication.
  • Network overlay encapsulation/decapsulation (e.g., VMware NSX Geneve headers).
  • Stateful firewall packet evaluation (NSX Distributed Firewall rules).
  • Quality of Service (QoS) scheduling and traffic shaping.

In high-throughput environments, this packet processing can consume a significant share of host CPU capacity. Cores that could run workload VMs are spent on hypervisor networking tasks instead.

+-------------------------------------------------------------------------+
| Standard vSphere Host (No DPU)       vSphere 8 with DSE (DPU Offload)   |
|                                                                         |
|   +--------------------------+         +--------------------------+     |
|   |    Workload VMs          |         |    Workload VMs          |     |
|   +--------------------------+         +--------------------------+     |
|   |  Infrastructure tax on   |         |  More x86 cycles left    |     |
|   |  x86 cores (vDS, NSX,    |         |  for workload VMs        |     |
|   |  DFW, Geneve)            |         |                          |     |
|   +--------------------------+         +--------------------------+     |
|   |      x86 Host CPU        |         |       x86 Host CPU       |     |
|   +--------------------------+         +--------------------------+     |
|                |                                     | PCIe             |
|                v                                     v                  |
|   +--------------------------+         +--------------------------+     |
|   |   Standard NIC           |         |  Data Processing Unit    |     |
|   |   (packet I/O only)      |         |  (SmartNIC, Arm cores)   |     |
|   +--------------------------+         |  - vDS datapath offload  |     |
|                                        |  - NSX services          |     |
|                                        |  - Geneve encap/decap    |     |
|                                        +--------------------------+     |
+-------------------------------------------------------------------------+

What is a Data Processing Unit (DPU)?

A Data Processing Unit (DPU) is an intelligent PCIe expansion card that combines:

  1. High-speed multi-port network controllers (dual 25 GbE, 50 GbE, or 100 GbE).
  2. Programmable multi-core embedded processors (for example, the NVIDIA BlueField-2 has 8 Arm Cortex-A72 cores).
  3. Its own onboard memory, separate from host RAM.
  4. Hardware acceleration engines for cryptography, compression, packet parsing (P4 engines), and direct memory access (DMA).

DSE Architectural Model: Dual-Hypervisor Pairing

The architectural breakthrough of vSphere Distributed Services Engine is that ESXi runs natively inside the DPU. DSE does not treat the DPU as a simple hardware offload ASIC; rather, it deploys a specialized instance of the ESXi hypervisor directly onto the card's ARM processor cores.

+---------------------------------------------------------------------------------+
| Physical Server Enclosure                                                       |
|                                                                                 |
|   +-------------------------------------------------------------------------+   |
|   | Host x86 Compute Domain                                                 |   |
|   |   [ VM 1 ]              [ VM 2 ]              [ VM 3 ]                  |   |
|   |   -------------------------------------------------------------------   |   |
|   |   ESXi 8.0 Hypervisor (x86-64 Kernel)                                   |   |
|   |   - User World (hostd, vpxa)                                            |   |
|   |   - VMkernel CPU & Memory Scheduler                                     |   |
|   +-------------------------------------------------------------------------+   |
|                                        |                                        |
|                                        | High-Speed PCIe Gen 4/5 Bus            |
|                                        v                                        |
|   +-------------------------------------------------------------------------+   |
|   | DPU Hardware Domain (PCIe SmartNIC)                                     |   |
|   |   ESXi 8.0 Hypervisor (ARM-compiled Kernel)                             |   |
|   |   - Embedded vDS Offload Engine                                         |   |
|   |   - NSX Distributed Firewall (DFW) In-Silicon Processing                 |   |
|   |   - Geneve Tunnel End Point (TEP) Hardware Offload                      |   |
|   |   -------------------------------------------------------------------   |   |
|   |   Physical Hardware: 8-16 ARM Cores | 16-32 GB RAM | Dual 25/100G Ports |   |
|   +-------------------------------------------------------------------------+   |
|                                        |                                        |
|                                        v                                        |
|                     Physical Data Center Network Fabric                         |
+---------------------------------------------------------------------------------+

Dual-Hypervisor Operational Mechanics

When a DSE-enabled host boots:

  1. The server hardware initializes the DPU PCIe device.
  2. The primary x86 ESXi image boots on the server's main x86 processor cores.
  3. Simultaneously, an ARM-compiled ESXi image boots on the DPU's onboard ARM cores.
  4. The host ESXi instance and the DPU ESXi instance establish an internal, secure control channel across the PCIe bus.
  5. In the vSphere Client, the administrator sees a single managed ESXi host. The presence of the DPU is reflected as a hardware accelerator component within the host configuration.

Offloaded Networking and Security Workloads

With DSE enabled, network packet processing is offloaded completely from the host CPU:

  • vDS Fast Path Offload: The vSphere Distributed Switch datapath executes directly inside the DPU. Network packets traveling from a virtual machine pass over the PCIe bus into the DPU, where switching decisions occur at line rate.
  • VMware NSX Services: With NSX, networking and security services such as the Distributed Firewall can also run on the DPU. Exactly which services are offloaded depends on the NSX release, so check the NSX documentation for your version.
  • Geneve Encap/Decap: Overlay encapsulation headers are stamped and stripped directly by the DPU hardware, eliminating software packet overhead.

Supported Hardware Platforms & Enterprise Benefits

vSphere Distributed Services Engine is engineered in close collaboration with leading silicon and server manufacturers. DSE cannot be deployed on generic, uncertified consumer PCIe network cards.

Certified Silicon & Server OEM Ecosystem

  • Supported DPU Silicon Vendors:
    • NVIDIA BlueField-2 DPU: Features dual 25 GbE or 100 GbE network interfaces, 8 ARM Cortex-A72 processor cores, and dedicated crypto accelerators.
    • AMD Pensando Distributed Services Card (DSC-25 / DSC-100): Features programmable P4 packet engines and ARM processing clusters.
  • Certified Server Platforms: DSE requires a server, DPU, and firmware combination listed on the VMware Compatibility Guide. Launch partners included Dell and HPE servers. The platform's BIOS, PCIe firmware, power, and cooling must be validated for the DPU card.

Quantifiable Architectural Advantages

Architectural DimensionTraditional Standard ESXi HostvSphere 8.0 with DSE and DPU Offload
Host CPU OverheadNetworking and security processing runs on host x86 coresMuch of that processing moves to the DPU's cores
Workload DensityInfrastructure services compete with VMs for CPUFreed x86 capacity can run more workload VMs
Network ThroughputLimited by software datapath processingHardware-accelerated datapath on the DPU
Packet Latency & JitterSubject to host CPU scheduling delaysDatapath runs outside the host CPU scheduler
Security ArchitectureDFW runs in the host VMkernelSecurity services can run in a separate DPU domain

Security Separation

Beyond performance, VMware positions DSE as a security improvement: the offloaded infrastructure datapath runs apart from the workload VMs.

In a standard hypervisor, if a rogue virtual machine successfully executes an advanced hypervisor escape exploit to achieve root control of the host x86 VMkernel, the attacker could theoretically manipulate local firewall rules and sniff traffic from adjacent VMs.

With vSphere Distributed Services Engine and NSX, infrastructure services can run in the DPU's own ESXi instance instead of on the host's x86 cores, so workload VMs never share processor cores with the offloaded networking and security datapath. Treat this as defense in depth, not a guarantee that a fully compromised host can never affect traffic.

Lifecycle Management of DPUs via vLCM

A critical concern for data center operators is operational complexity. Managing firmware, drivers, and operating systems across hundreds of auxiliary PCIe accelerators could quickly create management chaos. VMware addressed this by integrating DPU maintenance natively into vSphere Lifecycle Manager (vLCM).

+-------------------------------------------------------------------+
| vSphere Lifecycle Manager (vLCM) Cluster Image                    |
|                                                                   |
|   [ Base ESXi 8.0 Release (x86-64) ]                              |
|   [ Vendor Add-on (Server BIOS / CPLD / Platform Firmware) ]      |
|   [ DPU Component (ESXi on DPU ARM Image & DPU Firmware) ]        |
|                                                                   |
|                              |                                    |
|                              | Single Coordinated Remediation     |
|                              v                                    |
|   +-----------------------------------------------------------+   |
|   | Physical Host with DPU                                    |   |
|   |   1. Host Enters Maintenance Mode                         |   |
|   |   2. Host x86 ESXi Patched                                |   |
|   |   3. DPU ARM ESXi & Firmware Flashed                      |   |
|   |   4. Coordinated System Reboot                            |   |
|   +-----------------------------------------------------------+   |
+-------------------------------------------------------------------+

Unified vLCM Cluster Images

vLCM treats the server host and its attached DPU as a single, indivisible managed entity through declarative Cluster Images:

  • Single Image Definition: The vLCM cluster image specifies the base ESXi release, OEM vendor add-ons, and the exact DPU software component (containing the ARM ESXi build and DPU firmware).
  • Synchronized Remediation: When an administrator initiates cluster remediation, vLCM automatically places the host into Maintenance Mode, evacuating virtual machines via vSphere vMotion. It updates the host x86 hypervisor and simultaneously flashes the DPU firmware and ARM ESXi image. The host and DPU are then rebooted together.
  • Images Required: Clusters with DPU-backed hosts must be managed with vLCM images; baselines (the legacy vSphere Update Manager model) cannot lifecycle the DPU.

Configuring a VM Port Group for DPU Offload (Objective 5.6)

The blueprint asks you to configure a VM port group so its traffic is offloaded to a DPU. In vSphere 8 the offload decision is made at the distributed switch level, so the workflow is:

  1. Prepare DPU-backed hosts. Use servers, DPUs, and firmware listed together on the VMware Compatibility Guide. The ESXi installation also places ESXi on the DPU, and the cluster must be managed with a vLCM image.
  2. Create a vSphere Distributed Switch at version 8.0.0 or later. In the creation wizard, set Network Offloads compatibility to the DPU vendor: Pensando or NVIDIA BlueField. The default, None, means offloads are not supported on that switch.
  3. Note the Network I/O Control side effect. When you select Pensando or NVIDIA BlueField, Network I/O Control is disabled on the switch, so bandwidth management through NIOC shares and reservations is not available there.
  4. Add the DPU-backed hosts and map the DPU's physical ports to the switch uplinks.
  5. Activate the offloaded datapath with NSX Manager. The DPU datapath is NSX Enhanced Data Path, configured on the hosts as transport nodes. VMware's launch guidance was that NSX Manager is needed even for plain vDS offload (Enhanced Data Path Standard, which came with the vSphere Enterprise Plus entitlement), while overlay networking, the distributed firewall on the DPU, and UPT need NSX licensing.
  6. Create the distributed port groups for the VMs. VMs use one of two modes:
    • MUX mode (default): no guest requirements, and some processing stays on the x86 host.
    • Uniform Passthrough (UPTv2): near-passthrough performance with vMotion and HA preserved. It requires virtual hardware version 20, a VMXNET3 adapter with Use UPT Support enabled and a supported driver, and a full guest memory reservation.
SettingWhereExam point
vDS versionCreate Distributed Switch wizardMust be 8.0.0 or later
Network Offloads compatibilitySame wizardPensando or NVIDIA BlueField enables offload; None does not
Network I/O ControlSwitch propertiesDisabled when offload compatibility is selected
LifecycleCluster Updates tabvLCM image required for DPU hosts

Exam Trap: Enabling offload is not a per-VM checkbox on a standard switch. Port groups inherit the capability from a vDS created with the matching Network Offloads compatibility. A standard switch, or a vDS created with None, cannot offload traffic to the DPU.

Loading diagram...
vSphere Distributed Services Engine (DSE) Hardware & Software Boundary

Realistic Failure Scenarios & Troubleshooting

Scenario 1: Offload Configured on the Wrong Switch

Problem: A systems engineer adds new DPU-backed ESXi 8.0 hosts to an existing vSphere Distributed Switch and moves a latency-sensitive application's port group onto them. Host CPU usage for networking does not drop, and the switch reports no offload capability.

Root Cause Analysis:

  • The existing switch was created with Network Offloads compatibility: None, the default.
  • Offload capability comes from the distributed switch's compatibility setting. The switch must be version 8.0.0 or later and set to the DPU vendor (Pensando or NVIDIA BlueField).

Resolution:

  1. Create a vSphere Distributed Switch 8.0.0 or later and select the matching Network Offloads compatibility in the creation wizard.
  2. Add the DPU-backed hosts and assign the DPU ports as uplinks.
  3. Create or migrate the application's distributed port groups to the new switch.
  4. Plan bandwidth management without Network I/O Control, which is disabled on offload-compatible switches.

Scenario 2: vLCM Remediation Rejection

Problem: An administrator attempts to remediate a DSE-enabled cluster using vSphere Lifecycle Manager. The compliance check reports that the image cannot be applied because of a component dependency conflict involving the DPU software.

Root Cause Analysis:

  • The administrator manually modified the vLCM cluster image by selecting a newer ESXi 8.0 update release, but left the vendor add-on at an older version.
  • Because the DPU software component contains tight dependencies between the DPU firmware version and the ARM-compiled ESXi kernel, mismatched components violate vLCM dependency checks.

Resolution:

  1. In the vSphere Client, navigate to the Cluster -> Updates -> Image tab.
  2. Select Edit Image and update the OEM Vendor Add-On to match the target ESXi base release.
  3. Run Check Compliance. Once all dependency validations pass, initiate cluster remediation.
Test Your Knowledge

An administrator must offload the network datapath for a group of VMs to NVIDIA BlueField DPUs installed in new ESXi 8.0 hosts. Which configuration provides the offload for the VMs' port group?

A

A vSphere Standard Switch with Route Based on IP Hash teaming on the DPU ports

B

A vSphere Distributed Switch version 7.0.3 with Network I/O Control version 3 enabled

C

A vSphere Distributed Switch version 8.0.0 or later created with Network Offloads compatibility set to NVIDIA BlueField

D

Any vSphere Distributed Switch version with LACP configured on the DPU uplinks

Test Your Knowledge

Which vSphere operational management tool is mandatory for patching and maintaining hosts and their attached Data Processing Units (DPUs) in a unified, synchronized remediation workflow?

A

vSphere Lifecycle Manager (vLCM) using declarative Cluster Images

B

Legacy vSphere Update Manager (VUM) using attached baseline groups

C

ESXi Auto Deploy using stateless answer files

D

The Direct Console User Interface (DCUI) system recovery menu

Test Your Knowledge

What is the primary architectural mechanism through which vSphere Distributed Services Engine (DSE) delivers significant performance gains for compute-intensive workloads?

A

It converts virtual machine guest memory into non-volatile storage classes

B

It offloads network switching, overlay encapsulation, and security processing to the DPU's Arm cores and accelerators, returning x86 CPU cycles to workloads

C

It enables ESXi to run 32-bit legacy operating systems without binary translation overhead

D

It replaces virtual machine SCSI controllers with software emulation running in the vCenter Appliance

Sections you finish are checked off in the contents.