Created on Jul 2, 2026 (v 26.4 GA)
Scope
This Reference Deployment Guide (RDG) provides comprehensive instructions for deploying the NVIDIA DOCA Platform Framework (DPF) on high-performance, bare-metal infrastructure in Zero-Trust mode. The guide focuses on setting up an accelerated Host-Based Networking (HBN) service on NVIDIA® BlueField®-3 DPUs to deliver secure, isolated, and hardware-accelerated environments.
The guide is intended for experienced system administrators, systems engineers, and solution architects who build highly secure bare-metal environments with Host-Based Networking enabled using NVIDIA BlueField DPUs for acceleration, isolation, and infrastructure offload.
This document is an extension of the RDG for DPF Zero Trust (DPF-ZT) (referred to as the Baseline RDG). It details the additional steps and modifications required to deploy the HBN, Service into the Baseline RDG environment.
-
This reference implementation, as the name implies, is a specific, opinionated deployment example designed to address the use case described above.
-
Although other approaches may exist for implementing similar solutions, this document provides a detailed guide for this specific method.
Abbreviations and Acronyms
|
Term |
Definition |
Term |
Definition |
|---|---|---|---|
|
BFB |
BlueField Bootstream |
OOB |
Out-of-Band |
|
BGP |
Border Gateway Protocol |
PF |
Physical Function |
|
DOCA |
Data Center Infrastructure-on-a-Chip Architecture |
RDG |
Reference Deployment Guide |
|
DPF |
DOCA Platform Framework |
RDMA |
Remote Direct Memory Access |
|
DPU |
Data Processing Unit |
RoCE |
RDMA over Converged Ethernet |
|
HBN |
Host Based Networking |
SFC |
Service Function Chaining |
|
IPAM |
IP Address Management |
SR-IOV |
Single Root Input/Output Virtualization |
|
K8S |
Kubernetes |
VLAN |
Virtual LAN (Local Area Network) |
|
KVM |
Kernel-based Virtual Machine |
VNI |
Virtual Network Interface |
|
MAAS |
Metal as a Service |
VRF |
Virtual Router/Forwarder |
|
MTU |
Maximum Transmission Unit |
ZT |
Zero Trust |
|
NGC |
NVIDIA GPU Cloud |
|
|
Introduction
The NVIDIA BlueField-3 Data Processing Unit (DPU) is a 400 Gb/s infrastructure compute platform designed for line-rate processing of software-defined networking, storage, and cybersecurity workloads. It combines powerful compute resources, high-speed networking, and advanced programmability to deliver hardware-accelerated, software-defined solutions for modern data centers.
NVIDIA DOCA unleashes the full potential of the BlueField platform by enabling rapid development of applications and services that offload, accelerate, and isolate data center workloads.
One such service is Host-Based Networking (HBN) - a DOCA-enabled solution that allows network architects to design networks based on Layer 3 (L3) protocols. HBN enables routing on the server side by using BlueField as a BGP router. It encapsulates key networking functions in a containerized service pod, deployed directly on the BlueField’s Arm cores.
However, deploying and managing DPUs and their associated DOCA services, especially at scale, presents operational challenges. Without a robust provisioning and orchestration system, tasks such as lifecycle management, service deployment, and network configuration for service function chaining (SFC) can quickly become complex and error prone. This is where the DOCA Platform Framework (DPF) comes into play.
DPF automates the full DPU lifecycle, streamlines the deployment of DOCA services, and simplifies advanced network configurations. With DPF, services such as HBN can be deployed seamlessly, allowing for efficient offloading and intelligent routing of traffic through the DPU data plane.
By leveraging DPF, users can scale and automate DPU management across Bare Metal, Virtual, and Kubernetes customer environments - optimizing performance while simplifying operations.
DPF supports multiple deployment models. This guide focuses on the Zero Trust bare-metal deployment model. In this scenario:
-
The DPU is managed through its Baseboard Management Controller (BMC)
-
All management traffic occurs over the DPU's out-of-band (OOB) network
-
The host is considered as an untrusted entity towards the data center network. The DPU acts as a barrier between the host and the network.
-
The host sees the DPU as a standard NIC, with no access to the internal DPU management plane (Zero Trust Mode)
This Reference Deployment Guide (RDG) provides a step-by-step example for installing DPF in Zero-Trust mode and HBN. It also includes practical demonstrations of performance optimization, validated using standard RDMA and TCP workloads.
As part of the reference implementation, open-source components outside the scope of DPF (e.g., MAAS, pfSense, Kubespray) are used to simulate a realistic customer deployment environment. The guide includes the full end-to-end deployment process, including:
-
Infrastructure provisioning
-
DPF deployment
-
DPU provisioning (redfish)
-
Service configuration and deployment
-
Service chaining.
This document extends the capabilities of the DPF-managed Kubernetes cluster described in the RDG for DPF Zero Trust (DPF-ZT) (referred to as the Baseline RDG) by deploying the NVIDIA DOCA HBN Service within the existing DPF deployment to achieve a comprehensive, accelerated infrastructure.
References
Solution Architecture
Key Components and Technologies
-
NVIDIA BlueField® Data Processing Unit (DPU)
The NVIDIA® BlueField® data processing unit (DPU) ignites unprecedented innovation for modern data centers and supercomputing clusters. With its robust compute power and integrated software-defined hardware accelerators for networking, storage, and security, BlueField creates a secure and accelerated infrastructure for any workload in any environment, ushering in a new era of accelerated computing and AI.
-
NVIDIA DOCA Software Framework
NVIDIA DOCA™ unlocks the potential of the NVIDIA® BlueField® networking platform. By harnessing the power of BlueField DPUs and SuperNICs, DOCA enables the rapid creation of applications and services that offload, accelerate, and isolate data center workloads. It lets developers create software-defined, cloud-native, DPU- and SuperNIC-accelerated services with zero-trust protection, addressing the performance and security demands of modern data centers.
-
NVIDIA ConnectX SmartNICs
10/25/40/50/100/200 and 400G Ethernet Network Adapters
The industry-leading NVIDIA® ConnectX® family of smart network interface cards (SmartNICs) offer advanced hardware offloads and accelerations.
NVIDIA Ethernet adapters enable the highest ROI and lowest Total Cost of Ownership for hyperscale, public and private clouds, storage, machine learning, AI, big data, and telco platforms.
-
NVIDIA LinkX Cables
The NVIDIA® LinkX® product family of cables and transceivers provides the industry’s most complete line of 10, 25, 40, 50, 100, 200, and 400GbE in Ethernet and 100, 200 and 400Gb/s InfiniBand products for Cloud, HPC, hyperscale, Enterprise, telco, storage and artificial intelligence, data center applications.
-
NVIDIA Spectrum Ethernet Switches
Flexible form-factors with 16 to 128 physical ports, supporting 1GbE through 400GbE speeds.
Based on a ground-breaking silicon technology optimized for performance and scalability, NVIDIA Spectrum switches are ideal for building high-performance, cost-effective, and efficient Cloud Data Center Networks, Ethernet Storage Fabric, and Deep Learning Interconnects.
NVIDIA combines the benefits of NVIDIA Spectrum™ switches, based on an industry-leading application-specific integrated circuit (ASIC) technology, with a wide variety of modern network operating system choices, including NVIDIA Cumulus® Linux, SONiC and NVIDIA Onyx®.
-
NVIDIA Cumulus Linux
NVIDIA® Cumulus® Linux is the industry's most innovative open network operating system that allows you to automate, customize, and scale your data center network like no other.
-
Kubernetes
Kubernetes is an open-source container orchestration platform for deployment automation, scaling, and management of containerized applications.
-
Kubespray
Kubespray is a composition of Ansible playbooks, inventory, provisioning tools, and domain knowledge for generic OS/Kubernetes clusters configuration management tasks and provides:-
A highly available cluster
-
Composable attributes
-
Support for most popular Linux distributions
-
Solution Design
Solution Logical Design
The logical design includes the following components:
-
1 x Hypervisor node (KVM-based) with ConnectX-7:
-
1 x Firewall VM
-
1 x Jump Node VM
-
1 x MaaS VM
-
3 x K8s Master VMs running all K8s management components
-
-
4 x Worker nodes (PCI Gen5), each with a 1 x BlueField-3 NIC
-
Single High-Speed (HS) switch
-
1 Gb Host Management network
HBN service Logical Design
As part of this RDG, we will:
-
Create two fully isolated logical networks per bare-metal workload server using a single physical function (PF0).
-
Connect each network through the HBN service to a dedicated VLAN/VNI, mapped to separate VRFs (RED or BLUE).
-
-
Route all workload traffic through the HBN service, with routing and isolation enforced inside the DPU.
-
Assign PF0 as the sole network interface for each bare-metal workload server, with no networking configuration on the host.
-
Demonstrate accelerated RDMA and TCP traffic between workload servers running on different bare-metal hosts within the same network (for example, RED ↔ RED).
-
Validate strict network isolation by confirming that traffic between workloads in different networks (RED vs BLUE) is not permitted.
Firewall Design
The pfSense firewall in this solution serves a dual purpose:
-
Firewall—provides an isolated environment for the DPF system, ensuring secure operations
-
Router—enables Internet access for the management network
Port-forwarding rules for SSH and RDP are configured on the firewall to route traffic to the jump node’s IP address in the host management network. From the jump node, administrators can manage and access various devices in the setup, as well as handle the deployment of the Kubernetes (K8s) cluster and DPF components.
The following diagram illustrates the firewall design used in this solution:
Software Stack Components
Make sure to use the exact same versions for the software stack as described above.
Bill of Materials
Deployment and Configuration
Node and Switch Definitions
These are the definitions and parameters used for deploying the demonstrated fabric:
|
Switches Ports Usage |
||
|---|---|---|
|
Hostname |
Rack ID |
Ports |
|
|
1 |
swp1-5 |
|
|
1 |
swp1-9 |
|
Hosts |
|||||
|---|---|---|---|---|---|
|
Rack |
Server Type |
Server Name |
Switch Port |
IP and NICs |
Default Gateway |
|
Rack1
|
Hypervisor Node |
|
mgmt-switch: hs-switch: |
lab-br (interface eno1): Trusted LAN IP mgmt-br (interface eno2): - hs-br (interface enp1s0): - |
Trusted LAN GW |
|
Rack1 |
Firewall (Virtual) |
|
- |
WAN (lab-br): Trusted LAN IP LAN (mgmt-br): 10.0.110.254/24 OPT1(hs-br): 10.0.123.254/22 |
Trusted LAN GW |
|
Rack1 |
Jump Node (Virtual) |
|
- |
enp1s0: 10.0.110.253/24 |
10.0.110.254 |
|
Rack1 |
MaaS (Virtual) |
|
- |
enp1s0: 10.0.110.252/24 |
10.0.110.254 |
|
Rack1 |
Master Node
|
|
- |
enp1s0: 10.0.110.1/24 |
10.0.110.254 |
|
Rack1 |
Master Node
|
|
- |
enp1s0: 10.0.110.2/24 |
10.0.110.254 |
|
Rack1 |
Master Node
|
|
- |
enp1s0: 10.0.110.3/24 |
10.0.110.254 |
|
Rack1
|
Worker Node |
|
mgmt-switch: hs-switch: |
dpubmc: 10.0.110.21/24 ens1f0np0/ens1f1np1: 10.0.120.0/22 |
10.0.110.254 |
|
Rack1
|
Worker Node |
|
mgmt-switch: hs-switch: |
dpubmc: 10.0.110.22/24 ens1f0np0/ens1f1np1: 10.0.120.0/22 |
10.0.110.254 |
|
Rack1
|
Worker Node |
|
mgmt-switch: hs-switch: |
dpubmc: 10.0.110.23/24 ens1f0np0/ens1f1np1: 10.0.120.0/22 |
10.0.110.254 |
|
Rack1
|
Worker Node |
|
mgmt-switch: hs-switch: |
dpubmc: 10.0.110.24/24 ens1f0np0/ens1f1np1: 10.0.120.0/22 |
10.0.110.254 |
Note: On BlueField-3, the DPU BMC and DPU OOB management interfaces share a single 1G out-of-band link via an internal bridge (oob_net0 ↔ tmfifo_net0 on BMC side). Both IPs (.201/.211) reside on the same L2 segment of the management network (10.0.110.0/24) and are reached via a single switch port (swpN). It is necessary to set several environment variables before running this command.
$ source manifests/00-env-vars/envvars.env
Note: Workers' high-speed PFs (ens1f0np0, ens1f1np1) connect to hs-switch via 200GbE. No persistent host-side IP in Zero-Trust baseline mode — DPU acts as a transparent NIC. Per-tenant IP assignment (e.g. 10.0.121.x for HBN RED, 10.0.122.x for HBN BLUE) is configured by downstream services (see DPF-ZT with HBN sub-page). Subnet 10.0.120.0/22 is reserved for the high-speed fabric.
Wiring
Hypervisor Node
Bare Metal Worker Node
Fabric Configuration
Updating Cumulus Linux
As a best practice, make sure to use the latest released Cumulus Linux NOS version.
For information on how to upgrade Cumulus Linux, refer to the Cumulus Linux User Guide.
Configuring the Cumulus Linux Switch
The SN3700 switch (hs-switch), is configured as follows:
The SN2201 switch (mgmt-switch) is configured as follows:
Host Configuration
Make sure that the BIOS settings on the worker node servers have SR-IOV enabled and that the servers are tuned for maximum performance.
Required:
-
SR-IOV: Enabled
-
VT-d / AMD-Vi (IOMMU): Enabled
-
Above 4G Decoding: Enabled (mandatory for PCIe BAR sizes on BlueField-3)
Performance-recommended:
-
CPU C-states: Disabled (or up to C1 only)
-
Hyper-Threading: Enabled
-
Memory speed: Maximum supported
-
Power profile: Performance / Maximum Performance
All worker nodes must have the same PCIe placement for the BlueField-3 NIC and must display the same interface name.
Make sure that you have DPU BMC and OOB MAC addresses.
No change from the Reference Deployment Guide (Baseline RDG) (Section "Deployment and Configuration", Subsection "Host Configuration").
Hypervisor Installation and Configuration
No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Hypervisor Installation and Configuration").
Prepare Infrastructure Servers
No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Prepare Infrastructure Servers") regarding Firewall VM, Jump VM, MaaS VM.
Firewall VM – Bare Metal Server Outside Conection
To provide outside connection from Bare Metal Host via High Speed network, open Firefox web browser and go to the pfSense web UI (http://10.0.110.254).
System:
-
Routing → Static Routing → Add → “Destination network”: 10.0.125.0/24, “Gateway”: Switch - 172.169.50.2 → , “Description”: To DPU DHCP → Click "Save"→ Under "Default Gateway" - "Default gateway IPv4" choose WAN_DHCP → Click "Save"
Note that the IP addresses from the Trusted LAN network under "Gateway" and "Monitor IP" are blurred.
Provision Master VMs Using MaaS
No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Provision Master VMs Using MaaS").
K8s Cluster Deployment and Configuration
The procedures for initial Kubernetes cluster deployment using Kubespray for the master nodes, and subsequent verification, remain unchanged from the Baseline RDG (Section "K8s Cluster Deployment and Configuration", Subsections: "Kubespray Deployment and Configuration", "Deploying Cluster Using Kubespray Ansible Playbook","K8s Deployment Verification".
DPF Installation
The DPF installation process (Operator, System components) largely follows the Baseline RDG.
Software Prerequisites and Required Variables
-
Start by installing the remaining software perquisites.
Jump Node Console
## Connect to master1 to copy helm client utility that was installed during kubespray deployment $ depuser@jump:~$ ssh master1 depuser@master1:~$ cp /usr/local/bin/helm /tmp/ ## In another tab depuser@jump:~$ scp master1:/tmp/helm /tmp/ depuser@jump:~$ sudo chown root:root /tmp/helm depuser@jump:~$ sudo mv /tmp/helm /usr/local/bin/ ## Verify that envsubst utility is installed depuser@jump:~$ which envsubst /usr/bin/envsubst -
Proceed to clone the doca-platform Git repository:
Jump Node Console
$ git clone https://github.com/NVIDIA/doca-platform.git -
Change directory to doca-platform and checkout to tag v26.4.0:
Jump Node Console
$ cd doca-platform/ $ git checkout v26.4.0 -
Change directory to doca-platform/docs/public/user-guides/zero-trust/use-cases/hbn from where all the commands will be run:
Jump Node Console
$ cd doca-platform/docs/public/user-guides/zero-trust/use-cases/hbn -
Change the BMC root's password.
In Zero Trust mode, provisioning DPUs requires authentication with Redfish.
In order to do that, you must set the same root password to access the BMC for all DPUs DPF is going to manage.For more information on how to set the BMC root password refer to BlueField DPU Administrator Quick Start Guide.Connect to the first DPU BMC over SSH to change the BMC root's password:
Jump Node Console
$ ssh root@10.0.110.201 root@10.0.110.201's password: <BMC Root Password. Default root/0penBmc. need to change first time to $BMC_ROOT_PASSWORD in the manifests/00-env-vars/envvars.env file> -
Modify the variables in
manifests/00-env-vars/envvars.envto fit your environment, then source the file: -
Replace the values for the variables in the following file with the values that fit your setup. Specifically, pay attention to
DPUCLUSTER_INTERFACE,BMC_ROOT_PASSWORD, andDPU's serial number.
To get aDPU's serial numberyou can use following command. Sample:
$ curl -k -u root:'BMC root password' https://10.0.110.201/redfish/v1/Systems/Bluefield | jq -r '.SerialNumber | ascii_downcase'
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
100 4970 100 4970 0 0 4211 0 0:00:01 0:00:01 --:--:-- 4211
mt2402xz0f7xmanifests/00-env-vars/envvars.env
Bash## IP Address for the Kubernetes API server of the target cluster on which DPF is installed. ## This should never include a scheme or a port. ## e.g. 10.10.10.10 export TARGETCLUSTER_API_SERVER_HOST=10.0.110.10 ## Port for the Kubernetes API server of the target cluster on which DPF is installed. ## e.g. 6443 export TARGETCLUSTER_API_SERVER_PORT=6443 ## Virtual IP used by the load balancer for the DPU Cluster. Must be a reserved IP from the management subnet and not ## allocated by DHCP. export DPUCLUSTER_VIP=10.0.110.200 ## Interface on which the DPUCluster load balancer will listen. Should be the management interface of the control plane node. export DPUCLUSTER_INTERFACE=enp1s0 ## The repository URL for the NVIDIA Helm chart registry. ## Usually this is the NVIDIA Helm NGC registry. For development purposes, this can be set to a different repository. export HELM_REGISTRY_REPO_URL=https://helm.ngc.nvidia.com/nvidia/doca ## The repository URL for the HBN container image. ## Usually this is the NVIDIA NGC registry. For development purposes, this can be set to a different repository. export HBN_NGC_IMAGE_URL=nvcr.io/nvidia/doca/doca_hbn ## The DPF REGISTRY is the Helm repository URL where the DPF Operator Chart resides. ## Usually this is the NVIDIA Helm NGC registry. For development purposes, this can be set to a different repository. export REGISTRY=https://helm.ngc.nvidia.com/nvidia/doca ## The DPF TAG is the version of the DPF components which will be deployed in this guide. export TAG=v26.4.0 ## URL to the BFB used in the `bfb.yaml` and linked by the DPUSet. export BFB_URL="https://content.mellanox.com/BlueField/BFBs/Ubuntu24.04/bf-bundle-3.4.0-92_26.04_ubuntu-24.04_64k_prod.bfb" ## IP_RANGE_START and IP_RANGE_END ## These define the IP range for DPU discovery via Redfish/BMC interfaces ## Example: If your DPUs have BMC IPs in range 10.0.110.201-224 ## export IP_RANGE_START=10.0.110.201 ## export IP_RANGE_END=10.0.110.224 ## Start of DPUDiscovery IpRange export IP_RANGE_START=10.0.110.201 ## End of DPUDiscovery IpRange export IP_RANGE_END=10.0.110.204 # The password used for DPU BMC root login, must be the same for all DPUs # For more information on how to set the BMC root password refer to BlueField DPU Administrator Quick Start Guide. export BMC_ROOT_PASSWORD=<set your BMC_ROOT_PASSWORD> ## Serial number of DPUs. If you have more than 2 DPUs, you will need to parameterize the system accordingly and expose ## additional variables. ## All serial numbers must be in lowercase. ## Serial number of DPU1 export DPU1_SERIAL=mt2402xz0f7x ## Serial number of DPU2 export DPU2_SERIAL=mt2402xz0f80 ## Serial number of DPU3 export DPU3_SERIAL=mt2402xz0f9n ## Serial number of DPU4 export DPU4_SERIAL=mt2402xz0f8g -
Export environment variables for the installation:
Jump Node Console
$ source manifests/00-env-vars/envvars.env
DPF Operator Installation
No change from the Baseline RDG (Section "DPF Installation", Subsection "DPF Operator Installation").
DPF System Installation
No change from the Baseline RDG (Section "DPF Installation", Subsection "DPF System Installation").
DPU Services Installation
HBN DPU Service Installation
This section focuses on provisioning NVIDIA®BlueField®-3 DPUs using DPF, installing the HBN DPU Service on those DPUs and enabling workload traffic to pass through HBN before leaving the DPU.
-
Export environment variables for the installation:
Jump Node Console
$ source manifests/00-env-vars/envvars.env -
Use the following YAML to define a
BFBresource that downloads the Bluefield Bitstream to a shared volume:--- apiVersion: provisioning.dpu.nvidia.com/v1alpha1 kind: BFB metadata: name: bf-bundle-$TAG namespace: dpf-operator-system spec: url: $BFB_URL -
Change the DPUFlavor using the following YAML:
-
Change the
dpudeployment.yamlfile to reference the DPUFlavor.Please notice that with default nodeEffect above, DPU provisioning workflow will be paused and wait for an external signal (annotation) in order to proceed, as demonstrated in upcoming steps.
To implement a fully automated process that won’t require user intervention, see customAction option. -
Change the rest of the configuration files.
As explained in the introduction, these files create service chains that connect two physical functions PF0 or PF0 to the outer fabric through HBN, providing EVPN VXLAN overlay, VNI based isolation, and ECMP redundancy across both DPU uplinks (p0 and p1).
These are the configuration files.-
HBN DPUServiceConfig and DPUServiceTemplate to deploy HBN workloads to the DPUs.
-
Physical Interfaces for physical ports on the DPU.
-
DPU Service IPAM objects to set up IP Address Management on the DPUCluster.
--- apiVersion: svc.dpu.nvidia.com/v1alpha1 kind: DPUServiceIPAM metadata: name: pool1 namespace: dpf-operator-system spec: ipv4Network: network: "10.0.121.0/24" gatewayIndex: 2 prefixSize: 29 # These preallocations are not necessary. We specify them so that the validation commands are straightforward. allocations: dpu-node-${DPU1_SERIAL}-${DPU1_SERIAL}: 10.0.121.0/29 dpu-node-${DPU2_SERIAL}-${DPU2_SERIAL}: 10.0.121.8/29 --- apiVersion: svc.dpu.nvidia.com/v1alpha1 kind: DPUServiceIPAM metadata: name: pool2 namespace: dpf-operator-system spec: ipv4Network: network: "10.0.122.0/24" gatewayIndex: 2 prefixSize: 29 allocations: dpu-node-${DPU3_SERIAL}-${DPU3_SERIAL}: 10.0.122.0/29 dpu-node-${DPU4_SERIAL}-${DPU4_SERIAL}: 10.0.122.8/29--- apiVersion: svc.dpu.nvidia.com/v1alpha1 kind: DPUServiceIPAM metadata: name: loopback namespace: dpf-operator-system spec: ipv4Network: network: "11.0.0.0/24" prefixSize: 32It is necessary to set several environment variables before running this command.
$ source manifests/00-env-vars/envvars.env
-
-
Apply all of the YAML files mentioned above using the following command:
Jump Node Console
$ cat manifests/03.1-dpudeployment-installation-pf/*.yaml | envsubst | kubectl apply -f -Jump Node Console
$ kubectl wait --for=condition=ApplicationsReconciled --namespace dpf-operator-system dpuservices --all dpuservice.svc.dpu.nvidia.com/cni-installer condition met dpuservice.svc.dpu.nvidia.com/doca-hbn-x92vr condition met dpuservice.svc.dpu.nvidia.com/flannel condition met dpuservice.svc.dpu.nvidia.com/kube-state-metrics-05f12f695b condition met dpuservice.svc.dpu.nvidia.com/kube-state-metrics-rbac condition met dpuservice.svc.dpu.nvidia.com/multus condition met dpuservice.svc.dpu.nvidia.com/node-problem-detector condition met dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-05f12f695b condition met dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-node condition met dpuservice.svc.dpu.nvidia.com/ovs-cni condition met dpuservice.svc.dpu.nvidia.com/servicechainset-controller-05f12f695b condition met dpuservice.svc.dpu.nvidia.com/servicechainset-rbac-and-crds condition met dpuservice.svc.dpu.nvidia.com/sfc-controller condition met dpuservice.svc.dpu.nvidia.com/sriov-device-plugin condition met $ kubectl wait --for=condition=DPUIPAMObjectReconciled --namespace dpf-operator-system dpuserviceipam --all dpuserviceipam.svc.dpu.nvidia.com/loopback condition met dpuserviceipam.svc.dpu.nvidia.com/pool1 condition met $ kubectl wait --for=condition=ServiceInterfaceSetReconciled --namespace dpf-operator-system dpuserviceinterface --all dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-p0-if-f9hzk condition met dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-p1-if-2ld7q condition met dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-pf0hpf-if-gt8zw condition met dpuserviceinterface.svc.dpu.nvidia.com/p0 condition met dpuserviceinterface.svc.dpu.nvidia.com/p1 condition met dpuserviceinterface.svc.dpu.nvidia.com/pf0hpf condition met $ kubectl wait --for=condition=ServiceChainSetReconciled --namespace dpf-operator-system dpuservicechain --all dpuservicechain.svc.dpu.nvidia.com/hbn-only-cjpt5 condition met -
To follow the progress of DPU provisioning, run the following command to check its current phase:Jump Node Console
$ watch -n10 "kubectl describe dpu -n dpf-operator-system | grep 'Node Name\|Type\|Last\|Phase'"
-
Wait for the NodeEffect stage (at this point the provisioning is paused, waintig for external signal).
Run following command on all/specific DPU nodemaintanace object/s to proceed with provisioning:Jump Node Console
$ kubectl annotate dpunodemaintenances -n dpf-operator-system --all \ provisioning.dpu.nvidia.com/wait-for-external-nodeeffect=false \ maintenance.dpu.nvidia.com/wait-for-external-nodeeffect=false \ --overwrite -
To follow the progress of DPU provisioning, run the following command to check its current phase:
Jump Node Console
$ watch -n10 "kubectl -n dpf-operator-system get dpu,dpuset,dpudeployment,dpuservice,dpuserviceconfigurations,dpuservicetemplates" -
Wait for the Rebooted stage and then Power Cycle the bare-metal host manual.
After the DPU is up, run following command for each DPU worker:Jump Node Console
$ kubectl annotate dpunode -n dpf-operator-system --all provisioning.dpu.nvidia.com/dpunode-external-reboot-required- -
At this point, the DPU workers should be added to the cluster. As they being added to the cluster, the DPUs are provisioned.
Jump Node Console
$ watch -n 2 'kubectl -n dpf-operator-system get dpu,dpuset,dpudeployment,dpuservice,dpuserviceconfigurations,dpuservicetemplates' NAME READY OPERATIONAL PHASE AGE dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f7x True True Ready 114m dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f80 True True Ready 114m dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f8g True True Ready 114m dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f9n True True Ready 114m NAME READY AGE dpuset.provisioning.dpu.nvidia.com/hbn-only-dpuset1 True 114m NAME READY PHASE AGE dpudeployment.svc.dpu.nvidia.com/hbn-only True Success 115m NAME READY PHASE AGE dpuservice.svc.dpu.nvidia.com/cni-installer True Success 116m dpuservice.svc.dpu.nvidia.com/doca-hbn-x92vr True Success 114m dpuservice.svc.dpu.nvidia.com/flannel True Success 116m dpuservice.svc.dpu.nvidia.com/kube-state-metrics-05f12f695b True Success 116m dpuservice.svc.dpu.nvidia.com/kube-state-metrics-rbac True Success 116m dpuservice.svc.dpu.nvidia.com/multus True Success 116m dpuservice.svc.dpu.nvidia.com/node-problem-detector True Success 116m dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-05f12f695b True Success 116m dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-node True Success 116m dpuservice.svc.dpu.nvidia.com/ovs-cni True Success 116m dpuservice.svc.dpu.nvidia.com/servicechainset-controller-05f12f695b True Success 116m dpuservice.svc.dpu.nvidia.com/servicechainset-rbac-and-crds True Success 116m dpuservice.svc.dpu.nvidia.com/sfc-controller True Success 116m dpuservice.svc.dpu.nvidia.com/sriov-device-plugin True Success 116m NAME AGE dpuserviceconfiguration.svc.dpu.nvidia.com/doca-hbn 115m NAME AGE dpuservicetemplate.svc.dpu.nvidia.com/doca-hbn 115m -
Finally, validate that all the different DPU-related objects are now in the Ready state:
Jump Node Console
$ kubectl get secrets -n dpu-cplane-tenant1 dpu-cplane-tenant1-admin-kubeconfig -o json | jq -r '.data["admin.conf"]' | base64 --decode > /home/depuser/dpu-cluster.config $ echo "alias ki='KUBECONFIG=/home/depuser/dpu-cluster.config kubectl'" >> ~/.bashrc $ echo 'alias dpfctl="kubectl -n dpf-operator-system exec deploy/dpf-operator-controller-manager -- /dpfctl "' >> ~/.bashrc $ dpfctl describe dpudeployments NAME NAMESPACE STATUS REASON SINCE MESSAGE DPFOperatorConfig/dpfoperatorconfig dpf-operator-system Ready: True Success 3m3s └─DPUDeployments └─DPUDeployment/hbn dpf-operator-system Ready: True Success 22s ├─DPUServiceChains │ └─DPUServiceChain/hbn-wd7fs dpf-operator-system Ready: True Success 65s ├─DPUServiceInterfaces │ └─3 DPUServiceInterfaces... dpf-operator-system Ready: True Success 70s See doca-hbn-p0-if-749n9, doca-hbn-p1-if-fn8w5, doca-hbn-pf0hpf-if-9s8c6 ├─DPUSets │ └─DPUSet/hbn-dpuset1 dpf-operator-system Ready: True Success 71s │ ├─BFB/bf-bundle-v26.4.0 dpf-operator-system Ready: True Ready 39m File: 3.4.0-92_26.04_ubuntu-24.04_64k_prod.bfb, DOCA: 3.4.0 │ ├─DPUNodes │ │ └─4 DPUNodes... dpf-operator-system Ready: True Ready 98s See dpu-node-mt2402xz0f7x, dpu-node-mt2402xz0f80, dpu-node-mt2402xz0f8g, dpu-node-mt2402xz0f9n │ └─DPUs │ └─4 DPUs... dpf-operator-system Ready: True DPUReady 98s See dpu-node-mt2402xz0f7x-mt2402xz0f7x, dpu-node-mt2402xz0f80-mt2402xz0f80, │ dpu-node-mt2402xz0f8g-mt2402xz0f8g, dpu-node-mt2402xz0f9n-mt2402xz0f9n └─Services ├─DPUServiceTemplates │ └─DPUServiceTemplate/doca-hbn dpf-operator-system Ready: True Success 39m └─DPUServices └─1 DPUServices... dpf-operator-system Ready: True Success 50s See doca-hbn-jxkxw $ ki get node -A NAME STATUS ROLES AGE VERSION dpu-node-mt2402xz0f7x-mt2402xz0f7x Ready <none> 5m18s v1.34.8 dpu-node-mt2402xz0f80-mt2402xz0f80 Ready <none> 6m12s v1.34.8 dpu-node-mt2402xz0f8g-mt2402xz0f8g Ready <none> 6m14s v1.34.8 dpu-node-mt2402xz0f9n-mt2402xz0f9n Ready <none> 6m22s v1.34.8 $ kubectl get dpu -A NAMESPACE NAME READY PHASE AGE dpf-operator-system dpu-node-mt2402xz0f7x-mt2402xz0f7x True Ready 36m dpf-operator-system dpu-node-mt2402xz0f80-mt2402xz0f80 True Ready 36m dpf-operator-system dpu-node-mt2402xz0f8g-mt2402xz0f8g True Ready 36m dpf-operator-system dpu-node-mt2402xz0f9n-mt2402xz0f9n True Ready 36m $ kubectl wait --for=condition=ready --namespace dpf-operator-system dpu --all dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f7x condition met dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f80 condition met dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f8g-mt2402xz0f8g condition met dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f9n-mt2402xz0f9n condition met $ ki get pods -A -o wide NAMESPACE NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES dpf-operator-system dpu-cplane-tenant1-cni-installer-89kn4 1/1 Running 0 6m50s 10.244.2.3 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system dpu-cplane-tenant1-cni-installer-s8h4z 1/1 Running 0 7m1s 10.244.0.5 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-cni-installer-wb29j 1/1 Running 0 5m57s 10.244.3.2 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system dpu-cplane-tenant1-cni-installer-zhzqh 1/1 Running 0 6m53s 10.244.1.4 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-5sbzs 2/2 Running 0 2m54s 10.244.0.6 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-ftnpn 2/2 Running 0 2m54s 10.244.1.5 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq 2/2 Running 0 3m21s 10.244.3.4 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb 2/2 Running 0 2m54s 10.244.2.4 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system dpu-cplane-tenant1-nvidia-k8s-ipam-controller-5c77854fcc-grchr 1/1 Running 0 127m 10.244.0.3 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-krgzw 1/1 Running 0 6m53s 10.244.1.2 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-pr85m 1/1 Running 0 5m57s 10.244.3.3 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-x4lfs 1/1 Running 0 7m1s 10.244.0.2 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-zlzvf 1/1 Running 0 6m50s 10.244.2.2 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system dpu-cplane-tenant1-ovs-cni-arm64-bpljq 1/1 Running 0 7m1s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-ovs-cni-arm64-gls6h 1/1 Running 0 6m50s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system dpu-cplane-tenant1-ovs-cni-arm64-j8wr4 1/1 Running 0 5m57s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system dpu-cplane-tenant1-ovs-cni-arm64-kbrrn 1/1 Running 0 6m53s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system dpu-cplane-tenant1-sfc-controller-node-ds-vmfq4 1/1 Running 0 5m57s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system dpu-cplane-tenant1-sfc-controller-node-ds-x45nl 1/1 Running 0 6m53s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system dpu-cplane-tenant1-sfc-controller-node-ds-xskh9 1/1 Running 0 7m1s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system dpu-cplane-tenant1-sfc-controller-node-ds-zfmt5 1/1 Running 1 (5m46s ago) 6m50s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system kube-flannel-ds-2shh7 1/1 Running 0 7m2s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system kube-flannel-ds-42mlq 1/1 Running 0 6m54s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system kube-flannel-ds-m7xgt 1/1 Running 0 5m58s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system kube-flannel-ds-vd574 1/1 Running 0 6m52s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system kube-multus-ds-d5kb4 1/1 Running 0 6m53s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system kube-multus-ds-gnv88 1/1 Running 0 6m50s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system kube-multus-ds-l66tm 1/1 Running 0 7m1s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system kube-multus-ds-mh4cj 1/1 Running 0 5m57s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpf-operator-system kube-sriov-device-plugin-64c29 1/1 Running 0 7m1s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpf-operator-system kube-sriov-device-plugin-6js9j 1/1 Running 0 6m50s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> dpf-operator-system kube-sriov-device-plugin-g5gkx 1/1 Running 0 6m53s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpf-operator-system kube-sriov-device-plugin-lk4z7 1/1 Running 0 5m57s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> kube-system coredns-66bc5c9577-gqn8d 1/1 Running 0 127m 10.244.0.4 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> kube-system coredns-66bc5c9577-p2xnm 1/1 Running 0 127m 10.244.1.3 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> kube-system kube-proxy-64865 1/1 Running 0 5m58s 10.0.110.211 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> kube-system kube-proxy-hvjjp 1/1 Running 0 6m52s 10.0.110.212 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> kube-system kube-proxy-qfbwh 1/1 Running 0 6m54s 10.0.110.214 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> kube-system kube-proxy-w9gg4 1/1 Running 0 7m2s 10.0.110.213 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none>Congratulations! The DPF system with the HBN service has been successfully installed.
Zero-Trust Mode Checking
Here's a step-by-step procedure to check the Zero-Trust Mode on your NVIDIA BlueField DPU from the host server, including the installation of the Mellanox Firmware Tools (MFT).
Ubuntu 24.04 was installed on the servers.
-
Navigate to the NVIDIA Downloads Site: Open your web browser and go to the official NVIDIA Mellanox software downloads page.
-
Select the Latest Version for your OS:
-
Transfer and Extract MFT Tools on the Worker 1 BareMetal Host.
First Pod Console
root@worker1:~# tar -xvzf /tmp/mft-4.33.0-169-x86_64-deb.tgz -
Navigate into the Extracted Directory.
First Pod Console
root@worker1:~# cd mft-4.33.0-169-x86_64-deb/ -
Run following commands.
First Pod Console
root@worker1:~# apt-get install gcc make dkms root@worker1:~# ./install.sh -
Start MST (Mellanox Software Tools) Service and Identify DPU Device Name.
First Pod Console
root@worker1:~# mst start Starting MST (Mellanox Software Tools) driver set Loading MST PCI module - Success Loading MST PCI configuration module - Success Create devices Unloading MST PCI module (unused) - Success root@worker1:~# mst status MST modules: ------------ MST PCI module is not loaded MST PCI configuration module loaded MST devices: ------------ /dev/mst/mt41692_pciconf0 - PCI configuration cycles access. domain:bus:dev.fn=0000:2b:00.0 addr.reg=88 data.reg=92 cr_bar.gw_offset=-1 Chip revision is: 01 -
Perform Zero-Trust Checking.
First Pod Console
root@worker1:~# mlxprivhost -d 2b:00.0 q Host configurations ------------------- level : RESTRICTED Port functions status: ----------------------- disable_rshim : TRUE disable_tracer : TRUE disable_port_owner : TRUE disable_counter_rd : TRUE #Expected Zero-Trust Output.This is the most definitive confirmation.
level : RESTRICTEDmeans the host is in Zero-Trust Mode, and theTRUEflags confirm individual security restrictions are active. -
Verify the host cannot reset DPU firmware:
root@worker1:~# sudo mlxfwreset -d 2b:00.0 -y -l 3 resetExpected output on a Zero-Trust host:
-E- Failed to send Register MFRL: Register access Method not supported (264).The MFRL (Master Firmware Reset Lock) register is access-gated by RESTRICTED mode. Failure here confirms the host cannot trigger a DPU firmware reset.
-
Check Firmware Access with
mlxfwmanager:First Pod Console
root@worker1:~# mlxfwmanager -d 2b:00.0 --query Querying Mellanox devices firmware ... Device #1: ---------- Device Type: BlueField3 Part Number: -- Description: PSID: PCI Device Name: 2b:00.0 Base MAC: N/A Versions: Current Available FW -- Status: Failed to open deviceThe behaviour of
mlxfwmanager --querydepends on the MFT version installed:-
MFT < 4.33 returns
Status: Failed to open device— the host cannot read inventory at all in Zero-Trust mode. -
MFT 4.33+ returns inventory data (FW/PXE/UEFI versions, PSID, MAC) with
Status: No matching image found. The BMC fulfils the read from cached inventory; this is not a Zero-Trust failure. TheAvailable: N/Aand theNo matching image foundstatus both confirm the host has no path to upload firmware. Write-side operations (--update,mlxfwreset) remain blocked.
Either output is consistent with Zero-Trust mode. Use the
mlxprivhost -d <dev> pwrite probe in step 11 (below) for the definitive verification.
-
-
Check Device Configuration with
mlxconfig:First Pod Console
root@worker1:~# mlxconfig -d 2b:00.0 q Device #1: ---------- Device type: BlueField3 Name: 900-9D3B6-00CV-A_Ax Description: NVIDIA BlueField-3 B3220 P-Series FHHL DPU; 200GbE (default mode) / NDR200 IB; Dual-port QSFP112; PCIe Gen5.0 x16 with x16 PCIe extension option; 16 Arm cores; 32GB on-board DDR; integrated BMC; Crypto Enabled Device: 2b:00.0 Configurations: Next Boot ... ALLOW_RD_COUNTERS True(1) # No RO, but restricted by mlxprivhost ... PORT_OWNER True(1) # No RO, but restricted by mlxprivhost ... TRACER_ENABLE True(1) # No RO, but restricted by mlxprivhostMost configuration parameters are prefixed with
RO(Read-Only) — the host literally cannot change them, by design. A small number of parameters related to host-side control (PORT_OWNER,ALLOW_RD_COUNTERS,TRACER_ENABLE) are not markedRO, andmlxconfig setwill appear to succeed against them on a Zero-Trust host:
root@worker1:~# sudo mlxconfig -d 2b:00.0 set TRACER_ENABLE=0
...
Apply new Configuration? (y/n) [n] : y
Applying... Done!
-I- Please power cycle machine to load new configurations.However, the change is functionally a no-op. At runtime, the
mlxprivhostRESTRICTED layer (visible in step 7'sdisable_tracer: TRUE) overrides whatevermlxconfigsays. On the next DPU re-provision, the DPF operator'sDPUFlavor.spec.nvconfigre-applies the canonical settings anyway. The fact thatmlxconfig setreturns "Applying... Done!" does not indicate Zero-Trust is off.
-
Step 11. Verify the host cannot escape Zero-Trust mode (write probe with
mlxprivhost):
root@worker1:~# sudo mlxprivhost -d 2b:00.0 pExpected output on a Zero-Trust host:
-E- Operation is not permitted (refer to the DPU user manual)The
pargument asks the host to switch its privilege level back to PRIVILEGED. On a Zero-Trust DPU this must fail — that's the whole point of the RESTRICTED level. If this command succeeds and the level flips, the DPU is not in Zero-Trust mode and the DPUFlavor (spec.dpuMode) needs to be re-checked.This is the only command in this section that is fundamentally write-side and therefore unaffected by MFT-version changes to read-side behaviour.
-
Check Low-Level Hardware Access with
ethtool:First Pod Console
root@worker1:~# ethtool -d ens1f0np0 Cannot get register dump: Operation not supportedThis confirms the DPU is preventing deep, low-level hardware access from the host, aligning with Zero-Trust's isolation goals.
Conclusion
The host is operating in Zero-Trust Mode when ALL of the following are true:
|
# |
Test |
Expected on Zero-Trust host |
|---|---|---|
|
1 |
|
|
|
2 |
|
|
|
3 |
|
majority of parameters prefixed |
|
4 |
|
|
|
5 |
|
|
|
6 |
|
|
Tests 1–4 are read-side checks and confirm the configuration is in place. Test 5 (mlxprivhost p) is the authoritative proof: a Zero-Trust host cannot remove its own restrictions. Test 6 reinforces this for firmware-reset specifically.
This means the host has significantly restricted privileges and cannot perform sensitive operations on the DPU, ensuring its security and isolation.
Infrastructure Bandwidth & Latency Validation
Verify the deployment and confirm that the DPU system achieves link-speed performance and low latency by running various tests:
-
Iperf TCP—for bandwidth measurements
-
RDMA—for bandwidth and latency measurements
-
Network isolation
Each test is described in detail. At the end of each test, the achieved performance is displayed.
Notes
Make sure that the servers are tuned for maximum performance (not covered in this document).
Performance and Isolation Tests
Now that the test deployment is running, perform bandwidth and latency performance tests between two bare-metal workload servers.
Ubuntu 24.04 was installed on the servers.
-
Before running the tests, check the Gateway address on each HBN pod:
Jump Node Console
$ ki -n dpf-operator-system get pod -o wide | grep doca-hbn dpu-cplane-tenant1-doca-hbn-jxkxw-ds-5sbzs 2/2 Running 0 15m 10.244.0.6 dpu-node-mt2402xz0f9n-mt2402xz0f9n <none> <none> dpu-cplane-tenant1-doca-hbn-jxkxw-ds-ftnpn 2/2 Running 0 15m 10.244.1.5 dpu-node-mt2402xz0f8g-mt2402xz0f8g <none> <none> dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq 2/2 Running 0 16m 10.244.3.4 dpu-node-mt2402xz0f7x-mt2402xz0f7x <none> <none> dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb 2/2 Running 0 15m 10.244.2.4 dpu-node-mt2402xz0f80-mt2402xz0f80 <none> <none> $ ki exec -it -n dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq -- bash Defaulted container "doca-hbn" out of: doca-hbn, hbn-sidecar, hbn-init (init) root@dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq:/tmp# ip a s ... 9: vlan11@br_default: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 9216 qdisc noqueue master RED state UP group default qlen 1000 link/ether 0a:ff:4e:3e:99:24 brd ff:ff:ff:ff:ff:ff inet 10.0.121.2/29 scope global vlan11 valid_lft forever preferred_lft forever inet6 fe80::8ff:4eff:fe3e:9924/64 scope link valid_lft forever preferred_lft forever ... $ exit $ ki exec -it -n dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb -- bash Defaulted container "doca-hbn" out of: doca-hbn, hbn-sidecar, hbn-init (init) root@dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb:/tmp# ip a s ... 9: vlan11@br_default: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 9216 qdisc noqueue master RED state UP group default qlen 1000 link/ether 0e:7d:99:41:2e:11 brd ff:ff:ff:ff:ff:ff inet 10.0.121.10/29 scope global vlan11 valid_lft forever preferred_lft forever inet6 fe80::c7d:99ff:fe41:2e11/64 scope link valid_lft forever preferred_lft forever ... $ exit
-
Connect to a first Workload Server console, install iperf, perftest, check DPU Hight Speed Interfaces, set route to ethernet and identify the relevant RDMA device:
First Pod Console
root@worker1:~# apt install iperf3 root@worker1:~# apt install perftest root@worker1:~# ip a s ... 6: ens1f0np0: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000 link/ether 58:a2:e1:73:69:e6 brd ff:ff:ff:ff:ff:ff altname enp43s0f0np0 ... root@worker1:~# ip route add 10.0.123.0/22 via 10.0.121.2 depuser@worker1:~$ ping 8.8.8.8 PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data. 64 bytes from 8.8.8.8: icmp_seq=1 ttl=117 time=5.35 ms 64 bytes from 8.8.8.8: icmp_seq=2 ttl=117 time=5.10 ms 64 bytes from 8.8.8.8: icmp_seq=3 ttl=117 time=5.15 ms root@worker1:~# rdma link | grep ens1f0np0 link mlx5_0/1 state DOWN physical_state DISABLED netdev ens1f0np0 -
Configure the
ens1f0np0interface on Ubuntu 24.04 usingiproute2.
Configuration OverviewInterface
IP Address
Default Gateway
ens1f0np0
10.0.121.1/29
10.0.121.2/29
First Pod Console
# Bring up physical interfaces root@worker1:~# ip link set dev ens1f0np0 up # Assign IP addresses root@worker1:~# ip addr add 10.0.121.1/29 dev ens1f0np0 # Set default route root@worker1:~# ip route add default via 10.0.121.2 dev ens1f0np0
-
Using another console window, reconnect to the jump node and connect to a second Workload Server.
From within the servers, install iperf, perftest, check DPU Hight Speed Interfaces, set route to ethernet and identify the relevant RDMA device:First Pod Console
root@worker2:~# apt install iperf3 root@worker2:~# apt install perftest root@worker2:~# ip a s ... 6: ens1f0np0: <BROADCAST,MULTICAST> mtu 9000 qdisc noop state DOWN group default qlen 1000 link/ether 58:a2:e1:73:6a:58 brd ff:ff:ff:ff:ff:ff altname enp43s0f0np0 ... root@worker2:~# ip route add 10.0.123.0/22 via 10.0.121.10 depuser@worker2:~$ ping 8.8.8.8 PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data. 64 bytes from 8.8.8.8: icmp_seq=1 ttl=117 time=5.35 ms 64 bytes from 8.8.8.8: icmp_seq=2 ttl=117 time=5.10 ms 64 bytes from 8.8.8.8: icmp_seq=3 ttl=117 time=5.15 ms root@worker2:~# rdma link | grep ens1f0np0 link mlx5_0/1 state DOWN physical_state DISABLED netdev ens1f0np0 -
Configure the
ens1f0np0interface on Ubuntu 24.04 usingiproute2.Configuration Overview
Interface
IP Address
Default Gateway
ens1f0np0
10.0.121.9/29
10.0.121.10/29
First Pod Console
# Bring up physical interfaces
root@worker2:~# ip link set dev ens1f0np0 up
# Assign IP addresses
root@worker2:~# ip addr add 10.0.121.9/29 dev ens1f0np0
# Set default route
root@worker2:~# ip route add default via 10.0.121.10 dev ens1f0np0
iPerf TCP Bandwidth Test
Move back to the first server console.
-
Host NIC tuning (required for line-rate TCP on 200GbE) .
Without tuning, TCP throughput on Ubuntu 24.04 with stock NIC settings caps at ~167 Gbps. Apply these on every worker (sender + receiver) that will run iperf3:
First Pod Console (Sample)
# 1) Bump RX/TX ring buffers (default 1024 → max 8192)
root@worker1:~# ethtool -G ens1f0np0 rx 8192 tx 8192
# 2) Pin all mlx5_comp IRQs to the NIC's NUMA node
PCI=$(ethtool -i ens1f0np0 | awk '/bus-info/{print $2}')
NUMA=$(cat /sys/class/net/ens1f0np0/device/numa_node)
NUMA_CPUS=$(cat /sys/devices/system/node/node${NUMA}/cpulist | tr ',' '\n' | \
awk -F- '{if($2)print $2-$1+1; else print 1}' | paste -sd+ | bc)
for irq in $(awk -v p="$PCI" '$0~"mlx5_comp.*@pci:"p{gsub(":",""); print $1}' /proc/interrupts); do
q=$(awk -v i="$irq:" '$1==i{print $NF}' /proc/interrupts | grep -oP 'comp\K[0-9]+')
echo $((q % NUMA_CPUS)) | sudo tee /proc/irq/$irq/smp_affinity_list >/dev/null
done
# Expected gain on a 200 GbE BlueField-3 pair: 167 Gbps → 185 Gbps (8 TCP streams, retransmits drop from ~240K to ~120K).
# Persist via systemd unit if needed.
-
Start the
iperf3server side:First BM Server Console
root@worker1:~# iperf3 -s -p 5201 ----------------------------------------------------------- Server listening on 5201 (test #1) ------------------------------------------------------------ -
Move to the second server console.
Start theiperfclient side:Second BM Server Console
depuser@worker2:~$ iperf3 -c 10.0.121.1 -p 5201 -t 30 -P 8 -i 5 Connecting to host 10.0.121.1, port 5201 ... [ ID] Interval Transfer Bitrate Retr Cwnd [ 5] 0.00-5.01 sec 12.1 GBytes 20.7 Gbits/sec 1462 1.31 MBytes [ 7] 0.00-5.01 sec 14.3 GBytes 24.5 Gbits/sec 1620 883 KBytes [ 9] 0.00-5.01 sec 11.4 GBytes 19.6 Gbits/sec 1860 1.34 MBytes [ 11] 0.00-5.01 sec 12.7 GBytes 21.7 Gbits/sec 2204 743 KBytes [ 13] 0.00-5.01 sec 15.0 GBytes 25.8 Gbits/sec 2167 1.22 MBytes [ 15] 0.00-5.01 sec 12.9 GBytes 22.2 Gbits/sec 2251 926 KBytes [ 17] 0.00-5.01 sec 12.8 GBytes 22.0 Gbits/sec 1467 856 KBytes [ 19] 0.00-5.01 sec 16.7 GBytes 28.7 Gbits/sec 2032 1.37 MBytes [SUM] 0.00-5.01 sec 108 GBytes 185 Gbits/sec 15063 - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 5.01-10.00 sec 11.5 GBytes 19.7 Gbits/sec 2186 935 KBytes [ 7] 5.01-10.00 sec 13.6 GBytes 23.4 Gbits/sec 2263 900 KBytes [ 9] 5.01-10.00 sec 14.3 GBytes 24.5 Gbits/sec 3120 1.04 MBytes [ 11] 5.01-10.00 sec 12.9 GBytes 22.2 Gbits/sec 3150 1.05 MBytes [ 13] 5.01-10.00 sec 13.2 GBytes 22.7 Gbits/sec 2645 690 KBytes [ 15] 5.01-10.00 sec 13.8 GBytes 23.6 Gbits/sec 3126 1.91 MBytes [ 17] 5.01-10.01 sec 12.5 GBytes 21.4 Gbits/sec 2161 1.04 MBytes [ 19] 5.01-10.01 sec 15.6 GBytes 26.8 Gbits/sec 2769 1.28 MBytes [SUM] 5.01-10.00 sec 107 GBytes 184 Gbits/sec 21420 - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 10.00-15.01 sec 12.2 GBytes 21.0 Gbits/sec 1739 1.11 MBytes [ 7] 10.00-15.01 sec 14.1 GBytes 24.2 Gbits/sec 1835 1.89 MBytes [ 9] 10.00-15.01 sec 10.8 GBytes 18.5 Gbits/sec 2206 1.11 MBytes [ 11] 10.00-15.01 sec 14.5 GBytes 24.9 Gbits/sec 2888 1.37 MBytes [ 13] 10.00-15.01 sec 14.0 GBytes 24.0 Gbits/sec 2528 839 KBytes [ 15] 10.00-15.01 sec 15.0 GBytes 25.7 Gbits/sec 2798 1.80 MBytes [ 17] 10.01-15.01 sec 13.8 GBytes 23.7 Gbits/sec 2280 1.02 MBytes [ 19] 10.01-15.01 sec 13.6 GBytes 23.4 Gbits/sec 2205 1.25 MBytes [SUM] 10.00-15.01 sec 108 GBytes 185 Gbits/sec 18479 - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 15.01-20.01 sec 11.6 GBytes 19.9 Gbits/sec 2119 1.43 MBytes [ 7] 15.01-20.01 sec 14.2 GBytes 24.4 Gbits/sec 2457 1.01 MBytes [ 9] 15.01-20.01 sec 13.3 GBytes 22.8 Gbits/sec 2791 2.13 MBytes [ 11] 15.01-20.01 sec 13.6 GBytes 23.4 Gbits/sec 3383 1.65 MBytes [ 13] 15.01-20.01 sec 12.7 GBytes 21.9 Gbits/sec 2594 1.49 MBytes [ 15] 15.01-20.01 sec 14.3 GBytes 24.6 Gbits/sec 3274 2.30 MBytes [ 17] 15.01-20.01 sec 13.1 GBytes 22.6 Gbits/sec 2387 1.74 MBytes [ 19] 15.01-20.01 sec 14.8 GBytes 25.5 Gbits/sec 2714 1.20 MBytes [SUM] 15.01-20.01 sec 108 GBytes 185 Gbits/sec 21719 - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 20.01-25.01 sec 12.3 GBytes 21.2 Gbits/sec 2191 821 KBytes [ 7] 20.01-25.01 sec 15.2 GBytes 26.1 Gbits/sec 2159 1.74 MBytes [ 9] 20.01-25.01 sec 12.3 GBytes 21.2 Gbits/sec 2628 1.01 MBytes [ 11] 20.01-25.01 sec 14.5 GBytes 24.9 Gbits/sec 3484 1.14 MBytes [ 13] 20.01-25.01 sec 13.2 GBytes 22.6 Gbits/sec 2564 1.42 MBytes [ 15] 20.01-25.01 sec 13.5 GBytes 23.2 Gbits/sec 3048 1.07 MBytes [ 17] 20.01-25.01 sec 13.3 GBytes 22.8 Gbits/sec 2288 1.53 MBytes [ 19] 20.01-25.01 sec 13.4 GBytes 23.1 Gbits/sec 2559 1.23 MBytes [SUM] 20.01-25.01 sec 108 GBytes 185 Gbits/sec 20921 - - - - - - - - - - - - - - - - - - - - - - - - - [ 5] 25.01-30.01 sec 10.6 GBytes 18.2 Gbits/sec 1706 734 KBytes [ 7] 25.01-30.01 sec 15.8 GBytes 27.2 Gbits/sec 2563 1022 KBytes [ 9] 25.01-30.01 sec 10.5 GBytes 18.0 Gbits/sec 2482 874 KBytes [ 11] 25.01-30.01 sec 14.4 GBytes 24.8 Gbits/sec 3575 1.58 MBytes [ 13] 25.01-30.01 sec 15.6 GBytes 26.8 Gbits/sec 2679 1.75 MBytes [ 15] 25.01-30.01 sec 13.6 GBytes 23.3 Gbits/sec 3095 1022 KBytes [ 17] 25.01-30.01 sec 13.0 GBytes 22.3 Gbits/sec 2208 1.32 MBytes [ 19] 25.01-30.01 sec 15.4 GBytes 26.5 Gbits/sec 2771 2.59 MBytes [SUM] 25.01-30.01 sec 109 GBytes 187 Gbits/sec 21079 - - - - - - - - - - - - - - - - - - - - - - - - - [ ID] Interval Transfer Bitrate Retr [ 5] 0.00-30.01 sec 70.3 GBytes 20.1 Gbits/sec 11403 sender [ 5] 0.00-30.01 sec 70.3 GBytes 20.1 Gbits/sec receiver [ 7] 0.00-30.01 sec 87.2 GBytes 25.0 Gbits/sec 12897 sender [ 7] 0.00-30.01 sec 87.2 GBytes 25.0 Gbits/sec receiver [ 9] 0.00-30.01 sec 72.5 GBytes 20.8 Gbits/sec 15087 sender [ 9] 0.00-30.01 sec 72.5 GBytes 20.8 Gbits/sec receiver [ 11] 0.00-30.01 sec 82.6 GBytes 23.6 Gbits/sec 18684 sender [ 11] 0.00-30.01 sec 82.6 GBytes 23.6 Gbits/sec receiver [ 13] 0.00-30.01 sec 83.7 GBytes 24.0 Gbits/sec 15177 sender [ 13] 0.00-30.01 sec 83.7 GBytes 24.0 Gbits/sec receiver [ 15] 0.00-30.01 sec 83.1 GBytes 23.8 Gbits/sec 17592 sender [ 15] 0.00-30.01 sec 83.1 GBytes 23.8 Gbits/sec receiver [ 17] 0.00-30.01 sec 78.5 GBytes 22.5 Gbits/sec 12791 sender [ 17] 0.00-30.01 sec 78.5 GBytes 22.5 Gbits/sec receiver [ 19] 0.00-30.01 sec 89.6 GBytes 25.7 Gbits/sec 15050 sender [ 19] 0.00-30.01 sec 89.6 GBytes 25.7 Gbits/sec receiver [SUM] 0.00-30.01 sec 648 GBytes 185 Gbits/sec 118681 sender [SUM] 0.00-30.01 sec 647 GBytes 185 Gbits/sec receiver iperf Done.
RoCE Latency Test
Return to the first server console.
-
Start the
ib_read_latserver side:First BM Server Console
root@worker1:~# ib_read_lat -F -n 20000 -d mlx5_0 ************************************ * Waiting for client to connect... * ************************************ -
Move to the second server console.
Start theib_read_latclient side:
Second BM Server Console
root@worker2:~# ib_read_lat -F -n 20000 -d mlx5_0 10.0.121.1
---------------------------------------------------------------------------------------
RDMA_Read Latency Test
Dual-port : OFF Device : mlx5_0
Number of qps : 1 Transport type : IB
Connection type : RC Using SRQ : OFF
PCIe relax order: ON
ibv_wr* API : ON
TX depth : 1
Mtu : 1024[B]
Link type : Ethernet
GID index : 3
Outstand reads : 16
rdma_cm QPs : OFF
Data ex. method : Ethernet
---------------------------------------------------------------------------------------
local address: LID 0000 QPN 0x0048 PSN 0x77ae88 OUT 0x10 RKey 0x186ded VAddr 0x005fe0b3e3a000
GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:121:09
remote address: LID 0000 QPN 0x0048 PSN 0x51948d OUT 0x10 RKey 0x186ded VAddr 0x00577584a67000
GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:121:01
---------------------------------------------------------------------------------------
#bytes #iterations t_min[usec] t_max[usec] t_typical[usec] t_avg[usec] t_stdev[usec] 99% percentile[usec] 99.9% percentile[usec]
2 20000 3.98 65.30 4.08 7.89 7.17 31.51 36.33
---------------------------------------------------------------------------------------
RoCE Bandwidth Test
Return to the first server console.
-
Start the
ib_write_bwserver side:First BM Server Console
root@worker1:~# ib_write_bw -d mlx5_0 -F -a -q 4 --report_gbits ************************************ * Waiting for client to connect... * ************************************ -
Move to the second server console.
Start theib_write_bwclient side:Second BM Server Console
depuser@worker2:~$ ib_write_bw -d mlx5_0 -F -a -q 4 10.0.120.2 --report_gbits --------------------------------------------------------------------------------------- RDMA_Write BW Test Dual-port : OFF Device : mlx5_0 Number of qps : 4 Transport type : IB Connection type : RC Using SRQ : OFF PCIe relax order: ON ibv_wr* API : ON TX depth : 128 CQ Moderation : 100 Mtu : 4096[B] Link type : Ethernet GID index : 3 Max inline data : 0[B] rdma_cm QPs : OFF Data ex. method : Ethernet --------------------------------------------------------------------------------------- local address: LID 0000 QPN 0x0052 PSN 0x5b54f8 RKey 0x182e00 VAddr 0x0070e928a01000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10 local address: LID 0000 QPN 0x0053 PSN 0xa16782 RKey 0x182e00 VAddr 0x0070e929201000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10 local address: LID 0000 QPN 0x0054 PSN 0x15fa4 RKey 0x182e00 VAddr 0x0070e929a01000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10 local address: LID 0000 QPN 0x0055 PSN 0xd9b023 RKey 0x182e00 VAddr 0x0070e92a201000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10 remote address: LID 0000 QPN 0x0052 PSN 0xefbd15 RKey 0x182d00 VAddr 0x007ff2aa1d7000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02 remote address: LID 0000 QPN 0x0053 PSN 0x17c9db RKey 0x182d00 VAddr 0x007ff2aa9d7000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02 remote address: LID 0000 QPN 0x0054 PSN 0xd13589 RKey 0x182d00 VAddr 0x007ff2ab1d7000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02 remote address: LID 0000 QPN 0x0055 PSN 0x9f80a4 RKey 0x182d00 VAddr 0x007ff2ab9d7000 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02 --------------------------------------------------------------------------------------- #bytes #iterations BW peak[Gb/sec] BW average[Gb/sec] MsgRate[Mpps] 2 20000 0.022036 0.016848 1.053021 4 20000 0.25 0.25 7.739084 8 20000 0.50 0.49 7.721015 16 20000 0.99 0.99 7.728775 32 20000 1.98 1.97 7.692634 64 20000 3.96 3.96 7.728619 128 20000 7.90 7.86 7.675307 256 20000 15.81 15.77 7.702318 512 20000 31.51 31.39 7.663545 1024 20000 62.18 61.98 7.565755 2048 20000 121.66 121.25 7.400641 4096 20000 212.90 212.79 6.493855 8192 20000 228.04 164.11 2.504087 16384 20000 228.21 228.10 1.740301 32768 20000 229.78 229.36 0.874950 65536 20000 230.35 229.53 0.437792 131072 20000 230.52 229.68 0.219042 262144 20000 230.90 230.89 0.110097 524288 20000 186.92 186.91 0.044564 1048576 20000 179.16 179.16 0.021358 2097152 20000 182.22 182.22 0.010861 4194304 20000 181.55 181.52 0.005410 8388608 20000 181.72 181.72 0.002708 ---------------------------------------------------------------------------------------
Network Isolation Test
Finally, verify that workloads on different tenants (RED and BLUE) cannot communicate with each other. Both workloads use PF0 on their host, but isolation is enforced by HBN through separate VLAN, L2VNI, L3VNI, and VRF assignments.
Connect to the first workload server, with the PF0 network, and try to ping the PF0 on second node.
-
Run the
pingcommands from PF0 to PF0:First BM Server Console
root@worker1:~# ping -c 3 10.0.121.9 PING 10.0.121.9 (10.0.121.9) 56(84) bytes of data. 64 bytes from 10.0.121.9: icmp_seq=1 ttl=62 time=0.896 ms 64 bytes from 10.0.121.9: icmp_seq=2 ttl=62 time=0.241 ms 64 bytes from 10.0.121.9: icmp_seq=3 ttl=62 time=0.258 ms -
Try to ping the PF0 on nodes 3 and 4. Run the
pingcommands from PF0 to PF0:First BM Server Console
root@worker1:~# ping -c 3 10.0.122.1 PING 10.0.122.1 (10.0.122.1) 56(84) bytes of data. From 10.0.121.2 icmp_seq=1 Destination Host Unreachable From 10.0.121.2 icmp_seq=2 Destination Host Unreachable From 10.0.121.2 icmp_seq=3 Destination Host Unreachable --- 10.0.122.1 ping statistics --- 3 packets transmitted, 0 received, +3 errors, 100% packet loss, time 2045ms root@worker1:~# ping -c 3 10.0.122.9 PING 10.0.122.9 (10.0.122.9) 56(84) bytes of data. From 10.0.121.2 icmp_seq=1 Destination Host Unreachable From 10.0.121.2 icmp_seq=2 Destination Host Unreachable From 10.0.121.2 icmp_seq=3 Destination Host Unreachable --- 10.0.122.9 ping statistics --- 3 packets transmitted, 0 received, +3 errors, 100% packet loss, time 2067ms
This ping operation should fail due to the network isolation implemented in HBN using different VLANs, VNIs and VRFs.
Done.
Authors
|
|
Boris Kovalev
Boris Kovalev has worked for the past several years as a Solutions Architect, focusing on NVIDIA Networking/Mellanox technology, and is responsible for complex machine learning, Big Data and advanced VMware-based cloud research and design. Boris previously spent more than 20 years as a senior consultant and solutions architect at multiple companies, most recently at VMware. He has written multiple reference designs covering VMware, machine learning, Kubernetes, and container solutions which are available at the NVIDIA Documents website. |
NVIDIA, the NVIDIA logo, and BlueField are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries. Other company and product names may be trademarks of the respective companies with which they are associated.™
2025 NVIDIA Corporation. All rights reserved.©
Last updated: