Networking Solutions

RDG for DPF Zero Trust (DPF-ZT) with HBN DPU Service v26_4 multi VRF

 Created on Jul 2, 2026 (v 26.4 GA)


Scope

This Reference Deployment Guide (RDG) provides comprehensive instructions for deploying the NVIDIA DOCA Platform Framework (DPF) on high-performance, bare-metal infrastructure in Zero-Trust mode. The guide focuses on setting up an accelerated Host-Based Networking (HBN) service on NVIDIA® BlueField®-3 DPUs to deliver secure, isolated, and hardware-accelerated environments.

The guide is intended for experienced system administrators, systems engineers, and solution architects who build highly secure bare-metal environments with Host-Based Networking enabled using NVIDIA BlueField DPUs for acceleration, isolation, and infrastructure offload.

This document is an extension of the RDG for DPF Zero Trust (DPF-ZT) (referred to as the Baseline RDG). It details the additional steps and modifications required to deploy the HBN, Service into the Baseline RDG environment.

  • This reference implementation, as the name implies, is a specific, opinionated deployment example designed to address the use case described above. 

  • Although other approaches may exist for implementing similar solutions, this document provides a detailed guide for this specific method.

Abbreviations and Acronyms

Term

Definition

Term

Definition

BFB

BlueField Bootstream

OOB

Out-of-Band

BGP

Border Gateway Protocol

PF

Physical Function

DOCA

Data Center Infrastructure-on-a-Chip Architecture

RDG

Reference Deployment Guide

DPF

DOCA Platform Framework

RDMA

Remote Direct Memory Access

DPU

Data Processing Unit

RoCE

RDMA over Converged Ethernet

HBN

Host Based Networking

SFC

Service Function Chaining

IPAM

IP Address Management

SR-IOV

Single Root Input/Output Virtualization

K8S

Kubernetes

VLAN

Virtual LAN (Local Area Network)

KVM

Kernel-based Virtual Machine

VNI

Virtual Network Interface

MAAS

Metal as a Service

VRF

Virtual Router/Forwarder

MTU

Maximum Transmission Unit

ZT

Zero Trust

NGC

NVIDIA GPU Cloud



Introduction

The NVIDIA BlueField-3 Data Processing Unit (DPU) is a 400 Gb/s infrastructure compute platform designed for line-rate processing of software-defined networking, storage, and cybersecurity workloads. It combines powerful compute resources, high-speed networking, and advanced programmability to deliver hardware-accelerated, software-defined solutions for modern data centers.

NVIDIA DOCA unleashes the full potential of the BlueField platform by enabling rapid development of applications and services that offload, accelerate, and isolate data center workloads.

One such service is Host-Based Networking (HBN) - a DOCA-enabled solution that allows network architects to design networks based on Layer 3 (L3) protocols. HBN enables routing on the server side by using BlueField as a BGP router. It encapsulates key networking functions in a containerized service pod, deployed directly on the BlueField’s Arm cores.

However, deploying and managing DPUs and their associated DOCA services, especially at scale, presents operational challenges. Without a robust provisioning and orchestration system, tasks such as lifecycle management, service deployment, and network configuration for service function chaining (SFC) can quickly become complex and error prone. This is where the DOCA Platform Framework (DPF) comes into play.

DPF automates the full DPU lifecycle, streamlines the deployment of DOCA services, and simplifies advanced network configurations. With DPF, services such as HBN can be deployed seamlessly, allowing for efficient offloading and intelligent routing of traffic through the DPU data plane.

By leveraging DPF, users can scale and automate DPU management across Bare Metal, Virtual, and Kubernetes customer environments - optimizing performance while simplifying operations.

DPF supports multiple deployment models. This guide focuses on the Zero Trust bare-metal deployment model. In this scenario:

  • The DPU is managed through its Baseboard Management Controller (BMC)

  • All management traffic occurs over the DPU's out-of-band (OOB) network

  • The host is considered as an untrusted entity towards the data center network. The DPU acts as a barrier between the host and the network.

  • The host sees the DPU as a standard NIC, with no access to the internal DPU management plane (Zero Trust Mode)

This Reference Deployment Guide (RDG) provides a step-by-step example for installing DPF in Zero-Trust mode and HBN. It also includes practical demonstrations of performance optimization, validated using standard RDMA and TCP workloads.

As part of the reference implementation, open-source components outside the scope of DPF (e.g., MAAS, pfSense, Kubespray) are used to simulate a realistic customer deployment environment. The guide includes the full end-to-end deployment process, including:

  • Infrastructure provisioning

  • DPF deployment

  • DPU provisioning (redfish)

  • Service configuration and deployment

  • Service chaining.

This document extends the capabilities of the DPF-managed Kubernetes cluster described in the RDG for DPF Zero Trust (DPF-ZT) (referred to as the Baseline RDG) by deploying the NVIDIA DOCA HBN Service within the existing DPF deployment to achieve a comprehensive, accelerated infrastructure.

References


Solution Architecture

Key Components and Technologies


  • NVIDIA BlueField® Data Processing Unit (DPU)
    The NVIDIA® BlueField® data processing unit (DPU) ignites unprecedented innovation for modern data centers and supercomputing clusters. With its robust compute power and integrated software-defined hardware accelerators for networking, storage, and security, BlueField creates a secure and accelerated infrastructure for any workload in any environment, ushering in a new era of accelerated computing and AI.



  • NVIDIA DOCA Software Framework
    NVIDIA DOCA™ unlocks the potential of the NVIDIA® BlueField® networking platform. By harnessing the power of BlueField DPUs and SuperNICs, DOCA enables the rapid creation of applications and services that offload, accelerate, and isolate data center workloads. It lets developers create software-defined, cloud-native, DPU- and SuperNIC-accelerated services with zero-trust protection, addressing the performance and security demands of modern data centers.



  • NVIDIA ConnectX SmartNICs
    10/25/40/50/100/200 and 400G Ethernet Network Adapters
    The industry-leading NVIDIA® ConnectX® family of smart network interface cards (SmartNICs) offer advanced hardware offloads and accelerations.
    NVIDIA Ethernet adapters enable the highest ROI and lowest Total Cost of Ownership for hyperscale, public and private clouds, storage, machine learning, AI, big data, and telco platforms.


  • NVIDIA LinkX Cables 
    The NVIDIA® LinkX® product family of cables and transceivers provides the industry’s most complete line of 10, 25, 40, 50, 100, 200, and 400GbE in Ethernet and 100, 200 and 400Gb/s InfiniBand products for Cloud, HPC, hyperscale, Enterprise, telco, storage and artificial intelligence, data center applications.

  • NVIDIA Spectrum Ethernet Switches
    Flexible form-factors with 16 to 128 physical ports, supporting 1GbE through 400GbE speeds.
    Based on a ground-breaking silicon technology optimized for performance and scalability, NVIDIA Spectrum switches are ideal for building high-performance, cost-effective, and efficient Cloud Data Center Networks, Ethernet Storage Fabric, and Deep Learning Interconnects. 
    NVIDIA combines the benefits of NVIDIA Spectrum switches, based on an industry-leading application-specific integrated circuit (ASIC) technology, with a wide variety of modern network operating system choices, including NVIDIA Cumulus® LinuxSONiC and NVIDIA Onyx®.

  • NVIDIA Cumulus Linux 
    NVIDIA® Cumulus® Linux is the industry's most innovative open network operating system that allows you to automate, customize, and scale your data center network like no other.


  • Kubernetes
    Kubernetes is an open-source container orchestration platform for deployment automation, scaling, and management of containerized applications.



  • Kubespray 
    Kubespray is a composition of Ansible playbooks, inventory, provisioning tools, and domain knowledge for generic OS/Kubernetes clusters configuration management tasks and provides:

    • A highly available cluster

    • Composable attributes

    • Support for most popular Linux distributions


Solution Design

Solution Logical Design

The logical design includes the following components: 

  • 1 x Hypervisor node (KVM-based) with ConnectX-7:

    • 1 x Firewall VM

    • 1 x Jump Node VM

    • 1 x MaaS VM 

    • 3 x K8s Master VMs running all K8s management components

  • 4 x Worker nodes (PCI Gen5), each with a 1 x BlueField-3 NIC 

  • Single High-Speed (HS) switch

  • 1 Gb Host Management network

DPF_ZT_HBN_MULTI_VRFs.png


HBN service Logical Design

As part of this RDG, we will:

  • Create two fully isolated logical networks per bare-metal workload server using a single physical function (PF0).

    • Connect each network through the HBN service to a dedicated VLAN/VNI, mapped to separate VRFs (RED or BLUE).

  • Route all workload traffic through the HBN service, with routing and isolation enforced inside the DPU.

  • Assign PF0 as the sole network interface for each bare-metal workload server, with no networking configuration on the host.

  • Demonstrate accelerated RDMA and TCP traffic between workload servers running on different bare-metal hosts within the same network (for example, RED ↔ RED).

  • Validate strict network isolation by confirming that traffic between workloads in different networks (RED vs BLUE) is not permitted.

image-2026-2-17_13-18-41.png

Firewall Design

The pfSense firewall in this solution serves a dual purpose:

  • Firewall—provides an isolated environment for the DPF system, ensuring secure operations

  • Router—enables Internet access for the management network

Port-forwarding rules for SSH and RDP are configured on the firewall to route traffic to the jump node’s IP address in the host management network. From the jump node, administrators can manage and access various devices in the setup, as well as handle the deployment of the Kubernetes (K8s) cluster and DPF components.

The following diagram illustrates the firewall design used in this solution:

fw.png

Software Stack Components

SW stack 26_4_GA.png

Make sure to use the exact same versions for the software stack as described above.

Bill of Materials

image-2026-1-18_9-9-29-1.png

Deployment and Configuration

Node and Switch Definitions

These are the definitions and parameters used for deploying the demonstrated fabric:

Switches Ports Usage

Hostname

Rack ID

Ports

mgmt-switch

1

swp1-5

hs-switch

1

swp1-9

Hosts

Rack

Server Type

Server Name

Switch Port

IP and NICs

Default Gateway

Rack1


Hypervisor Node

hypervisor

mgmt-switch: swp1

hs-switch: swp1

lab-br (interface eno1): Trusted LAN IP

mgmt-br (interface eno2): -

hs-br (interface enp1s0): -

Trusted LAN GW

Rack1

Firewall (Virtual)

fw

-

WAN (lab-br): Trusted LAN IP

LAN (mgmt-br): 10.0.110.254/24

    OPT1(hs-br): 10.0.123.254/22

Trusted LAN GW

Rack1

Jump Node (Virtual)

jump

-

enp1s0: 10.0.110.253/24

10.0.110.254

Rack1

MaaS (Virtual)

maas

-

enp1s0: 10.0.110.252/24

10.0.110.254

Rack1

Master Node
(Virtual) 

master1

-

enp1s0: 10.0.110.1/24

10.0.110.254

Rack1

Master Node
(Virtual)

master2

-

enp1s0: 10.0.110.2/24

10.0.110.254

Rack1

Master Node
(Virtual)

master3

-

enp1s0: 10.0.110.3/24

10.0.110.254

Rack1


Worker Node

worker1

mgmt-switch: swp2(DPU OOB) 

hs-switch: swp2-swp3

dpubmc: 10.0.110.21/24

ens1f0np0/ens1f1np1: 10.0.120.0/22

10.0.110.254

Rack1


Worker Node

worker2

mgmt-switch: swp3(DPU OOB)

hs-switch: swp4-swp5

dpubmc: 10.0.110.22/24

ens1f0np0/ens1f1np1: 10.0.120.0/22

10.0.110.254

Rack1


Worker Node

worker3

mgmt-switch: swp4(DPU OOB) 

hs-switch: swp6-swp7

dpubmc: 10.0.110.23/24

ens1f0np0/ens1f1np1: 10.0.120.0/22

10.0.110.254

Rack1


Worker Node

worker4

mgmt-switch: swp5(DPU OOB)

hs-switch: swp8-swp9

dpubmc: 10.0.110.24/24

ens1f0np0/ens1f1np1: 10.0.120.0/22

10.0.110.254

Note: On BlueField-3, the DPU BMC and DPU OOB management interfaces share a single 1G out-of-band link via an internal bridge (oob_net0tmfifo_net0 on BMC side). Both IPs (.201/.211) reside on the same L2 segment of the management network (10.0.110.0/24) and are reached via a single switch port (swpN). It is necessary to set several environment variables before running this command.

$ source manifests/00-env-vars/envvars.env

Note: Workers' high-speed PFs (ens1f0np0, ens1f1np1) connect to hs-switch via 200GbE. No persistent host-side IP in Zero-Trust baseline mode — DPU acts as a transparent NIC. Per-tenant IP assignment (e.g. 10.0.121.x for HBN RED, 10.0.122.x for HBN BLUE) is configured by downstream services (see DPF-ZT with HBN sub-page). Subnet 10.0.120.0/22 is reserved for the high-speed fabric.

Wiring

Hypervisor Node 

HW node.png

Bare Metal Worker Node

image-2026-2-17_13-20-44-1.png

Fabric Configuration

Updating Cumulus Linux

As a best practice, make sure to use the latest released Cumulus Linux NOS version.

For information on how to upgrade Cumulus Linux, refer to the Cumulus Linux User Guide.

Configuring the Cumulus Linux Switch

The SN3700 switch (hs-switch), is configured as follows:

SN3700 Switch Console
nv set evpn state enable
nv set interface eth0 ip address dhcp
nv set interface eth0 ip vrf mgmt
nv set interface eth0 type eth
nv set interface lo ipv4 address 11.0.0.101/32
nv set interface lo type loopback
nv set interface swp1-9 link state up
nv set interface swp1-9 type swp
nv set interface swp1 ipv4 address 172.169.50.2/30
nv set qos roce mode lossless
nv set qos roce state enabled
nv set router bgp autonomous-system 65001
nv set router bgp state enabled
nv set router bgp graceful-restart mode full
nv set router bgp router-id 11.0.0.101
nv set vrf default router bgp address-family ipv4-unicast state enabled
nv set vrf default router bgp address-family ipv4-unicast redistribute connected state enabled
nv set vrf default router bgp address-family ipv4-unicast redistribute static state enabled
nv set vrf default router bgp address-family ipv6-unicast state enabled
nv set vrf default router bgp address-family ipv6-unicast redistribute connected state enabled
nv set vrf default router bgp address-family l2vpn-evpn state enabled
nv set vrf default router bgp state enabled
nv set vrf default router bgp neighbor swp2-9 peer-group hbn
nv set vrf default router bgp neighbor swp2-9 type unnumbered
nv set vrf default router bgp path-selection multipath aspath-ignore enabled
nv set vrf default router bgp peer-group hbn address-family l2vpn-evpn state enabled
nv set vrf default router bgp peer-group hbn remote-as external
nv set vrf default router static 0.0.0.0/0 address-family ipv4-unicast
nv set vrf default router static 0.0.0.0/0 via 172.169.50.1 type ipv4-address
nv config apply -y
nv config save

The SN2201 switch (mgmt-switch) is configured as follows:

SN2201 Switch Console
nv set interface swp1-5 link state up
nv set interface swp1-5 type swp
nv set interface swp1-5 bridge domain br_default
nv set bridge domain br_default untagged 1
nv config apply
nv config save -y

Host Configuration

Make sure that the BIOS settings on the worker node servers have SR-IOV enabled and that the servers are tuned for maximum performance.
Required:

  • SR-IOV: Enabled

  • VT-d / AMD-Vi (IOMMU): Enabled

  • Above 4G Decoding: Enabled (mandatory for PCIe BAR sizes on BlueField-3)

Performance-recommended:

  • CPU C-states: Disabled (or up to C1 only)

  • Hyper-Threading: Enabled

  • Memory speed: Maximum supported

  • Power profile: Performance / Maximum Performance

All worker nodes must have the same PCIe placement for the BlueField-3 NIC and must display the same interface name.

Make sure that you have DPU BMC and OOB MAC addresses.

No change from the Reference Deployment Guide (Baseline RDG) (Section "Deployment and Configuration", Subsection "Host Configuration").

Hypervisor Installation and Configuration

No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Hypervisor Installation and Configuration").  

Prepare Infrastructure Servers

No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Prepare Infrastructure Servers") regarding Firewall VM, Jump VM, MaaS VM.

Firewall VM – Bare Metal Server Outside Conection 

To provide outside connection from Bare Metal Host via High Speed network, open Firefox web browser and go to the pfSense web UI (http://10.0.110.254).

System:

  • Routing → Static Routing → Add → “Destination network”: 10.0.125.0/24, “Gateway”: Switch - 172.169.50.2 → , “Description”: To DPU DHCP → Click "Save"→ Under "Default Gateway" - "Default gateway IPv4" choose WAN_DHCP → Click "Save"

    PFsense_route.png

Note that the IP addresses from the Trusted LAN network under "Gateway" and "Monitor IP" are blurred.

Provision Master VMs Using MaaS

No change from the Baseline RDG (Section "Deployment and Configuration", Subsection "Provision Master VMs Using MaaS").

K8s Cluster Deployment and Configuration

The procedures for initial Kubernetes cluster deployment using Kubespray for the master nodes, and subsequent verification, remain unchanged from the Baseline RDG (Section "K8s Cluster Deployment and Configuration", Subsections: "Kubespray Deployment and Configuration", "Deploying Cluster Using Kubespray Ansible Playbook","K8s Deployment Verification".

DPF Installation

The DPF installation process (Operator, System components) largely follows the Baseline RDG. 

Software Prerequisites and Required Variables

  1. Start by installing the remaining software perquisites.

    Jump Node Console

    ## Connect to master1 to copy helm client utility that was installed during kubespray deployment
    $ depuser@jump:~$ ssh master1
    depuser@master1:~$ cp /usr/local/bin/helm /tmp/
    
    ## In another tab 
    depuser@jump:~$ scp master1:/tmp/helm /tmp/
    depuser@jump:~$ sudo chown root:root /tmp/helm
    depuser@jump:~$ sudo mv /tmp/helm /usr/local/bin/
    
    ## Verify that envsubst utility is installed 
    depuser@jump:~$ which envsubst
    /usr/bin/envsubst
    
  2. Proceed to clone the doca-platform Git repository:

    Jump Node Console

    $ git clone https://github.com/NVIDIA/doca-platform.git
    
  3. Change directory to doca-platform and checkout to tag v26.4.0

    Jump Node Console

    $ cd doca-platform/
    $ git checkout v26.4.0
    
  4. Change directory to doca-platform/docs/public/user-guides/zero-trust/use-cases/hbn from where all the commands will be run:

    Jump Node Console

    $ cd doca-platform/docs/public/user-guides/zero-trust/use-cases/hbn
    
  5. Change the BMC root's password.
    In Zero Trust mode, provisioning DPUs requires authentication with Redfish.
    In order to do that, you must set the same root password to access the BMC for all DPUs DPF is going to manage.For more information on how to set the BMC root password refer to BlueField DPU Administrator Quick Start Guide

    Connect to the first DPU BMC over SSH to change the BMC root's password:

    Jump Node Console

    $ ssh root@10.0.110.201
    root@10.0.110.201's password: <BMC Root Password. Default root/0penBmc. need to change first time to $BMC_ROOT_PASSWORD in the manifests/00-env-vars/envvars.env file>
    
  6. Modify the variables in manifests/00-env-vars/envvars.env to fit your environment, then source the file: 

  7. Replace the values for the variables in the following file with the values that fit your setup. Specifically, pay attention to DPUCLUSTER_INTERFACEBMC_ROOT_PASSWORD, and DPU's serial number.
    To get a DPU's serial number you can use following command. Sample:
    $ curl -k -u root:'BMC root password' https://10.0.110.201/redfish/v1/Systems/Bluefield | jq -r '.SerialNumber | ascii_downcase'
      % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                     Dload  Upload   Total   Spent    Left  Speed
    100  4970  100  4970    0     0   4211      0  0:00:01  0:00:01 --:--:--  4211
    mt2402xz0f7x

    manifests/00-env-vars/envvars.env

    Bash
    ## IP Address for the Kubernetes API server of the target cluster on which DPF is installed.
    ## This should never include a scheme or a port.
    ## e.g. 10.10.10.10
    export TARGETCLUSTER_API_SERVER_HOST=10.0.110.10
    
     ## Port for the Kubernetes API server of the target cluster on which DPF is installed.
     ## e.g. 6443
     export TARGETCLUSTER_API_SERVER_PORT=6443
     
    ## Virtual IP used by the load balancer for the DPU Cluster. Must be a reserved IP from the management subnet and not
    ## allocated by DHCP.
    export DPUCLUSTER_VIP=10.0.110.200
     
    ## Interface on which the DPUCluster load balancer will listen. Should be the management interface of the control plane node.
    export DPUCLUSTER_INTERFACE=enp1s0
     
    ## The repository URL for the NVIDIA Helm chart registry.
    ## Usually this is the NVIDIA Helm NGC registry. For development purposes, this can be set to a different repository.
    export HELM_REGISTRY_REPO_URL=https://helm.ngc.nvidia.com/nvidia/doca
     
    ## The repository URL for the HBN container image.
    ## Usually this is the NVIDIA NGC registry. For development purposes, this can be set to a different repository.
    export HBN_NGC_IMAGE_URL=nvcr.io/nvidia/doca/doca_hbn
     
    ## The DPF REGISTRY is the Helm repository URL where the DPF Operator Chart resides.
    ## Usually this is the NVIDIA Helm NGC registry. For development purposes, this can be set to a different repository.
    export REGISTRY=https://helm.ngc.nvidia.com/nvidia/doca
      
    ## The DPF TAG is the version of the DPF components which will be deployed in this guide.
    export TAG=v26.4.0
      
    ## URL to the BFB used in the `bfb.yaml` and linked by the DPUSet.
    export BFB_URL="https://content.mellanox.com/BlueField/BFBs/Ubuntu24.04/bf-bundle-3.4.0-92_26.04_ubuntu-24.04_64k_prod.bfb"
    
    
    ## IP_RANGE_START and IP_RANGE_END
    ## These define the IP range for DPU discovery via Redfish/BMC interfaces
    ## Example: If your DPUs have BMC IPs in range 10.0.110.201-224
    ## export IP_RANGE_START=10.0.110.201
    ## export IP_RANGE_END=10.0.110.224
     
    ## Start of DPUDiscovery IpRange
    export IP_RANGE_START=10.0.110.201
     
    ## End of DPUDiscovery IpRange
    export IP_RANGE_END=10.0.110.204
     
    # The password used for DPU BMC root login, must be the same for all DPUs
    # For more information on how to set the BMC root password refer to BlueField DPU Administrator Quick Start Guide. 
    export BMC_ROOT_PASSWORD=<set your BMC_ROOT_PASSWORD>
     
    ## Serial number of DPUs. If you have more than 2 DPUs, you will need to parameterize the system accordingly and expose
    ## additional variables.
    ## All serial numbers must be in lowercase.
     
    ## Serial number of DPU1
    export DPU1_SERIAL=mt2402xz0f7x
     
    ## Serial number of DPU2
    export DPU2_SERIAL=mt2402xz0f80
     
    ## Serial number of DPU3
    export DPU3_SERIAL=mt2402xz0f9n
     
    ## Serial number of DPU4
    export DPU4_SERIAL=mt2402xz0f8g
    
  8. Export environment variables for the installation:

    Jump Node Console

    $ source manifests/00-env-vars/envvars.env
    

DPF Operator Installation

No change from the Baseline RDG (Section "DPF Installation", Subsection "DPF Operator Installation").

DPF System Installation

No change from the Baseline RDG (Section "DPF Installation", Subsection "DPF System Installation").

DPU Services Installation 

HBN DPU Service Installation

This section focuses on provisioning NVIDIA®BlueField®-3 DPUs using DPF, installing the HBN DPU Service on those DPUs and enabling workload traffic to pass through HBN before leaving the DPU.

  1. Export environment variables for the installation:

    Jump Node Console

    $ source manifests/00-env-vars/envvars.env
    
  2. Use the following YAML to define a BFB resource that downloads the Bluefield Bitstream to a shared volume:

    ---
    apiVersion: provisioning.dpu.nvidia.com/v1alpha1
    kind: BFB
    metadata:
      name: bf-bundle-$TAG
      namespace: dpf-operator-system
    spec:
      url: $BFB_URL
    
  3. Change the DPUFlavor using the following YAML:

    manifests/03.1-dpudeployment-installation-pf/hbn-dpuflavor.yaml
    ---
    apiVersion: provisioning.dpu.nvidia.com/v1alpha1
    kind: DPUFlavor
    metadata:
      name: hbn-$TAG
      namespace: dpf-operator-system
    spec:
      bfcfgParameters:
      - UPDATE_ATF_UEFI=yes
      - UPDATE_DPU_OS=yes
      - WITH_NIC_FW_UPDATE=yes
      configFiles:
      - operation: override
        path: /etc/mellanox/mlnx-bf.conf
        permissions: "0644"
        raw: |
          ALLOW_SHARED_RQ="no"
          IPSEC_FULL_OFFLOAD="no"
          ENABLE_ESWITCH_MULTIPORT="yes"
      - operation: override
        path: /etc/mellanox/mlnx-ovs.conf
        permissions: "0644"
        raw: |
          CREATE_OVS_BRIDGES="no"
          OVS_DOCA="yes"
      - operation: override
        path: /etc/mellanox/mlnx-sf.conf
        permissions: "0644"
        raw: ""
      grub:
        kernelParameters:
        - console=hvc0
        - console=ttyAMA0
        - earlycon=pl011,0x13010000
        - fixrttc
        - net.ifnames=0
        - biosdevname=0
        - iommu.passthrough=1
        - cgroup_no_v1=net_prio,net_cls
        - hugepagesz=2048kB
        - hugepages=3072
      nvconfig:
      - device: '*'
        parameters:
        - PF_BAR2_ENABLE=0
        - PER_PF_NUM_SF=1
        - PF_TOTAL_SF=20
        - PF_SF_BAR_SIZE=10
        - NUM_PF_MSIX_VALID=0
        - PF_NUM_PF_MSIX_VALID=1
        - PF_NUM_PF_MSIX=228
        - INTERNAL_CPU_MODEL=1
        - INTERNAL_CPU_OFFLOAD_ENGINE=0
        - SRIOV_EN=1
        - NUM_OF_VFS=46
        - LAG_RESOURCE_ALLOCATION=1
        - LINK_TYPE_P1=ETH
        - LINK_TYPE_P2=ETH
    	- EXP_ROM_UEFI_x86_ENABLE=1
      ovs:
        rawConfigScript: |
          _ovs-vsctl() {
            ovs-vsctl --timeout 15 "$@"
          }
    
          # Remove default OVS configuration on the DPU and ensure no leftovers on the OVS kernel side
          _ovs-vsctl --if-exists del-br ovsbr1
          _ovs-vsctl --if-exists del-br ovsbr2
          ovs-appctl --timeout 15 dpctl/del-dp system@ovs-system || true
    
          _ovs-vsctl set Open_vSwitch . other_config:doca-init=true
          _ovs-vsctl set Open_vSwitch . other_config:dpdk-max-memzones=50000
          _ovs-vsctl set Open_vSwitch . other_config:hw-offload=true
          _ovs-vsctl set Open_vSwitch . other_config:pmd-quiet-idle=true
          _ovs-vsctl set Open_vSwitch . other_config:max-idle=20000
          _ovs-vsctl set Open_vSwitch . other_config:max-revalidator=5000
          _ovs-vsctl remove Open_vSwitch . other_config default-datapath-type || true
    
          if systemctl list-unit-files openvswitch-switch.service &>/dev/null; then
            systemctl restart openvswitch-switch
          elif systemctl list-unit-files openvswitch.service &>/dev/null; then
            systemctl restart openvswitch
          fi
          _ovs-vsctl --may-exist add-br br-sfc
          _ovs-vsctl set bridge br-sfc datapath_type=netdev
          _ovs-vsctl set bridge br-sfc fail_mode=secure
          _ovs-vsctl --may-exist add-port br-sfc p0
          _ovs-vsctl set Interface p0 type=dpdk
          _ovs-vsctl set Interface p0 mtu_request=9216
          _ovs-vsctl set Port p0 external_ids:dpf-type=physical
          _ovs-vsctl --may-exist add-port br-sfc p1
          _ovs-vsctl set Interface p1 type=dpdk
          _ovs-vsctl set Interface p1 mtu_request=9216
          _ovs-vsctl set Port p1 external_ids:dpf-type=physical
          _ovs-vsctl --may-exist add-br br-hbn
          _ovs-vsctl set bridge br-hbn datapath_type=netdev
          _ovs-vsctl set bridge br-hbn fail_mode=secure
    
  4. Change the dpudeployment.yaml file to reference the DPUFlavor.

    manifests/03.1-dpudeployment-installation-pf/dpudeployment.yaml
    ---
    apiVersion: svc.dpu.nvidia.com/v1alpha1
    kind: DPUDeployment
    metadata:
      name: hbn-only
      namespace: dpf-operator-system
    spec:
      dpus:
        bfb: bf-bundle-$TAG
        flavor: hbn-$TAG
        nodeEffect:
          hold: true
        dpuSetStrategy:
          type: OnDelete
        dpuSets:
        - nameSuffix: "dpuset1"
          dpuNodeSelector:
            matchLabels:
              feature.node.kubernetes.io/dpu-enabled: "true"
      services:
        doca-hbn:
          serviceTemplate: doca-hbn
          serviceConfiguration: doca-hbn
      serviceChains:
        switches:
          - ports:
            serviceMTU: 9000 
            - serviceInterface:
                matchLabels:
                  interface: p0
            - service:
                name: doca-hbn
                interface: p0_if
          - ports:
            serviceMTU: 9000
            - serviceInterface:
                matchLabels:
                  interface: p1
            - service:
                name: doca-hbn
                interface: p1_if
          - ports:
            serviceMTU: 9000
            - serviceInterface:
                matchLabels:
                  interface: pf0hpf
            - service:
                interface: pf0hpf_if
                name: doca-hbn
    

    Please notice that with default nodeEffect above, DPU provisioning workflow will be paused and wait for an external signal (annotation) in order to proceed, as demonstrated in upcoming steps.
    To implement a fully automated process that won’t require user intervention, see customAction option.

  5. Change the rest of the configuration files.

    As explained in the introduction, these files create service chains that connect two physical functions PF0  or PF0 to the outer fabric through HBN, providing EVPN VXLAN overlay, VNI based isolation, and ECMP redundancy across both DPU uplinks (p0 and p1).
    These are the configuration files.

    • HBN DPUServiceConfig and DPUServiceTemplate to deploy HBN workloads to the DPUs.

      manifests/03.1-dpudeployment-installation-pf/hbn-dpuserviceconfig.yaml
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceConfiguration
      metadata:
        name: doca-hbn
        namespace: dpf-operator-system
      spec:
        deploymentServiceName: "doca-hbn"
        serviceConfiguration:
          serviceDaemonSet:
            annotations:
              k8s.v1.cni.cncf.io/networks: |-
                [
                  {"name": "iprequest", "interface": "ip_lo", "cni-args": {"poolNames": ["loopback"], "poolType": "cidrpool"}},
                  {"name": "iprequest", "interface": "ip_pf0hpf_red", "cni-args": {"poolNames": ["pool1"], "poolType": "cidrpool", "allocateDefaultGateway": true}},
                  {"name": "iprequest", "interface": "ip_pf0hpf_blue", "cni-args": {"poolNames": ["pool2"], "poolType": "cidrpool", "allocateDefaultGateway": true}}
                ]
      
          helmChart:
            values:
              configuration:
                perDPUValuesYAML: |
                  - hostnamePattern: "*"
                    values:
                      bgp_peer_group: hbn
      
                  # ---- DPU1, DPU2 => RED only ----
                  - hostnamePattern: "dpu-node-${DPU1_SERIAL}-${DPU1_SERIAL}*"
                    values:
                      role: RED
                      vrf: RED
                      vlan: 11
                      l2vni: 10010
                      l3vni: 100001
                      bgp_autonomous_system: 65101
      
                  - hostnamePattern: "dpu-node-${DPU2_SERIAL}-${DPU2_SERIAL}*"
                    values:
                      role: RED
                      vrf: RED
                      vlan: 11
                      l2vni: 10010
                      l3vni: 100001
                      bgp_autonomous_system: 65201
      
                  # ---- DPU3, DPU4 => BLUE only ----
                  - hostnamePattern: "dpu-node-${DPU3_SERIAL}-${DPU3_SERIAL}*"
                    values:
                      role: BLUE
                      vrf: BLUE
                      vlan: 21
                      l2vni: 10020
                      l3vni: 100002
                      bgp_autonomous_system: 65301
      
                  - hostnamePattern: "dpu-node-${DPU4_SERIAL}-${DPU4_SERIAL}*"
                    values:
                      role: BLUE
                      vrf: BLUE
                      vlan: 21
                      l2vni: 10020
                      l3vni: 100002
                      bgp_autonomous_system: 65401
      
                startupYAMLJ2: |
                  - header:
                      model: bluefield
                      nvue-api-version: nvue_v1
                      rev-id: 1.0
                      version: HBN 2.4.0
      
                  - set:
                      bridge:
                        domain:
                          br_default:
                            vlan:
                              {{ config.vlan }}:
                                vni:
                                  {{ config.l2vni }}: {}
      
                      evpn:
                        enable: on
                        route-advertise: {}
      
                      interface:
                        lo:
                          ip:
                            address:
                              {{ ipaddresses.ip_lo.ip }}/32: {}
                          type: loopback
      
                        p0_if,p1_if,pf0hpf_if:
                          type: swp
                          link:
                            mtu: 9216
      
                        pf0hpf_if:
                          bridge:
                            domain:
                              br_default:
                                access: {{ config.vlan }}
      
                        vlan{{ config.vlan }}:
                          type: svi
                          vlan: {{ config.vlan }}
      					# SVI MTU must match swp uplinks; default 1500 caps cross-VRF routed traffic.
                          link:
                            mtu: 9216
                          ip:
                            address:
                              {% if config.role == "RED" %}
                              {{ ipaddresses.ip_pf0hpf_red.cidr }}: {}
                              {% else %}
                              {{ ipaddresses.ip_pf0hpf_blue.cidr }}: {}
                              {% endif %}
                            vrf: {{ config.vrf }}
      
                      nve:
                        vxlan:
                          arp-nd-suppress: on
                          enable: on
                          source:
                            address: {{ ipaddresses.ip_lo.ip }}
      
                      router:
                        bgp:
                          enable: on
                          graceful-restart:
                            mode: full
      
                      vrf:
                        default:
                          router:
                            bgp:
                              address-family:
                                ipv4-unicast:
                                  enable: on
                                  redistribute:
                                    connected:
                                      enable: on
                                l2vpn-evpn:
                                  enable: on
                              autonomous-system: {{ config.bgp_autonomous_system }}
                              enable: on
                              neighbor:
                                p0_if:
                                  peer-group: {{ config.bgp_peer_group }}
                                  type: unnumbered
                                p1_if:
                                  peer-group: {{ config.bgp_peer_group }}
                                  type: unnumbered
                              path-selection:
                                multipath:
                                  aspath-ignore: on
                              peer-group:
                                {{ config.bgp_peer_group }}:
                                  address-family:
                                    ipv4-unicast:
                                      enable: on
                                    l2vpn-evpn:
                                      enable: on
                                  remote-as: external
                              router-id: {{ ipaddresses.ip_lo.ip }}
      
                        {{ config.vrf }}:
                          evpn:
                            enable: on
                            vni:
                              {{ config.l3vni }}: {}
                          loopback:
                            ip:
                              address:
                                {{ ipaddresses.ip_lo.ip }}/32: {}
                          router:
                            bgp:
                              address-family:
                                ipv4-unicast:
                                  enable: on
                                  redistribute:
                                    connected:
                                      enable: on
                                  route-export:
                                    to-evpn:
                                      enable: on
                              autonomous-system: {{ config.bgp_autonomous_system }}
                              enable: on
                              router-id: {{ ipaddresses.ip_lo.ip }}
      
        interfaces:
          - name: p0_if
            network: mybrhbn
          - name: p1_if
            network: mybrhbn
          - name: pf0hpf_if
            network: mybrhbn
      
      manifests/03.1-dpudeployment-installation-pf/hbn-dpuservicetemplate.yaml
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceTemplate
      metadata:
        name: doca-hbn
        namespace: dpf-operator-system
      spec:
        deploymentServiceName: "doca-hbn"
        helmChart:
          source:
            repoURL: $HELM_REGISTRY_REPO_URL
            version: 3.4.0-0
            chart: doca-hbn
          values:
            image:
              repository: $HBN_NGC_IMAGE_URL
              tag: release-3.4.0-doca3.4.0
            resources:
              memory: 6Gi
              nvidia.com/bf_sf: 3
      
    • Physical Interfaces for physical ports on the DPU.

      manifests/03.1-dpudeployment-installation-pf/physical-ifaces.yaml
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceInterface
      metadata:
        name: p0
        namespace: dpf-operator-system
      spec:
        template:
          spec:
            template:
              metadata:
                labels:
                  interface: "p0"
              spec:
                interfaceType: physical
                physical:
                  interfaceName: p0
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceInterface
      metadata:
        name: p1
        namespace: dpf-operator-system
      spec:
        template:
          spec:
            template:
              metadata:
                labels:
                  interface: "p1"
              spec:
                interfaceType: physical
                physical:
                  interfaceName: p1
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceInterface
      metadata:
        name: pf0hpf
        namespace: dpf-operator-system
      spec:
        template:
          spec:
            template:
              metadata:
                labels:
                  interface: "pf0hpf"
              spec:
                interfaceType: pf
                pf:
                  pfID: 0
      
      
    • DPU Service IPAM objects to set up IP Address Management on the DPUCluster.

      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceIPAM
      metadata:
        name: pool1
        namespace: dpf-operator-system
      spec:
        ipv4Network:
          network: "10.0.121.0/24"
          gatewayIndex: 2
          prefixSize: 29
          # These preallocations are not necessary. We specify them so that the validation commands are straightforward.
          allocations:
            dpu-node-${DPU1_SERIAL}-${DPU1_SERIAL}: 10.0.121.0/29
            dpu-node-${DPU2_SERIAL}-${DPU2_SERIAL}: 10.0.121.8/29
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceIPAM
      metadata:
        name: pool2
        namespace: dpf-operator-system
      spec:
        ipv4Network:
          network: "10.0.122.0/24"
          gatewayIndex: 2
          prefixSize: 29
          allocations:
            dpu-node-${DPU3_SERIAL}-${DPU3_SERIAL}: 10.0.122.0/29
            dpu-node-${DPU4_SERIAL}-${DPU4_SERIAL}: 10.0.122.8/29  
      
      ---
      apiVersion: svc.dpu.nvidia.com/v1alpha1
      kind: DPUServiceIPAM
      metadata:
        name: loopback
        namespace: dpf-operator-system
      spec:
        ipv4Network:
          network: "11.0.0.0/24"
          prefixSize: 32
      

      It is necessary to set several environment variables before running this command.

      $ source manifests/00-env-vars/envvars.env

  6. Apply all of the YAML files mentioned above using the following command:

    Jump Node Console

    $ cat manifests/03.1-dpudeployment-installation-pf/*.yaml | envsubst | kubectl apply -f -
    

     

    Jump Node Console

    $ kubectl wait --for=condition=ApplicationsReconciled --namespace dpf-operator-system dpuservices --all
    dpuservice.svc.dpu.nvidia.com/cni-installer condition met
    dpuservice.svc.dpu.nvidia.com/doca-hbn-x92vr condition met
    dpuservice.svc.dpu.nvidia.com/flannel condition met
    dpuservice.svc.dpu.nvidia.com/kube-state-metrics-05f12f695b condition met
    dpuservice.svc.dpu.nvidia.com/kube-state-metrics-rbac condition met
    dpuservice.svc.dpu.nvidia.com/multus condition met
    dpuservice.svc.dpu.nvidia.com/node-problem-detector condition met
    dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-05f12f695b condition met
    dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-node condition met
    dpuservice.svc.dpu.nvidia.com/ovs-cni condition met
    dpuservice.svc.dpu.nvidia.com/servicechainset-controller-05f12f695b condition met
    dpuservice.svc.dpu.nvidia.com/servicechainset-rbac-and-crds condition met
    dpuservice.svc.dpu.nvidia.com/sfc-controller condition met
    dpuservice.svc.dpu.nvidia.com/sriov-device-plugin condition met
    
    $ kubectl wait --for=condition=DPUIPAMObjectReconciled --namespace dpf-operator-system dpuserviceipam --all
    dpuserviceipam.svc.dpu.nvidia.com/loopback condition met
    dpuserviceipam.svc.dpu.nvidia.com/pool1 condition met
    
    $ kubectl wait --for=condition=ServiceInterfaceSetReconciled --namespace dpf-operator-system dpuserviceinterface --all
    dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-p0-if-f9hzk condition met
    dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-p1-if-2ld7q condition met
    dpuserviceinterface.svc.dpu.nvidia.com/doca-hbn-pf0hpf-if-gt8zw condition met
    dpuserviceinterface.svc.dpu.nvidia.com/p0 condition met
    dpuserviceinterface.svc.dpu.nvidia.com/p1 condition met
    dpuserviceinterface.svc.dpu.nvidia.com/pf0hpf condition met
    
    $ kubectl wait --for=condition=ServiceChainSetReconciled --namespace dpf-operator-system dpuservicechain --all
    dpuservicechain.svc.dpu.nvidia.com/hbn-only-cjpt5 condition met
    
  7. To follow the progress of DPU provisioning, run the following command to check its current phase:

    Jump Node Console

    $ watch -n10 "kubectl describe dpu -n dpf-operator-system | grep 'Node Name\|Type\|Last\|Phase'"
    
    


  8. Wait for the NodeEffect stage (at this point the provisioning is paused, waintig for external signal).
    Run following command on all/specific DPU nodemaintanace object/s to proceed with provisioning:

    Jump Node Console

    $ kubectl annotate dpunodemaintenances -n dpf-operator-system --all \
      provisioning.dpu.nvidia.com/wait-for-external-nodeeffect=false \
      maintenance.dpu.nvidia.com/wait-for-external-nodeeffect=false \
      --overwrite
    
  9. To follow the progress of DPU provisioning, run the following command to check its current phase:

    Jump Node Console

    $ watch -n10 "kubectl -n dpf-operator-system get dpu,dpuset,dpudeployment,dpuservice,dpuserviceconfigurations,dpuservicetemplates"
    
  10. Wait for the Rebooted stage and then Power Cycle the bare-metal host manual.
    After the DPU is up, run following command for each DPU worker:

    Jump Node Console

    $ kubectl annotate dpunode -n dpf-operator-system --all provisioning.dpu.nvidia.com/dpunode-external-reboot-required-
    
  11. At this point, the DPU workers should be added to the cluster. As they being added to the cluster, the DPUs are provisioned.

    Jump Node Console

    $ watch -n 2 'kubectl -n dpf-operator-system get dpu,dpuset,dpudeployment,dpuservice,dpuserviceconfigurations,dpuservicetemplates'
    
    NAME                                                                 READY   OPERATIONAL   PHASE   AGE
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f7x   True    True          Ready   114m
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f80   True    True          Ready   114m
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f8g   True    True          Ready   114m
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f9n   True    True          Ready   114m
    
    NAME                                                  READY   AGE
    dpuset.provisioning.dpu.nvidia.com/hbn-only-dpuset1   True    114m
    
    NAME                                        READY   PHASE     AGE
    dpudeployment.svc.dpu.nvidia.com/hbn-only   True    Success   115m
    
    NAME                                                                  READY   PHASE     AGE
    dpuservice.svc.dpu.nvidia.com/cni-installer                           True    Success   116m
    dpuservice.svc.dpu.nvidia.com/doca-hbn-x92vr                          True    Success   114m
    dpuservice.svc.dpu.nvidia.com/flannel                                 True    Success   116m
    dpuservice.svc.dpu.nvidia.com/kube-state-metrics-05f12f695b           True    Success   116m
    dpuservice.svc.dpu.nvidia.com/kube-state-metrics-rbac                 True    Success   116m
    dpuservice.svc.dpu.nvidia.com/multus                                  True    Success   116m
    dpuservice.svc.dpu.nvidia.com/node-problem-detector                   True    Success   116m
    dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-05f12f695b              True    Success   116m
    dpuservice.svc.dpu.nvidia.com/nvidia-k8s-ipam-node                    True    Success   116m
    dpuservice.svc.dpu.nvidia.com/ovs-cni                                 True    Success   116m
    dpuservice.svc.dpu.nvidia.com/servicechainset-controller-05f12f695b   True    Success   116m
    dpuservice.svc.dpu.nvidia.com/servicechainset-rbac-and-crds           True    Success   116m
    dpuservice.svc.dpu.nvidia.com/sfc-controller                          True    Success   116m
    dpuservice.svc.dpu.nvidia.com/sriov-device-plugin                     True    Success   116m
    
    NAME                                                  AGE
    dpuserviceconfiguration.svc.dpu.nvidia.com/doca-hbn   115m
    
    NAME                                             AGE
    dpuservicetemplate.svc.dpu.nvidia.com/doca-hbn   115m
    
  12.  Finally, validate that all the different DPU-related objects are now in the Ready state:

    Jump Node Console

    $ kubectl get secrets -n dpu-cplane-tenant1 dpu-cplane-tenant1-admin-kubeconfig -o json | jq -r '.data["admin.conf"]' | base64 --decode > /home/depuser/dpu-cluster.config
     
    $ echo "alias ki='KUBECONFIG=/home/depuser/dpu-cluster.config kubectl'" >> ~/.bashrc
    $ echo 'alias dpfctl="kubectl -n dpf-operator-system exec deploy/dpf-operator-controller-manager -- /dpfctl "' >> ~/.bashrc
     
    $ dpfctl describe dpudeployments
    NAME                                   NAMESPACE            STATUS       REASON    SINCE  MESSAGE
    DPFOperatorConfig/dpfoperatorconfig    dpf-operator-system  Ready: True  Success   3m3s
    └─DPUDeployments
      └─DPUDeployment/hbn                  dpf-operator-system  Ready: True  Success   22s
        ├─DPUServiceChains
        │ └─DPUServiceChain/hbn-wd7fs      dpf-operator-system  Ready: True  Success   65s
        ├─DPUServiceInterfaces
        │ └─3 DPUServiceInterfaces...      dpf-operator-system  Ready: True  Success   70s    See doca-hbn-p0-if-749n9, doca-hbn-p1-if-fn8w5, doca-hbn-pf0hpf-if-9s8c6
        ├─DPUSets
        │ └─DPUSet/hbn-dpuset1             dpf-operator-system  Ready: True  Success   71s
        │   ├─BFB/bf-bundle-v26.4.0       dpf-operator-system  Ready: True  Ready     39m    File: 3.4.0-92_26.04_ubuntu-24.04_64k_prod.bfb, DOCA: 3.4.0
        │   ├─DPUNodes
        │   │ └─4 DPUNodes...              dpf-operator-system  Ready: True  Ready     98s    See dpu-node-mt2402xz0f7x, dpu-node-mt2402xz0f80, dpu-node-mt2402xz0f8g, dpu-node-mt2402xz0f9n
        │   └─DPUs
        │     └─4 DPUs...                  dpf-operator-system  Ready: True  DPUReady  98s    See dpu-node-mt2402xz0f7x-mt2402xz0f7x, dpu-node-mt2402xz0f80-mt2402xz0f80,
        │                                                                                     dpu-node-mt2402xz0f8g-mt2402xz0f8g, dpu-node-mt2402xz0f9n-mt2402xz0f9n
        └─Services
          ├─DPUServiceTemplates
          │ └─DPUServiceTemplate/doca-hbn  dpf-operator-system  Ready: True  Success   39m
          └─DPUServices
            └─1 DPUServices...             dpf-operator-system  Ready: True  Success   50s    See doca-hbn-jxkxw
    
    
    $ ki get node -A
    NAME                                 STATUS   ROLES    AGE     VERSION
    dpu-node-mt2402xz0f7x-mt2402xz0f7x   Ready    <none>   5m18s   v1.34.8
    dpu-node-mt2402xz0f80-mt2402xz0f80   Ready    <none>   6m12s   v1.34.8
    dpu-node-mt2402xz0f8g-mt2402xz0f8g   Ready    <none>   6m14s   v1.34.8
    dpu-node-mt2402xz0f9n-mt2402xz0f9n   Ready    <none>   6m22s   v1.34.8
     
    $ kubectl get dpu -A
    NAMESPACE             NAME                                 READY   PHASE   AGE
    dpf-operator-system   dpu-node-mt2402xz0f7x-mt2402xz0f7x   True    Ready   36m
    dpf-operator-system   dpu-node-mt2402xz0f80-mt2402xz0f80   True    Ready   36m
    dpf-operator-system   dpu-node-mt2402xz0f8g-mt2402xz0f8g   True    Ready   36m
    dpf-operator-system   dpu-node-mt2402xz0f9n-mt2402xz0f9n   True    Ready   36m
    
    $ kubectl wait --for=condition=ready --namespace dpf-operator-system dpu --all
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f7x-mt2402xz0f7x condition met
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f80-mt2402xz0f80 condition met
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f8g-mt2402xz0f8g condition met
    dpu.provisioning.dpu.nvidia.com/dpu-node-mt2402xz0f9n-mt2402xz0f9n condition met
    
    $ ki get pods -A -o wide
    NAMESPACE             NAME                                                             READY   STATUS    RESTARTS      AGE     IP             NODE                                 NOMINATED NODE   READINESS GATES
    dpf-operator-system   dpu-cplane-tenant1-cni-installer-89kn4                           1/1     Running   0               6m50s   10.244.2.3     dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-cni-installer-s8h4z                           1/1     Running   0               7m1s    10.244.0.5     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-cni-installer-wb29j                           1/1     Running   0               5m57s   10.244.3.2     dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-cni-installer-zhzqh                           1/1     Running   0               6m53s   10.244.1.4     dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-doca-hbn-jxkxw-ds-5sbzs                       2/2     Running   0               2m54s   10.244.0.6     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-doca-hbn-jxkxw-ds-ftnpn                       2/2     Running   0               2m54s   10.244.1.5     dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq                       2/2     Running   0               3m21s   10.244.3.4     dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb                       2/2     Running   0               2m54s   10.244.2.4     dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-nvidia-k8s-ipam-controller-5c77854fcc-grchr   1/1     Running   0               127m    10.244.0.3     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-krgzw                 1/1     Running   0               6m53s   10.244.1.2     dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-pr85m                 1/1     Running   0               5m57s   10.244.3.3     dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-x4lfs                 1/1     Running   0               7m1s    10.244.0.2     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-nvidia-k8s-ipam-node-ds-zlzvf                 1/1     Running   0               6m50s   10.244.2.2     dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-ovs-cni-arm64-bpljq                           1/1     Running   0               7m1s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-ovs-cni-arm64-gls6h                           1/1     Running   0               6m50s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-ovs-cni-arm64-j8wr4                           1/1     Running   0               5m57s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-ovs-cni-arm64-kbrrn                           1/1     Running   0               6m53s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-sfc-controller-node-ds-vmfq4                  1/1     Running   0               5m57s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-sfc-controller-node-ds-x45nl                  1/1     Running   0               6m53s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-sfc-controller-node-ds-xskh9                  1/1     Running   0               7m1s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   dpu-cplane-tenant1-sfc-controller-node-ds-zfmt5                  1/1     Running   1 (5m46s ago)   6m50s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   kube-flannel-ds-2shh7                                            1/1     Running   0               7m2s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   kube-flannel-ds-42mlq                                            1/1     Running   0               6m54s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   kube-flannel-ds-m7xgt                                            1/1     Running   0               5m58s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   kube-flannel-ds-vd574                                            1/1     Running   0               6m52s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   kube-multus-ds-d5kb4                                             1/1     Running   0               6m53s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   kube-multus-ds-gnv88                                             1/1     Running   0               6m50s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   kube-multus-ds-l66tm                                             1/1     Running   0               7m1s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   kube-multus-ds-mh4cj                                             1/1     Running   0               5m57s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpf-operator-system   kube-sriov-device-plugin-64c29                                   1/1     Running   0               7m1s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpf-operator-system   kube-sriov-device-plugin-6js9j                                   1/1     Running   0               6m50s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    dpf-operator-system   kube-sriov-device-plugin-g5gkx                                   1/1     Running   0               6m53s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpf-operator-system   kube-sriov-device-plugin-lk4z7                                   1/1     Running   0               5m57s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    kube-system           coredns-66bc5c9577-gqn8d                                         1/1     Running   0               127m    10.244.0.4     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    kube-system           coredns-66bc5c9577-p2xnm                                         1/1     Running   0               127m    10.244.1.3     dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    kube-system           kube-proxy-64865                                                 1/1     Running   0               5m58s   10.0.110.211   dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    kube-system           kube-proxy-hvjjp                                                 1/1     Running   0               6m52s   10.0.110.212   dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    kube-system           kube-proxy-qfbwh                                                 1/1     Running   0               6m54s   10.0.110.214   dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    kube-system           kube-proxy-w9gg4                                                 1/1     Running   0               7m2s    10.0.110.213   dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    

    Congratulations! The DPF system with the HBN service has been successfully installed.

Zero-Trust Mode Checking

Here's a step-by-step procedure to check the Zero-Trust Mode on your NVIDIA BlueField DPU from the host server, including the installation of the Mellanox Firmware Tools (MFT).

Ubuntu 24.04 was installed on the servers.

  1. Navigate to the NVIDIA Downloads Site: Open your web browser and go to the official NVIDIA Mellanox software downloads page.

  2. Select the Latest Version for your OS: image-2025-9-9_12-24-17.png

  3. Transfer and Extract MFT Tools on the Worker 1 BareMetal Host.

    First Pod Console

    root@worker1:~# tar -xvzf /tmp/mft-4.33.0-169-x86_64-deb.tgz
    
  4. Navigate into the Extracted Directory.

    First Pod Console

    root@worker1:~# cd mft-4.33.0-169-x86_64-deb/
    
  5. Run following commands.

    First Pod Console

    root@worker1:~# apt-get install gcc make dkms
    root@worker1:~# ./install.sh
    
  6. Start MST (Mellanox Software Tools) Service and Identify DPU Device Name.

    First Pod Console

    root@worker1:~# mst start
    
    Starting MST (Mellanox Software Tools) driver set
    Loading MST PCI module - Success
    Loading MST PCI configuration module - Success
    Create devices
    Unloading MST PCI module (unused) - Success
    
    root@worker1:~# mst status
    
    MST modules:
    ------------
        MST PCI module is not loaded
        MST PCI configuration module loaded
    
    MST devices:
    ------------
    /dev/mst/mt41692_pciconf0        - PCI configuration cycles access.
                                       domain:bus:dev.fn=0000:2b:00.0 addr.reg=88 data.reg=92 cr_bar.gw_offset=-1
                                       Chip revision is: 01
    
    
  7. Perform Zero-Trust Checking.

    First Pod Console

    root@worker1:~# mlxprivhost -d 2b:00.0 q
    Host configurations
    -------------------
    level                         : RESTRICTED
    
    Port functions status:
    -----------------------
    disable_rshim                 : TRUE
    disable_tracer                : TRUE
    disable_port_owner            : TRUE
    disable_counter_rd            : TRUE
    
    #Expected Zero-Trust Output.
    

    This is the most definitive confirmation. level : RESTRICTED means the host is in Zero-Trust Mode, and the TRUE flags confirm individual security restrictions are active.

  8. Verify the host cannot reset DPU firmware:
    root@worker1:~# sudo mlxfwreset -d 2b:00.0 -y -l 3 reset

    Expected output on a Zero-Trust host:
    -E- Failed to send Register MFRL: Register access Method not supported (264).

    The MFRL (Master Firmware Reset Lock) register is access-gated by RESTRICTED mode. Failure here confirms the host cannot trigger a DPU firmware reset.

  9. Check Firmware Access with mlxfwmanager:

    First Pod Console

    root@worker1:~# mlxfwmanager -d 2b:00.0 --query
    Querying Mellanox devices firmware ...
    
    Device #1:
    ----------
    
      Device Type:      BlueField3
      Part Number:      --
      Description:
      PSID:
      PCI Device Name:  2b:00.0
      Base MAC:         N/A
      Versions:         Current        Available
         FW             --
    
      Status:           Failed to open device
    

    The behaviour of mlxfwmanager --query depends on the MFT version installed:

    • MFT < 4.33 returns Status: Failed to open device — the host cannot read inventory at all in Zero-Trust mode.

    • MFT 4.33+ returns inventory data (FW/PXE/UEFI versions, PSID, MAC) with Status: No matching image found. The BMC fulfils the read from cached inventory; this is not a Zero-Trust failure. The Available: N/A and the No matching image found status both confirm the host has no path to upload firmware. Write-side operations (--updatemlxfwreset) remain blocked.

    Either output is consistent with Zero-Trust mode. Use the mlxprivhost -d <dev> p write probe in step 11 (below) for the definitive verification.


  10. Check Device Configuration with mlxconfig:

    First Pod Console

    root@worker1:~# mlxconfig -d 2b:00.0 q
    
    Device #1:
    ----------
    
    Device type:        BlueField3
    Name:               900-9D3B6-00CV-A_Ax
    Description:        NVIDIA BlueField-3 B3220 P-Series FHHL DPU; 200GbE (default mode) / NDR200 IB; Dual-port QSFP112; PCIe Gen5.0 x16 with x16 PCIe extension option; 16 Arm cores; 32GB on-board DDR; integrated BMC; Crypto Enabled
    Device:             2b:00.0
    
    Configurations:                                          Next Boot
    ...
            ALLOW_RD_COUNTERS                           True(1)   # No RO, but restricted by mlxprivhost
    ...
            PORT_OWNER                                  True(1)   # No RO, but restricted by mlxprivhost
    ...        
            TRACER_ENABLE                               True(1)   # No RO, but restricted by mlxprivhost
    

    Most configuration parameters are prefixed with RO (Read-Only) — the host literally cannot change them, by design. A small number of parameters related to host-side control (PORT_OWNERALLOW_RD_COUNTERSTRACER_ENABLE) are not marked RO, and mlxconfig set will appear to succeed against them on a Zero-Trust host:
    root@worker1:~# sudo mlxconfig -d 2b:00.0 set TRACER_ENABLE=0
    ...
    Apply new Configuration? (y/n) [n] : y
    Applying... Done!
    -I- Please power cycle machine to load new configurations.

    However, the change is functionally a no-op. At runtime, the mlxprivhost RESTRICTED layer (visible in step 7's disable_tracer: TRUE) overrides whatever mlxconfig says. On the next DPU re-provision, the DPF operator's DPUFlavor.spec.nvconfig re-applies the canonical settings anyway. The fact that mlxconfig set returns "Applying... Done!" does not indicate Zero-Trust is off.


  11. Step 11. Verify the host cannot escape Zero-Trust mode (write probe with mlxprivhost):
    root@worker1:~# sudo mlxprivhost -d 2b:00.0 p

    Expected output on a Zero-Trust host:
    -E- Operation is not permitted (refer to the DPU user manual)

    The p argument asks the host to switch its privilege level back to PRIVILEGED. On a Zero-Trust DPU this must fail — that's the whole point of the RESTRICTED level. If this command succeeds and the level flips, the DPU is not in Zero-Trust mode and the DPUFlavor (spec.dpuMode) needs to be re-checked.

    This is the only command in this section that is fundamentally write-side and therefore unaffected by MFT-version changes to read-side behaviour.

  12. Check Low-Level Hardware Access with ethtool:

    First Pod Console

    root@worker1:~# ethtool -d ens1f0np0
    Cannot get register dump: Operation not supported
    

     This confirms the DPU is preventing deep, low-level hardware access from the host, aligning with Zero-Trust's isolation goals.


Conclusion

The host is operating in Zero-Trust Mode when ALL of the following are true:

#

Test

Expected on Zero-Trust host

1

mlxprivhost q -> level line

RESTRICTED

2

mlxprivhost q -> disable_* flags

disable_rshimdisable_tracerdisable_port_ownerdisable_counter_rd all TRUE

3

mlxconfig q -> count of RO rows

majority of parameters prefixed RO

4

ethtool -d <iface>

Operation not supported

5

mlxprivhost p (write probe — definitive)

Operation is not permitted

6

mlxfwreset reset (optional write check)

Method not supported

Tests 1–4 are read-side checks and confirm the configuration is in place. Test 5 (mlxprivhost p) is the authoritative proof: a Zero-Trust host cannot remove its own restrictions. Test 6 reinforces this for firmware-reset specifically.


This means the host has significantly restricted privileges and cannot perform sensitive operations on the DPU, ensuring its security and isolation.

Infrastructure Bandwidth & Latency Validation 

Verify the deployment and confirm that the DPU system achieves link-speed performance and low latency by running various tests:

  1. Iperf TCP—for bandwidth measurements 

  2. RDMA—for bandwidth and latency measurements 

  3. Network isolation

Each test is described in detail. At the end of each test, the achieved performance is displayed. 

Notes

Make sure that the servers are tuned for maximum performance (not covered in this document).  

Performance and Isolation Tests

Now that the test deployment is running, perform bandwidth and latency performance tests between two bare-metal workload servers.

Ubuntu 24.04 was installed on the servers.

  1. Before running the tests, check the Gateway address on each HBN pod:

    Jump Node Console

    $ ki -n dpf-operator-system get pod -o wide | grep doca-hbn
    dpu-cplane-tenant1-doca-hbn-jxkxw-ds-5sbzs                       2/2     Running   0             15m    10.244.0.6     dpu-node-mt2402xz0f9n-mt2402xz0f9n   <none>           <none>
    dpu-cplane-tenant1-doca-hbn-jxkxw-ds-ftnpn                       2/2     Running   0             15m    10.244.1.5     dpu-node-mt2402xz0f8g-mt2402xz0f8g   <none>           <none>
    dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq                       2/2     Running   0             16m    10.244.3.4     dpu-node-mt2402xz0f7x-mt2402xz0f7x   <none>           <none>
    dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb                       2/2     Running   0             15m    10.244.2.4     dpu-node-mt2402xz0f80-mt2402xz0f80   <none>           <none>
    
    
    $ ki exec -it -n dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq -- bash
    Defaulted container "doca-hbn" out of: doca-hbn, hbn-sidecar, hbn-init (init)
    
    root@dpu-cplane-tenant1-doca-hbn-jxkxw-ds-gjsqq:/tmp# ip a s
    ...
    9: vlan11@br_default: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 9216 qdisc noqueue master RED state UP group default qlen 1000
        link/ether 0a:ff:4e:3e:99:24 brd ff:ff:ff:ff:ff:ff
        inet 10.0.121.2/29 scope global vlan11
           valid_lft forever preferred_lft forever
        inet6 fe80::8ff:4eff:fe3e:9924/64 scope link
           valid_lft forever preferred_lft forever
    ...
    
    $ exit
    
    $  ki exec -it -n dpf-operator-system dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb -- bash
    Defaulted container "doca-hbn" out of: doca-hbn, hbn-sidecar, hbn-init (init)
    
    root@dpu-cplane-tenant1-doca-hbn-jxkxw-ds-k78vb:/tmp# ip a s
    ...
    9: vlan11@br_default: <BROADCAST,MULTICAST,UP,LOWER_UP> mtu 9216 qdisc noqueue master RED state UP group default qlen 1000
        link/ether 0e:7d:99:41:2e:11 brd ff:ff:ff:ff:ff:ff
        inet 10.0.121.10/29 scope global vlan11
           valid_lft forever preferred_lft forever
        inet6 fe80::c7d:99ff:fe41:2e11/64 scope link
           valid_lft forever preferred_lft forever
    ...
    
    $ exit
    


  2. Connect to a first Workload Server console, install iperf, perftest, check DPU Hight Speed Interfaces, set route to ethernet and identify the relevant RDMA device:

    First Pod Console

    root@worker1:~# apt install iperf3
    root@worker1:~# apt install perftest
    root@worker1:~# ip a s
    ...
    6: ens1f0np0: <BROADCAST,MULTICAST> mtu 1500 qdisc noop state DOWN group default qlen 1000
        link/ether 58:a2:e1:73:69:e6 brd ff:ff:ff:ff:ff:ff
        altname enp43s0f0np0
    ...
    
    root@worker1:~# ip route add 10.0.123.0/22 via 10.0.121.2
    
    depuser@worker1:~$ ping 8.8.8.8
    PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data.
    64 bytes from 8.8.8.8: icmp_seq=1 ttl=117 time=5.35 ms
    64 bytes from 8.8.8.8: icmp_seq=2 ttl=117 time=5.10 ms
    64 bytes from 8.8.8.8: icmp_seq=3 ttl=117 time=5.15 ms
    
    root@worker1:~#  rdma link | grep ens1f0np0
    link mlx5_0/1 state DOWN physical_state DISABLED netdev ens1f0np0
    
  3. Configure the ens1f0np0 interface on Ubuntu 24.04 using iproute2 .
    Configuration Overview

    Interface

    IP Address

    Default Gateway

    ens1f0np0

    10.0.121.1/29

    10.0.121.2/29


    First Pod Console

    # Bring up physical interfaces
    root@worker1:~# ip link set dev ens1f0np0 up
    
    # Assign IP addresses
    root@worker1:~# ip addr add 10.0.121.1/29 dev ens1f0np0
    
    # Set default route
    root@worker1:~# ip route add default via 10.0.121.2 dev ens1f0np0
    


  4. Using another console window, reconnect to the jump node and connect to a second Workload Server.
    From within the servers, install iperf, perftest, check DPU Hight Speed Interfaces, set route to ethernet and identify the relevant RDMA device:

    First Pod Console

    root@worker2:~# apt install iperf3
    root@worker2:~# apt install perftest
    root@worker2:~# ip a s
    ...
    6: ens1f0np0: <BROADCAST,MULTICAST> mtu 9000 qdisc noop state DOWN group default qlen 1000
        link/ether 58:a2:e1:73:6a:58 brd ff:ff:ff:ff:ff:ff
        altname enp43s0f0np0
    ...
    
    root@worker2:~# ip route add 10.0.123.0/22 via 10.0.121.10
    
    depuser@worker2:~$ ping 8.8.8.8
    PING 8.8.8.8 (8.8.8.8) 56(84) bytes of data.
    64 bytes from 8.8.8.8: icmp_seq=1 ttl=117 time=5.35 ms
    64 bytes from 8.8.8.8: icmp_seq=2 ttl=117 time=5.10 ms
    64 bytes from 8.8.8.8: icmp_seq=3 ttl=117 time=5.15 ms
    
    
    root@worker2:~# rdma link | grep ens1f0np0
    link mlx5_0/1 state DOWN physical_state DISABLED netdev ens1f0np0
    
    

     

  5. Configure the ens1f0np0 interface on Ubuntu 24.04 using iproute2.

    Configuration Overview

    Interface

    IP Address

    Default Gateway

    ens1f0np0

    10.0.121.9/29

    10.0.121.10/29

First Pod Console
# Bring up physical interfaces
root@worker2:~# ip link set dev ens1f0np0 up

# Assign IP addresses
root@worker2:~# ip addr add 10.0.121.9/29 dev ens1f0np0

# Set default route
root@worker2:~# ip route add default via 10.0.121.10 dev ens1f0np0
iPerf TCP Bandwidth Test

Move back to the first server console.

  1. Host NIC tuning (required for line-rate TCP on 200GbE) .
    Without tuning, TCP throughput on Ubuntu 24.04 with stock NIC settings caps at ~167 Gbps. Apply these on every worker (sender + receiver) that will run iperf3:

First Pod Console (Sample)
# 1) Bump RX/TX ring buffers (default 1024 → max 8192)
root@worker1:~# ethtool -G ens1f0np0 rx 8192 tx 8192

# 2) Pin all mlx5_comp IRQs to the NIC's NUMA node
PCI=$(ethtool -i ens1f0np0 | awk '/bus-info/{print $2}')
NUMA=$(cat /sys/class/net/ens1f0np0/device/numa_node)
NUMA_CPUS=$(cat /sys/devices/system/node/node${NUMA}/cpulist | tr ',' '\n' | \
            awk -F- '{if($2)print $2-$1+1; else print 1}' | paste -sd+ | bc)
for irq in $(awk -v p="$PCI" '$0~"mlx5_comp.*@pci:"p{gsub(":",""); print $1}' /proc/interrupts); do
  q=$(awk -v i="$irq:" '$1==i{print $NF}' /proc/interrupts | grep -oP 'comp\K[0-9]+')
  echo $((q % NUMA_CPUS)) | sudo tee /proc/irq/$irq/smp_affinity_list >/dev/null
done

# Expected gain on a 200 GbE BlueField-3 pair: 167 Gbps → 185 Gbps (8 TCP streams, retransmits drop from ~240K to ~120K). 
# Persist via systemd unit if needed.
  1. Start the iperf3 server side:

    First BM Server Console

    root@worker1:~# iperf3 -s -p 5201
    -----------------------------------------------------------
    Server listening on 5201 (test #1)
    ------------------------------------------------------------
    
  2. Move to the second server console.
    Start the iperf client side:

    Second BM Server Console

    depuser@worker2:~$ iperf3 -c 10.0.121.1 -p 5201 -t 30 -P 8 -i 5
    Connecting to host 10.0.121.1, port 5201
    ...
    [ ID] Interval           Transfer     Bitrate         Retr  Cwnd
    [  5]   0.00-5.01   sec  12.1 GBytes  20.7 Gbits/sec  1462   1.31 MBytes
    [  7]   0.00-5.01   sec  14.3 GBytes  24.5 Gbits/sec  1620    883 KBytes
    [  9]   0.00-5.01   sec  11.4 GBytes  19.6 Gbits/sec  1860   1.34 MBytes
    [ 11]   0.00-5.01   sec  12.7 GBytes  21.7 Gbits/sec  2204    743 KBytes
    [ 13]   0.00-5.01   sec  15.0 GBytes  25.8 Gbits/sec  2167   1.22 MBytes
    [ 15]   0.00-5.01   sec  12.9 GBytes  22.2 Gbits/sec  2251    926 KBytes
    [ 17]   0.00-5.01   sec  12.8 GBytes  22.0 Gbits/sec  1467    856 KBytes
    [ 19]   0.00-5.01   sec  16.7 GBytes  28.7 Gbits/sec  2032   1.37 MBytes
    [SUM]   0.00-5.01   sec   108 GBytes   185 Gbits/sec  15063
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [  5]   5.01-10.00  sec  11.5 GBytes  19.7 Gbits/sec  2186    935 KBytes
    [  7]   5.01-10.00  sec  13.6 GBytes  23.4 Gbits/sec  2263    900 KBytes
    [  9]   5.01-10.00  sec  14.3 GBytes  24.5 Gbits/sec  3120   1.04 MBytes
    [ 11]   5.01-10.00  sec  12.9 GBytes  22.2 Gbits/sec  3150   1.05 MBytes
    [ 13]   5.01-10.00  sec  13.2 GBytes  22.7 Gbits/sec  2645    690 KBytes
    [ 15]   5.01-10.00  sec  13.8 GBytes  23.6 Gbits/sec  3126   1.91 MBytes
    [ 17]   5.01-10.01  sec  12.5 GBytes  21.4 Gbits/sec  2161   1.04 MBytes
    [ 19]   5.01-10.01  sec  15.6 GBytes  26.8 Gbits/sec  2769   1.28 MBytes
    [SUM]   5.01-10.00  sec   107 GBytes   184 Gbits/sec  21420
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [  5]  10.00-15.01  sec  12.2 GBytes  21.0 Gbits/sec  1739   1.11 MBytes
    [  7]  10.00-15.01  sec  14.1 GBytes  24.2 Gbits/sec  1835   1.89 MBytes
    [  9]  10.00-15.01  sec  10.8 GBytes  18.5 Gbits/sec  2206   1.11 MBytes
    [ 11]  10.00-15.01  sec  14.5 GBytes  24.9 Gbits/sec  2888   1.37 MBytes
    [ 13]  10.00-15.01  sec  14.0 GBytes  24.0 Gbits/sec  2528    839 KBytes
    [ 15]  10.00-15.01  sec  15.0 GBytes  25.7 Gbits/sec  2798   1.80 MBytes
    [ 17]  10.01-15.01  sec  13.8 GBytes  23.7 Gbits/sec  2280   1.02 MBytes
    [ 19]  10.01-15.01  sec  13.6 GBytes  23.4 Gbits/sec  2205   1.25 MBytes
    [SUM]  10.00-15.01  sec   108 GBytes   185 Gbits/sec  18479
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [  5]  15.01-20.01  sec  11.6 GBytes  19.9 Gbits/sec  2119   1.43 MBytes
    [  7]  15.01-20.01  sec  14.2 GBytes  24.4 Gbits/sec  2457   1.01 MBytes
    [  9]  15.01-20.01  sec  13.3 GBytes  22.8 Gbits/sec  2791   2.13 MBytes
    [ 11]  15.01-20.01  sec  13.6 GBytes  23.4 Gbits/sec  3383   1.65 MBytes
    [ 13]  15.01-20.01  sec  12.7 GBytes  21.9 Gbits/sec  2594   1.49 MBytes
    [ 15]  15.01-20.01  sec  14.3 GBytes  24.6 Gbits/sec  3274   2.30 MBytes
    [ 17]  15.01-20.01  sec  13.1 GBytes  22.6 Gbits/sec  2387   1.74 MBytes
    [ 19]  15.01-20.01  sec  14.8 GBytes  25.5 Gbits/sec  2714   1.20 MBytes
    [SUM]  15.01-20.01  sec   108 GBytes   185 Gbits/sec  21719
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [  5]  20.01-25.01  sec  12.3 GBytes  21.2 Gbits/sec  2191    821 KBytes
    [  7]  20.01-25.01  sec  15.2 GBytes  26.1 Gbits/sec  2159   1.74 MBytes
    [  9]  20.01-25.01  sec  12.3 GBytes  21.2 Gbits/sec  2628   1.01 MBytes
    [ 11]  20.01-25.01  sec  14.5 GBytes  24.9 Gbits/sec  3484   1.14 MBytes
    [ 13]  20.01-25.01  sec  13.2 GBytes  22.6 Gbits/sec  2564   1.42 MBytes
    [ 15]  20.01-25.01  sec  13.5 GBytes  23.2 Gbits/sec  3048   1.07 MBytes
    [ 17]  20.01-25.01  sec  13.3 GBytes  22.8 Gbits/sec  2288   1.53 MBytes
    [ 19]  20.01-25.01  sec  13.4 GBytes  23.1 Gbits/sec  2559   1.23 MBytes
    [SUM]  20.01-25.01  sec   108 GBytes   185 Gbits/sec  20921
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [  5]  25.01-30.01  sec  10.6 GBytes  18.2 Gbits/sec  1706    734 KBytes
    [  7]  25.01-30.01  sec  15.8 GBytes  27.2 Gbits/sec  2563   1022 KBytes
    [  9]  25.01-30.01  sec  10.5 GBytes  18.0 Gbits/sec  2482    874 KBytes
    [ 11]  25.01-30.01  sec  14.4 GBytes  24.8 Gbits/sec  3575   1.58 MBytes
    [ 13]  25.01-30.01  sec  15.6 GBytes  26.8 Gbits/sec  2679   1.75 MBytes
    [ 15]  25.01-30.01  sec  13.6 GBytes  23.3 Gbits/sec  3095   1022 KBytes
    [ 17]  25.01-30.01  sec  13.0 GBytes  22.3 Gbits/sec  2208   1.32 MBytes
    [ 19]  25.01-30.01  sec  15.4 GBytes  26.5 Gbits/sec  2771   2.59 MBytes
    [SUM]  25.01-30.01  sec   109 GBytes   187 Gbits/sec  21079
    - - - - - - - - - - - - - - - - - - - - - - - - -
    [ ID] Interval           Transfer     Bitrate         Retr
    [  5]   0.00-30.01  sec  70.3 GBytes  20.1 Gbits/sec  11403             sender
    [  5]   0.00-30.01  sec  70.3 GBytes  20.1 Gbits/sec                  receiver
    [  7]   0.00-30.01  sec  87.2 GBytes  25.0 Gbits/sec  12897             sender
    [  7]   0.00-30.01  sec  87.2 GBytes  25.0 Gbits/sec                  receiver
    [  9]   0.00-30.01  sec  72.5 GBytes  20.8 Gbits/sec  15087             sender
    [  9]   0.00-30.01  sec  72.5 GBytes  20.8 Gbits/sec                  receiver
    [ 11]   0.00-30.01  sec  82.6 GBytes  23.6 Gbits/sec  18684             sender
    [ 11]   0.00-30.01  sec  82.6 GBytes  23.6 Gbits/sec                  receiver
    [ 13]   0.00-30.01  sec  83.7 GBytes  24.0 Gbits/sec  15177             sender
    [ 13]   0.00-30.01  sec  83.7 GBytes  24.0 Gbits/sec                  receiver
    [ 15]   0.00-30.01  sec  83.1 GBytes  23.8 Gbits/sec  17592             sender
    [ 15]   0.00-30.01  sec  83.1 GBytes  23.8 Gbits/sec                  receiver
    [ 17]   0.00-30.01  sec  78.5 GBytes  22.5 Gbits/sec  12791             sender
    [ 17]   0.00-30.01  sec  78.5 GBytes  22.5 Gbits/sec                  receiver
    [ 19]   0.00-30.01  sec  89.6 GBytes  25.7 Gbits/sec  15050             sender
    [ 19]   0.00-30.01  sec  89.6 GBytes  25.7 Gbits/sec                  receiver
    [SUM]   0.00-30.01  sec   648 GBytes   185 Gbits/sec  118681             sender
    [SUM]   0.00-30.01  sec   647 GBytes   185 Gbits/sec                  receiver
    
    iperf Done.
    
RoCE Latency Test 

Return to the first server console.

  1. Start the ib_read_lat server side:

    First BM Server Console

    root@worker1:~# ib_read_lat -F -n 20000 -d mlx5_0
    
    ************************************
    * Waiting for client to connect... *
    ************************************
    
  2. Move to the second server console.
    Start the ib_read_lat client side:

Second BM Server Console
root@worker2:~# ib_read_lat -F -n 20000 -d mlx5_0 10.0.121.1

---------------------------------------------------------------------------------------
                    RDMA_Read Latency Test
 Dual-port       : OFF          Device         : mlx5_0
 Number of qps   : 1            Transport type : IB
 Connection type : RC           Using SRQ      : OFF
 PCIe relax order: ON
 ibv_wr* API     : ON
 TX depth        : 1
 Mtu             : 1024[B]
 Link type       : Ethernet
 GID index       : 3
 Outstand reads  : 16
 rdma_cm QPs     : OFF
 Data ex. method : Ethernet
---------------------------------------------------------------------------------------
 local address: LID 0000 QPN 0x0048 PSN 0x77ae88 OUT 0x10 RKey 0x186ded VAddr 0x005fe0b3e3a000
 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:121:09
 remote address: LID 0000 QPN 0x0048 PSN 0x51948d OUT 0x10 RKey 0x186ded VAddr 0x00577584a67000
 GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:121:01
---------------------------------------------------------------------------------------
 #bytes #iterations    t_min[usec]    t_max[usec]  t_typical[usec]    t_avg[usec]    t_stdev[usec]   99% percentile[usec]   99.9% percentile[usec]
 2       20000          3.98           65.30        4.08               7.89             7.17            31.51                   36.33
---------------------------------------------------------------------------------------
RoCE Bandwidth Test

Return to the first server console.

  1. Start the ib_write_bw server side:

    First BM Server Console

    root@worker1:~# ib_write_bw -d mlx5_0 -F -a -q 4 --report_gbits
    
    ************************************
    * Waiting for client to connect... *
    ************************************
    
  2. Move to the second server console.
    Start the ib_write_bw client side:

    Second BM Server Console

    depuser@worker2:~$ ib_write_bw -d mlx5_0 -F -a -q 4 10.0.120.2 --report_gbits
    ---------------------------------------------------------------------------------------
                        RDMA_Write BW Test
     Dual-port       : OFF          Device         : mlx5_0
     Number of qps   : 4            Transport type : IB
     Connection type : RC           Using SRQ      : OFF
     PCIe relax order: ON
     ibv_wr* API     : ON
     TX depth        : 128
     CQ Moderation   : 100
     Mtu             : 4096[B]
     Link type       : Ethernet
     GID index       : 3
     Max inline data : 0[B]
     rdma_cm QPs     : OFF
     Data ex. method : Ethernet
    ---------------------------------------------------------------------------------------
     local address: LID 0000 QPN 0x0052 PSN 0x5b54f8 RKey 0x182e00 VAddr 0x0070e928a01000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10
     local address: LID 0000 QPN 0x0053 PSN 0xa16782 RKey 0x182e00 VAddr 0x0070e929201000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10
     local address: LID 0000 QPN 0x0054 PSN 0x15fa4 RKey 0x182e00 VAddr 0x0070e929a01000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10
     local address: LID 0000 QPN 0x0055 PSN 0xd9b023 RKey 0x182e00 VAddr 0x0070e92a201000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:10
     remote address: LID 0000 QPN 0x0052 PSN 0xefbd15 RKey 0x182d00 VAddr 0x007ff2aa1d7000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02
     remote address: LID 0000 QPN 0x0053 PSN 0x17c9db RKey 0x182d00 VAddr 0x007ff2aa9d7000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02
     remote address: LID 0000 QPN 0x0054 PSN 0xd13589 RKey 0x182d00 VAddr 0x007ff2ab1d7000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02
     remote address: LID 0000 QPN 0x0055 PSN 0x9f80a4 RKey 0x182d00 VAddr 0x007ff2ab9d7000
     GID: 00:00:00:00:00:00:00:00:00:00:255:255:10:00:120:02
    ---------------------------------------------------------------------------------------
     #bytes     #iterations    BW peak[Gb/sec]    BW average[Gb/sec]   MsgRate[Mpps]
     2          20000           0.022036            0.016848            1.053021
     4          20000            0.25               0.25               7.739084
     8          20000            0.50               0.49               7.721015
     16         20000            0.99               0.99               7.728775
     32         20000            1.98               1.97               7.692634
     64         20000            3.96               3.96               7.728619
     128        20000            7.90               7.86               7.675307
     256        20000            15.81              15.77              7.702318
     512        20000            31.51              31.39              7.663545
     1024       20000            62.18              61.98              7.565755
     2048       20000            121.66             121.25             7.400641
     4096       20000            212.90             212.79             6.493855
     8192       20000            228.04             164.11             2.504087
     16384      20000            228.21             228.10             1.740301
     32768      20000            229.78             229.36             0.874950
     65536      20000            230.35             229.53             0.437792
     131072     20000            230.52             229.68             0.219042
     262144     20000            230.90             230.89             0.110097
     524288     20000            186.92             186.91             0.044564
     1048576    20000            179.16             179.16             0.021358
     2097152    20000            182.22             182.22             0.010861
     4194304    20000            181.55             181.52             0.005410
     8388608    20000            181.72             181.72             0.002708
    ---------------------------------------------------------------------------------------
    

Network Isolation Test

Finally, verify that workloads on different tenants (RED and BLUE) cannot communicate with each other. Both workloads use PF0 on their host, but isolation is enforced by HBN through separate VLAN, L2VNI, L3VNI, and VRF assignments.

Connect to the first workload server, with the PF0 network, and try to ping the PF0 on second node.

  1. Run the ping commands from PF0 to PF0:

    First BM Server Console

    root@worker1:~# ping -c 3 10.0.121.9
    PING 10.0.121.9 (10.0.121.9) 56(84) bytes of data.
    64 bytes from 10.0.121.9: icmp_seq=1 ttl=62 time=0.896 ms
    64 bytes from 10.0.121.9: icmp_seq=2 ttl=62 time=0.241 ms
    64 bytes from 10.0.121.9: icmp_seq=3 ttl=62 time=0.258 ms
    
  2. Try to ping the PF0 on nodes 3 and 4. Run the ping commands from PF0 to PF0:

    First BM Server Console

    root@worker1:~# ping -c 3 10.0.122.1
    PING 10.0.122.1 (10.0.122.1) 56(84) bytes of data.
    From 10.0.121.2 icmp_seq=1 Destination Host Unreachable
    From 10.0.121.2 icmp_seq=2 Destination Host Unreachable
    From 10.0.121.2 icmp_seq=3 Destination Host Unreachable
    
    --- 10.0.122.1 ping statistics ---
    3 packets transmitted, 0 received, +3 errors, 100% packet loss, time 2045ms
    
    root@worker1:~# ping -c 3 10.0.122.9
    PING 10.0.122.9 (10.0.122.9) 56(84) bytes of data.
    From 10.0.121.2 icmp_seq=1 Destination Host Unreachable
    From 10.0.121.2 icmp_seq=2 Destination Host Unreachable
    From 10.0.121.2 icmp_seq=3 Destination Host Unreachable
    
    --- 10.0.122.9 ping statistics ---
    3 packets transmitted, 0 received, +3 errors, 100% packet loss, time 2067ms
    
    

     

This ping operation should fail due to the network isolation implemented in HBN using different VLANs, VNIs and VRFs.

 Done.

Authors


BK.jpg

Boris Kovalev

Boris Kovalev has worked for the past several years as a Solutions Architect, focusing on NVIDIA Networking/Mellanox technology, and is responsible for complex machine learning, Big Data and advanced VMware-based cloud research and design. Boris previously spent more than 20 years as a senior consultant and solutions architect at multiple companies, most recently at VMware. He has written multiple reference designs covering VMware, machine learning, Kubernetes, and container solutions which are available at the NVIDIA Documents website.




NVIDIA, the NVIDIA logo, and BlueField are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries. Other company and product names may be trademarks of the respective companies with which they are associated.
2025 NVIDIA Corporation. All rights reserved.©













Last updated: