GPU-Live/network-operator
root 59fbb40e28 net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
..
charts net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
crds net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
hostdev net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
ipiob net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
sriov net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
templates net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
.helmignore net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
Chart.lock net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
Chart.yaml net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
README.md net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00
values.yaml net_operator, nccl, nerdctl init 2025-12-02 17:20:54 +09:00

README.md

Nvidia Network Operator Helm Chart

Nvidia Network Operator Helm Chart provides an easy way to install and manage the lifecycle of Nvidia network operator.

Nvidia Network Operator

Nvidia Network Operator leverages Kubernetes CRDs and Operator SDK to manage Networking related Components in order to enable Fast networking, RDMA and GPUDirect for workloads in a Kubernetes cluster. Network Operator works in conjunction with GPU-Operator to enable GPU-Direct RDMA on compatible systems.

The Goal of Network Operator is to manage all networking related components to enable execution of RDMA and GPUDirect RDMA workloads in a kubernetes cluster including:

  • Mellanox Networking drivers to enable advanced features
  • Kubernetes device plugins to provide hardware resources for fast network
  • Kubernetes secondary network for Network intensive workloads

Documentation

For more information please visit the official documentation.

Additional components

Node Feature Discovery

Nvidia Network Operator relies on the existance of specific node labels to operate properly. e.g label a node as having Nvidia networking hardware available. This can be achieved by either manually labeling Kubernetes nodes or using Node Feature Discovery to perform the labeling.

To allow zero touch deployment of the Operator we provide a helm chart to be used to optionally deploy Node Feature Discovery in the cluster. This is enabled via nfd.enabled chart parameter.

SR-IOV Network Operator

Nvidia Network Operator can operate in unison with SR-IOV Network Operator to enable SR-IOV workloads in a Kubernetes cluster. We provide a helm chart to be used to optionally deploy SR-IOV Network Operator in the cluster. This is enabled via sriovNetworkOperator.enabled chart parameter.

SR-IOV Network Operator can work in conjuction with IB Kubernetes to use InfiniBand PKEY Membership Types

For more information on how to configure SR-IOV in your Kubernetes cluster using SR-IOV Network Operator refer to the project's github.

QuickStart

System Requirements

  • RDMA capable hardware: Mellanox ConnectX-5 NIC or newer.
  • NVIDIA GPU and driver supporting GPUDirect e.g Quadro RTX 6000/8000 or Tesla T4 or Tesla V100 or Tesla V100. (GPU-Direct only)
  • Operating Systems: Ubuntu 20.04 LTS.

NOTE: ConnectX-6 Lx is not supported.

Tested Network Adapters

The following Network Adapters have been tested with network-operator:

  • ConnectX-5
  • ConnectX-6 Dx

Prerequisites

  • Kubernetes v1.17+
  • Helm v3.5.3+
  • Ubuntu 20.04 LTS

Install Helm

Helm provides an install script to copy helm binary to your system:

$ curl -fsSL -o get_helm.sh https://raw.githubusercontent.com/helm/helm/master/scripts/get-helm-3
$ chmod 500 get_helm.sh
$ ./get_helm.sh

For additional information and methods for installing Helm, refer to the official helm website

Deploy Network Operator

# Add Repo
$ helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
$ helm repo update

# Install Operator
$ helm install -n network-operator --create-namespace --wait network-operator nvidia/network-operator

# View deployed resources
$ kubectl -n network-operator get pods

Deploy Network Operator without Node Feature Discovery

By default the network operator deploys Node Feature Discovery (NFD) in order to perform node labeling in the cluster to allow proper scheduling of Network Operator resources. If the nodes where already labeled by other means (either deployed from upstream or deployed within another deployment), it is possible to disable the deployment of NFD by setting nfd.enabled=false chart parameter and make sure that the installed version is v0.13.2 or newer and has NodeFeatureApi enabled.

Deploy NFD from upstream with NodeFeatureApi enabled
$ export NFD_NS=node-feature-discovery
$ helm repo add nfd https://kubernetes-sigs.github.io/node-feature-discovery/charts
$ helm repo update
$ helm install nfd/node-feature-discovery --namespace $NFD_NS --create-namespace --generate-name --set enableNodeFeatureApi='true'

For additional information , refer to the official NVD deployment with Helm

Deploy Network Operator without Node Feature Discovery
$ helm install --set nfd.enabled=false -n network-operator --create-namespace --wait network-operator nvidia/network-operator
Currently the following NFD labels are used:
Label Where
feature.node.kubernetes.io/pci-15b3.present Nodes bearing Nvidia Mellanox Networking hardware
nvidia.com/gpu.present Nodes bearing Nvidia GPU hardware

Note: The labels which Network Operator depends on may change between releases.

Note: By default the operator is deployed without an instance of NicClusterPolicy and MacvlanNetwork custom resources. The user is required to create it later with configuration matching the cluster or use chart parameters to deploy it together with the operator.

Deploy development version of Network Operator

To install development version of Network Operator you need to clone repository first and install helm chart from the local directory:

# Clone Network Operator Repository
$ git clone https://github.com/Mellanox/network-operator.git

# Update chart dependencies
$ cd network-operator/deployment/network-operator && helm dependency update

# Install Operator
$ helm install -n network-operator --create-namespace --wait network-operator ./

# View deployed resources
$ kubectl -n network-operator get pods

Deploy Network Operator with Admission Controller

The Admission Controller can be optionally included as part of the Network Operator installation process.
It has the capability to validate supported Custom Resource Definitions (CRDs), which currently include NicClusterPolicy and HostDeviceNetwork.
By default, the deployment of the admission controller is disabled. To enable it, you must set operator.admissionController.enabled to true.

Enabling the admission controller provides you with two options for managing certificates.
You can either utilize cert-manager for generating a self-signed certificate automatically, or you can provide your own self-signed certificate.

To use cert-manager, ensure that operator.admissionController.useCertManager is set to true. Additionally, make sure that you deploy cert-manager before initiating the Network Operator deployment.

If you prefer not to use cert-manager, set operator.admissionController.useCertManager to false, and then provide your custom certificate and key using operator.admissionController.certificate.tlsCrt and operator.admissionController.certificate.tlsKey.

NOTE: When using your own certificate, the certificate must be valid for <Release_Name>-webhook-service.< Release_Namespace>.svc, e.g. network-operator-webhook-service.network-operator.svc

NOTE: When deploying network operator with admission controller using helm, you need to append --wait to helm install and helm upgrade commands

Generating self-signed certificate using OpenSSL

To generate a self-signed SSL certificate valid for a specific hostname, you can use the openssl command-line tool. First, navigate to the directory where you want to store your certificate and key files. Then, run the following command:

SVCNAME="network-operator-webhook-service.network-operator.svc"
openssl req -x509 -nodes -batch -newkey rsa:2048 -keyout server.key -out server.crt -days 365 -addext "subjectAltName=DNS:$SVCNAME"

Replace SVCNAME with the SVC name follows this convention <Release_Name>-webhook-service.<Release_Namespace>.svc. This command will generate a new RSA key pair with 2048 bits and create a self-signed certificate (server.crt) and private key (server.key) that are valid for 365 days.

Upgrade

NOTE: Upgrade capabilities are limited now. Additional manual actions required when containerized OFED driver is used

Before starting the upgrade to a specific release version, please, check release notes for this version to ensure that no additional actions are required.

Check available releases

helm search repo nvidia/network-operator -l

NOTE: add --devel option if you want to list beta releases as well

Upgrade CRDs to compatible version

The network-operator helm chart contains a hook(pre-install, pre-upgrade) that will automatically upgrade required CRDs in the cluster. The hook is enabled by default. If you don't want to upgrade CRDs with helm automatically, you can disable auto upgrade by setting upgradeCRDs: false in the helm chart values. Then you can follow the guide below to download and apply CRDs for the concrete version of the network-operator.

It is possible to retrieve updated CRDs from the Helm chart or from the release branch on GitHub. Example bellow show how to download and unpack Helm chart for specified release and then apply CRDs update from it.

helm pull nvidia/network-operator --version <VERSION> --untar --untardir network-operator-chart

NOTE: --devel option required if you want to use the beta release

kubectl apply -f network-operator-chart/network-operator/crds \
              -f network-operator-chart/network-operator/charts/sriov-network-operator/crds

Prepare Helm values for the new release

Download Helm values for the specific release

helm show values nvidia/network-operator --version=<VERSION> > values-<VERSION>.yaml

Edit values-<VERSION>.yaml file as required for your cluster.

Apply Helm chart update

helm upgrade -n network-operator  network-operator nvidia/network-operator --version=<VERSION> -f values-<VERSION>.yaml --force

NOTE: --devel option required if you want to use the beta release

Chart parameters

In order to tailor the deployment of the network operator to your cluster needs, Chart parameters are available. See official documentation.