Skip to content
K8sNet course / Module 3

Kubernetes networking for AI clusters

0%

0 of 5 cases closed

Operators

Deploy and version the NVIDIA operator stack correctly, push RoCE and GPUDirect settings into NIC firmware safely, and align NIC to GPU so a training job runs at full speed.

  1. 3.1 Network Operator and the NicClusterPolicyThe policy that was applied and never became anything
  2. 3.2 Version matrices, upgrades and heterogeneous clustersThe bookmark that says 1.31 is fine
  3. 3.3 NIC firmware config in Kubernetes: RoCE, PFC and rebootsThe change request that says no outage expected
  4. 3.4 GPUDirect RDMA: dma-buf, nvidia-peermem and GDSlsmod says nothing, so GPUDirect must be broken
  5. 3.5 NUMA alignment with Topology Manager and CPU ManagerSix nodes are slower and nobody can say which
  6. Checkpoint