K8sNet course / Module 3
Kubernetes networking for AI clusters
0%
0 of 5 cases closed
Operators
Deploy and version the NVIDIA operator stack correctly, push RoCE and GPUDirect settings into NIC firmware safely, and align NIC to GPU so a training job runs at full speed.
- 3.1 Network Operator and the NicClusterPolicyThe policy that was applied and never became anything
- 3.2 Version matrices, upgrades and heterogeneous clustersThe bookmark that says 1.31 is fine
- 3.3 NIC firmware config in Kubernetes: RoCE, PFC and rebootsThe change request that says no outage expected
- 3.4 GPUDirect RDMA: dma-buf, nvidia-peermem and GDSlsmod says nothing, so GPUDirect must be broken
- 3.5 NUMA alignment with Topology Manager and CPU ManagerSix nodes are slower and nobody can say which
- Checkpoint