2.4 Administering Kubernetes for GPUs
Checking the GPU Operator, NFD and GPU time-slicing.
Key points
The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.
What NVIDIA says (1)
“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”
NFD (Node Feature Discovery) labels nodes with their hardware features. The Operator needs it on every node. If NFD already runs, deploying a second copy must be turned off.
What NVIDIA says (1)
“If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”
Time-slicing lets several pods take turns on one GPU. It shares the GPU but does not partition memory. MIG does.
What NVIDIA says (1)
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”
The device plugin tells Kubernetes how many GPUs a node has. With time-slicing, it advertises several replicas of each GPU.
What NVIDIA says (1)
“The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”
Key terms
- NVIDIA GPU Operator: A Kubernetes operator that installs and manages the GPU driver, container toolkit, device plugin and related parts.
- Node Feature Discovery: A Kubernetes add-on that labels nodes with their hardware features; the GPU Operator can deploy it.
- GPU time-slicing: Sharing a GPU between pods by taking turns, with no memory or fault isolation.
Try it
Sample question
After installing the GPU Operator, how do you confirm it is healthy before running GPU workloads?
Show the answer
Answer: kubectl get pods -n gpu-operator and check all pods are Running
The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.
What NVIDIA says (1)
“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”