2.4 Administering Kubernetes for GPUs

NCP-AIO · Administration (23% of the exam) · Official objective: “Administer Kubernetes”

Checking the GPU Operator, NFD and GPU time-slicing.

Key points

  1. The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.

    What NVIDIA says (1)

    “Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”

    — GPU Operator: Installing the NVIDIA GPU Operator

  2. NFD (Node Feature Discovery) labels nodes with their hardware features. The Operator needs it on every node. If NFD already runs, deploying a second copy must be turned off.

    What NVIDIA says (1)

    “If NFD is already running in the cluster, then you must disable deploying NFD when you install the Operator.”

    — GPU Operator: Installing the NVIDIA GPU Operator

  3. Time-slicing lets several pods take turns on one GPU. It shares the GPU but does not partition memory. MIG does.

    What NVIDIA says (1)

    “Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”

    — Time-Slicing GPUs in Kubernetes

  4. The device plugin tells Kubernetes how many GPUs a node has. With time-slicing, it advertises several replicas of each GPU.

    What NVIDIA says (1)

    “The NVIDIA GPU Operator enables oversubscription of GPUs through a set of extended options for the NVIDIA Kubernetes Device Plugin .”

    — Time-Slicing GPUs in Kubernetes

Key terms

Try it

Sample question

After installing the GPU Operator, how do you confirm it is healthy before running GPU workloads?

Show the answer

Answer: kubectl get pods -n gpu-operator and check all pods are Running

The GPU Operator runs several pods: driver, toolkit, device plugin, feature discovery and validators. All should be Running or Completed.

What NVIDIA says (1)

“Check that all GPU Operator pods are running: $ kubectl get pods -n gpu-operator”

— GPU Operator: Installing the NVIDIA GPU Operator

Practice 2.4 (4 questions) Full Administration guide

← 2.3 Administering Run:ai · 2.5 Configuring MIG →