4.6 Troubleshooting NGC container deployment
GPU Operator pod failures, lost GPU access and readiness checks.
Key points
Operand pods wait for the driver and toolkit pods. Fix the driver pod first.
What NVIDIA says (2)
“Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”
“kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”
The container's device access was removed. A fresh container gets it back. Then apply a lasting fix.
What NVIDIA says (1)
“You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”
Triton is NVIDIA's inference server. Its readiness endpoint returns 200 when ready.
What NVIDIA says (1)
“The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”
CrashLoopBackOff means Kubernetes keeps restarting a failing container. A hook error points to the container toolkit.
What NVIDIA says (2)
“State: Waiting Reason: CrashLoopBackOff”
“When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”
Key terms
- NVIDIA GPU Operator: A Kubernetes operator that installs and manages the GPU driver, container toolkit, device plugin and related parts.
- NGC: NVIDIA's catalog of GPU-optimized software; its container registry is nvcr.io.
- NVIDIA Container Toolkit: Software whose runtime hook gives containers access to host GPUs.
- NVML: The NVIDIA Management Library, the API behind nvidia-smi; NVML errors in a container mean GPU access is broken.
- Triton Inference Server: NVIDIA's open-source inference server, shipped as an NGC container.
Try it
Sample question
A GPU Operator driver pod is not Ready and other operand pods are stuck in Init. What do you check first?
Show the answer
Answer: The nvidia-driver-daemonset pod logs (kubectl logs -n gpu-operator ... -c nvidia-driver-ctr)
Operand pods wait for the driver and toolkit pods. Fix the driver pod first.
What NVIDIA says (2)
“Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”
“kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”
Practice 4.6 (4 questions) Full Troubleshooting and Optimization guide