4.6 Troubleshooting NGC container deployment

NCP-AIO · Troubleshooting and Optimization (23% of the exam) · Official objective: “Troubleshoot the deployment of a container from NGC”

GPU Operator pod failures, lost GPU access and readiness checks.

Key points

  1. Operand pods wait for the driver and toolkit pods. Fix the driver pod first.

    What NVIDIA says (2)

    “Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”

    — NVIDIA GPU Operator Troubleshooting

    “kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”

    — NVIDIA GPU Operator Troubleshooting

  2. The container's device access was removed. A fresh container gets it back. Then apply a lasting fix.

    What NVIDIA says (1)

    “You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”

    — NVIDIA Container Toolkit: Troubleshooting

  3. Triton is NVIDIA's inference server. Its readiness endpoint returns 200 when ready.

    What NVIDIA says (1)

    “The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”

    — Quickstart — NVIDIA Triton Inference Server

  4. CrashLoopBackOff means Kubernetes keeps restarting a failing container. A hook error points to the container toolkit.

    What NVIDIA says (2)

    “State: Waiting Reason: CrashLoopBackOff”

    — NVIDIA GPU Operator Troubleshooting

    “When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”

    — NVIDIA Container Toolkit: Troubleshooting

Key terms

Try it

Sample question

A GPU Operator driver pod is not Ready and other operand pods are stuck in Init. What do you check first?

Show the answer

Answer: The nvidia-driver-daemonset pod logs (kubectl logs -n gpu-operator ... -c nvidia-driver-ctr)

Operand pods wait for the driver and toolkit pods. Fix the driver pod first.

What NVIDIA says (2)

“Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”

— NVIDIA GPU Operator Troubleshooting

“kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”

— NVIDIA GPU Operator Troubleshooting

Practice 4.6 (4 questions) Full Troubleshooting and Optimization guide

← 4.5 Storage performance