Troubleshooting and Optimization

23% of the NCP-AIO exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Installation and Deployment · Administration · Workload Management · Troubleshooting and Optimization

4.1 Troubleshooting Docker with GPUs

Official objective: “Troubleshoot Docker”

Selecting GPUs, runtime debug logs and lost GPU access.

Key points

  1. The device list is quoted twice: single quotes outside, double quotes inside.

    What NVIDIA says (1)

    “For example, '"device=2,3"' enumerates GPUs 2 and 3 to the container.”

    — NVIDIA Container Toolkit: Specialized Configurations with Docker

  2. NVIDIA_VISIBLE_DEVICES is the environment variable the NVIDIA Container Runtime reads to choose GPUs.

    What NVIDIA says (1)

    “none No GPU is accessible, but driver capabilities are enabled.”

    — NVIDIA Container Toolkit: Specialized Configurations with Docker

  3. Debug logs show what the runtime hook did, which usually reveals the root cause.

    What NVIDIA says (1)

    “Edit your runtime configuration under /etc/nvidia-container-runtime/config.toml and uncomment the debug=... line.”

    — NVIDIA Container Toolkit: Troubleshooting

  4. NVML (NVIDIA Management Library) is the API nvidia-smi uses. With systemd managing cgroups, a reload can strip the container's device access.

    What NVIDIA says (2)

    “For Docker, use cgroupfs as the cgroup driver for containers.”

    — NVIDIA Container Toolkit: Troubleshooting

    “On systems where systemd is used to manage the cgroups of the container, reloading the systemd unit files ( systemctl daemon-reload ) is sufficient to trigger container updates and cause a loss of GPU access.”

    — NVIDIA Container Toolkit: Troubleshooting

  5. SELinux (Security-Enhanced Linux) labels files and processes. This option stops labeling for the container.

    What NVIDIA says (2)

    “you might have to specify --security-opt=label=disable on the Docker or Podman command line”

    — NVIDIA Container Toolkit: Troubleshooting

    “However, using this option disables SELinux separation in the container, and the container runs in an unconfined type.”

    — NVIDIA Container Toolkit: Troubleshooting

Key terms: NVIDIA Container Toolkit NVML

Try it: NGC Container Flow

Practice 4.1 (5 questions)

4.2 Troubleshooting Fabric Manager

Official objective: “Troubleshoot the fabric manager service for NVLink and NVSwitch systems”

Service status, version match, logs and the not-ready error.

Key points

  1. Fabric Manager (FM) configures the NVSwitch fabric so GPUs can talk over NVLink. It runs as a systemd service.

    What NVIDIA says (1)

    “To check FM, for Linux based OS distributions, run the following command: sudo systemctl status”

    — NVIDIA Fabric Manager User Guide

  2. Without FM, the NVSwitch fabric is not set up, so new CUDA jobs are refused.

    What NVIDIA says (1)

    “A new CUDA job launch will fail with a cudaErrorSystemNotReady error.”

    — NVIDIA Fabric Manager User Guide

  3. A mismatch between FM and the driver is a common cause of FM start failures.

    What NVIDIA says (1)

    “NVIDIA Fabric Manager Package (same version as the Driver package).”

    — NVIDIA Fabric Manager User Guide

  4. Check this log first when FM fails to start or reports NVSwitch errors.

    What NVIDIA says (1)

    “Default Value LOG_FILE_NAME=/var/log/fabricmanager.log”

    — NVIDIA Fabric Manager User Guide

Key terms: Fabric Manager

Practice 4.2 (4 questions)

4.3 Troubleshooting BCM

Official objective: “Troubleshoot Base Command Manager”

CMDaemon logs and debug mode, provisioning states and status.

Key points

  1. CMDaemon is the BCM service on every node that carries out management commands.

    What NVIDIA says (1)

    “CMDaemon generates log messages in /var/log/cmdaemon from specific internal subsystems”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  2. Debug output is off by default because it grows the log quickly. Turn it off when done.

    What NVIDIA says (1)

    “A global debug mode can be enabled in CMDaemon using cmdaemonctl”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  3. The node-installer is the small system that provisions a node's image at boot.

    What NVIDIA says (1)

    “This state is entered from the INSTALLING state when the head node CMDaemon can no longer ping the node. It could indicate the node has crashed while running the node-installer.”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

  4. Provisioning copies software images to nodes. This command shows its queue and state.

    What NVIDIA says (1)

    “provisioningstatus Provisioning subsystem status Pending request: node001, node002”

    — NVIDIA Base Command Manager 11 Administrator Manual (PDF)

Key terms: Base Command Manager CMDaemon

Practice 4.3 (4 questions)

4.4 Troubleshooting Magnum IO

Official objective: “Troubleshoot Magnum IO components”

NCCL debug and interface selection, peer-to-peer tests and ACS for GDS.

Key points

  1. Magnum IO is NVIDIA's set of data-movement libraries, including NCCL and GPUDirect Storage. NCCL_DEBUG controls NCCL's debug output.

    What NVIDIA says (1)

    “The NCCL_DEBUG variable controls the debug information that is displayed from NCCL. This variable is commonly used for debugging.”

    — NCCL Environment Variables

  2. It takes a list of name prefixes. A ^ excludes interfaces.

    What NVIDIA says (1)

    “The NCCL_SOCKET_IFNAME variable specifies which IP interfaces to use for communication.”

    — NCCL Environment Variables

  3. Peer-to-peer (P2P) is a direct GPU-to-GPU path. Turning it off is a test, not a fix.

    What NVIDIA says (1)

    “The NCCL_P2P_DISABLE variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”

    — NCCL Environment Variables

  4. GDS moves data straight between storage and GPU memory. ACS forces that traffic up through the CPU.

    What NVIDIA says (1)

    “For optimal GDS performance, disable ACS. Note To list all of the PCI switches that have ACS enabled, issue /usr/local/cuda/gds/tools/gdscheck -p .”

    — GPUDirect Storage Best Practices Guide

Key terms: Magnum IO NCCL GPUDirect Storage Access Control Services

Try it: AllReduce Deep Dive GPUDirect Storage

Practice 4.4 (4 questions)

4.5 Storage performance

Official objective: “Troubleshoot storage performance”

GDS compatibility mode, IOMMU, PCIe topology and benchmarking.

Key points

  1. cuFile is the GDS API. Compatibility mode works without the direct path but is slower.

    What NVIDIA says (1)

    “In the /etc/cufile.json file, verify that allow_compat_mode is set to true . gdscheck -p displays whether the allow_compat_mode property is set to true .”

    — GPUDirect Storage Troubleshooting Guide

  2. IOMMU (input-output memory management unit) translates device addresses. Here it forces a slower path.

    What NVIDIA says (2)

    “When the IOMMU setting is enabled, PCIe traffic will be routed through the CPU root ports.”

    — GPUDirect Storage Best Practices Guide

    “Before you install GDS, you must disable IOMMU.”

    — GPUDirect Storage Best Practices Guide

  3. Fewer hops between GPU and NIC means higher throughput for GDS.

    What NVIDIA says (1)

    “The optimal path between a GPU and NIC will be one PCIe switch path designated as PIX. The least optimal path is designated as SYS”

    — GPUDirect Storage Configuration and Benchmarking Guide

  4. Benchmark with several thread counts. The number of I/O threads strongly affects results.

    What NVIDIA says (1)

    “the number of processes/threads generating IO is critical to determining maximum performance. The gdsio tool provides for specifying random reads or random writes”

    — GPUDirect Storage Configuration and Benchmarking Guide

Key terms: High-speed storage GPUDirect Storage IOMMU gdsio

Try it: GPUDirect Storage Storage Bottleneck

Practice 4.5 (4 questions)

4.6 Troubleshooting NGC container deployment

Official objective: “Troubleshoot the deployment of a container from NGC”

GPU Operator pod failures, lost GPU access and readiness checks.

Key points

  1. Operand pods wait for the driver and toolkit pods. Fix the driver pod first.

    What NVIDIA says (2)

    “Note that the operand pods will only come up when the driver daemonset and toolkit pods come up successfully.”

    — NVIDIA GPU Operator Troubleshooting

    “kubectl logs -n gpu-operator nvidia-driver-daemonset-p97x5 -c nvidia-driver-ctr”

    — NVIDIA GPU Operator Troubleshooting

  2. The container's device access was removed. A fresh container gets it back. Then apply a lasting fix.

    What NVIDIA says (1)

    “You need to delete the container after the issue occurs. When it is restarted, either manually or automatically depending on whether you use a container orchestration platform, it regains access to the GPU.”

    — NVIDIA Container Toolkit: Troubleshooting

  3. Triton is NVIDIA's inference server. Its readiness endpoint returns 200 when ready.

    What NVIDIA says (1)

    “The HTTP request returns status 200 if Triton is ready and non-200 if it is not ready.”

    — Quickstart — NVIDIA Triton Inference Server

  4. CrashLoopBackOff means Kubernetes keeps restarting a failing container. A hook error points to the container toolkit.

    What NVIDIA says (2)

    “State: Waiting Reason: CrashLoopBackOff”

    — NVIDIA GPU Operator Troubleshooting

    “When using the NVIDIA Container Runtime Hook (that is, the Docker --gpus flag or the NVIDIA Container Runtime in legacy mode) to inject requested GPUs and driver libraries into a container”

    — NVIDIA Container Toolkit: Troubleshooting

Key terms: NVIDIA GPU Operator NGC NVIDIA Container Toolkit NVML Triton Inference Server

Try it: Kubernetes GPU Ops

Practice 4.6 (4 questions)