4.1 Troubleshooting Docker with GPUs
Selecting GPUs, runtime debug logs and lost GPU access.
Key points
The device list is quoted twice: single quotes outside, double quotes inside.
What NVIDIA says (1)
“For example, '"device=2,3"' enumerates GPUs 2 and 3 to the container.”
NVIDIA_VISIBLE_DEVICES is the environment variable the NVIDIA Container Runtime reads to choose GPUs.
What NVIDIA says (1)
“none No GPU is accessible, but driver capabilities are enabled.”
Debug logs show what the runtime hook did, which usually reveals the root cause.
What NVIDIA says (1)
“Edit your runtime configuration under /etc/nvidia-container-runtime/config.toml and uncomment the debug=... line.”
NVML (NVIDIA Management Library) is the API nvidia-smi uses. With systemd managing cgroups, a reload can strip the container's device access.
What NVIDIA says (2)
“For Docker, use cgroupfs as the cgroup driver for containers.”
“On systems where systemd is used to manage the cgroups of the container, reloading the systemd unit files ( systemctl daemon-reload ) is sufficient to trigger container updates and cause a loss of GPU access.”
SELinux (Security-Enhanced Linux) labels files and processes. This option stops labeling for the container.
What NVIDIA says (2)
“you might have to specify --security-opt=label=disable on the Docker or Podman command line”
“However, using this option disables SELinux separation in the container, and the container runs in an unconfined type.”
Key terms
- NVIDIA Container Toolkit: Software whose runtime hook gives containers access to host GPUs.
- NVML: The NVIDIA Management Library, the API behind nvidia-smi; NVML errors in a container mean GPU access is broken.
Try it
Sample question
Which Docker option starts a container that can see only GPUs 2 and 3?
Show the answer
Answer: --gpus '"device=2,3"'
The device list is quoted twice: single quotes outside, double quotes inside.
What NVIDIA says (1)
“For example, '"device=2,3"' enumerates GPUs 2 and 3 to the container.”
Practice 4.1 (5 questions) Full Troubleshooting and Optimization guide
← 3.7 Containers from NGC · 4.2 Troubleshooting Fabric Manager →