3.4 Virtualizing GPUs
MIG, time-slicing and vGPU: how each shares a GPU and what isolation it gives.
Key points
NVIDIA says MIG lets supported GPUs be securely partitioned into up to seven separate GPU instances. MIG mode is not enabled by default. An administrator enables it, then creates instances from profiles such as 1g.10gb on H100. A profile name gives the instance's share of compute (1g = one compute slice) and its memory size (10gb).
What NVIDIA says (3)
“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”
“By default, MIG mode is not enabled on the GPU.”
“Table 10 GPU Instance Profiles on H100”
A virtual machine (VM) is a software-defined computer that runs on a hypervisor. NVIDIA vGPU software is installed on a physical GPU in a server and creates virtual GPUs that can be shared across multiple VMs. NVIDIA says this raises GPU utilization while keeping management and security centralized.
What NVIDIA says (2)
“NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”
“Installed on a physical GPU in a cloud or enterprise data center server, NVIDIA vGPU software creates virtual GPUs that can be shared across multiple virtual machines, scaling GPU utilization and enhancing ROI while centralizing manageability and security for enterprise IT.”
Multi-Instance GPU (MIG) partitions one physical GPU into separate GPU instances. NVIDIA says each instance gets dedicated compute and memory resources. Each instance's processors have isolated paths through the whole memory system. MIG also provides fault isolation between clients such as VMs, containers or processes.
What NVIDIA says (2)
“The Multi-Instance GPU (MIG) User Guide explains how to partition supported NVIDIA GPUs into multiple isolated instances, each with dedicated compute and memory resources.”
“MIG can partition available GPU compute resources (including streaming multiprocessors or SMs, and GPU engines such as copy engines or decoders), to provide a defined quality of service (QoS) with fault isolation for different clients such as VMs, containers or processes.”
Time-slicing lets workloads on an oversubscribed GPU take turns. Oversubscribed means more workloads than GPUs. NVIDIA says that, unlike MIG, there is no memory or fault isolation between the shared replicas. Time-slicing trades MIG's isolation for the ability to share a GPU among more users.
What NVIDIA says (2)
“GPU time-slicing enables workloads that are scheduled on oversubscribed GPUs to interleave with one another.”
“Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”
Key terms
- Multi-Instance GPU: Splitting one GPU into up to seven isolated instances, each with its own compute and memory.
- Time-slicing: Sharing a GPU by letting workloads take turns, without the memory and fault isolation of MIG.
- Virtual GPU: NVIDIA software that creates virtual GPUs shared across virtual machines.
- Streaming multiprocessor: The GPU compute unit that runs threads; a GPU has many SMs, and MIG divides them between instances.
Try it
Sample question
Up to how many separate GPU instances can MIG create on one supported GPU, and what must you do first?
Show the answer
Answer: Up to seven. First enable MIG mode on the GPU (it is off by default).
NVIDIA says MIG lets supported GPUs be securely partitioned into up to seven separate GPU instances. MIG mode is not enabled by default. An administrator enables it, then creates instances from profiles such as 1g.10gb on H100. A profile name gives the instance's share of compute (1g = one compute slice) and its memory size (10gb).
What NVIDIA says (3)
“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”
“By default, MIG mode is not enabled on the GPU.”
“Table 10 GPU Instance Profiles on H100”