3.4 Virtualizing GPUs

NCA-AIIO · AI Operations (22% of the exam) · Official objective: “Identify the key considerations for virtualizing accelerated infrastructure.”

MIG, time-slicing and vGPU: how each shares a GPU and what isolation it gives.

Key points

  1. NVIDIA says MIG lets supported GPUs be securely partitioned into up to seven separate GPU instances. MIG mode is not enabled by default. An administrator enables it, then creates instances from profiles such as 1g.10gb on H100. A profile name gives the instance's share of compute (1g = one compute slice) and its memory size (10gb).

    What NVIDIA says (3)

    “allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”

    — MIG User Guide: Introduction

    “By default, MIG mode is not enabled on the GPU.”

    — MIG User Guide: Getting Started with MIG

    “Table 10 GPU Instance Profiles on H100”

    — MIG User Guide: Supported MIG Profiles

  2. A virtual machine (VM) is a software-defined computer that runs on a hypervisor. NVIDIA vGPU software is installed on a physical GPU in a server and creates virtual GPUs that can be shared across multiple VMs. NVIDIA says this raises GPU utilization while keeping management and security centralized.

    What NVIDIA says (2)

    “NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”

    — NVIDIA Virtual GPU Software Documentation

    “Installed on a physical GPU in a cloud or enterprise data center server, NVIDIA vGPU software creates virtual GPUs that can be shared across multiple virtual machines, scaling GPU utilization and enhancing ROI while centralizing manageability and security for enterprise IT.”

    — NVIDIA Virtual GPU Solutions

  3. Multi-Instance GPU (MIG) partitions one physical GPU into separate GPU instances. NVIDIA says each instance gets dedicated compute and memory resources. Each instance's processors have isolated paths through the whole memory system. MIG also provides fault isolation between clients such as VMs, containers or processes.

    What NVIDIA says (2)

    “The Multi-Instance GPU (MIG) User Guide explains how to partition supported NVIDIA GPUs into multiple isolated instances, each with dedicated compute and memory resources.”

    — NVIDIA Multi-Instance GPU User Guide

    “MIG can partition available GPU compute resources (including streaming multiprocessors or SMs, and GPU engines such as copy engines or decoders), to provide a defined quality of service (QoS) with fault isolation for different clients such as VMs, containers or processes.”

    — MIG User Guide: Introduction

  4. Time-slicing lets workloads on an oversubscribed GPU take turns. Oversubscribed means more workloads than GPUs. NVIDIA says that, unlike MIG, there is no memory or fault isolation between the shared replicas. Time-slicing trades MIG's isolation for the ability to share a GPU among more users.

    What NVIDIA says (2)

    “GPU time-slicing enables workloads that are scheduled on oversubscribed GPUs to interleave with one another.”

    — Time-Slicing GPUs in Kubernetes

    “Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”

    — Time-Slicing GPUs in Kubernetes

Key terms

Try it

Sample question

Up to how many separate GPU instances can MIG create on one supported GPU, and what must you do first?

Show the answer

Answer: Up to seven. First enable MIG mode on the GPU (it is off by default).

NVIDIA says MIG lets supported GPUs be securely partitioned into up to seven separate GPU instances. MIG mode is not enabled by default. An administrator enables it, then creates instances from profiles such as 1g.10gb on H100. A profile name gives the instance's share of compute (1g = one compute slice) and its memory size (10gb).

What NVIDIA says (3)

“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”

— MIG User Guide: Introduction

“By default, MIG mode is not enabled on the GPU.”

— MIG User Guide: Getting Started with MIG

“Table 10 GPU Instance Profiles on H100”

— MIG User Guide: Supported MIG Profiles

Practice 3.4 (4 questions) Full AI Operations guide

← 3.3 Monitoring GPUs