NCA-AIIO hands-on labs
Guided labs in the Aegis lab simulator. They use realistic simulated command output, not real hardware. Each lab is tagged to official objectives and cites the NVIDIA source for the concept it practises.
NVLink Topology: simulator lab
Inspect the GPU-to-GPU topology of an 8-GPU server and see which pairs talk over NVLink.
Try this
- Read the topology matrix.
- Find the fastest path between two GPUs.
What NVIDIA says (10)
“is a direct GPU-to-GPU interconnect that scales multi-GPU input/output (IO) in the server.”
“NV# = Connection traversing a bonded set of # NVLinks”
“every GPU has four NVLinks that connect to two of the NVSwitches, and five NVLinks that connect to the remaining two NVSwitches”
“nvidia-smi nvlink --errorcounters Get the Nvlink error counters”
“./build/all_reduce_perf -b 8 -e 128M -f 2 -g 8”
“Using this bus bandwidth, we can compare it with the hardware peak bandwidth, independently of the number of ranks used.”
“PHB = Connection traversing PCIe as well as a PCIe Host Bridge”
“INFO - Prints debug information.”
“variable disables the peer to peer (P2P) transport, which uses CUDA direct access between GPUs, using NVLink or PCI.”
“The nvidia-smi nvlink command can provide additional details on NVLink errors, and connection information on the links.”
MIG Partitioning: simulator lab
Enable Multi-Instance GPU and create isolated GPU instances.
Try this
- Enable MIG mode.
- Create instances and list them.
What NVIDIA says (9)
“allows GPUs (starting with NVIDIA Ampere architecture) to be securely partitioned into up to seven separate GPU Instances for CUDA applications”
“By default, MIG mode is not enabled on the GPU.”
“Without creating GPU instances (and corresponding compute instances), CUDA workloads cannot be run on the GPU.”
“MIG 1g.10gb 1/8 1/7 1 NVDEC /1 JPEG /0 OFA 1/8 1 7”
“Once the GPU instances are created, you need to create the corresponding Compute Instances (CI).”
“MIG 3g.20gb Device 0: (UUID: MIG-c7384736-a75d-5afc-978f-d2f1294409fd)”
“CUDA_VISIBLE_DEVICES=MIG-c7384736-a75d-5afc-978f-d2f1294409fd ./BlackScholes”
“sudo nvidia-smi mig -dci && sudo nvidia-smi mig -dgi”
“the created MIG devices are not persistent across system reboots”
ECC Error Lifecycle: simulator lab
Follow an ECC error from DCGM counters to the Xid 48 recovery workflow.
Try this
- Read field 311 in DCGM.
- Pick the recovery action for Xid 48.
What NVIDIA says (8)
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
“DCGM_FI_DEV_ECC_SBE_VOL_TOTAL 310 Total single bit volatile ECC errors”
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors”
“On GPUs that support row remapping, starting with NVIDIA® Ampere archtecture GPUs, these events provide details on row remapper activity.”
“48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU. This is also reported back to the user application. A GPU reset or node reboot is needed to clear this error.”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
XID Fault Drill: simulator lab
Diagnose an NVLink fault from Xid 74 and decide the recovery step.
Try this
- Read the Xid in the kernel log.
- Decide whether to re-seat links and run diagnostics.
What NVIDIA says (11)
“This event is logged when the GPU detects that a problem with a connection from the GPU to another GPU or NVSwitch over NVLink.”
“Check link mechanical connections and re-seat if a field resolution is required. Run diags if issue persists.”
“48 ROBUST_CHANNEL_GPU_ECC_DBE Double Bit ECC Error”
“A GPU reset or node reboot is needed to clear this error.”
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors”
“This event is logged when the GPU driver attempts to access the GPU over its PCI Express connection and finds that the GPU is not accessible.”
“Reviewing system event logs and kernel PCI event logs may provide additional indications of the source of the link failures.”
“Trigger a reset of one or more GPUs.”
“79 ROBUST_CHANNEL_GPU_HAS_FALLEN_OFF_THE_BUS GPU has fallen off the bus YES YES YES YES RESTART_BM”
“For example, if a GPU fails, another GPU connected to it over NVLink may report an Xid 74 simply because the link went down as a result.”
“79 ROBUST_CHANNEL_GPU_HAS_FALLEN_OFF_THE_BUS”
CUDA Stack Verification: simulator lab
Check that the driver and CUDA versions on a node can run a given application, using CUDA compatibility rules.
Try this
- Read the driver and CUDA versions from simulated nvidia-smi output.
- Decide whether forward compatibility is needed.
What NVIDIA says (8)
“CUDA Compatibility helps bridge that gap by defining supported ways to run newer CUDA software on existing driver installations, within documented limits.”
“NVIDIA-SMI 535.86.10 Driver Version: 535.86.10 CUDA Version: 12.2”
“Backwards compatibility ensures that a newer NVIDIA driver can be used with an older CUDA Toolkit.”
“Minor version and forward compatibility ensure that an older NVIDIA driver can be used with a newer CUDA Toolkit.”
“Drivers have always been backwards compatible with CUDA.”
“It’s mainly intended to support applications built on newer CUDA Toolkits to run on systems installed with an older NVIDIA Linux GPU driver from different major release families.”
“NVIDIA creates an updated set of Docker containers for the frameworks monthly.”
“where deep learning frameworks are tuned, optimized, tested, and containerized for your use.”
NGC Container Flow: simulator lab
Pull a GPU-optimized container from the NGC catalog and run it with the NVIDIA Container Toolkit.
Try this
- Pull a framework container.
- Run it with GPU access and confirm the GPUs are visible.
What NVIDIA says (6)
“The NGC Catalog consists of containers, pretrained models, Helm charts for Kubernetes deployments, and industry-specific AI toolkits with software development kits (SDKs).”
“It currently includes: The NVIDIA Container Runtime ( nvidia-container-runtime ) The NVIDIA Container Toolkit CLI ( nvidia-ctk )”
“sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi”
“As of Docker release 19.03, NVIDIA GPUs are natively supported as devices in the Docker runtime.”
“docker run --gpus all -it --rm –v local_dir:container_dir nvcr.io/nvidia/tensorflow:<xx.xx>-tf2-py3”
“nvidia-smi dmon”
Distributed Training (DDP): simulator lab
Run a simulated multi-GPU training job and watch weights being adjusted over many iterations.
Try this
- Start a distributed training run.
- Compare throughput as GPUs are added.
What NVIDIA says (9)
“Instead, the error is propagated back through the network’s layers, and the model must adjust its weights and try again.”
“Data Parallelism (DP) replicates the model across multiple GPUs.”
“is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”
“NCCL implements both collective communication and point-to-point send/receive primitives.”
“Data batches are evenly distributed between GPUs and the data-parallel GPUs process them independently.”
“it sums the gradients of all model copies using all-reduce communication collectives.”
“The AllReduce operation performs reductions on data (for example, sum, min, max) across devices and stores the result in the receive buffer of every rank.”
“Distributed Data Parallelism (DDP) keeps the model copies consistent by synchronizing parameter gradients across data-parallel GPUs before each parameter update.”
“bottleneck, limiting the performance and scalability of training and inference.”
AllReduce Deep Dive: simulator lab
Run an all-reduce across GPUs and nodes and see how NCCL uses NVLink inside a node and the network between nodes.
Try this
- Run all-reduce on one node, then two.
- Compare bandwidth for each path.
What NVIDIA says (11)
“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”
“NCCL implements both collective communication and point-to-point send/receive primitives.”
“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”
“INFO - Prints debug information.”
“A ring would do that operation in an order which follows the ring”
“variable defines which algorithms NCCL will use.”
“we need 2(n-1) data transfers (x number of elements) to perform an allReduce operation.”
“B = S/t * (2*(n-1)/n) = algbw * (2*(n-1)/n)”
“Using this bus bandwidth, we can compare it with the hardware peak bandwidth, independently of the number of ranks used.”
“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”
“NCCL will instead fall back to another available transport such as IP sockets.”
InfiniBand Fabric: simulator lab
Check InfiniBand port state and fabric health, including the Subnet Manager.
Try this
- Read simulated ibstat output.
- Find the port that is down.
What NVIDIA says (10)
“InfiniBand is a high-performance, low latency, RDMA capable networking technology”
“All InfiniBand-compliant ULPs require a proper operation of a Subnet Manager (SM) running on the InfiniBand fabric, at all times.”
“ibstat is a binary which displays basic information obtained from the local IB driver. Output includes LID, SMLID, port state, link width active, and port physical state.”
“Queries InfiniBand ports’ performance and error counters.”
“Calculates the BW of RDMA write between a pair of machines. One acts as a server and the other as a client.”
“Output includes LID, SMLID, port state, link width active, and port physical state.”
“Scans the fabric using directed route packets and extracts all the available information regarding its connectivity and devices.”
“Links in INIT state and unresponsive links detection Counters fetch Error counters check Routing checks Link width and speed checks”
“Link width and speed checks”
“ibstat is a binary which displays basic information obtained from the local IB driver.”
RoCEv2 + PFC/ECN: simulator lab
Configure lossless RoCE: MTU, Priority Flow Control (PFC) and ECN, then diagnose a PFC pause storm.
Try this
- Enable PFC on the RoCE priority.
- Check ECN.
- Read pause counters during a storm.
What NVIDIA says (13)
“the BlueField-3 SuperNIC provides best-in-class remote direct-memory access over converged Ethernet (RoCE) network connectivity between GPU servers”
“The regular Ethernet MTU applies on the RoCE frame.”
“RDMA over Converged Ethernet (RoCE) is a mechanism to provide this efficient data transfer with very low latencies on lossless Ethernet networks.”
“In order to function reliably, RoCE requires a form of flow control.”
“The normal and optimal way to use RoCE is to use Priority Flow Control (PFC). To use PFC, it must be enabled on all endpoints and switches in the flow path.”
“For example, PFC can provide lossless service for the RoCE traffic and best-effort service for the standard Ethernet traffic.”
“It allows reliable communication by notifying all ends of communication when congestion occurs. This is done without dropping packets.”
“cat /sys/class/net/<interface>/ecn/<protocol>/enable/X”
“Calculates the BW of RDMA write between a pair of machines.”
“Several ingress and egress counters per priority are supported. Run ethtool -S to get the full list of port counters.”
“prio4_tx_pause: 26832”
“PFC storm prevention enables toggling between default and auto modes.”
“The stall prevention timeout is configured to 8 seconds by default. Auto mode sets the stall prevention timeout to be 100 msec.”
NCCL Fallback Drill: simulator lab
Find out why NCCL fell back from InfiniBand to IP sockets, and fix it.
Try this
- Read the NCCL debug output.
- Find the setting that disabled the IB/RoCE transport.
What NVIDIA says (7)
“The NCCL_IB_DISABLE variable prevents the IB/RoCE transport from being used by NCCL.”
“NCCL will instead fall back to another available transport such as IP sockets.”
“It supports a variety of interconnect technologies including PCIe, NVLINK, InfiniBand Verbs, and IP sockets.”
“INFO - Prints debug information.”
“ibstat is a binary which displays basic information obtained from the local IB driver.”
“The NCCL_IB_HCA variable specifies which Host Channel Adapter (RDMA) interfaces to use for communication.”
“Using this bus bandwidth, we can compare it with the hardware peak bandwidth, independently of the number of ranks used.”
Storage Bottleneck: simulator lab
Find a storage bottleneck that starves GPUs, cache the dataset on local NVMe, and offload preprocessing with DALI.
Try this
- Spot the I/O bottleneck in the simulated metrics.
- Stage data to local NVMe and compare.
What NVIDIA says (8)
“The key I/O operation in DL training is re-read.”
“In addition, the DGX H100 system provides local NVMe storage that can also be used for caching or staging data.”
“bottleneck, limiting the performance and scalability of training and inference.”
“It is not just that data is read, but it must be reused again and again due to the iterative nature of DL training.”
“Ideally, data is cached during the first read of the dataset, so data does not have to be retrieved across the network.”
“Reading files from cache can be an order of magnitude faster than from remote storage.”
“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the”
“NVIDIA GPUDirect Storage® (GDS) provides a way to read data from the remote filesystem or local NVMe directly into GPU memory providing higher sustained I/O performance”
GPUDirect Storage: simulator lab
Compare a normal file read through CPU memory with a GPUDirect Storage read straight into GPU memory.
Try this
- Run both read paths.
- Compare CPU load and throughput.
What NVIDIA says (11)
“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which avoids a bounce buffer through the CPU.”
“an extra copy through a bounce buffer in the CPU is necessary, which introduces latency and lowers effective bandwidth.”
“GPUDirect® Storage (GDS) enables a direct data path for direct memory access (DMA) transfers between GPU memory and storage, which”
“The direct data path that GDS provides relies on the availability of file system drivers that are enabled with”
“/usr/local/cuda-<x>.<y>/gds/tools/gdscheck.py -p”
“Creating this direct path involves distributed file systems such as NFSoRDMA, DDN EXAScaler parallel file system solutions (based on the Lustre file system), Amazon FSx for Lustre, and WekaFS”
“compatibility mode is available for unsupported configurations that maps IO operations to a fallback path.”
“tool included when GDS is installed, gdsio, is covered and its use demonstrated.”
“xfer_type : 0 - Storage -> GPU ( GDS ) 1 - Storage -> CPU 2 - Storage -> CPU -> GPU”
“xfer_type : 0 - Storage -> GPU ( GDS )”
“avoids a bounce buffer through the CPU. Using this direct path can relieve system bandwidth bottlenecks and decrease”
DCGM Monitoring: simulator lab
Collect GPU metrics with DCGM Exporter and Prometheus, and set an alert.
Try this
- Scrape the /metrics endpoint.
- Write an alert on double-bit ECC errors.
What NVIDIA says (10)
“DCGM Exporter is written in Go and exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”
“Address of listening http server. Default: “:9400””
“curl localhost:9400/metrics”
“DCGM_FI_DEV_ECC_DBE_VOL_TOTAL 311 Total double bit volatile ECC errors”
“exposes GPU metrics at an HTTP endpoint ( /metrics ) for monitoring solutions such as Prometheus.”
“serviceMonitorSelectorNilUsesHelmValues: false”
“To add a dashboard for DCGM, you can use a standard dashboard that NVIDIA has made available, which can also be customized.”
“https://grafana.com/grafana/dashboards/12239”
“This event is logged when the GPU detects that an uncorrectable error occurs on the GPU.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
Slurm Scheduler: simulator lab
Submit and inspect GPU jobs on a Slurm cluster managed with Base Command Manager.
Try this
- Submit a containerized GPU job with Pyxis.
- Validate nodes with an NCCL test.
- Drain and resume a node.
What NVIDIA says (10)
“NVIDIA Base Command Manager streamlines cluster provisioning, workload management, and infrastructure monitoring.”
“Integrate with tools, like Slurm or NVIDIA Run:ai , to support traditional HPC or AI and analytics workloads across bare-metal and containerized environments.”
“Slurm is a classic workload manager used to orchestrate complex workloads in a multi-node, batch-style, compute environment”
“#SBATCH --container-image nvcr.io\#nvidia/pytorch:21.12-py3”
“the srun command will be in the queue until the nodes become available”
“--container-image=[USER@][REGISTRY#]IMAGE[:TAG]|PATH”
“srun_exports: NCCL_DEBUG=INFO”
“The validation playbook will verify that Pyxis and Enroot can run GPU jobs”
“Nodes which fail this check will be automatically drained in Slurm to prevent jobs run”
“This tool will run periodically on idle nodes to validate that the hardware and software is set up as expected.”
Kubernetes GPU Ops: simulator lab
Check the GPU Operator, request GPUs in Kubernetes, debug a Pending pod, and spot time-slicing oversubscription.
Try this
- Request nvidia.com/gpu in a pod spec.
- Read why a pod is Pending.
- Run the CUDA sample workload after maintenance.
What NVIDIA says (7)
“With the daemonset deployed, NVIDIA GPUs can now be requested by a container using the nvidia.com/gpu resource type”
“Check that all GPU Operator pods are running:”
“within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU.”
“nvidia.com/gpu: 1 # requesting 1 GPU”
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas, but for some workloads this is better than not being able to share at all.”
“WORKFLOW_XID_48 Data Center Recovery Action Solo: RESET_GPU”
“Verification: Running Sample GPU Applications”
AI, ML & DL Foundations: simulator lab
Sort examples into AI, machine learning and deep learning, and see why GPUs suit deep learning.
Try this
- Classify five workloads.
- Explain why a deep learning workload benefits from a GPU.
What NVIDIA says (9)
“Deep learning is a subset of machine learning, with the difference that DL algorithms can automatically learn representations from data such as images, video, or text, without introducing human domain knowledge.”
“a GPU is designed to excel at executing thousands of threads in parallel, trading off lower single-thread performance to achieve much greater total throughput.”
“In contrast, a GPU is composed of hundreds of cores that can handle thousands of threads simultaneously.”
“As a subset of AI, machine learning in its most elemental form uses algorithms to parse data, learn from it, and then make predictions or determinations about something in the real world.”
“GPUs are specialized for highly parallel computations and devote more transistors to data processing units, while CPUs dedicate more transistors to data caching and flow control.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“Architecturally, the CPU is composed of just a few cores with lots of cache memory that can handle a few software threads at a time.”
“This parallelism maps naturally to GPUs , providing a significant computation speedup over CPU-only training”
“Tensor Cores were introduced in the NVIDIA Volta™ GPU architecture to accelerate matrix multiply and accumulate operations for machine learning and scientific applications.”
Training vs Inference Serving: simulator lab
Compare a training job with an inference service: what each needs and how latency and throughput are traded.
Try this
- Profile a training step and an inference request.
- Change batch size and watch latency and throughput.
What NVIDIA says (11)
“That’s why inference optimization techniques are so critical: they allow enterprises to optimize throughput and latency so they can meet their service level agreements across a variety of use cases.”
“Over many iterations, it converges on the correct weights, enabling it to reliably suggest accurate, useful code completions or translations.”
“Inference is the process where a trained AI model generates new outputs by reasoning and making predictions on new data — classifying inputs and applying learned knowledge in real time.”
“Triton Inference Server enables teams to deploy any AI model from multiple deep learning and machine learning frameworks, including TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL, and more.”
“Concurrent model execution Dynamic batching”
“is an SDK for optimizing deep learning inference on NVIDIA GPUs.”
“optimizes inference using quantization, layer and tensor fusion, and kernel tuning techniques.”
“It takes trained models from frameworks such as PyTorch and ONNX and compiles them into engines, which are optimized executable artifacts for a specific deployment configuration.”
“TensorRT supports mixed precision”
“TensorRT includes libraries that optimize neural network models trained on all major frameworks, calibrate them for lower precision with high accuracy”
“Inference can’t happen without training.”
NVIDIA AI Software Stack: simulator lab
Walk the NVIDIA AI software stack from driver and CUDA up to libraries, containers and inference services.
Try this
- List which layer each tool belongs to.
- Match a workload to the NVIDIA solution that serves it.
What NVIDIA says (13)
“to enable any computational workload to use the throughput capability of GPUs independent of graphics APIs.”
“Libraries like cuBLAS, cuFFT, cuDNN, and CUTLASS are just a few examples of libraries that help developers avoid reimplementing well-established algorithms.”
“NVIDIA CUDA-X™, built on CUDA, is a collection of libraries that deliver dramatically higher performance across application domains, including AI and HPC.”
“The NVIDIA CUDA Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives for deep neural networks.”
“NCCL provides the following collective communication primitives: AllReduce Broadcast Reduce AllGather ReduceScatter”
“NCCL implements both collective communication and point-to-point send/receive primitives.”
“All of them also include the needed GPU libraries, configuration files, and tools to rebuild the container.”
“NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI”
“NVIDIA NIM microservices are a set of easy-to-use microservices for accelerating the deployment of foundation models on any cloud or data center”
“NVIDIA NIM™ provides prebuilt, optimized inference microservices for rapidly deploying the latest AI models on any NVIDIA-accelerated infrastructure—cloud, data center, workstation, and edge.”
“is a GPU-accelerated library for tabular data processing.”
“throughout the entire ML lifecycle—from exploratory analysis, data preparation, and model development to deployment, monitoring, and ongoing optimization.”
“DALI addresses the problem of the CPU bottleneck by offloading data preprocessing to the”
AI Infrastructure Planning: simulator lab
Size racks for DGX H100 systems: power per rack, rack density, and the extra racks needed when power or cooling is limited.
Try this
- Work out the power for one, two and four systems per rack.
- Decide how to deploy one SU when cooling is oversubscribed.
What NVIDIA says (13)
“System power consumption 10.2 kW max”
“However, rack densities can be customized to fit within the available power and cooling capacities at the data center.”
“The system is built upon building blocks of scalable units (SU), each containing 32 DGX H100 systems, which provides for rapid deployment of systems of multiple sizes.”
“DGX H100 systems are optimally deployed at a rack density of four systems per rack.”
“4 8 326.4 kW 40.8 kW”
“Oversubscribed cooling can sometimes be mitigated by lowering the rack density”
“2 16 326.4 kW 20.4 kW”
“In this example, the resource constraint in cooling is “paid for” using floor space.”
“In either case, the main benefit of aisle containment is the prevention of air recirculation from the hot aisle to the cold aisle”
“Four of the six power supplies must be energized for the system to operate.”
“Due to this requirement, the data center must minimally provide N+1 power, where N equals two power sources.”
“Each power source must be sized to support 50% of the total peak load.”
“The three main resource constraints in an air-cooled data center environment are power, cooling, and space.”
DPU Offload & Cloud vs On-Prem: simulator lab
Decide when to offload infrastructure work to a DPU, and compare cloud and on-prem costs over time.
Try this
- Identify host CPU work a DPU can take over.
- Compare up-front and ongoing costs for a multi-year workload.
What NVIDIA says (13)
“The NVIDIA DOCA™ Framework enables rapidly creating and managing applications and services on top of the BlueField networking platform, leveraging industry-standard APIs.”
“by harnessing the power of NVIDIA's BlueField data-processing units (DPUs) and SuperNICs”
“IT leaders should evaluate the total cost of ownership (TCO) over time and consider factors such as data storage, compute resources, and ongoing maintenance.”
“a CPU is designed to excel at executing a serial sequence of operations (called a thread) as fast as possible and can execute a few tens of these threads in parallel”
“It offloads demanding work that can bog down CPUs, processors that typically execute tasks in serial fashion.”
“The BlueField-3 platforms integrate x8 / x16 Armv8.2+ A78 Hercules cores (64-bit)”
“A DPU is a system on a chip, or SoC, that combines: An industry-standard, high-performance, software-programmable, multi-core CPU, typically based on the widely used Arm architecture, tightly coupled to the other SoC components.”
“Data packet parsing, matching and manipulation to implement an open virtual switch (OVS)”
“BlueField-3 DPUs offload, accelerate, and isolate software-defined networking, storage, security, and management functions, significantly enhancing data center performance, efficiency, and security.”
“By decoupling data center infrastructure from business applications, BlueField-3 creates a secure, zero-trust data center infrastructure”
“Cloud-based solutions offer a cost-effective way to start AI initiatives by reducing acquisition costs and shifting capital expenditures (CapEx) to operational expenditures (OpEx).”
“Yet, while cloud solutions may have lower initial costs, long-term expenses can add up.”
“Specialists in moving data in data centers, DPUs, or data processing units, are a new class of programmable processor and will join CPUs and GPUs as one of the three pillars of computing.”
GPU Virtualization: simulator lab
Share GPUs across virtual machines with NVIDIA vGPU and compare with MIG and time-slicing.
Try this
- Pick a vGPU profile.
- Compare isolation with time-slicing.
What NVIDIA says (6)
“NVIDIA virtual GPU software enables multiple virtual machines (VMs) to have simultaneous, direct access to a single physical GPU”
“Time-slicing trades the memory and fault-isolation that is provided by MIG for the ability to share a GPU by a larger number of users.”
“Display information on GRID virtual GPUs.”
“With MIG, each instance’s processors have separate and isolated paths through the entire memory system”
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas, but for some workloads this is better than not being able to share at all.”
“Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas”