3.2 Inference with Run:ai

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Deploy inference workloads with Run:ai”

Run:ai inference workloads, their prerequisites and custom servers.

Key points

  1. Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.

    What NVIDIA says (1)

    “Make sure Knative is properly installed by your administrator.”

    — NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

  2. Every Run:ai workload belongs to a project. A project's quota is the GPU share it is guaranteed.

    What NVIDIA says (1)

    “The inference workload is assigned to a project and is affected by the project’s quota.”

    — NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

  3. The Custom server option lets you bring your own container image and server configuration.

    What NVIDIA says (1)

    “allows you to bring your own container image and server configuration - for example, when using an inference server not natively supported by NVIDIA Run:ai, such as SGLang.”

    — NVIDIA Run:ai: Deploy Inference Workloads with a Custom Server

  4. Run:ai inference covers small and very large models. Large LLMs can span several nodes.

    What NVIDIA says (1)

    “The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”

    — NVIDIA Run:ai Inference Overview

Key terms

Try it

Sample question

Before a user deploys a Run:ai inference workload with NIM, what must the administrator have installed?

Show the answer

Answer: Knative

Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.

What NVIDIA says (1)

“Make sure Knative is properly installed by your administrator.”

— NVIDIA Run:ai: Deploy Inference Workloads with NVIDIA NIM

Practice 3.2 (4 questions) Full Workload Management guide

← 3.1 Inference on Kubernetes · 3.3 Training with Slurm →