3.2 Inference with Run:ai
Run:ai inference workloads, their prerequisites and custom servers.
Key points
Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.
What NVIDIA says (1)
“Make sure Knative is properly installed by your administrator.”
Every Run:ai workload belongs to a project. A project's quota is the GPU share it is guaranteed.
What NVIDIA says (1)
“The inference workload is assigned to a project and is affected by the project’s quota.”
The Custom server option lets you bring your own container image and server configuration.
What NVIDIA says (1)
“allows you to bring your own container image and server configuration - for example, when using an inference server not natively supported by NVIDIA Run:ai, such as SGLang.”
Run:ai inference covers small and very large models. Large LLMs can span several nodes.
What NVIDIA says (1)
“The platform supports both single-node and multi-node architectures and is compatible with NVIDIA NIM, vLLM, and custom inference servers.”
Key terms
- Run:ai project: The main Run:ai unit for a team: it holds a GPU quota and maps to a Kubernetes namespace.
- NVIDIA NIM: Containerized NVIDIA inference microservices that serve AI models.
- Knative: A Kubernetes add-on for request-driven serving that Run:ai inference workloads need.
Try it
Sample question
Before a user deploys a Run:ai inference workload with NIM, what must the administrator have installed?
Show the answer
Answer: Knative
Knative is a Kubernetes add-on for serving and autoscaling request-driven workloads. Run:ai inference workloads depend on it.
What NVIDIA says (1)
“Make sure Knative is properly installed by your administrator.”