3.1 Inference on Kubernetes

NCP-AIO · Workload Management (23% of the exam) · Official objective: “Deploy inference workloads with Kubernetes”

Deploying NIM with the NIM Operator, caching models and autoscaling.

Key points

  1. The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.

    What NVIDIA says (1)

    “The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

    — NVIDIA NIM Operator

  2. Models are large. The NIM Operator can pre-cache them on cluster storage. New pods then start from the cache instead of downloading again.

    What NVIDIA says (2)

    “One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”

    — NVIDIA NIM Operator

    “Models, their many available profiles, and training datasets are large and can take a long time to download.”

    — NVIDIA NIM Operator

  3. A NIMService is a custom resource. kubectl lists it like any other resource. -A means all namespaces.

    What NVIDIA says (1)

    “View the NIM services custom resources: $ kubectl get nimservices.apps.nvidia.com -A”

    — NVIDIA NIM Operator: NIM Service

  4. Horizontal Pod Autoscaling adds or removes pod replicas based on metrics. The NIM Operator docs list Prometheus as the prerequisite.

    What NVIDIA says (1)

    “Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”

    — NVIDIA NIM Operator: NIM Service

Key terms

Try it

Sample question

What does the NVIDIA NIM Operator do in a Kubernetes cluster?

Show the answer

Answer: It deploys and manages the lifecycle of NIM inference microservices through custom resources

The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.

What NVIDIA says (1)

“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”

— NVIDIA NIM Operator

Practice 3.1 (4 questions) Full Workload Management guide

← 2.5 Configuring MIG · 3.2 Inference with Run:ai →