3.1 Inference on Kubernetes
Deploying NIM with the NIM Operator, caching models and autoscaling.
Key points
The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.
What NVIDIA says (1)
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”
Models are large. The NIM Operator can pre-cache them on cluster storage. New pods then start from the cache instead of downloading again.
What NVIDIA says (2)
“One key benefit of using the NIM Operator is its ability to pre-cache models and datasets.”
“Models, their many available profiles, and training datasets are large and can take a long time to download.”
A NIMService is a custom resource. kubectl lists it like any other resource. -A means all namespaces.
What NVIDIA says (1)
“View the NIM services custom resources: $ kubectl get nimservices.apps.nvidia.com -A”
Horizontal Pod Autoscaling adds or removes pod replicas based on metrics. The NIM Operator docs list Prometheus as the prerequisite.
What NVIDIA says (1)
“Configuring Horizontal Pod Autoscaling # Prerequisites # Prometheus installed on your cluster.”
Key terms
- NVIDIA NIM: Containerized NVIDIA inference microservices that serve AI models.
- NIM Operator: A Kubernetes operator that deploys NIM services and caches their models.
- Horizontal Pod Autoscaling: Kubernetes adding or removing pod replicas based on metrics.
Try it
Sample question
What does the NVIDIA NIM Operator do in a Kubernetes cluster?
Show the answer
Answer: It deploys and manages the lifecycle of NIM inference microservices through custom resources
The NIM Operator is a Kubernetes operator for NIM. You describe the NIM deployment you want in a custom resource, and the Operator makes it so.
What NVIDIA says (1)
“The NVIDIA NIM Operator enables Kubernetes cluster administrators to operate the software components and services necessary to deploy NVIDIA NIMs microservices in Kubernetes.”