3.7 Writing software components and scripts
Calling a NIM VLM through its OpenAI-compatible API and starting Triton with a model repository.
Key points
OpenAI-compatible means existing OpenAI client code can call it by changing the base URL. Chat completions take a message history and return the model's reply. VLM means vision language model. NIM means NVIDIA Inference Microservice.
What NVIDIA says (2)
“NIM VLM exposes an OpenAI-compatible inference API backed by vLLM”
“POST /v1/chat/completions Multi-turn chat completions with message history.”
A model repository is a folder with each model and its configuration. Triton loads models from it at startup.
What NVIDIA says (1)
“tritonserver --model-repository=/models”
Key terms
- Triton Inference Server: NVIDIA's open-source server for serving trained models.
- NVIDIA NIM: Containerized NVIDIA inference microservices; NIM for VLMs exposes an OpenAI-compatible API.
Sample question
You write a script to send an image question to a NIM VLM. Which API style does it expose?
Show the answer
Answer: An OpenAI-compatible API, with POST /v1/chat/completions
OpenAI-compatible means existing OpenAI client code can call it by changing the base URL. Chat completions take a message history and return the model's reply. VLM means vision language model. NIM means NVIDIA Inference Microservice.
What NVIDIA says (2)
“NIM VLM exposes an OpenAI-compatible inference API backed by vLLM”
“POST /v1/chat/completions Multi-turn chat completions with message history.”
Practice 3.7 (2 questions) Full Multimodal Data guide
← 3.6 Python packages for traditional ML · 4.1 Working with clients on requirements →