3.1 Scalability, performance and reliability

NCA-GENM · Multimodal Data (15% of the exam) · Official objective: “Assist in the deployment and evaluations of model scalability, performance, and reliability under the supervision of a senior team member.”

Dynamic batching, load tests versus benchmarks, and readiness probes.

Key points

  1. Triton Inference Server is NVIDIA's model serving software. Dynamic batching combines separate requests into one batch on the server. Bigger batches use the GPU better.

    What NVIDIA says (1)

    “Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”

    — NVIDIA Triton Inference Server: Batchers

  2. By default the batcher does not wait for more requests. You can set a queue delay if you prefer bigger batches over latency.

    What NVIDIA says (1)

    “By default the dynamic batcher will create batches as large as possible up to the maximum batch size and will not delay when forming batches.”

    — NVIDIA Triton Inference Server: Batchers

  3. Scalability is how well a service handles more traffic. Load tests probe scale. Benchmarks such as GenAI-Perf measure raw model performance. NVIDIA recommends both.

    What NVIDIA says (1)

    “Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”

    — LLM Inference Benchmarking: Fundamental Concepts

  4. A readiness probe checks whether a service can take traffic. NIM for VLMs returns 200 on /v1/health/ready once the model is loaded. VLM means vision language model. VLMs means vision language models. NIM means NVIDIA Inference Microservice.

    What NVIDIA says (1)

    “GET /v1/health/ready Readiness probe. Returns 200 when the model is loaded and inference is available.”

    — NVIDIA NIM for Vision Language Models: API Reference

Key terms

Sample question

Your multimodal model on Triton gets many small requests and GPU use is low. Which Triton feature raises throughput?

Show the answer

Answer: Dynamic batching

Triton Inference Server is NVIDIA's model serving software. Dynamic batching combines separate requests into one batch on the server. Bigger batches use the GPU better.

What NVIDIA says (1)

“Dynamic batching is a feature of Triton that allows inference requests to be combined by the server, so that a batch is created dynamically. Creating a batch of requests typically results in increased throughput.”

— NVIDIA Triton Inference Server: Batchers

Practice 3.1 (4 questions) Full Multimodal Data guide

← 2.10 TensorFlow and PyTorch · 3.2 RAG, chatbots and summarizers →