2.7 Writing components and scripts
Configuring and checking serving and safety components.
Key points
NVIDIA's Triton quickstart says all models should show READY status. If a model fails to load, the status reports the failure and a reason.
What NVIDIA says (2)
“All the models should show “READY” status to indicate that they loaded correctly.”
“If a model fails to load the status will report the failure and a reason for the failure.”
NeMo Guardrails is an open-source Python package for adding programmable guardrails to large language model (LLM) applications. It can block, alter or validate unsafe, off-topic or policy-violating inputs and responses.
What NVIDIA says (1)
“is an open-source Python package for adding programmable guardrails to LLM-based applications. Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”
Triton is NVIDIA's inference server; each model has a configuration. The docs say max_batch_size should be set to a value greater than or equal to 1 indicating the maximum batch size Triton should use.
What NVIDIA says (1)
“should be set to a value greater-or-equal-to 1 that indicates the maximum batch size that Triton should use with the model.”
NVIDIA's quickstart says running Triton revolves around building model repositories. You start the server with the --model-repository option pointing at that folder.
What NVIDIA says (2)
“Launching and maintaining Triton Inference Server revolves around the use of building model repositories.”
“tritonserver --model-repository=/models”
Key terms
- Triton Inference Server: NVIDIA's inference server, which loads models from a model repository.
- NeMo Guardrails: An open-source library that adds programmable safety and topic rules around an LLM app.
Sample question
After starting Triton, how do you confirm models loaded correctly?
Show the answer
Answer: Check that every model shows READY status
NVIDIA's Triton quickstart says all models should show READY status. If a model fails to load, the status reports the failure and a reason.
What NVIDIA says (2)
“All the models should show “READY” status to indicate that they loaded correctly.”
“If a model fails to load the status will report the failure and a reason for the failure.”
Practice 2.7 (4 questions) Full Software Development guide
← 2.6 Traditional ML packages in code · 3.1 Training and training optimization →