2.7 Writing components and scripts

NCA-GENL · Software Development (24% of the exam) · Official objective: “Write software components or scripts under the supervision of a senior team member.”

Configuring and checking serving and safety components.

Key points

  1. NVIDIA's Triton quickstart says all models should show READY status. If a model fails to load, the status reports the failure and a reason.

    What NVIDIA says (2)

    “All the models should show “READY” status to indicate that they loaded correctly.”

    — Quickstart — NVIDIA Triton Inference Server

    “If a model fails to load the status will report the failure and a reason for the failure.”

    — Quickstart — NVIDIA Triton Inference Server

  2. NeMo Guardrails is an open-source Python package for adding programmable guardrails to large language model (LLM) applications. It can block, alter or validate unsafe, off-topic or policy-violating inputs and responses.

    What NVIDIA says (1)

    “is an open-source Python package for adding programmable guardrails to LLM-based applications. Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”

    — Overview | NVIDIA NeMo Guardrails Library Developer Guide

  3. Triton is NVIDIA's inference server; each model has a configuration. The docs say max_batch_size should be set to a value greater than or equal to 1 indicating the maximum batch size Triton should use.

    What NVIDIA says (1)

    “should be set to a value greater-or-equal-to 1 that indicates the maximum batch size that Triton should use with the model.”

    — Model Configuration — NVIDIA Triton Inference Server

  4. NVIDIA's quickstart says running Triton revolves around building model repositories. You start the server with the --model-repository option pointing at that folder.

    What NVIDIA says (2)

    “Launching and maintaining Triton Inference Server revolves around the use of building model repositories.”

    — Quickstart — NVIDIA Triton Inference Server

    “tritonserver --model-repository=/models”

    — Quickstart — NVIDIA Triton Inference Server

Key terms

Sample question

After starting Triton, how do you confirm models loaded correctly?

Show the answer

Answer: Check that every model shows READY status

NVIDIA's Triton quickstart says all models should show READY status. If a model fails to load, the status reports the failure and a reason.

What NVIDIA says (2)

“All the models should show “READY” status to indicate that they loaded correctly.”

— Quickstart — NVIDIA Triton Inference Server

“If a model fails to load the status will report the failure and a reason for the failure.”

— Quickstart — NVIDIA Triton Inference Server

Practice 2.7 (4 questions) Full Software Development guide

← 2.6 Traditional ML packages in code · 3.1 Training and training optimization →