Software Development

24% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI

2.1 Measuring LLM performance in software

Official objective: “Assist in the deployment and evaluations of model scalability, performance, and reliability under the supervision of a senior team member.”

Latency, throughput and testing approaches for LLM services.

Key points

  1. Time to first token (TTFT) is how long a user waits before the first output token appears. NVIDIA says TTFT generally includes request queuing, prefill and network latency. A longer prompt gives a larger TTFT because attention needs the whole input sequence.

    What NVIDIA says (2)

    “TTFT generally includes both request queuing time, prefill time, and network latency. The longer the prompt, the larger the TTFT.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “This is because the attention mechanism requires the whole input sequence to compute”

    — LLM Inference Benchmarking: Fundamental Concepts

  2. NVIDIA describes load testing as simulating many concurrent requests to assess handling of real-world traffic at scale. Performance benchmarking, for example with the NVIDIA GenAI-Perf tool, measures the model's own throughput, latency and token-level metrics. NVIDIA recommends combining both.

    What NVIDIA says (3)

    “Load testing focuses on simulating a large number of concurrent requests to a model to assess its ability to handle real-world traffic at scale.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “In contrast, performance benchmarking, as demonstrated by the NVIDIA GenAI-Perf tool, is concerned with measuring the actual performance of the model itself, such as its throughput, latency, and token-level metrics.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “By combining both approaches, developers can gain a comprehensive understanding”

    — LLM Inference Benchmarking: Fundamental Concepts

  3. NVIDIA's benchmarking post says the cost of a large language model (LLM) deployment depends on how many queries it can process per second while staying responsive. So serving performance drives cost.

    What NVIDIA says (1)

    “The cost of an LLM application deployment depends on how many queries it can process per second while being responsive”

    — LLM Inference Benchmarking: Fundamental Concepts

Key terms: Time to first token

Try it: KV cache lab

Practice 2.1 (3 questions)

2.2 Building LLM features in software

Official objective: “Build LLM use cases such as RAGs, chatbots, and summarizers.”

Agents, reranking and context limits when you build LLM features.

Key points

  1. A large language model (LLM)-powered agent uses an LLM to reason through a problem, create a plan and carry it out with tools. NVIDIA's example question needs planning, memory and several tools, which a simple lookup cannot provide.

    What NVIDIA says (2)

    “they can be described as a system that can use an LLM to reason through a problem, create a plan to solve the problem, and execute the plan with the help of a set of tools.”

    — Introduction to LLM Agents

    “This inquiry requires planning, tailored focus, memory, using different tools”

    — Introduction to LLM Agents

  2. Reranking is a second stage after fast retrieval such as Best Matching 25, a keyword-ranking function (BM25) or vector search. The candidates go to a large language model (LLM) that judges how relevant each one is to the query, so the best passages reach the generator.

    What NVIDIA says (3)

    “Initially, a set of candidate documents or passages is retrieved using traditional information retrieval methods like BM25 or vector similarity search.”

    — Enhancing RAG Pipelines with Re-Ranking

    “Re-ranking is typically used as a second stage after an initial fast retrieval step”

    — Enhancing RAG Pipelines with Re-Ranking

    “These candidates are then fed into an LLM that analyzes the semantic relevance between the query and each document.”

    — Enhancing RAG Pipelines with Re-Ranking

  3. The context window is the most text the large language model (LLM) can take in at once. NVIDIA notes the whole prompt, retrieved chunks plus the query, must fit in it, so chunk sizes should not be too big.

    What NVIDIA says (2)

    “The entire prompt (retrieved chunks plus the user query) must fit within the LLM’s context window.”

    — Enhancing RAG Pipelines with Re-Ranking

    “specify chunk sizes too big”

    — Enhancing RAG Pipelines with Re-Ranking

  4. NVIDIA's LLMOps post says that once a customized model is used alone or in a chain of models and application programming interfaces (APIs), you must test the complete AI system for accuracy, speed and vulnerabilities, and add guardrails.

    What NVIDIA says (2)

    “At this point, it is crucial to test the complete AI system for accuracy, speed”

    — Mastering LLM Techniques: LLMOps

    “it is crucial to test the complete AI system for accuracy, speed, and vulnerabilities, and add guardrails”

    — Mastering LLM Techniques: LLMOps

Key terms: Large language model Retrieval-augmented generation Reranking Context window LLM agent

Practice 2.2 (4 questions)

2.3 Python language packages in practice

Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

Tokenizing, entity extraction and DataFrames in real pipelines.

Key points

  1. Subword tokenization splits words into smaller, reusable pieces, for example “anyplace” into “any” and “place”. NVIDIA's RAPIDS natural language processing (NLP) post describes a graphics processing unit (GPU) Bidirectional Encoder Representations from Transformers (BERT) subword tokenizer, so tokenization runs on the GPU too.

    What NVIDIA says (3)

    “We first introduced the GPU BERT subword tokenizer in a previous”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “To solve these problems, we use subword tokenization.”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “For example, the word “anyplace” can be broken down into “any” and “place,””

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  2. The post explains that spaCy needed its inputs on the central processing unit (CPU), so the graphics processing unit (GPU) pipeline had to copy data to CPU memory and back, which slowed it down.

    What NVIDIA says (1)

    “spaCy currently needs your inputs on CPU and thus was slow as it required a copy to CPU memory and back to GPU memory.”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  3. NVIDIA's pandas glossary describes the DataFrame as a two-dimensional, array-like table where each column represents a variable and each row a set of values.

    What NVIDIA says (1)

    “A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable”

    — What Is Pandas and Why Does it Matter?

Key terms: cuDF pandas / DataFrame spaCy

Practice 2.3 (3 questions)

2.4 Picking the right components

Official objective: “Identify system data, hardware, or software components required to meet user needs.”

Which NVIDIA software components fit which need.

Key points

  1. NIM (NVIDIA Inference Microservices) are containerized solutions with industry-standard application programming interfaces (APIs) and Helm charts to scale. They let developers focus on application logic instead of serving infrastructure.

    What NVIDIA says (2)

    “NIMs are containerized solutions, which come with industry-standard APIs and Helm charts to scale.”

    — Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

    “developers to focus on working on their application logic rather than having to spend cycles on building and scaling out the infrastructure.”

    — Develop Production-Grade Text Retrieval Pipelines for RAG with NVIDIA NeMo Retriever

  2. NVIDIA Collective Communications Library (NCCL) provides inter-GPU communication (communication between graphics processing units, or GPUs) primitives that are topology-aware. Its AllReduce collective is heavily used in neural-network training.

    What NVIDIA says (2)

    “is a library providing inter-GPU communication primitives that are topology-aware and can be easily integrated into applications.”

    — Overview of NCCL — NCCL 2.32.3 documentation

    “NCCL has found great application in Deep Learning Frameworks, where the AllReduce collective is heavily used for neural network training.”

    — Overview of NCCL — NCCL 2.32.3 documentation

  3. NVIDIA NeMo Retriever is a collection of microservices with a single application programming interface (API) for indexing and querying user data. It covers extraction, embedding and reranking pipelines.

    What NVIDIA says (2)

    “NVIDIA NeMo Retriever (NeMo Retriever) is a collection of microservices that present a single API for indexing and querying of user data.”

    — NVIDIA NeMo Retriever

    “for building and scaling multimodal data extraction, embedding, and reranking pipelines”

    — NVIDIA NeMo Retriever

  4. NVIDIA NeMo Framework is a scalable, cloud-native generative AI framework for large language models (LLMs), multimodal and speech models. It lets users create, customize and deploy models from existing code and pre-trained checkpoints.

    What NVIDIA says (2)

    “NVIDIA NeMo Framework is a scalable and cloud-native generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and”

    — Overview — NVIDIA NeMo Framework User Guide

    “It enables users to efficiently create, customize, and deploy new generative AI models by leveraging existing code and pre-trained model checkpoints.”

    — Overview — NVIDIA NeMo Framework User Guide

  5. FP16 (half precision) uses 16 bits instead of 32 bits for FP32 (single precision). NVIDIA's mixed-precision guide says lowering memory enables larger models or larger mini-batches.

    What NVIDIA says (1)

    “Half-precision floating point format (FP16) uses 16 bits, compared to 32 bits for single precision (FP32). Lowering the required memory enables training of larger models or training with larger mini-batches.”

    — Train With Mixed Precision

Key terms: FP16 / FP32 NVIDIA NIM NVIDIA NeMo Framework NeMo Retriever NCCL

Practice 2.4 (5 questions)

2.5 Monitoring data, experiments and processes

Official objective: “Monitor functioning of data collection, experiments, and other software processes.”

How MLOps keeps models and data under control over time.

Key points

  1. MLOps (machine learning operations) is a set of best practices for running AI successfully. It is modeled on DevOps and adds data scientists to the team.

    What NVIDIA says (3)

    “A shorthand for machine learning operations, MLOps is a set of best practices for businesses to run AI successfully.”

    — What is MLOps?

    “MLOps is modeled on the existing discipline of DevOps”

    — What is MLOps?

    “MLOps adds to the team the data scientists, who curate datasets and build AI models that analyze them.”

    — What is MLOps?

  2. NVIDIA notes that datasets are massive and can change in real time. Models need careful tracking through cycles of experiments, tuning and retraining.

    What NVIDIA says (2)

    “AI models require careful tracking through cycles of experiments, tuning and retraining.”

    — What is MLOps?

    “Datasets are massive and growing, and can change in real time.”

    — What is MLOps?

  3. NVIDIA defines an AI data flywheel as a self-improving loop. Data from AI interactions refines the models, which produces better outcomes and more valuable data.

    What NVIDIA says (2)

    “An AI data flywheel is a self-improving loop where data collected from AI interactions or processes is used to continuously refine AI models”

    — Data flywheel: What it is and how it works

    “generating better outcomes and more valuable data for continued improvement.”

    — Data flywheel: What it is and how it works

  4. A foundation model only knows its pretraining and fine-tuning data, which becomes outdated over time. NVIDIA says retrieval-augmented generation (RAG) keeps the model grounded with external knowledge at query time.

    What NVIDIA says (2)

    “The knowledge of a foundation model is limited to the pretraining and fine-tuning data, becoming outdated over time”

    — Mastering LLM Techniques: LLMOps

    “workflow is used to maintain freshness and keep the model grounded with external knowledge during query time.”

    — Mastering LLM Techniques: LLMOps

Key terms: MLOps AI data flywheel

Practice 2.5 (4 questions)

2.6 Traditional ML packages in code

Official objective: “Use Python packages (spaCy, NumPy, Keras, etc.) to implement specific traditional machine learning analyses.”

Tuning, vectorizing and manipulating data with Python packages.

Key points

  1. Hyperparameters are the settings you tune to get the best model. NVIDIA's scikit-learn glossary says grid search automates this by testing various configurations, and is used with cross-validation to evaluate models.

    What NVIDIA says (2)

    “scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”

    — What is scikit-learn?

    “Grid search further automates hyperparameter optimization, testing various configurations for improved accuracy.”

    — What is scikit-learn?

  2. TF-IDF (term frequency–inverse document frequency) is a way to vectorize text, that is, turn it into numbers. NVIDIA's RAPIDS text post describes Count and TF-IDF vectorizers in cuML that can scale across multiple graphics processing units (GPUs) and machines.

    What NVIDIA says (3)

    “subpackage in cuML by adding Count and TF-IDF vectorizer”

    — NLP and Text Processing with RAPIDS: Now Simpler and Faster

    “You can also scale your TF-IDF workflow to multiple GPUs and machines using cuml”

    — NLP and Text Processing with RAPIDS: Now Simpler and Faster

    “by first vectorizing them using TF-IDF”

    — NLP and Text Processing with RAPIDS: Now Simpler and Faster

  3. NVIDIA's pandas glossary lists data manipulation and cleaning operations such as selecting a subset, derived columns, sorting, joining, filling, replacing, summary statistics and plotting.

    What NVIDIA says (1)

    “pandas also allows for various data manipulation operations and data cleaning features, including selecting a subset, creating derived columns, sorting, joining, filling, replacing, summary statistics, and plotting.”

    — What Is Pandas and Why Does it Matter?

Key terms: cuML scikit-learn TF-IDF Grid search

Practice 2.6 (3 questions)

2.7 Writing components and scripts

Official objective: “Write software components or scripts under the supervision of a senior team member.”

Configuring and checking serving and safety components.

Key points

  1. NVIDIA's Triton quickstart says all models should show READY status. If a model fails to load, the status reports the failure and a reason.

    What NVIDIA says (2)

    “All the models should show “READY” status to indicate that they loaded correctly.”

    — Quickstart — NVIDIA Triton Inference Server

    “If a model fails to load the status will report the failure and a reason for the failure.”

    — Quickstart — NVIDIA Triton Inference Server

  2. NeMo Guardrails is an open-source Python package for adding programmable guardrails to large language model (LLM) applications. It can block, alter or validate unsafe, off-topic or policy-violating inputs and responses.

    What NVIDIA says (1)

    “is an open-source Python package for adding programmable guardrails to LLM-based applications. Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”

    — Overview | NVIDIA NeMo Guardrails Library Developer Guide

  3. Triton is NVIDIA's inference server; each model has a configuration. The docs say max_batch_size should be set to a value greater than or equal to 1 indicating the maximum batch size Triton should use.

    What NVIDIA says (1)

    “should be set to a value greater-or-equal-to 1 that indicates the maximum batch size that Triton should use with the model.”

    — Model Configuration — NVIDIA Triton Inference Server

  4. NVIDIA's quickstart says running Triton revolves around building model repositories. You start the server with the --model-repository option pointing at that folder.

    What NVIDIA says (2)

    “Launching and maintaining Triton Inference Server revolves around the use of building model repositories.”

    — Quickstart — NVIDIA Triton Inference Server

    “tritonserver --model-repository=/models”

    — Quickstart — NVIDIA Triton Inference Server

Key terms: Triton Inference Server NeMo Guardrails

Practice 2.7 (4 questions)