Software Development

15% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI

4.1 Working with clients on requirements

Official objective: “Collaborate with the client during requirements acquisition, data gathering, progress reporting, deployment, and integration.”

Aligning stakeholders, writing for nontechnical readers, and turning usage into sizing and design choices.

Key points

  1. A stakeholder is anyone affected by the project, such as the client, users and operators. Agreeing on goals first sets the requirements for everything else. Then you define who owns what, and the roles. MLOps means machine learning operations.

    What NVIDIA says (1)

    “Align stakeholders on the goals, create an organizational structure that defines who owns what, then define responsibilities and roles”

    — What is MLOps?

  2. A model card summarizes what a model does and its limits. Writing for clients and users helps agree on fitness for use.

    What NVIDIA says (1)

    “They are built to communicate to nontechnical stakeholders, customers using the software, and even students.”

    — Enhancing AI Transparency and Ethical Considerations with Model Card++

  3. ISL and OSL mean input and output sequence length. Longer inputs raise time to first token (TTFT). Longer outputs raise the time between tokens. Knowing the mix lets you size the deployment.

    What NVIDIA says (2)

    “A longer ISL will increase the memory requirement for the prefill stage and thus increase the TTFT.”

    — LLM Inference Benchmarking: Fundamental Concepts

    “It is important to understand the distribution of inputs and outputs in your LLM deployment to best optimize your hardwar”

    — LLM Inference Benchmarking: Fundamental Concepts

  4. The primary modality is chosen from the focus of the application. Here the focus is text Q&A, so images get text descriptions. RAG means retrieval-augmented generation.

    What NVIDIA says (1)

    “Another option is to pick a primary modality based on the focus of the application and ground all other modalities in the primary modality.”

    — An Easy Introduction to Multimodal Retrieval-Augmented Generation

Key terms: Model card MLOps Stakeholder

Practice 4.1 (4 questions)

4.2 Best practices and software quality

Official objective: “Ensure adherence to best practices and maintain high standards of software quality and reliability.”

Import guarding, readiness checks, guardrails and dataset licence checks.

Key points

  1. Import guarding lets code run with or without an optional dependency. safe_import returns the module (or a placeholder) and a boolean saying whether it loaded.

    What NVIDIA says (1)

    “In either of these cases, it’s important to guard the optional imports.”

    — NeMo Framework 24.09: Best Practices

  2. A quality check confirms a service is really working before use. Triton prints a status for each model and a reason if loading failed.

    What NVIDIA says (1)

    “All the models should show “READY” status to indicate that they loaded correctly.”

    — Quickstart — NVIDIA Triton Inference Server

  3. A guardrail is a rule that checks what goes into or comes out of a model. NeMo Guardrails is an open-source Python package for this. DALI means the NVIDIA Data Loading Library. LLM means large language model.

    What NVIDIA says (1)

    “Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”

    — Overview | NVIDIA NeMo Guardrails Library Developer Guide

  4. Good practice includes checking legal and content fitness of data, not just code quality.

    What NVIDIA says (1)

    “It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”

    — NeMo Framework 24.09: Text to Image Datasets

Key terms: NeMo Guardrails Import guarding

Practice 4.2 (4 questions)

4.3 Prompting generative models

Official objective: “Use prompt engineering to better influence the output of generative AI models.”

Text-to-image prompts, DreamBooth identifiers, top-k, system prompts and Perfusion gating.

Key points

  1. A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”

    — NeMo Framework 24.09: Stable Diffusion

  2. DreamBooth teaches a model a new subject from a few images. It binds a unique identifier token to the subject. Using that token in a prompt recalls it in new scenes.

    What NVIDIA says (1)

    “so that it learns to bind a unique identifier with a special subject.”

    — NeMo Framework 24.09: DreamBooth

  3. Top-k limits the next-token choice to the k most likely tokens. Top-p instead keeps tokens whose probabilities add up to p. LLM means large language model.

    What NVIDIA says (1)

    “Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”

    — How to Get Better Outputs from Your Large Language Model

  4. A system prompt is hidden setup text that frames every conversation. It sets role, tone and rules. LLM means large language model.

    What NVIDIA says (1)

    “This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”

    — Mastering LLM Techniques: Customization

  5. Personalization teaches a model a specific concept, such as your own teddy bear. Perfusion adds a gate that controls how strongly the concept shows, without retraining. VAE means variational autoencoder.

    What NVIDIA says (1)

    “We also added a gating mechanism to regulate how strongly the learned concept is considered”

    — NVIDIA Technical Blog: Personalizing Text-to-Image Models

Key terms: Text encoder DreamBooth Prompt engineering Top-k sampling System prompt

Try it: Sampling lab

Practice 4.3 (5 questions)

4.4 U-Nets: from noise, and as an autoencoder

Official objective: “Build a U-Net to generate images from pure noise and as a type of autoencoder.”

How denoising diffusion works, what the U-Net predicts and the VAE autoencoder.

Key points

  1. A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.

    What NVIDIA says (1)

    “first draw a random image of pure white noise, and then chip away at the noise level”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  2. The timestep tells the U-Net how noisy the input is. Subtracting the predicted noise moves the latent toward a clean image. VAE means variational autoencoder.

    What NVIDIA says (1)

    “The Unet processes the noisy latents (x) to predict the noise, utilizing a conditional model which also incorporates the timestep (t) and text embedding for guidance.”

    — NeMo Framework 24.09: Stable Diffusion

  3. At high noise, many clean images are possible, so the best guess is their average. Repeated small steps sharpen this into one image.

    What NVIDIA says (1)

    “the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”

    — NVIDIA Technical Blog: Demystifying Diffusion-Based Models

  4. An autoencoder compresses data with an encoder and rebuilds it with a decoder. In Stable Diffusion the VAE plays this role around the U-Net. VAE means variational autoencoder.

    What NVIDIA says (1)

    “Subsequently, during inference, the decoder reverses this process by transforming denoised latent representations back into their original, tangible image forms.”

    — NeMo Framework 24.09: Stable Diffusion

  5. This is a cascade: one base model and super-resolution models in sequence. Each stage is a diffusion U-Net. GAN means generative adversarial network.

    What NVIDIA says (1)

    “Imagen first generates an image at a 64x64 resolution and then upsamples the generated image to 256x256 and 1024x1024 resolutions, all using diffusion models.”

    — NeMo Framework 24.09: Imagen

Key terms: Stable Diffusion Denoising diffusion U-Net Variational autoencoder Latent space Imagen

Try it: Diffusion lab

Practice 4.4 (5 questions)

4.5 CLIP for text-to-image

Official objective: “Generate images from English text prompts using CLIP, and use CLIP to train a text-to-image diffusion model.”

CLIP's role in Stable Diffusion, matching dimensions, and other uses of CLIP.

Key points

  1. CLIP learned a shared space for images and text. Its text encoder gives diffusion training a strong prompt signal. It stays frozen during SD training. CLIP means Contrastive Language-Image Pre-training. VAE means variational autoencoder.

    What NVIDIA says (1)

    “Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”

    — NeMo Framework 24.09: Stable Diffusion

  2. context_dim is the width of the text embeddings that the U-Net's cross-attention reads. A mismatch makes the shapes incompatible.

    What NVIDIA says (1)

    “context_dim : Must be adjusted to match the text encoder’s output dimension.”

    — NeMo Framework 24.09: Stable Diffusion

  3. A Vision Transformer splits an image into patches and processes them like tokens. CLIP means Contrastive Language-Image Pre-training. ML means machine learning.

    What NVIDIA says (1)

    “CLIP’s vision model is based on the Vision Transformer (ViT) architecture.”

    — NeMo Framework 24.09: CLIP

  4. CLIP embeddings are a general-purpose image representation. They feed vision-language models and data-quality tools. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (2)

    “NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”

    — NeMo Framework 24.09: NeVA

    “a linear classifier that takes OpenAI CLIP ViT-L/14 image embeddings as input.”

    — NeMo Curator (NeMo Framework 24.09): Aesthetic Classifier

  5. Both encoders must output the same size so their embeddings can be compared. CLIP means Contrastive Language-Image Pre-training.

    What NVIDIA says (1)

    “output_dim : Represents the dimensionality of the output embeddings for both the text and vision models.”

    — NeMo Framework 24.09: CLIP

Key terms: Vision encoder CLIP Vision Transformer Stable Diffusion Text encoder

Try it: CLIP lab

Practice 4.5 (5 questions)