Software Development
15% of the NCA-GENM exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Experimentation · Core Machine Learning and AI Knowledge · Multimodal Data · Software Development · Data Analysis and Visualization · Performance Optimization · Trustworthy AI
4.1 Working with clients on requirements
Aligning stakeholders, writing for nontechnical readers, and turning usage into sizing and design choices.
Key points
A stakeholder is anyone affected by the project, such as the client, users and operators. Agreeing on goals first sets the requirements for everything else. Then you define who owns what, and the roles. MLOps means machine learning operations.
What NVIDIA says (1)
“Align stakeholders on the goals, create an organizational structure that defines who owns what, then define responsibilities and roles”
A model card summarizes what a model does and its limits. Writing for clients and users helps agree on fitness for use.
What NVIDIA says (1)
“They are built to communicate to nontechnical stakeholders, customers using the software, and even students.”
ISL and OSL mean input and output sequence length. Longer inputs raise time to first token (TTFT). Longer outputs raise the time between tokens. Knowing the mix lets you size the deployment.
What NVIDIA says (2)
“A longer ISL will increase the memory requirement for the prefill stage and thus increase the TTFT.”
“It is important to understand the distribution of inputs and outputs in your LLM deployment to best optimize your hardwar”
The primary modality is chosen from the focus of the application. Here the focus is text Q&A, so images get text descriptions. RAG means retrieval-augmented generation.
What NVIDIA says (1)
“Another option is to pick a primary modality based on the focus of the application and ground all other modalities in the primary modality.”
Key terms: Model card MLOps Stakeholder
4.2 Best practices and software quality
Import guarding, readiness checks, guardrails and dataset licence checks.
Key points
Import guarding lets code run with or without an optional dependency. safe_import returns the module (or a placeholder) and a boolean saying whether it loaded.
What NVIDIA says (1)
“In either of these cases, it’s important to guard the optional imports.”
A quality check confirms a service is really working before use. Triton prints a status for each model and a reason if loading failed.
What NVIDIA says (1)
“All the models should show “READY” status to indicate that they loaded correctly.”
A guardrail is a rule that checks what goes into or comes out of a model. NeMo Guardrails is an open-source Python package for this. DALI means the NVIDIA Data Loading Library. LLM means large language model.
What NVIDIA says (1)
“Use it to block, alter, or validate unsafe, off-topic, malicious, or policy-violating user inputs and model responses.”
Good practice includes checking legal and content fitness of data, not just code quality.
What NVIDIA says (1)
“It is the responsibility of each user to check the content of the dataset, review the applicable licenses, and determine if it is suitable for their intended use.”
Key terms: NeMo Guardrails Import guarding
4.3 Prompting generative models
Text-to-image prompts, DreamBooth identifiers, top-k, system prompts and Perfusion gating.
Key points
A text-to-image prompt is a description of the image you want. Its embeddings condition each denoising step. Clear, specific prompts give better guidance. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“The text-encoder, typically a simple transformer like CLIP, converts input prompts into embeddings, which guides the U-Net’s denoising process.”
DreamBooth teaches a model a new subject from a few images. It binds a unique identifier token to the subject. Using that token in a prompt recalls it in new scenes.
What NVIDIA says (1)
“so that it learns to bind a unique identifier with a special subject.”
Top-k limits the next-token choice to the k most likely tokens. Top-p instead keeps tokens whose probabilities add up to p. LLM means large language model.
What NVIDIA says (1)
“Top-k tells the model that it has to keep the top k highest probability tokens, from which the next token is selected at random. Lower values reduce randomness”
A system prompt is hidden setup text that frames every conversation. It sets role, tone and rules. LLM means large language model.
What NVIDIA says (1)
“This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”
Personalization teaches a model a specific concept, such as your own teddy bear. Perfusion adds a gate that controls how strongly the concept shows, without retraining. VAE means variational autoencoder.
What NVIDIA says (1)
“We also added a gating mechanism to regulate how strongly the learned concept is considered”
Key terms: Text encoder DreamBooth Prompt engineering Top-k sampling System prompt
Try it: Sampling lab
4.4 U-Nets: from noise, and as an autoencoder
How denoising diffusion works, what the U-Net predicts and the VAE autoencoder.
Key points
A denoiser removes noise. Applied many times to pure noise, it reveals a new image. The training data decides what kinds of images appear.
What NVIDIA says (1)
“first draw a random image of pure white noise, and then chip away at the noise level”
The timestep tells the U-Net how noisy the input is. Subtracting the predicted noise moves the latent toward a clean image. VAE means variational autoencoder.
What NVIDIA says (1)
“The Unet processes the noisy latents (x) to predict the noise, utilizing a conditional model which also incorporates the timestep (t) and text embedding for guidance.”
At high noise, many clean images are possible, so the best guess is their average. Repeated small steps sharpen this into one image.
What NVIDIA says (1)
“the denoiser must output the blurry average of all possible clean images that could have been hiding under the noise.”
An autoencoder compresses data with an encoder and rebuilds it with a decoder. In Stable Diffusion the VAE plays this role around the U-Net. VAE means variational autoencoder.
What NVIDIA says (1)
“Subsequently, during inference, the decoder reverses this process by transforming denoised latent representations back into their original, tangible image forms.”
This is a cascade: one base model and super-resolution models in sequence. Each stage is a diffusion U-Net. GAN means generative adversarial network.
What NVIDIA says (1)
“Imagen first generates an image at a 64x64 resolution and then upsamples the generated image to 256x256 and 1024x1024 resolutions, all using diffusion models.”
Key terms: Stable Diffusion Denoising diffusion U-Net Variational autoencoder Latent space Imagen
Try it: Diffusion lab
4.5 CLIP for text-to-image
CLIP's role in Stable Diffusion, matching dimensions, and other uses of CLIP.
Key points
CLIP learned a shared space for images and text. Its text encoder gives diffusion training a strong prompt signal. It stays frozen during SD training. CLIP means Contrastive Language-Image Pre-training. VAE means variational autoencoder.
What NVIDIA says (1)
“Stable diffusion has three main components: A U-Net, an image encoder(Variational Autoencoder, VAE) and a text-encoder(CLIP).”
context_dim is the width of the text embeddings that the U-Net's cross-attention reads. A mismatch makes the shapes incompatible.
What NVIDIA says (1)
“context_dim : Must be adjusted to match the text encoder’s output dimension.”
A Vision Transformer splits an image into patches and processes them like tokens. CLIP means Contrastive Language-Image Pre-training. ML means machine learning.
What NVIDIA says (1)
“CLIP’s vision model is based on the Vision Transformer (ViT) architecture.”
CLIP embeddings are a general-purpose image representation. They feed vision-language models and data-quality tools. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (2)
“NeVA harnesses the power of the pre-trained CLIP visual encoder, ViT-L/14”
“a linear classifier that takes OpenAI CLIP ViT-L/14 image embeddings as input.”
Both encoders must output the same size so their embeddings can be compared. CLIP means Contrastive Language-Image Pre-training.
What NVIDIA says (1)
“output_dim : Represents the dimensionality of the output embeddings for both the text and vision models.”
Key terms: Vision encoder CLIP Vision Transformer Stable Diffusion Text encoder
Try it: CLIP lab