Core Machine Learning and AI Knowledge
30% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.
Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI
1.1 Serving LLMs at scale
How to judge whether an LLM deployment is fast, scalable and reliable, and what limits it.
Key points
NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.
What NVIDIA says (2)
“In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”
“Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”
Requests in a batch generate different numbers of tokens, so with static batching every request waits for the longest one to finish. NVIDIA names in-flight batching as one way to reduce this waste.
What NVIDIA says (2)
“As a result, all requests in the batch must wait until the longest request is finished, which can be exacerbated by a large variance in the generation lengths.”
“There are methods to mitigate this, such as in-flight batching.”
Model parallelism means splitting one model across several graphics processing units (GPUs). NVIDIA notes this reduces the per-device memory footprint of the weights. Tensor parallelism splits individual layers into blocks on different devices. Pipeline parallelism places groups of layers on separate devices.
What NVIDIA says (3)
“One way to reduce the per-device memory footprint of the model weights is to distribute the model over several GPUs.”
“Pipeline parallelism involves sharding the model (vertically) into chunks, where each chunk comprises a subset of layers that is executed on a separate device.”
“Tensor parallelism involves sharding (horizontally) individual layers of the model into smaller, independent blocks of computation that can be executed on different devices.”
Key terms: KV cache In-flight batching Tensor / pipeline parallelism
Try it: KV cache lab
1.2 Finding insights in large datasets
How exploration, mining and visualization turn raw data into findings.
Key points
NVIDIA's exploratory data analysis (EDA) tutorial says you must explore whether the data has major gaps, either missing or invalid inputs, because those issues affect whether the data can be used as a reliable source.
What NVIDIA says (1)
“However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”
NVIDIA's machine-learning glossary describes unsupervised learning as working without labeled data to find previously unknown patterns, with clustering (for example customer segmentation) as a common task.
What NVIDIA says (2)
“Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”
“An example of clustering is a company that wants to segment its customers in order to better tailor products and offerings.”
NVIDIA's RAPIDS visualization guide says visuals should be used throughout exploration, not just at the end. Visualization is good at finding outliers, anomalies and patterns.
What NVIDIA says (2)
“While data visuals are an effective tool for explaining data insights at the end of a project, they should ideally be used throughout the data exploration and enriching process.”
“Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”
Key terms: Unsupervised learning
1.3 Building LLM applications: RAG, chatbots, summarizers
How the common LLM application patterns fit together.
Key points
RAG (retrieval-augmented generation) retrieves relevant passages at query time and gives them to the large language model (LLM). NVIDIA notes it gives models sources they can cite, is faster and cheaper than retraining, and lets you hot-swap new sources.
What NVIDIA says (3)
“Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”
“That makes the method faster and less expensive than retraining a model with additional datasets. And it lets users hot-swap new sources on the fly.”
“when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
NVIDIA calls a very plausible but incorrect answer a hallucination. It says retrieval-augmented generation (RAG) reduces the chance of hallucination.
What NVIDIA says (1)
“It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”
NVIDIA's retrieval-augmented generation (RAG) 101 post explains that document ingestion happens offline, while retrieval and response generation happen when an online query comes in.
What NVIDIA says (1)
“The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”
In retrieval-augmented generation (RAG), the large language model (LLM) is the generative component: it writes the answer using the user query and the contextual information retrieved from the vector database.
What NVIDIA says (1)
“In the context of RAG, LLMs are used to generate fully formed responses based on the user query and contextual information retrieved from the vector DBs during user queries.”
Key terms: Large language model Retrieval-augmented generation Hallucination
Try it: RAG lab
1.4 Preparing content for RAG
How documents are split, embedded and stored so retrieval works.
Key points
NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.
What NVIDIA says (2)
“One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”
“The retrieval step typically examines smaller chunks of the original text rather than all documents.”
NVIDIA's chunking study says poor chunking can lead to irrelevant or incomplete responses. A smart strategy improves retrieval precision and contextual coherence.
What NVIDIA says (2)
“When done poorly, chunking can lead to irrelevant or incomplete responses, frustrating users and undermining trust in the system.”
“a smart chunking strategy improves retrieval precision and contextual coherence”
NVIDIA's vector-database glossary: ingested data is chunked, a vector is created to represent each chunk, and chunks and vectors are stored together with optional metadata.
What NVIDIA says (1)
“When private enterprise data is ingested, it’s chunked, a vector is created to represent it, and the data chunks with their corresponding vectors are stored in a vector database along with optional metadata for later retrieval.”
Key terms: Vector database Retrieval-augmented generation Chunking
Try it: RAG lab Chunking lab
1.5 Machine-learning fundamentals
Core ideas you need before any LLM work: task types, model comparison and validation.
Key points
Cross-validation splits the training data into several subsets so the model can be tested and validated in turns. scikit-learn pairs it with grid search to find the best hyperparameters and evaluate models.
What NVIDIA says (2)
“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”
“scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”
Regression predicts a continuous numeric value from one or more features. NVIDIA's glossary uses linear regression to estimate house price (the label) from house size (the feature).
What NVIDIA says (2)
“Regression estimates the relationship between a target outcome label and one or more feature variables to predict a continuous numeric value.”
“linear regression is used to estimate the house price (the label) based on the house size (the feature).”
NVIDIA's XGBoost glossary: random-forest bagging minimizes variance and overfitting, while gradient-boosted decision trees (GBDT) boosting minimizes bias and underfitting.
What NVIDIA says (1)
“Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”
NVIDIA's scikit-learn glossary says feature scaling or encoding categorical variables prepares input data for optimal model performance.
What NVIDIA says (1)
“For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”
Key terms: Supervised learning Unsupervised learning Cross-validation Bagging / boosting
1.6 Python NLP and vector tools
What the main Python language libraries and vector databases do.
Key points
A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.
What NVIDIA says (2)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
“Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
NeMo's token-classification guide defines named entity recognition (NER) as detecting and classifying key information (entities) in text, with the example that “Mary” is a person, “Santa Clara” a location and “NVIDIA” a company. Libraries such as spaCy also provide NER pipelines.
What NVIDIA says (3)
“NER, also referred to as entity chunking, identification, or extraction, is the task of detecting and classifying key information (entities) in text.”
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
“is a person, “Santa Clara” is a location, and “NVIDIA” is a company.”
NVIDIA's scikit-learn glossary states that scikit-learn is built on NumPy, which is optimized for high-performance linear algebra and array operations.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
Key terms: Vector database NumPy spaCy Named-entity recognition
1.7 Reading research and spotting trends
The ideas behind modern LLMs and the trends NVIDIA highlights.
Key points
Attention (self-attention) lets a model learn how even distant elements in a sequence relate to each other. NVIDIA's explainer credits the 2017 paper with defining transformers, and NVIDIA's exam page lists “Attention Is All You Need” as suggested reading.
What NVIDIA says (3)
“Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”
“Attention Is All You Need”
“one of eight co-authors of the 2017 paper that defined transformers.”
NVIDIA describes foundation models as neural networks trained on massive unlabeled datasets that, with a little fine-tuning, handle jobs from translating text to analyzing medical images.
What NVIDIA says (2)
“Foundation models are AI neural networks trained on massive unlabeled datasets to handle a wide variety of jobs from translating text to analyzing medical images.”
“With a little fine-tuning, foundation models can handle jobs from translating text to analyzing medical images to performing agent-based behaviors.”
NVIDIA's MoE glossary explains that a learned routing mechanism sparsely selects which subnetworks participate instead of running every parameter on every step.
What NVIDIA says (1)
“Instead of running every parameter on every step, a learned routing mechanism sparsely selects which subnetworks should participate, allowing the model to grow capacity without paying the full compute cost.”
Key terms: Transformer Attention Foundation model Mixture of experts
1.8 Text embeddings
What embeddings are, how similarity is measured, and how to choose a model.
Key points
Embedding generation converts data into high-dimensional vectors representing text numerically; the query's vector is compared with stored vectors to find relevant information.
What NVIDIA says (2)
“Generating embeddings involves converting data into high-dimensional vectors, which represent text in a numerical format.”
“The system identifies relevant information by comparing the query vector with the stored vectors in the vector DBs.”
Cosine similarity focuses on the angle between vectors, so it captures semantic similarity by orientation. NVIDIA's glossary calls it ideal for text processing and information retrieval.
What NVIDIA says (2)
“Focuses on the angle between vectors. Ideal for text processing and information retrieval, capturing semantic similarities based on orientation rather than traditional distance.”
“Cosine similarity: Focuses on the angle between vectors.”
NVIDIA's vector-database glossary states the selection among embedding techniques depends on application needs, balancing semantic depth, computational efficiency, data types and dimensionality.
What NVIDIA says (1)
“The selection among embedding techniques depends on application needs, balancing factors like semantic depth, computational efficiency, the types of data to be encoded, and dimensionality.”
Key terms: Embedding Cosine similarity
Try it: RAG lab
1.9 Prompt engineering
How to word prompts and set sampling options to get the result you want.
Key points
NVIDIA describes chain-of-thought prompting as providing few-shot examples where the reasoning process is explained, so the large language model (LLM) shows its reasoning when it answers.
What NVIDIA says (1)
“Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”
At a lower temperature the model is more conservative and only picks higher-probability tokens. NVIDIA notes lower temperatures suit definitive tasks like question answering or summarization.
What NVIDIA says (2)
“Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”
“At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”
NVIDIA defines zero-shot as prompting the model without any example of the expected behavior. For example, a zero-shot prompt simply asks a question.
What NVIDIA says (2)
“Zero-shot means prompting the model without any example of expected behavior from the model.”
“For example, a zero-shot prompt asks a question.”
System prompting adds a system-level prompt to the user prompt. It gives the large language model (LLM) specific, detailed instructions on how to behave.
What NVIDIA says (1)
“This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”
Key terms: Prompt engineering Zero-shot prompt Few-shot prompt Chain-of-thought prompting Temperature System prompt
Try it: Sampling lab
1.10 Traditional ML with Python packages
How scikit-learn style workflows are built.
Key points
NVIDIA's scikit-learn glossary defines an estimator as the core machine-learning algorithm that fits the training data to produce a model.
What NVIDIA says (1)
“An estimator is the core machine learning algorithm that fits the training data to produce a model.”
Pipelines chain transformers and estimators so preprocessing, training and prediction stay consistent, which makes workflows reproducible.
What NVIDIA says (1)
“Pipelines in scikit-learn chain transformers and estimators into a cohesive workflow, ensuring consistent preprocessing, training, and prediction steps.”
NVIDIA's scikit-learn glossary names principal component analysis (PCA) as a dimensionality-reduction technique that reduces the number of variables while retaining meaningful patterns.
What NVIDIA says (1)
“For datasets with many features, dimensionality reduction techniques, such as PCA, can simplify the input data by reducing the number of variables while retaining meaningful patterns.”
Key terms: scikit-learn Principal component analysis