Core Machine Learning and AI Knowledge

30% of the NCA-GENL exam. Read each objective's key points, open “What NVIDIA says” to see the source, then practise.

Core Machine Learning and AI Knowledge · Software Development · Experimentation · Data Analysis and Visualization · Trustworthy AI

1.1 Serving LLMs at scale

Official objective: “Assist in deployment and evaluation of model scalability, performance, and reliability under the supervision of senior team members.”

How to judge whether an LLM deployment is fast, scalable and reliable, and what limits it.

Key points

  1. NVIDIA's inference-optimization post states that the two main contributors to graphics processing unit (GPU) memory during large language model (LLM) inference are the model weights and the key-value (KV) cache. The KV cache grows with batch size and sequence length, which is why bigger batches eventually overflow memory.

    What NVIDIA says (2)

    “In effect, the two main contributors to the GPU LLM memory requirement are model weights and the KV cache.”

    — Mastering LLM Techniques: Inference Optimization

    “Growing linearly with batch size and sequence length, the memory requirement can quickly scale.”

    — Mastering LLM Techniques: Inference Optimization

  2. Requests in a batch generate different numbers of tokens, so with static batching every request waits for the longest one to finish. NVIDIA names in-flight batching as one way to reduce this waste.

    What NVIDIA says (2)

    “As a result, all requests in the batch must wait until the longest request is finished, which can be exacerbated by a large variance in the generation lengths.”

    — Mastering LLM Techniques: Inference Optimization

    “There are methods to mitigate this, such as in-flight batching.”

    — Mastering LLM Techniques: Inference Optimization

  3. Model parallelism means splitting one model across several graphics processing units (GPUs). NVIDIA notes this reduces the per-device memory footprint of the weights. Tensor parallelism splits individual layers into blocks on different devices. Pipeline parallelism places groups of layers on separate devices.

    What NVIDIA says (3)

    “One way to reduce the per-device memory footprint of the model weights is to distribute the model over several GPUs.”

    — Mastering LLM Techniques: Inference Optimization

    “Pipeline parallelism involves sharding the model (vertically) into chunks, where each chunk comprises a subset of layers that is executed on a separate device.”

    — Mastering LLM Techniques: Inference Optimization

    “Tensor parallelism involves sharding (horizontally) individual layers of the model into smaller, independent blocks of computation that can be executed on different devices.”

    — Mastering LLM Techniques: Inference Optimization

Key terms: KV cache In-flight batching Tensor / pipeline parallelism

Try it: KV cache lab

Practice 1.1 (3 questions)

1.2 Finding insights in large datasets

Official objective: “Awareness of the process of extracting insights from large datasets using data mining, data visualization, and similar techniques.”

How exploration, mining and visualization turn raw data into findings.

Key points

  1. NVIDIA's exploratory data analysis (EDA) tutorial says you must explore whether the data has major gaps, either missing or invalid inputs, because those issues affect whether the data can be used as a reliable source.

    What NVIDIA says (1)

    “However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”

    — Accelerated Data Analytics: Speed Up Data Exploration with RAPIDS cuDF

  2. NVIDIA's machine-learning glossary describes unsupervised learning as working without labeled data to find previously unknown patterns, with clustering (for example customer segmentation) as a common task.

    What NVIDIA says (2)

    “Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”

    — What is Machine Learning and Why Does It Matter?

    “An example of clustering is a company that wants to segment its customers in order to better tailor products and offerings.”

    — What is Machine Learning and Why Does It Matter?

  3. NVIDIA's RAPIDS visualization guide says visuals should be used throughout exploration, not just at the end. Visualization is good at finding outliers, anomalies and patterns.

    What NVIDIA says (2)

    “While data visuals are an effective tool for explaining data insights at the end of a project, they should ideally be used throughout the data exploration and enriching process.”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

    “Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”

    — Accelerated Data Analytics: A Guide to Data Visualization with RAPIDS

Key terms: Unsupervised learning

Practice 1.2 (3 questions)

1.3 Building LLM applications: RAG, chatbots, summarizers

Official objective: “Build LLM use cases such as retrieval-augmented generation (RAG), chatbots, and summarizers.”

How the common LLM application patterns fit together.

Key points

  1. RAG (retrieval-augmented generation) retrieves relevant passages at query time and gives them to the large language model (LLM). NVIDIA notes it gives models sources they can cite, is faster and cheaper than retraining, and lets you hot-swap new sources.

    What NVIDIA says (3)

    “Retrieval-augmented generation gives models sources they can cite, like footnotes in a research paper, so users can check any claims.”

    — What Is Retrieval-Augmented Generation aka RAG

    “That makes the method faster and less expensive than retraining a model with additional datasets. And it lets users hot-swap new sources on the fly.”

    — What Is Retrieval-Augmented Generation aka RAG

    “when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

  2. NVIDIA calls a very plausible but incorrect answer a hallucination. It says retrieval-augmented generation (RAG) reduces the chance of hallucination.

    What NVIDIA says (1)

    “It also reduces the possibility that a model will give a very plausible but incorrect answer, a phenomenon called hallucination.”

    — What Is Retrieval-Augmented Generation aka RAG

  3. NVIDIA's retrieval-augmented generation (RAG) 101 post explains that document ingestion happens offline, while retrieval and response generation happen when an online query comes in.

    What NVIDIA says (1)

    “The process of document ingestion occurs offline, and when an online query comes in, the retrieval of relevant documents and the generation of a response occurs.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

  4. In retrieval-augmented generation (RAG), the large language model (LLM) is the generative component: it writes the answer using the user query and the contextual information retrieved from the vector database.

    What NVIDIA says (1)

    “In the context of RAG, LLMs are used to generate fully formed responses based on the user query and contextual information retrieved from the vector DBs during user queries.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

Key terms: Large language model Retrieval-augmented generation Hallucination

Try it: RAG lab

Practice 1.3 (4 questions)

1.4 Preparing content for RAG

Official objective: “Curate and embed content datasets for RAGs.”

How documents are split, embedded and stored so retrieval works.

Key points

  1. NVIDIA notes that text splitting breaks long text into smaller segments, which is necessary for fitting the text into the embedding model; the retrieval step then works on these smaller chunks rather than whole documents.

    What NVIDIA says (2)

    “One transformation method is text-splitting, which breaks down long text into smaller segments. This is necessary for fitting the text into the embedding model,”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

    “The retrieval step typically examines smaller chunks of the original text rather than all documents.”

    — Enhancing RAG Pipelines with Re-Ranking

  2. NVIDIA's chunking study says poor chunking can lead to irrelevant or incomplete responses. A smart strategy improves retrieval precision and contextual coherence.

    What NVIDIA says (2)

    “When done poorly, chunking can lead to irrelevant or incomplete responses, frustrating users and undermining trust in the system.”

    — Finding the Best Chunking Strategy for Accurate AI Responses

    “a smart chunking strategy improves retrieval precision and contextual coherence”

    — Finding the Best Chunking Strategy for Accurate AI Responses

  3. NVIDIA's vector-database glossary: ingested data is chunked, a vector is created to represent each chunk, and chunks and vectors are stored together with optional metadata.

    What NVIDIA says (1)

    “When private enterprise data is ingested, it’s chunked, a vector is created to represent it, and the data chunks with their corresponding vectors are stored in a vector database along with optional metadata for later retrieval.”

    — What is a Vector Database and How Does it Work?

Key terms: Vector database Retrieval-augmented generation Chunking

Try it: RAG lab Chunking lab

Practice 1.4 (3 questions)

1.5 Machine-learning fundamentals

Official objective: “Familiarity with the fundamentals of machine learning (e.g., feature engineering, model comparison, cross validation).”

Core ideas you need before any LLM work: task types, model comparison and validation.

Key points

  1. Cross-validation splits the training data into several subsets so the model can be tested and validated in turns. scikit-learn pairs it with grid search to find the best hyperparameters and evaluate models.

    What NVIDIA says (2)

    “Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

    — What is scikit-learn?

    “scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”

    — What is scikit-learn?

  2. Regression predicts a continuous numeric value from one or more features. NVIDIA's glossary uses linear regression to estimate house price (the label) from house size (the feature).

    What NVIDIA says (2)

    “Regression estimates the relationship between a target outcome label and one or more feature variables to predict a continuous numeric value.”

    — What is Machine Learning and Why Does It Matter?

    “linear regression is used to estimate the house price (the label) based on the house size (the feature).”

    — What is Machine Learning and Why Does It Matter?

  3. NVIDIA's XGBoost glossary: random-forest bagging minimizes variance and overfitting, while gradient-boosted decision trees (GBDT) boosting minimizes bias and underfitting.

    What NVIDIA says (1)

    “Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

    — What Is XGBoost and Why Does It Matter?

  4. NVIDIA's scikit-learn glossary says feature scaling or encoding categorical variables prepares input data for optimal model performance.

    What NVIDIA says (1)

    “For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”

    — What is scikit-learn?

Key terms: Supervised learning Unsupervised learning Cross-validation Bagging / boosting

Practice 1.5 (4 questions)

1.6 Python NLP and vector tools

Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

What the main Python language libraries and vector databases do.

Key points

  1. A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.

    What NVIDIA says (2)

    “A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

    — What is a Vector Database and How Does it Work?

    “Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

    — What is a Vector Database and How Does it Work?

  2. NeMo's token-classification guide defines named entity recognition (NER) as detecting and classifying key information (entities) in text, with the example that “Mary” is a person, “Santa Clara” a location and “NVIDIA” a company. Libraries such as spaCy also provide NER pipelines.

    What NVIDIA says (3)

    “NER, also referred to as entity chunking, identification, or extraction, is the task of detecting and classifying key information (entities) in text.”

    — Token Classification Model with Named Entity Recognition (NER) — NVIDIA NeMo Framework User Guide

    “relied on spaCy for NER but, spaCy currently needs your inputs on CPU”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “is a person, “Santa Clara” is a location, and “NVIDIA” is a company.”

    — Token Classification Model with Named Entity Recognition (NER) — NVIDIA NeMo Framework User Guide

  3. NVIDIA's scikit-learn glossary states that scikit-learn is built on NumPy, which is optimized for high-performance linear algebra and array operations.

    What NVIDIA says (1)

    “scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”

    — What is scikit-learn?

Key terms: Vector database NumPy spaCy Named-entity recognition

Practice 1.6 (3 questions)

1.7 Reading research and spotting trends

Official objective: “Read research papers (articles, conference papers, etc.) to identify emerging LLM trends and technologies.”

The ideas behind modern LLMs and the trends NVIDIA highlights.

Key points

  1. Attention (self-attention) lets a model learn how even distant elements in a sequence relate to each other. NVIDIA's explainer credits the 2017 paper with defining transformers, and NVIDIA's exam page lists “Attention Is All You Need” as suggested reading.

    What NVIDIA says (3)

    “Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.”

    — What Is a Transformer Model?

    “Attention Is All You Need”

    — NVIDIA-Certified Associate Generative AI LLMs Certification Exam

    “one of eight co-authors of the 2017 paper that defined transformers.”

    — What Is a Transformer Model?

  2. NVIDIA describes foundation models as neural networks trained on massive unlabeled datasets that, with a little fine-tuning, handle jobs from translating text to analyzing medical images.

    What NVIDIA says (2)

    “Foundation models are AI neural networks trained on massive unlabeled datasets to handle a wide variety of jobs from translating text to analyzing medical images.”

    — What Are Foundation Models?

    “With a little fine-tuning, foundation models can handle jobs from translating text to analyzing medical images to performing agent-based behaviors.”

    — What Are Foundation Models?

  3. NVIDIA's MoE glossary explains that a learned routing mechanism sparsely selects which subnetworks participate instead of running every parameter on every step.

    What NVIDIA says (1)

    “Instead of running every parameter on every step, a learned routing mechanism sparsely selects which subnetworks should participate, allowing the model to grow capacity without paying the full compute cost.”

    — What Is Mixture of Experts (MoE) and How It Works?

Key terms: Transformer Attention Foundation model Mixture of experts

Practice 1.7 (3 questions)

1.8 Text embeddings

Official objective: “Select and use models to create text embeddings.”

What embeddings are, how similarity is measured, and how to choose a model.

Key points

  1. Embedding generation converts data into high-dimensional vectors representing text numerically; the query's vector is compared with stored vectors to find relevant information.

    What NVIDIA says (2)

    “Generating embeddings involves converting data into high-dimensional vectors, which represent text in a numerical format.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

    “The system identifies relevant information by comparing the query vector with the stored vectors in the vector DBs.”

    — RAG 101: Demystifying Retrieval-Augmented Generation Pipelines

  2. Cosine similarity focuses on the angle between vectors, so it captures semantic similarity by orientation. NVIDIA's glossary calls it ideal for text processing and information retrieval.

    What NVIDIA says (2)

    “Focuses on the angle between vectors. Ideal for text processing and information retrieval, capturing semantic similarities based on orientation rather than traditional distance.”

    — What is a Vector Database and How Does it Work?

    “Cosine similarity: Focuses on the angle between vectors.”

    — What is a Vector Database and How Does it Work?

  3. NVIDIA's vector-database glossary states the selection among embedding techniques depends on application needs, balancing semantic depth, computational efficiency, data types and dimensionality.

    What NVIDIA says (1)

    “The selection among embedding techniques depends on application needs, balancing factors like semantic depth, computational efficiency, the types of data to be encoded, and dimensionality.”

    — What is a Vector Database and How Does it Work?

Key terms: Embedding Cosine similarity

Try it: RAG lab

Practice 1.8 (3 questions)

1.9 Prompt engineering

Official objective: “Use prompt engineering principles to create prompts to achieve desired results.”

How to word prompts and set sampling options to get the result you want.

Key points

  1. NVIDIA describes chain-of-thought prompting as providing few-shot examples where the reasoning process is explained, so the large language model (LLM) shows its reasoning when it answers.

    What NVIDIA says (1)

    “Do this by providing some few-shot examples, where the reasoning process is explained. When the LLM answers the prompt, it shows its reasoning process as well.”

    — An Introduction to Large Language Models: Prompt Engineering and P-Tuning

  2. At a lower temperature the model is more conservative and only picks higher-probability tokens. NVIDIA notes lower temperatures suit definitive tasks like question answering or summarization.

    What NVIDIA says (2)

    “Lower temperatures are suitable for more definitive tasks like question-answering or summarization.”

    — How to Get Better Outputs from Your Large Language Model

    “At a lower temperature, the model is more conservative and is limited to choosing tokens with higher probabilities.”

    — How to Get Better Outputs from Your Large Language Model

  3. NVIDIA defines zero-shot as prompting the model without any example of the expected behavior. For example, a zero-shot prompt simply asks a question.

    What NVIDIA says (2)

    “Zero-shot means prompting the model without any example of expected behavior from the model.”

    — An Introduction to Large Language Models: Prompt Engineering and P-Tuning

    “For example, a zero-shot prompt asks a question.”

    — An Introduction to Large Language Models: Prompt Engineering and P-Tuning

  4. System prompting adds a system-level prompt to the user prompt. It gives the large language model (LLM) specific, detailed instructions on how to behave.

    What NVIDIA says (1)

    “This approach involves adding a system-level prompt in addition to the user prompt to provide specific and detailed instructions to the LLMs to behave as intended.”

    — Mastering LLM Techniques: Customization

Key terms: Prompt engineering Zero-shot prompt Few-shot prompt Chain-of-thought prompting Temperature System prompt

Try it: Sampling lab

Practice 1.9 (4 questions)

1.10 Traditional ML with Python packages

Official objective: “Use Python packages (spaCy, NumPy, Keras, etc.) to implement specific traditional machine learning analyses.”

How scikit-learn style workflows are built.

Key points

  1. NVIDIA's scikit-learn glossary defines an estimator as the core machine-learning algorithm that fits the training data to produce a model.

    What NVIDIA says (1)

    “An estimator is the core machine learning algorithm that fits the training data to produce a model.”

    — What is scikit-learn?

  2. Pipelines chain transformers and estimators so preprocessing, training and prediction stay consistent, which makes workflows reproducible.

    What NVIDIA says (1)

    “Pipelines in scikit-learn chain transformers and estimators into a cohesive workflow, ensuring consistent preprocessing, training, and prediction steps.”

    — What is scikit-learn?

  3. NVIDIA's scikit-learn glossary names principal component analysis (PCA) as a dimensionality-reduction technique that reduces the number of variables while retaining meaningful patterns.

    What NVIDIA says (1)

    “For datasets with many features, dimensionality reduction techniques, such as PCA, can simplify the input data by reducing the number of variables while retaining meaningful patterns.”

    — What is scikit-learn?

Key terms: scikit-learn Principal component analysis

Practice 1.10 (3 questions)