1.5 Machine-learning fundamentals

NCA-GENL · Core Machine Learning and AI Knowledge (30% of the exam) · Official objective: “Familiarity with the fundamentals of machine learning (e.g., feature engineering, model comparison, cross validation).”

Core ideas you need before any LLM work: task types, model comparison and validation.

Key points

  1. Cross-validation splits the training data into several subsets so the model can be tested and validated in turns. scikit-learn pairs it with grid search to find the best hyperparameters and evaluate models.

    What NVIDIA says (2)

    “Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

    — What is scikit-learn?

    “scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”

    — What is scikit-learn?

  2. Regression predicts a continuous numeric value from one or more features. NVIDIA's glossary uses linear regression to estimate house price (the label) from house size (the feature).

    What NVIDIA says (2)

    “Regression estimates the relationship between a target outcome label and one or more feature variables to predict a continuous numeric value.”

    — What is Machine Learning and Why Does It Matter?

    “linear regression is used to estimate the house price (the label) based on the house size (the feature).”

    — What is Machine Learning and Why Does It Matter?

  3. NVIDIA's XGBoost glossary: random-forest bagging minimizes variance and overfitting, while gradient-boosted decision trees (GBDT) boosting minimizes bias and underfitting.

    What NVIDIA says (1)

    “Random forest “bagging” minimizes the variance and overfitting, while GBDT “boosting” minimizes the bias and underfitting.”

    — What Is XGBoost and Why Does It Matter?

  4. NVIDIA's scikit-learn glossary says feature scaling or encoding categorical variables prepares input data for optimal model performance.

    What NVIDIA says (1)

    “For example, feature scaling or encoding categorical variables prepares the subset of input data for optimal performance.”

    — What is scikit-learn?

Key terms

Sample question

What does cross-validation do during model selection?

Show the answer

Answer: Splits the training data into multiple subsets so the model is trained and validated iteratively

Cross-validation splits the training data into several subsets so the model can be tested and validated in turns. scikit-learn pairs it with grid search to find the best hyperparameters and evaluate models.

What NVIDIA says (2)

“Cross-validation splits the training data into multiple subsets, allowing iterative testing and validation.”

— What is scikit-learn?

“scikit-learn incorporates tools like grid search and cross-validation to identify the best hyperparameters and evaluate model performance.”

— What is scikit-learn?

Practice 1.5 (4 questions) Full Core Machine Learning and AI Knowledge guide

← 1.4 Preparing content for RAG · 1.6 Python NLP and vector tools →