3.6 Python packages for traditional ML

NCA-GENM · Multimodal Data (15% of the exam) · Official objective: “Use Python packages (spaCy, NumPy, Keras, etc.) to implement specific traditional machine learning analyses.”

XGBoost, pandas and scikit-learn estimators.

Key points

  1. Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.

    What NVIDIA says (1)

    “is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

    — What Is XGBoost and Why Does It Matter?

  2. pandas provides DataFrames, which are tables with labeled rows and columns. It is a common first step before traditional ML. ML means machine learning.

    What NVIDIA says (1)

    “pandas is the most popular software library for data manipulation and data analysis for the Python programming languages.”

    — What Is Pandas and Why Does it Matter?

  3. Every scikit-learn model, such as a random forest, is an estimator with fit and predict methods.

    What NVIDIA says (1)

    “An estimator is the core machine learning algorithm that fits the training data to produce a model.”

    — What is scikit-learn?

Key terms

Sample question

What is XGBoost?

Show the answer

Answer: A scalable, distributed gradient-boosted decision tree library

Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.

What NVIDIA says (1)

“is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”

— What Is XGBoost and Why Does It Matter?

Practice 3.6 (3 questions) Full Multimodal Data guide

← 3.5 Monitoring data collection and experiments · 3.7 Writing software components and scripts →