3.6 Python packages for traditional ML
XGBoost, pandas and scikit-learn estimators.
Key points
Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.
What NVIDIA says (1)
“is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”
pandas provides DataFrames, which are tables with labeled rows and columns. It is a common first step before traditional ML. ML means machine learning.
What NVIDIA says (1)
“pandas is the most popular software library for data manipulation and data analysis for the Python programming languages.”
Every scikit-learn model, such as a random forest, is an estimator with fit and predict methods.
What NVIDIA says (1)
“An estimator is the core machine learning algorithm that fits the training data to produce a model.”
Key terms
- scikit-learn: A Python library for traditional machine learning built on NumPy.
- Estimator: A scikit-learn algorithm object that fits data to produce a model.
- XGBoost: A library for gradient-boosted decision trees.
- pandas: The most popular Python library for working with tables of data.
Sample question
What is XGBoost?
Show the answer
Answer: A scalable, distributed gradient-boosted decision tree library
Gradient boosting builds many small decision trees, each fixing the errors of the last. XGBoost is a popular library for it on tabular data.
What NVIDIA says (1)
“is a scalable, distributed gradient-boosted decision tree (GBDT) machine learning library.”
Practice 3.6 (3 questions) Full Multimodal Data guide
← 3.5 Monitoring data collection and experiments · 3.7 Writing software components and scripts →