1.2 Finding insights in large datasets
How exploration, mining and visualization turn raw data into findings.
Key points
NVIDIA's exploratory data analysis (EDA) tutorial says you must explore whether the data has major gaps, either missing or invalid inputs, because those issues affect whether the data can be used as a reliable source.
What NVIDIA says (1)
“However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”
NVIDIA's machine-learning glossary describes unsupervised learning as working without labeled data to find previously unknown patterns, with clustering (for example customer segmentation) as a common task.
What NVIDIA says (2)
“Unsupervised learning, also called descriptive analytics, doesn’t have labeled data provided in advance, and can aid data scientists in finding previously unknown patterns in data.”
“An example of clustering is a company that wants to segment its customers in order to better tailor products and offerings.”
NVIDIA's RAPIDS visualization guide says visuals should be used throughout exploration, not just at the end. Visualization is good at finding outliers, anomalies and patterns.
What NVIDIA says (2)
“While data visuals are an effective tool for explaining data insights at the end of a project, they should ideally be used throughout the data exploration and enriching process.”
“Visualization excels at enhancing data understanding by finding outliers, anomalies, and patterns”
Key terms
- Unsupervised learning: Finding patterns in data without labels, as in clustering.
Sample question
You start exploratory data analysis (EDA) on a large sensor dataset. Per NVIDIA's cuDF EDA walkthrough, what should you check early because it decides whether the data can be relied on?
Show the answer
Answer: Whether the data has major gaps from missing or invalid values
NVIDIA's exploratory data analysis (EDA) tutorial says you must explore whether the data has major gaps, either missing or invalid inputs, because those issues affect whether the data can be used as a reliable source.
What NVIDIA says (1)
“However, you must still explore whether this data has major gaps, either with missing or invalid data inputs. These issues affect whether this data can be used as a reliable source on its own.”
Practice 1.2 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 1.1 Serving LLMs at scale · 1.3 Building LLM applications: RAG, chatbots, summarizers →