1.2 Cleaning data and handling quality and governance

NCA-ADS · Data Manipulation and Preparation (23% of the exam) · Official objective: “Data cleaning, quality handling, and governance compliance”

Nulls in cuDF, dropna and fillna, duplicates, and protecting personal data.

Key points

  1. A missing value is a cell with no data, also called a null. cuDF allows nulls in every data type and shows them as <NA>. NA means not available (missing).

    What NVIDIA says (2)

    “cudf supports having missing values in all dtypes.”

    — cuDF: Working with missing data

    “To detect missing values, you can use isna() and notna() functions.”

    — cuDF: Working with missing data

  2. dropna removes rows or columns with nulls. how='any' drops a row with at least one null; how='all' drops only fully empty rows.

    What NVIDIA says (1)

    “any (default) drops rows (or columns) containing at least one null value. all drops only rows (or columns) containing all null values.”

    — cuDF API: DataFrame.dropna

  3. Imputation means filling missing values with a chosen value. fillna takes a scalar, a Series, or a dict of per-column values.

    What NVIDIA says (1)

    “A dict can be used to provide different values to fill nulls in different columns.”

    — cuDF API: DataFrame.fillna

  4. Data governance means rules about who may use which data, and how. Compliance means following those rules and the law, such as privacy rules for personal data.

    What NVIDIA says (2)

    “In addition to complying with privacy and consumer protection laws, trustworthy AI models are tested for safety, security and mitigation of unwanted bias.”

    — What Is Trustworthy AI?

    “Ensuring that this data is free from duplicates, personal identifiable information (PII), and toxic content is crucial.”

    — Enhancing Generative AI Model Accuracy with NVIDIA NeMo Curator

  5. Data quality means data is correct, complete and free of junk such as duplicates. drop_duplicates() keeps the first copy by default and drops the rest.

    What NVIDIA says (2)

    “Training models on these datasets without proper processing can result in higher training time and lower model quality.”

    — Mastering LLM Techniques: Text Data Processing

    “Determines which duplicates (if any) to keep. - ‘first’ : Drop duplicates except for the first occurrence.”

    — cuDF API: DataFrame.drop_duplicates

Key terms

Sample question

How does cuDF represent a missing value, and how do you find missing values?

Show the answer

Answer: As <NA> (a null) in any data type; find them with isna() and notna()

A missing value is a cell with no data, also called a null. cuDF allows nulls in every data type and shows them as <NA>. NA means not available (missing).

What NVIDIA says (2)

“cudf supports having missing values in all dtypes.”

— cuDF: Working with missing data

“To detect missing values, you can use isna() and notna() functions.”

— cuDF: Working with missing data

Practice 1.2 (5 questions) Full Data Manipulation and Preparation guide

← 1.1 Joining and manipulating data with cuDF and pandas · 1.3 GPU-accelerated ETL with RAPIDS, Dask and Spark →