3.2 Preparing large datasets

NCA-GENL · Experimentation (22% of the exam) · Official objective: “Assist in preparing (e.g., scraping, tokenization) large datasets for pretraining, fine-tuning, and RLHF.”

Cleaning, deduplicating and tokenizing data for pretraining and fine-tuning.

Key points

  1. Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.

    What NVIDIA says (5)

    “Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”

    — Mastering LLM Techniques: Training

    “Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”

    — Mastering LLM Techniques: Training

    “Character-based tokenizers lead to long sequences”

    — Mastering LLM Techniques: Training

    “The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”

    — Mastering LLM Techniques: Training

    “the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”

    — Mastering LLM Techniques: Training

  2. Exact deduplication hashes each document and keeps one per hash bucket, so it only removes identical copies. Fuzzy deduplication finds near-duplicates using MinHash signatures and Locality-Sensitive Hashing (LSH).

    What NVIDIA says (2)

    “Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH) to identify similar documents.”

    — Mastering LLM Techniques: Text Data Processing

    “This method generates hash signatures for each document and groups documents by their hashes into buckets”

    — Mastering LLM Techniques: Text Data Processing

  3. NVIDIA's data-processing guide says deduplication improves training efficiency, cuts compute cost and keeps data diverse. It also helps prevent overfitting to repeated content.

    What NVIDIA says (2)

    “Deduplication is essential for improving model training efficiency, reducing computational costs, and ensuring data diversity.”

    — Mastering LLM Techniques: Text Data Processing

    “It helps prevent models from overfitting to repeated content and improves generalization.”

    — Mastering LLM Techniques: Text Data Processing

  4. NVIDIA lists Unicode fixing and language identification as crucial early steps when curating large web-scraped text corpora.

    What NVIDIA says (1)

    “Unicode fixing and language identification represent crucial early steps in the data curation pipeline, particularly when dealing with large-scale web-scraped text corpora.”

    — Mastering LLM Techniques: Text Data Processing

  5. Heuristic filtering uses rule-based metrics and statistics to remove low-quality content. It checks things like document length, repetition patterns and punctuation distribution.

    What NVIDIA says (2)

    “Heuristic filtering employs rule-based metrics and statistical measures to identify and remove low-quality content.”

    — Mastering LLM Techniques: Text Data Processing

    “such as document length, repetition patterns, punctuation distribution”

    — Mastering LLM Techniques: Text Data Processing

  6. Large language models (LLMs) are evaluated on unseen test data. Decontamination deals with test data leaking into the training data, which would make evaluation results unreliable.

    What NVIDIA says (2)

    “Downstream task decontamination is a step that addresses the potential leakage of test data into training datasets”

    — Mastering LLM Techniques: Text Data Processing

    “After training, LLMs are usually evaluated by their performance on downstream tasks consisting of unseen test data.”

    — Mastering LLM Techniques: Text Data Processing

  7. Personally identifiable information (PII) ranges from direct identifiers like names and social security numbers to indirect ones. NVIDIA says PII redaction removes it to protect privacy and comply with data-protection regulations.

    What NVIDIA says (2)

    “Personally Identifiable Information (PII) redaction involves identifying and removing sensitive information from datasets to protect individual privacy and ensure compliance with data protection regulations.”

    — Mastering LLM Techniques: Text Data Processing

    “from direct identifiers like names and social security numbers to indirect identifiers”

    — Mastering LLM Techniques: Text Data Processing

Key terms

Try it

Sample question

Why do most modern large language models (LLMs) use subword tokenizers such as Byte Pair Encoding (BPE) or WordPiece?

Show the answer

Answer: Word tokenizers give huge vocabularies and out-of-vocabulary words, character tokenizers give long sequences; subwords split rare words into meaningful pieces

Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.

What NVIDIA says (5)

“Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”

— Mastering LLM Techniques: Training

“Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”

— Mastering LLM Techniques: Training

“Character-based tokenizers lead to long sequences”

— Mastering LLM Techniques: Training

“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”

— Mastering LLM Techniques: Training

“the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”

— Mastering LLM Techniques: Training

Practice 3.2 (7 questions) Full Experimentation guide

← 3.1 Training and training optimization · 3.3 Testing LLM applications →