3.2 Preparing large datasets
Cleaning, deduplicating and tokenizing data for pretraining and fine-tuning.
Key points
Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.
What NVIDIA says (5)
“Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”
“Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”
“Character-based tokenizers lead to long sequences”
“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”
“the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”
Exact deduplication hashes each document and keeps one per hash bucket, so it only removes identical copies. Fuzzy deduplication finds near-duplicates using MinHash signatures and Locality-Sensitive Hashing (LSH).
What NVIDIA says (2)
“Fuzzy deduplication addresses near-duplicate content using MinHash signatures and Locality-Sensitive Hashing (LSH) to identify similar documents.”
“This method generates hash signatures for each document and groups documents by their hashes into buckets”
NVIDIA's data-processing guide says deduplication improves training efficiency, cuts compute cost and keeps data diverse. It also helps prevent overfitting to repeated content.
What NVIDIA says (2)
“Deduplication is essential for improving model training efficiency, reducing computational costs, and ensuring data diversity.”
“It helps prevent models from overfitting to repeated content and improves generalization.”
NVIDIA lists Unicode fixing and language identification as crucial early steps when curating large web-scraped text corpora.
What NVIDIA says (1)
“Unicode fixing and language identification represent crucial early steps in the data curation pipeline, particularly when dealing with large-scale web-scraped text corpora.”
Heuristic filtering uses rule-based metrics and statistics to remove low-quality content. It checks things like document length, repetition patterns and punctuation distribution.
What NVIDIA says (2)
“Heuristic filtering employs rule-based metrics and statistical measures to identify and remove low-quality content.”
“such as document length, repetition patterns, punctuation distribution”
Large language models (LLMs) are evaluated on unseen test data. Decontamination deals with test data leaking into the training data, which would make evaluation results unreliable.
What NVIDIA says (2)
“Downstream task decontamination is a step that addresses the potential leakage of test data into training datasets”
“After training, LLMs are usually evaluated by their performance on downstream tasks consisting of unseen test data.”
Personally identifiable information (PII) ranges from direct identifiers like names and social security numbers to indirect ones. NVIDIA says PII redaction removes it to protect privacy and comply with data-protection regulations.
What NVIDIA says (2)
“Personally Identifiable Information (PII) redaction involves identifying and removing sensitive information from datasets to protect individual privacy and ensure compliance with data protection regulations.”
“from direct identifiers like names and social security numbers to indirect identifiers”
Key terms
- Token: A small unit of text, such as a word or part of a word, that a model processes.
- Tokenization: Splitting text into tokens and mapping them to numeric IDs the model can use.
- Byte Pair Encoding: A subword tokenizer that starts from characters and repeatedly merges the most frequent adjacent pairs.
- NeMo Curator: NVIDIA's GPU-accelerated toolkit for curating LLM training data.
- Deduplication (exact / fuzzy / semantic): Removing duplicate or near-duplicate documents from training data.
- Personally identifiable information: Data that can identify a person, such as a name or social security number.
- Decontamination: Removing test data that leaked into training data.
Try it
Sample question
Why do most modern large language models (LLMs) use subword tokenizers such as Byte Pair Encoding (BPE) or WordPiece?
Show the answer
Answer: Word tokenizers give huge vocabularies and out-of-vocabulary words, character tokenizers give long sequences; subwords split rare words into meaningful pieces
Tokenization splits text into smaller units called tokens. NVIDIA explains that word tokenizers create huge vocabularies and out-of-vocabulary tokens, while character tokenizers create long sequences. Subword tokenizers split rare words into meaningful subwords. Popular ones are byte-pair encoding (BPE), WordPiece, Unigram and SentencePiece.
What NVIDIA says (5)
“Word-based tokenizers lead to a large vocabulary size and words not seen during the tokenizer training process cause many out-of-vocabulary tokens.”
“Popular subword tokenization algorithms include Byte Pair Encoding (BPE), WordPiece, Unigram, and SentencePiece.”
“Character-based tokenizers lead to long sequences”
“The focus of subword tokenization algorithms is to split rare words into smaller, meaningful subwords”
“the process of tokenization assumes a pivotal role in fragmenting text into smaller units known as tokens.”
Practice 3.2 (7 questions) Full Experimentation guide
← 3.1 Training and training optimization · 3.3 Testing LLM applications →