2.3 Python language packages in practice
Tokenizing, entity extraction and DataFrames in real pipelines.
Key points
Subword tokenization splits words into smaller, reusable pieces, for example “anyplace” into “any” and “place”. NVIDIA's RAPIDS natural language processing (NLP) post describes a graphics processing unit (GPU) Bidirectional Encoder Representations from Transformers (BERT) subword tokenizer, so tokenization runs on the GPU too.
What NVIDIA says (3)
“We first introduced the GPU BERT subword tokenizer in a previous”
“To solve these problems, we use subword tokenization.”
“For example, the word “anyplace” can be broken down into “any” and “place,””
The post explains that spaCy needed its inputs on the central processing unit (CPU), so the graphics processing unit (GPU) pipeline had to copy data to CPU memory and back, which slowed it down.
What NVIDIA says (1)
“spaCy currently needs your inputs on CPU and thus was slow as it required a copy to CPU memory and back to GPU memory.”
NVIDIA's pandas glossary describes the DataFrame as a two-dimensional, array-like table where each column represents a variable and each row a set of values.
What NVIDIA says (1)
“A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable”
Key terms
- cuDF: RAPIDS GPU DataFrame library with a pandas-like API.
- pandas / DataFrame: Python library for tabular data; a DataFrame is a table of rows and columns.
- spaCy: Python NLP library; used here for named-entity recognition on the CPU.
Sample question
Which NVIDIA-documented capability does cuDF provide for text work before feeding a Bidirectional Encoder Representations from Transformers (BERT)-style model?
Show the answer
Answer: GPU subword tokenization that keeps text processing on the GPU
Subword tokenization splits words into smaller, reusable pieces, for example “anyplace” into “any” and “place”. NVIDIA's RAPIDS natural language processing (NLP) post describes a graphics processing unit (GPU) Bidirectional Encoder Representations from Transformers (BERT) subword tokenizer, so tokenization runs on the GPU too.
What NVIDIA says (3)
“We first introduced the GPU BERT subword tokenizer in a previous”
“To solve these problems, we use subword tokenization.”
“For example, the word “anyplace” can be broken down into “any” and “place,””
Practice 2.3 (3 questions) Full Software Development guide
← 2.2 Building LLM features in software · 2.4 Picking the right components →