2.3 Python language packages in practice

NCA-GENL · Software Development (24% of the exam) · Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

Tokenizing, entity extraction and DataFrames in real pipelines.

Key points

  1. Subword tokenization splits words into smaller, reusable pieces, for example “anyplace” into “any” and “place”. NVIDIA's RAPIDS natural language processing (NLP) post describes a graphics processing unit (GPU) Bidirectional Encoder Representations from Transformers (BERT) subword tokenizer, so tokenization runs on the GPU too.

    What NVIDIA says (3)

    “We first introduced the GPU BERT subword tokenizer in a previous”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “To solve these problems, we use subword tokenization.”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “For example, the word “anyplace” can be broken down into “any” and “place,””

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  2. The post explains that spaCy needed its inputs on the central processing unit (CPU), so the graphics processing unit (GPU) pipeline had to copy data to CPU memory and back, which slowed it down.

    What NVIDIA says (1)

    “spaCy currently needs your inputs on CPU and thus was slow as it required a copy to CPU memory and back to GPU memory.”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  3. NVIDIA's pandas glossary describes the DataFrame as a two-dimensional, array-like table where each column represents a variable and each row a set of values.

    What NVIDIA says (1)

    “A pandas DataFrame is a two-dimensional, array-like table where each column represents values of a specific variable”

    — What Is Pandas and Why Does it Matter?

Key terms

Sample question

Which NVIDIA-documented capability does cuDF provide for text work before feeding a Bidirectional Encoder Representations from Transformers (BERT)-style model?

Show the answer

Answer: GPU subword tokenization that keeps text processing on the GPU

Subword tokenization splits words into smaller, reusable pieces, for example “anyplace” into “any” and “place”. NVIDIA's RAPIDS natural language processing (NLP) post describes a graphics processing unit (GPU) Bidirectional Encoder Representations from Transformers (BERT) subword tokenizer, so tokenization runs on the GPU too.

What NVIDIA says (3)

“We first introduced the GPU BERT subword tokenizer in a previous”

— Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

“To solve these problems, we use subword tokenization.”

— Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

“For example, the word “anyplace” can be broken down into “any” and “place,””

— Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

Practice 2.3 (3 questions) Full Software Development guide

← 2.2 Building LLM features in software · 2.4 Picking the right components →