1.6 Python NLP and vector tools

NCA-GENL · Core Machine Learning and AI Knowledge (30% of the exam) · Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

What the main Python language libraries and vector databases do.

Key points

  1. A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.

    What NVIDIA says (2)

    “A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

    — What is a Vector Database and How Does it Work?

    “Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

    — What is a Vector Database and How Does it Work?

  2. NeMo's token-classification guide defines named entity recognition (NER) as detecting and classifying key information (entities) in text, with the example that “Mary” is a person, “Santa Clara” a location and “NVIDIA” a company. Libraries such as spaCy also provide NER pipelines.

    What NVIDIA says (3)

    “NER, also referred to as entity chunking, identification, or extraction, is the task of detecting and classifying key information (entities) in text.”

    — Token Classification Model with Named Entity Recognition (NER) — NVIDIA NeMo Framework User Guide

    “relied on spaCy for NER but, spaCy currently needs your inputs on CPU”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

    “is a person, “Santa Clara” is a location, and “NVIDIA” is a company.”

    — Token Classification Model with Named Entity Recognition (NER) — NVIDIA NeMo Framework User Guide

  3. NVIDIA's scikit-learn glossary states that scikit-learn is built on NumPy, which is optimized for high-performance linear algebra and array operations.

    What NVIDIA says (1)

    “scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”

    — What is scikit-learn?

Key terms

Sample question

What is the main job of a vector database in a large language model (LLM) application?

Show the answer

Answer: Storing vector embeddings and retrieving those most similar to a query's embedding

A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.

What NVIDIA says (2)

“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

— What is a Vector Database and How Does it Work?

“Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

— What is a Vector Database and How Does it Work?

Practice 1.6 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 1.5 Machine-learning fundamentals · 1.7 Reading research and spotting trends →