3.3 Python natural-language packages

NCA-GENM · Multimodal Data (15% of the exam) · Official objective: “Familiarity with the capabilities of Python natural language packages (spaCy, NumPy, vector databases, etc.).”

Vector databases, similarity search, spaCy and NumPy.

Key points

  1. An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.

    What NVIDIA says (1)

    “A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

    — What is a Vector Database and How Does it Work?

  2. Similarity search is also called vector or semantic search. It ranks stored vectors by closeness to the query vector.

    What NVIDIA says (1)

    “refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”

    — What is a Vector Database and How Does it Work?

  3. spaCy is a popular Python library for natural-language processing, such as named-entity recognition (NER). Moving data back to the CPU slows a GPU pipeline.

    What NVIDIA says (1)

    “relied on spaCy for NER but, spaCy currently needs your inputs on CPU”

    — Run State of the Art NLP Workloads at Scale with RAPIDS, HuggingFace, and Dask

  4. NumPy is the core Python library for arrays and linear algebra. Many data and NLP libraries build on it. NLP means natural-language processing.

    What NVIDIA says (1)

    “scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”

    — What is scikit-learn?

Key terms

Try it

Sample question

What is a vector database?

Show the answer

Answer: An organized collection of vector embeddings that can be created, read, updated and deleted

An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.

What NVIDIA says (1)

“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”

— What is a Vector Database and How Does it Work?

Practice 3.3 (4 questions) Full Multimodal Data guide

← 3.2 RAG, chatbots and summarizers · 3.4 Identifying data, hardware and software components →