3.3 Python natural-language packages
Vector databases, similarity search, spaCy and NumPy.
Key points
An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.
What NVIDIA says (1)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
Similarity search is also called vector or semantic search. It ranks stored vectors by closeness to the query vector.
What NVIDIA says (1)
“refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
spaCy is a popular Python library for natural-language processing, such as named-entity recognition (NER). Moving data back to the CPU slows a GPU pipeline.
What NVIDIA says (1)
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
NumPy is the core Python library for arrays and linear algebra. Many data and NLP libraries build on it. NLP means natural-language processing.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
Key terms
- Vector database: A store of embeddings that can find the vectors closest to a query.
- Embedding: A list of numbers that represents the meaning of an input such as text or an image.
- spaCy: A Python library for natural-language processing, such as named-entity recognition.
- NumPy: The core Python library for arrays and linear algebra.
Try it
Sample question
What is a vector database?
Show the answer
Answer: An organized collection of vector embeddings that can be created, read, updated and deleted
An embedding is a list of numbers that captures meaning. A vector database stores embeddings and finds the ones closest to a query.
What NVIDIA says (1)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
Practice 3.3 (4 questions) Full Multimodal Data guide
← 3.2 RAG, chatbots and summarizers · 3.4 Identifying data, hardware and software components →