1.6 Python NLP and vector tools
What the main Python language libraries and vector databases do.
Key points
A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.
What NVIDIA says (2)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
“Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
NeMo's token-classification guide defines named entity recognition (NER) as detecting and classifying key information (entities) in text, with the example that “Mary” is a person, “Santa Clara” a location and “NVIDIA” a company. Libraries such as spaCy also provide NER pipelines.
What NVIDIA says (3)
“NER, also referred to as entity chunking, identification, or extraction, is the task of detecting and classifying key information (entities) in text.”
“relied on spaCy for NER but, spaCy currently needs your inputs on CPU”
“is a person, “Santa Clara” is a location, and “NVIDIA” is a company.”
NVIDIA's scikit-learn glossary states that scikit-learn is built on NumPy, which is optimized for high-performance linear algebra and array operations.
What NVIDIA says (1)
“scikit-learn is a versatile Python library built on NumPy, optimized for high-performance linear algebra and array operations.”
Key terms
- Vector database: A database that stores embeddings and finds the ones most similar to a query.
- NumPy: Python library for fast array and linear-algebra operations.
- spaCy: Python NLP library; used here for named-entity recognition on the CPU.
- Named-entity recognition: Finding and labeling entities in text, such as people, places and companies.
Sample question
What is the main job of a vector database in a large language model (LLM) application?
Show the answer
Answer: Storing vector embeddings and retrieving those most similar to a query's embedding
A vector database is an organized collection of vector embeddings; similarity (vector) search retrieves vectors semantically similar to a query's embedding.
What NVIDIA says (2)
“A vector database is an organized collection of vector embeddings that can be created, read, updated, and deleted at any point in time.”
“Similarity search, also known as vector search, vector similarity, or semantic search, refers to the process when an AI application efficiently retrieves vectors from the database that are semantically similar to a given query’s vector embeddings”
Practice 1.6 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 1.5 Machine-learning fundamentals · 1.7 Reading research and spotting trends →