Chapter 4: Vector Databases and ChromaDB [chapter] deterministic
## Chapter Objectives - Set up ChromaDB for local vector storage - Implement efficient document chunking - Build a complete RAG pipeline In modern AI applications, the ability to retrieve relevant inf
Chapter Objectives - Set up ChromaDB for local vector storage - Implement efficient document chunking - Build a complete RAG pipeline In modern AI applications, the ability to retrieve relevant information quickly and accurately is paramount. Traditional relational databases excel at structured data but falter when dealing with unstructured text, embeddings, or semantic similarity searches. This chapter introduces vector databases as a solution, focusing on ChromaDB as a lightweight, local-first option. We will explore how to set up ChromaDB, chunk documents efficiently, and integrate it into a retrieval-augmented generation (RAG) pipeline.
Understanding Vector Databases
A vector database is a specialized storage designed to manage high-dimensional vectors, typically generated by embedding models such as SentenceTransformers or OpenAI’s text-embedding-ada-002. These vectors represent the semantic meaning of text snippets. Vector databases enable similarity searches, allowing applications to retrieve the most relevant documents based on a query’s vector representation.
Unlike traditional databases that rely on exact matches or keyword searches, vector databases use distance metrics such as cosine similarity or Euclidean distance to find the closest vectors. This approach is particularly powerful in scenarios where the ’s intent is not explicitly stated, but can be inferred from the context of a query.
Key Concepts in Vector Databases - **Embeddings**: Numerical representations of text, images, or other data. - **Similarity Search**: Finding the most similar vectors to a query vector. - **Indexing**: Storing vectors in a way that optimizes retrieval speed. - **Metadata**: Additional information stored alongside vectors, such as document titles or timestamps.
Setting Up ChromaDB ChromaDB is a popular open-source vector database that runs locally, making it ideal for developers who want to keep data private and avoid cloud dependencies. It integrates seamlessly with Python and supports fast vector similarity searches.
Installation and Basic Setup To get started, install ChromaDB via pip:
```bash pip install chromadb
`
Once installed, you can create a client and a collection:
```python import chromadb client = chromadb.Client() collection = client.create_collection(name="my_collection")
`
The create_collection method initializes a new collection where you will store vectors and associated metadata. You can specify additional parameters such as the distance metric (e.g., cosine, euclidean) and the embedding function.
Adding Documents To add documents to the collection, you first need to generate embeddings for each document. ChromaDB provides built-in embedding functions, but you can also use custom ones.
```python import chromadb client = chromadb.Client() collection = client.create_collection(name="my_collection", embedding_function=chromadb.Embeddings()) documents = [ "The quick brown fox jumps over the lazy dog.", "A fast brown fox leaps over a sleeping dog.", "The lazy dog slept under the tree." ] embeddings = client.get_embedding_function()(documents) collection.add( ids=["doc1", "doc2", "doc3"], documents=documents, embeddings=embeddings )
`
In this example, we use the default embedding function provided by ChromaDB. The add method stores the documents along with their embeddings and unique IDs.
Document Chunking Large documents often exceed the token limits of embedding models and can dilute the semantic meaning of a text snippet. To address this, we need to split documents into smaller, coherent chunks. This process is called **chunking** and is a critical part of any RAG pipeline.
Chunking Strategies There are several strategies for chunking: - **Fixed-size chunking**: Split the document into chunks of a fixed number of tokens or characters. - **Paragraph-based chunking**: Split the document by paragraphs, ensuring that each chunk contains a complete thought. - **Semantic chunking**: Use a model to detect natural breakpoints in the text, such as sentence boundaries or topic shifts. For most use cases, paragraph-based chunking provides a good balance between granularity and coherence.
Implementing Chunking in Python
Here’s a simple implementation of paragraph-based chunking using the re module:
```python import re from typing import List def chunk_by_paragraphs(text: str, max_chunk_size: int = 512) -> List[str]: """Split text into chunks by paragraphs, respecting max_chunk_size.""" paragraphs = re.split(r'\n\s*\n', text.strip()) chunks = [] current_chunk = [] current_size = 0 for paragraph in paragraphs: paragraph_size = len(paragraph) if current_size + paragraph_size > max_chunk_size: if current_chunk: chunks.append(' '.join(current_chunk)) current_chunk = [] current_size = 0 current_chunk.append(paragraph) current_size += paragraph_size if current_chunk: chunks.append(' '.join(current_chunk)) return chunks
`
This function splits the text by double newlines and ensures that each chunk does not exceed max_chunk_size characters.
Building a RAG Pipeline With ChromaDB set up and documents chunked, we can now build a retrieval-augmented generation (RAG) pipeline. RAG combines retrieval from a vector database with generation from a language model to answer questions with contextually relevant information.
Retrieval Step The retrieval step involves converting the ’s query into an embedding and searching the vector database for the most similar documents.
```python def retrieve_relevant_documents(query: str, collection, top_k: int = 3) -> List[str]: """Retrieve top_k relevant documents for a given query.""" query_embedding = client.get_embedding_function()([query])[0] results = collection.query(query_embeddings=[query_embedding], n_results=top_k) return results['documents'][0]
`
This function uses the same embedding function as before to convert the query into a vector and then queries the collection for the top k most similar documents.
Generation Step Once we have retrieved relevant documents, we pass them along with the query to a language model to generate a response.
```python from transformers import pipeline generator = pipeline("text2text-generation", model="t5-small") def generate_response(query: str, context: List[str]) -> str: """Generate a response using retrieved context.""" context_str = "\n".join(context) prompt = f"Context: {context_str}\nQuestion: {query}\nAnswer:" response = generator(prompt, max_length=150, num_return_sequences=1) return response[0]['generated_text']
`
This function uses the T5 model to generate a response based on the query and the retrieved context.
Putting It All Together Finally, we combine the retrieval and generation steps into a single function:
```python def rag_pipeline(query: str, collection, top_k: int = 3) -> str: """Full RAG pipeline: retrieve and generate.""" context = retrieve_relevant_documents(query, collection, top_k) return generate_response(query, context)
`
This function demonstrates how ChromaDB integrates into a RAG pipeline, enabling the to answer questions with up-to-date, contextually relevant information.
Advanced Topics
Metadata Filtering ChromaDB supports metadata filtering, allowing you to restrict the search to documents that match certain criteria. For example, you can filter by document type, author, or date.
```python results = collection.query( query_embeddings=[query_embedding], where={"date": {"$gte": "2023-01-01"}}, n_results=5 )
`
This query retrieves documents published after January 1, 2023, ensuring that the search is limited to recent content.
Hybrid Search
Hybrid search combines vector similarity with keyword-based search, leveraging the strengths of both approaches. ChromaDB supports hybrid search through the HybridSearch class, which allows you to specify both vector and keyword search criteria.
```python from chromadb import HybridSearch hybri