Sovereign AI Ecosystem

Understanding RAG Systems [chapter] deterministic

Retrieval-Augmented Generation (RAG) has emerged as a cornerstone technique for building AI systems that can draw on external knowledge while retaining the flexibility of large language models. At its

knowledge_systemsovereignty

Retrieval-Augmented Generation (RAG) has emerged as a cornerstone technique for building AI systems that can draw on external knowledge while retaining the flexibility of large language models. At its core, RAG augments a generative model with retrieved facts, allowing the model to produce responses grounded in up‑to‑date or domain‑specific information without retraining the underlying parameters. In local‑first applications, where privacy and data sovereignty are paramount, RAG enables developers to keep sensitive corpora on‑premises while still benefiting from state‑of‑the‑art generation capabilities. This chapter walks you through the architecture of RAG pipelines, the role of embedding models and vector search, and the most useful patterns for structuring retrieval‑generation workflows. We will also examine practical considerations such as handling unsupported patterns, scaling with volume patterns, and preserving conversation history for multi‑turn interactions.

Fundamentals of RAG Architecture A RAG is built from three tightly coupled stages: retrieval, augmentation, and generation. The retrieval stage fetches relevant documents or snippets from an external knowledge base, usually organized as a vector store or a hybrid index. The augmentation stage combines the retrieved material with the original query, often inserting it into a prompt template that instructs the language model to answer based on the provided context. Finally, the generation stage passes the augmented prompt to a language model, which produces the final response. By decoupling knowledge from model weights, RAG offers a modular architecture that can be iterated on independently: you can swap out the retrieval engine, adjust the embedding model, or tune the generator without rebuilding the entire .

Retrieval Engine The retrieval engine is responsible for turning a query into a set of candidate documents. In practice, this involves encoding the query into a dense representation (embedding) and then searching a corpus of similarly encoded documents. The engine must balance recall (finding all relevant items) with precision (avoiding irrelevant noise). Indexing strategies such as HNSW or IVF‑PQ, combined with similarity metrics like cosine or inner product, are common choices for dense retrieval. For hybrid approaches, the engine may also maintain a lexical index (e.g., BM25) to capture exact keyword matches. A simple illustration of a retrieval step uses a vector store that exposes a search method:

```python import chromadb client = chromadb.Client() collection = client.get_collection("documents") def retrieve(query: str, top_k: int = 5) -> list[dict]: results = collection.query( query_texts=[query], n_results=top_k, ) return [ {"id": doc_id, "text": text} for doc_id, text in zip(results["ids"][0], results["documents"][0]) ]

`

The function above demonstrates how a single query can be transformed into a list of candidate documents. In a production , you would also apply metadata filters, re‑ranking, or query expansion to improve relevance.

Augmentation Step Once the retrieval engine has produced a set of candidates, the augmentation step constructs a prompt that blends the original query with the retrieved context. The goal is to give the language model enough information to answer accurately while constraining it to the provided facts. A typical augmentation template looks like this:

```python def augment(query: str, docs: list[dict]) -> str: context = "\n".join(f"[{i}] {doc['text']}" for i, doc in enumerate(docs, 1)) return f"Answer the following question using the provided context.\n\nContext:\n{context}\n\nQuestion: {query}"

`

This straightforward construction works well for single‑turn queries. When the supports multi‑turn dialogue, the augmentation step must also incorporate conversation history, as discussed later in the chapter.

Generation Phase The final stage hands the augmented prompt to a language model. In a local‑first setup, you might use an open‑source model such as Llama‑2 or Mistral, deployed via an inference server like vLLM or Ollama. The model receives the augmented prompt and generates a response that should be faithful to the retrieved context. Because the model is not retrained on the external corpus, it relies on the prompt to supply the necessary facts. This reliance makes prompt design and retrieval quality critical. A minimal generation call looks like this:

```python import requests def generate(prompt: str) -> str: response = requests.post( "http://localhost:8080/generate", json={"prompt": prompt, "max_tokens": 500}, ) return response.json()["text"]

`

In production, you would add temperature control, stop sequences, and possibly a post‑processing step to enforce formatting or extract structured data.

Embedding Models and Vector Search Embedding models are the backbone of dense retrieval in RAG. They map text into a high‑dimensional vector space where semantically similar passages are close together. The choice of embedding model influences both the quality of retrieval and the latency of the . Open‑source models such as Sentence‑Transformers, BGE, or GTE provide a good balance of performance and accessibility, while proprietary APIs like OpenAI embeddings offer higher accuracy at the cost of external dependencies.

Choosing an Embedding Model When selecting an embedding model, consider the following factors: - **Accuracy**: Measured by benchmark suites such as MTEB. Higher accuracy typically yields better retrieval precision. - **Speed**: Inference time per token matters for real‑time applications. Smaller models (e.g., 110M parameters) are faster but may sacrifice recall. - **Size**: Model weights determine storage requirements. Local deployments often favor models under 1 GB. - **License**: Ensure compliance with your distribution model and any data‑privacy policies. A practical way to evaluate candidates is to run a small benchmark on a representative subset of your corpus and compare the Mean Reciprocal Rank (MRR) or Recall@K across models.

Vector Search Implementation Vector search is typically implemented using libraries such as FAISS, HNSWLib, or Chroma. These libraries provide efficient indexing structures and similarity search routines. Below is an example of building an index with FAISS and performing a query:

```python import numpy as np import faiss def build_index(vectors: np.ndarray, dim: int) -> faiss.Index: index = faiss.IndexFlatL2(dim) index.add(vectors) return index def search(index: faiss.Index, query_vec: np.ndarray, top_k: int) -> tuple[np.ndarray, np.ndarray]: distances, indices = index.search(query_vec.reshape(1, -1), top_k) return distances, indices

`

The build_index function creates a flat L2 index, while search returns the nearest neighbors. For large corpora, you would replace the flat index with a hierarchical or product‑quantized index to reduce memory usage and improve query speed.

RAG Patterns RAG is not a single monolithic architecture; it is a family of patterns that can be combined and extended to suit different use cases. Understanding these patterns helps you choose the right retrieval strategy for a given scenario.

Naive RAG Naive RAG is the simplest form: retrieve a top‑k set of documents, augment the prompt, and generate. It works well when the corpus is small and the queries are straightforward. However, it struggles with queries that require multiple pieces of information spread across different documents.

Hybrid RAG Hybrid RAG combines dense retrieval with lexical search. The runs both a vector search and a keyword search (e.g., BM25), then fuses the results using a scoring function such as reciprocal rank fusion. This approach captures both semantic relevance and exact matches, improving overall retrieval quality.

```python def hybrid_search(query: str, vector_results: list[dict], keyword_results: list[dict]) -> list[dict]:

Sources

Sovereign AI: Building Local-First Intelligent Systems (book) · source

Related (2)

discusses Knowledge Systems conf=0.6
discusses Local-First / Sovereignty conf=0.6

← all Book