Sovereign AI Ecosystem

AI-Powered Study Systems [chapter] deterministic

The rise of AI in education has fundamentally changed how we approach learning, and the most impactful systems are those that respect the learner's privacy and data sovereignty. In this chapter, we'll

sovereignty

The rise of AI in education has fundamentally changed how we approach learning, and the most impactful systems are those that respect the learner's privacy and data sovereignty. In this chapter, we'll build a complete study system that parses academic PDFs, extracts citations, and provides intelligent Q&A capabilities—all running locally on your machine.

Parsing Academic PDFs with Docling

Academic PDFs present unique challenges for parsing: they often contain complex layouts, footnotes, tables, and embedded figures that traditional parsers struggle with. Docling, an open-source PDF parser developed by IBM, addresses these challenges by using a combination of layout analysis and deep learning to extract structured content from PDFs.

The first step is to install Docling and its dependencies:

```python pip install docling docling-core

`

Once installed, we can initialize the parser:

```python from docling.document_converter import DocumentConverter converter = DocumentConverter()

`

The DocumentConverter class is the main entry point for parsing documents. It supports various file formats, including PDF, EPUB, and HTML. For academic PDFs, we'll focus on the PDF format.

To parse a PDF, we simply call the convert method:

```python result = converter.convert("academic_paper.pdf")

`

The result object contains the parsed document, including its text, metadata, and other information. We can access the text using the text attribute:

```python text = result.document.text

`

However, academic PDFs often contain more than just plain text. They may include footnotes, tables, and figures that need to be extracted separately. Docling provides a DocumentConverter class that can handle these cases.

One of the key features of Docling is its ability to extract citations from the parsed text. This is crucial for building citation-aware study systems, as we'll see later.

Handling Different File Types

While PDFs are the most common format for academic papers, we may also encounter other formats, such as EPUB or HTML. Docling supports these formats as well. To parse an EPUB file, we simply pass the file path to the convert method:

```python result = converter.convert("ebook.epub")

`

Similarly, for HTML files, we can use the convert method:

```python result = converter.convert("article.html")

`

Docling's ability to handle multiple file types makes it a versatile tool for building study systems that can ingest a variety of academic materials.

Extracting Metadata

In addition to text, Docling also extracts metadata from the parsed documents. This metadata includes information such as the author, title, publication date, and DOI. We can access this metadata using the metadata attribute of the result object:

```python metadata = result.document.metadata

`

This metadata can be useful for organizing and searching the parsed documents.

Performance Considerations

When parsing large PDFs, performance can become a concern. Docling provides a DocumentConverter class that can handle large documents efficiently. However, if we need to parse multiple PDFs in parallel, we can use the asyncio library to run the parsing tasks concurrently.

For example, we can define an async function to parse a PDF:

```python import asyncio

async def parse_pdf(file_path): converter = DocumentConverter() result = await converter.convert_async(file_path) return result

`

Then, we can run multiple parsing tasks concurrently using asyncio.gather:

```python pdf_files = ["paper1.pdf", "paper2.pdf", "paper3.pdf"] results = await asyncio.gather(*[parse_pdf(file) for file in pdf_files])

`

This approach can significantly speed up the parsing process when dealing with large collections of academic papers.

Building Citation-Aware RAG

One of the key challenges in building a study system is ensuring that the system can provide accurate and cited responses to user queries. Traditional RAG systems often lack this capability, as they don't track citations or provide traceability to the original sources.

Citation-aware RAG addresses this by extracting citations from the parsed documents and storing them in a vector database. When a user queries the system, the RAG engine retrieves the most relevant citations and returns them along with the answer.

Extracting Citations

The first step in building a citation-aware RAG system is to extract citations from the parsed documents. Docling provides a DocumentConverter class that can extract citations from the parsed text.

To extract citations, we can use the extract_citations method of the DocumentConverter class:

```python citations = converter.extract_citations(result.document)

`

This method returns a list of citation objects, each containing the citation text, the source document, and the citation context.

Storing Citations in a Vector Database

Once we have extracted the citations, we need to store them in a vector database. This allows us to perform fast and efficient retrieval of citations based on user queries.

We can use the chromadb library to create a vector database:

```python import chromadb from chromadb.utils import embedding_functions

embedding_function = embedding_functions.SentenceTransformerEmbeddingFunction() client = chromadb.Client() collection = client.create_collection("citations", embedding_function=embedding_function)

`

We can then add the extracted citations to the collection:

```python for citation in citations: collection.add( ids=[citation.id], documents=[citation.text], metadatas=[{"source": citation.source, "context": citation.context}] )

`

This approach allows us to store the citations in a vector database and perform fast retrieval based on user queries.

Retrieving Citations

When a user queries the system, we can retrieve the most relevant citations using the query method of the collection:

```python query = "What is the main contribution of the paper?" results = collection.query(query_texts=[query], n_results=5)

`

The results object contains the most relevant citations, along with their IDs, documents, and metadata. We can then use this information to provide a cited response to the user.

Challenges in Citation Extraction

Extracting citations from academic PDFs can be challenging due to the variety of citation formats and the presence of noisy text. Docling's extract_citations method addresses these challenges by using a combination of regex and machine learning to extract citations accurately.

However, there are still some challenges to consider. For example, some citations may be embedded in footnotes or endnotes, which can be difficult to extract. Additionally, some papers may use non-standard citation formats, which can make extraction more difficult.

To address these challenges, we can use a combination of regex and machine learning to extract citations more accurately. For example, we can use regex to extract citations that follow a specific format, and then use a machine learning model to extract citations that don't follow the format.

Building Intelligent Study Assistants

With the parsed documents and citation-aware RAG system in place, we can now build an intelligent study assistant that can answer user queries, generate summaries, and provide explanations.

Setting Up the Assistant

We can use the LangChain library to build the study assistant. LangChain provides a set of tools for building AI-powered applications, including tools for building chatbots and question-answering systems.

First, we need to install LangChain and its dependencies:

```python pip install langchain langchain-openai

`

Then, we can initialize the assistant:

```python from langchain.chains import RetrievalQA from langchain.embeddings import OpenAIEmbeddings from langchain.vectorstores import Chroma from langchain.llms import OpenAI from langchain.memory

Sources

Sovereign AI: Building Local-First Intelligent Systems (book) · source

Related (1)

discusses Local-First / Sovereignty conf=0.96

← all Book