"Knowledge Compiler: Why I'm Building a Compiler for Human Knowledge Instead of Another RAG System" [post] deterministic
"Knowledge Compiler transforms collections of Markdown documents into statically-deployable semantic artifacts — knowledge graphs, concept hierarchies, vector embeddings, and cluster maps — using a multi-pass compilation
> **Repository:** [github.com/kliewerdaniel/knowledge-compiler](https://github.com/kliewerdaniel/knowledge-compiler) > **Overview video:** [NotebookLM](https://notebooklm.google.com/notebook/1833f401-8f66-466b-9794-e2669107ab41/artifact/0a6cbfd3-dfb3-4316-b7b9-b8005397f7fc) > > *This is not a chatbot. It is a compiler.*
---
**Why does an AI system need to rediscover the same knowledge every time it answers a question?**
This is the fundamental inefficiency that most knowledge systems silently accept. Every query against a RAG pipeline pays the full cost of retrieval, context assembly, and generation — even when the knowledge domain is static. Even when the question has been asked before. Even when the answer could have been precomputed.
Knowledge Compiler is an exploration of a different tradeoff: what if semantic understanding is performed at compile time instead of runtime? What if the artifacts of that compilation — knowledge graphs, concept hierarchies, vector embeddings, cluster maps — are themselves the deployable unit?
What if a knowledge application can be *compiled* like software, served from a static CDN, and never touch an LLM at inference time?
I built this to find out.
---
I. The Runtime Tax
Let's be concrete about the costs that current architectures accept as unavoidable.
Retrieval-Augmented Generation (RAG)
Every query in a standard RAG pipeline:
1. Embeds the query (one API call, ~100-500ms) 2. Searches a vector index (one ANN search, ~10-100ms) 3. Retrieves context chunks (one or more document lookups) 4. Constructs a prompt with the retrieved context 5. Sends the prompt to an LLM (one generation call, ~500ms-5s depending on output length)
**Per-query cost:** ~1-6 seconds of latency, $0.001-$0.01 in API fees, and one round of GPU inference.
Scale this to an organization processing thousands of queries per day against a stable knowledge base — documentation, legal archives, medical literature, scientific papers — and you are paying the same tax for every single query, even though the underlying knowledge has not changed.
GraphRAG
GraphRAG improves retrieval quality by organizing documents into a graph structure, enabling multi-hop reasoning and community detection. Microsoft's GraphRAG paper demonstrated that graph-based retrieval significantly outperforms naive vector search on complex, sensemaking queries.
But GraphRAG introduces its own runtime costs:
- Query expansion to identify graph-relevant entities
- Graph traversal across multiple hops
- Community summarization at query time (often requiring additional LLM calls)
- Secondary retrieval to fetch supporting evidence
The architectural assumption is the same: intelligence happens at query time.
Agentic Knowledge Systems
The current frontier — multi-agent systems that navigate knowledge bases, break down queries, and synthesize answers — multiplies these costs further. Each agent in the swarm may independently retrieve, reason, and generate. Task decomposition, tool selection, and result synthesis each require LLM calls.
The result is a system that is powerful but expensive, both in latency and in compute.
---
II. The Compiler Alternative
There is a well-understood precedent for this class of problem.
Software compilers transform source code (human-readable, expressive, redundant) into optimized executables (machine-efficient, pre-analyzed, deployable). The compilation step is expensive. The runtime step is cheap. The fundamental insight is that analysis can be *amortized* across all executions.
| Software Compiler | Knowledge Compiler | |---|---| | Source code | Markdown documents | | Lexical analysis | Markdown parsing (MDAST) | | Abstract Syntax Tree | Document AST with position tracking | | Intermediate Representation | Semantic IR (knowledge graphs, concept hierarchies, vectors) | | Optimization passes | Pruning, deduplication, quantization | | Object files | JSON artifacts + binary embedding store | | Executable | Static Next.js application |
Knowledge Compiler applies this same amortization strategy to knowledge. Instead of analyzing documents at query time, it performs a complete semantic analysis during a build step, producing artifacts that encode the full relational and semantic structure of the knowledge base.
**The compiled artifacts are not an index into the source documents. They are a self-contained reasoning substrate.**
---
III. Architecture of the Knowledge Compiler
The compiler is organized as a monorepo with seven packages and one application:
packages/
ir/ — Intermediate Representation types (Zod schemas)
config/ — Configuration system (cosmiconfig + Zod validation)
cache/ — Two-level cache (L1 memory, L2 disk, XXH3 hashing)
artifacts/ — Artifact serialization (binary embeddings, atomic writes)
plugins/ — Plugin registry and pass lifecycle interfaces
core/ — Pipeline engine, scheduler, 23 built-in passes
cli/ — CLI tool (cac-based), binary: kc
apps/
web/ — Next.js app for browsing compiled knowledge
The Compilation Pipeline
The pipeline executes **9 phases** in sequence, each containing one or more compiler passes. Passes declare dependencies (hard and optional), and the scheduler resolves them via topological sort (Kahn's algorithm).
Source (Markdown)
│
▼
1. PARSING Glob resolution → File reading → Frontmatter extraction → MDAST parsing
│
▼
2. ANALYSIS Link extraction → Named entity recognition → TF-IDF keywords → Concept hierarchy
│
▼
3. GRAPH Knowledge graph construction → PageRank → Graph statistics
│
▼
4. EMBEDDING Sentence-level chunking → Vector embedding → Dimensionality reduction
│
▼
5. CLUSTERING Similarity matrix → Connected-component clustering → Centroid computation
│
▼
6. OPTIMIZATION Edge pruning → SimHash near-duplicate detection → Int8 quantization
│
▼
7. GENERATION Artifact serialization → Manifest building
│
▼
8. COMPLETE Report aggregation
I'll walk through each phase.
Phase 1: Parsing (4 passes)
**GlobResolverPass** uses fast-glob to resolve user-specified patterns (default **/*.md) against the base directory, with a manual recursive-walk fallback.
**FileReaderPass** reads each file asynchronously with SHA-256 content hashing.
**FrontmatterParserPass** extracts YAML frontmatter using js-yaml, with a hand-written parseSimpleYaml() fallback.
**MDASTParserPass** parses markdown into an MDAST (Markdown Abstract Syntax Tree) using unified/remark with GFM and frontmatter support. The resulting AST is stored in the IR store as a DocAST:
typescript
interface DocAST extends IRGraph<DocNode> {
sourcePath: string;
sourceHash: string;
rootNodeId: UUID;
totalTokens: number;
statistics: DocStatistics;
}
Each DocNode tracks its source position (start/end line and column), parent-child relationships, and node-type-specific metadata (heading levels, code language, link URLs, etc.).
**Token estimation** uses a simple heuristic: Math.ceil(words.length * 1.3). In practice this correlates well with actual token counts for technical prose.
Phase 2: Analysis (4 passes)
**LinkExtractorPass** walks the AST recursively, classifying links as internal (matching *.md patterns) or external. Internal links become candidates for knowledge graph edges between documents.
**EntityExtractorPass** performs regex-based named entity recognition against 13 patterns:
- PERSON (with honorific prefixes: Dr., Prof., Sen., etc.)
- ORG (with suffixes: Inc., Corp., LLC, Ltd.)
- LOCATION (known US cities and common locations)
- DATE (full date formats and ISO dates)
- MONEY, EMAIL, URL, PHONE, CODE (constant identifiers)
Entities are deduplicated and ranked by frequency across the document.
**KeywordExtractorPass** implements classic TF-IDF: