Sovereign AI Ecosystem

"Compile-Time AI: Why the Industry Is Quietly Building an LLVM for Knowledge" [post] deterministic

"A taxonomy of the emerging Compile-Time AI movement — kib, Kompile, Brian Letort's Context Compilation Theory, llm-wiki-compiler, OVIR, and the SkCC paper — and why moving reasoning from runtime to compile time is the n

knowledge_systemsovereigntycompile_time_aimcp

Compile-Time AI: Why the Industry Is Quietly Building an LLVM for Knowledge

Every few years, systems programming rediscovers the same lesson: expensive work done once, ahead of time, beats expensive work done repeatedly, on demand. That lesson is why we have compilers instead of interpreting source code line-by-line on every execution. It's why we have query planners instead of re-deriving an execution strategy for every SQL statement. And it is, I'd argue, why a scattered but growing set of teams — with no coordination between them — have independently started describing their AI systems using the vocabulary of compilers: intermediate representations, lowering passes, optimizers, emitters, static analysis.

I don't think this is a coincidence, and I don't think it's marketing convergence. I think it's the AI industry rediscovering a systems-design pattern that is older than AI itself, applied to a resource that changed the economics: tokens.

This piece is a survey and a taxonomy, not a pitch for any one project. I looked closely at six efforts — kib, Kompile, Brian Letort's Context Compilation Theory, llm-wiki-compiler, OVIR, and the SkCC academic paper on skill compilation — plus the compiler literature they explicitly or implicitly borrow from (LLVM, MLIR). I want to show you the pattern underneath all of them, where they diverge, and where I think the analogy to traditional compilers breaks down.

The Core Thesis: Runtime AI vs. Compile-Time AI

Most AI systems built since 2023 follow the same shape. A question arrives. The system searches, retrieves, assembles a prompt, and asks a large model to reason over it — from scratch, every single time.

``` Traditional RAG (Runtime AI)

Documents | v Embeddings | v Vector DB | v LLM <----- every query pays full reasoning cost | v Answer ```

This works. It is also, structurally, an interpreter. Every query re-derives meaning from raw material. Nothing compounds. If you ask the same conceptual question twice, phrased two different ways, the system does the same expensive work twice, with no memory that it already did it once.

Compile-Time AI proposes a different shape: do the expensive reasoning once, offline, and compile it into a structured artifact that cheap runtime processes can consume.

``` Compile-Time AI

Documents | v IR1 (extraction / parsing) | v IR2 (concept / entity normalization) | v Semantic Passes (dedup, contradiction detection, linking) | v Knowledge Graph | v Optimization (pruning, confidence scoring, compaction) | v Static Application / Runtime Artifact | v Deployment <----- queries hit compiled artifact, not raw reasoning ```

The reasoning still happens — this isn't "avoid LLMs." It happens once, offline, at compile time, and the *output* of that reasoning becomes the thing that gets served. Compare that to a traditional compiler pipeline, and the analogy holds up better than you'd expect:

Source Code Human Knowledge | | v v AST Document IR | | v v IR Concept IR | | v v Optimization Relationship IR | | v v Machine Code Heuristic IR | v Application IR | v Website

A traditional compiler frontend turns source text into an AST, lowers it into one or more intermediate representations, runs optimization passes, and emits machine code that a much dumber, much faster CPU can execute directly. Compile-Time AI systems turn raw documents into structured semantic objects, lower them through progressively more typed representations, run passes that deduplicate and validate and score confidence, and emit an artifact — a graph, a wiki, a static app — that a cheap runtime process can serve without re-reasoning.

The question worth asking isn't "is this a real trend." It's "why now." Two forces are pushing simultaneously: token economics (frontier-model reasoning is expensive at the volumes production systems now operate at, so amortizing it across many future queries is financially rational), and reliability (a system that reasons fresh every time is also nondeterministic every time — compiling a decision once and auditing it once is a fundamentally different governance posture than re-deriving it under time pressure on every request).

Six Independent Efforts, One Pattern

kib — the headless knowledge compiler

kib is the most literal instance of the pattern. It's a CLI-first tool, built by Keegan Thompson, that ingests URLs, PDFs, YouTube transcripts, GitHub repos, and images, then runs an explicit compile step that turns those raw sources into a structured, queryable markdown wiki. The workflow is unapologetically compiler-shaped: kib init, kib ingest, kib compile, kib query. Search is BM25 full-text over the compiled artifact, not embedding search over raw chunks, and the whole thing ships as an MCP server so agents in Claude Code, Cursor, or Claude Desktop can drive the pipeline directly. Output is plain markdown files under version control — no proprietary database, no lock-in.

What's notable architecturally is the framing on their own site: the tool doesn't call itself a RAG system or a note-taking app. It calls itself a compiler, and the CLI verbs mirror that self-description precisely.

Kompile — enterprise-scale, three-pillar compilation

Kompile is the most ambitious of the group and the one furthest from a single-developer tool. It frames itself around three simultaneous compilation targets: models, knowledge, and applications. On the model side, it runs models through a 25-pass fixed-point graph optimizer — documented to reduce LLaMA cast operations from 668 down to 108 — doing fusion, dead-code elimination, constant folding, and hardware targeting that will look immediately familiar to anyone who has read an LLVM or XLA paper. On the knowledge side, it crawls an organization's data estate (Confluence, Jira, Slack, databases, email) through an eight-phase pipeline — load, classify, route, chunk, extract, resolve, compute edges, index — and compiles it into a typed, hierarchical knowledge graph with seven node levels, full provenance on every mutation, and support for Multi-Entity Bayesian Networks for causal and probabilistic reasoning over the graph. On the application side, it presents one unified interface so a business can swap LLM providers, vector stores, or embedding models without rewriting application logic.

Kompile is the clearest expression of "sovereign AI architecture" among the projects surveyed here: everything runs on the customer's own infrastructure, with air-gapped model archives (.karch files) explicitly built for regulated, disconnected environments. It's early access only as of this writing, so the production-scale evidence isn't public yet, but the architectural ambition is the most complete instance of "compile the whole stack" I found.

Brian Letort's Context Compilation Theory — the missing layer, formalized

Of everything surveyed, Brian Letort's work is the most rigorous attempt to give this pattern a formal theory rather than just an implementation. His argument, laid out across a series of posts, starts from a measurement problem: existing AI benchmarks evaluate answer quality but don't expose the compilation decisions that produced the context a model reasoned over. That

Sources

DanielKliewer.com blog · source

Related (4)

discusses Knowledge Systems conf=0.96
discusses Local-First / Sovereignty conf=0.96
discusses Compile-Time AI conf=0.96
discusses Model Context Protocol conf=0.5

← all Blog