Sovereign AI Ecosystem

'The Loop Is the Product: Inside the Sovereign Intelligence Observatory' [post] deterministic

"A technical deep dive into the Sovereign Intelligence Observatory: a six-component, local-first pipeline that turns every agent run into a versioned recipe, routes evaluation by confidence tier, detects capability drift

sovereign-intelligence-observatoryagent-recipesdrift-detectionexpert-signal-routingtacit-knowledge-extractionautonomy-laddersknowledge-graphslocal-aisovereign-aisovereign-memory-banksovereignspecsynthintobservabilityrecipesignal_routerevaluation_loopobservatoryapprenticeshipsovereigntytacit_judgmentmcp

The Loop Is the Product: Inside the Sovereign Intelligence Observatory

**July 3, 2026**

---

Every agent framework on the market answers the same question: how do you get a model to do a task. Almost none of them answer the question that actually determines whether your system gets better over time: what happened, in what order, under what confidence, judged by whom, and is that judgment still valid six months later.

The [Sovereign Intelligence Observatory](https://github.com/kliewerdaniel/sovereign-intelligence-observatory) is a six-component, local-first Python system built to answer that second question. It doesn't wrap an LLM. It doesn't compete with LangGraph or CrewAI for orchestration mindshare. It sits downstream of whatever agent runtime you're already using and treats every decision that runtime makes as a first-class, versioned, queryable artifact. The project's own framing is blunt about it: intelligence isn't the weights, it's the accumulated decisions that shaped them, and if the loop is the product, observability of the loop is the operating system.

This post walks through the architecture at the code level -- the drift statistics, the sandboxing model, the ledger chain, the concurrency guarantees -- and makes the case for why this pattern matters for anyone building agents they intend to keep improving rather than keep re-prompting.

The core insight: recipes, not logs

Most agent systems produce logs. Logs are append-only text optimized for a human to read once, during an incident, and then forget. The Observatory instead produces **recipes**: structured, versioned artifacts that capture the complete decision context of a single agent run.

json { "recipe_id": "recipe-20240101-120000-abc123", "objective": "classify_ai_paper", "model": "qwen3.5", "prompt_version": 5, "memory_version": 12, "retrieved_docs": ["doc_1", "doc_2"], "reasoning_patterns": ["compare", "retrieve", "synthesize"], "evaluation": {"score": 0.95, "reviewed_by": "expert"}, "outcome": "accepted" }

The distinction matters because a log is write-once and a recipe is a **row in a schema**. Once your agent's behavior has a schema, it can be indexed (SQLite FTS5 full-text search), diffed across prompt or memory versions, embedded and searched semantically (optional ChromaDB), streamed out as training data, and — critically — fed back into the system that decides whether your agent is getting better or worse. The Agent Recipe Compiler is the component that does the capturing; everything downstream consumes its output. The system frames this as the missing primitive most agent stacks never build, and the framing holds up: without it, "improving the agent" means eyeballing transcripts.

Architecture: six layers, one feedback loop

Agent | v Recipe Compiler ----------------------------------------+ | | v | Expert Signal Router | | | v | Autonomous Evaluation Loop | | | v | Tacit Judgment Extractor | | | v | Sovereign Apprenticeship Engine | | | v | Intelligence Observatory <------------------------------+ | v Intelligence Timeline -> Actionable Insights

Each layer produces the input for the next, and the Observatory at the bottom folds everything back into a timeline that determines whether the whole loop is compounding or decaying. Six components, six SQLite databases in WAL mode, one FastAPI surface per component, 176 tests across 8 suites. Let's go through them in the order data actually flows.

1. Agent Recipe Compiler — the ledger of what happened

Every run gets ingested through POST /api/recipes, indexed with SQLite FTS5, and made available for full-text and (optionally) semantic search. It supports chunked streaming JSON export specifically so recipe history can be turned into fine-tuning data later without loading the whole table into memory. This is the layer everything else is built on top of, and it's deliberately boring: SQLite, JSON, HTTP. No vector database is required to get started; ChromaDB is dependency-injected and the system falls back to FTS5 silently if it isn't installed.

2. Expert Signal Router — deciding who judges the output

Recipes tell you what happened. They don't tell you if it was any good, and worse, they don't tell you who should be bothered to find out. The router implements a tiered confidence gate:

Agent Output | v Confidence >= 0.95? --YES--> Auto-accepted | NO v Confidence >= 0.80? --YES--> Cheap evaluation | NO v Expert review required

The thresholds aren't fixed. A dynamic calibration matrix adjusts them per objective based on historical error rate, so a task class that keeps fooling the cheap evaluator gets escalated more aggressively over time, and one that experts keep rubber-stamping gets cheaper to clear. Every expert decision the router captures becomes a labeled training example for the next tier down — this is the mechanism that lets human judgment gradually get absorbed into the automated evaluation layer instead of staying a permanent cost center.

3. Autonomous Evaluation Loop — catching drift before it becomes an outage

This is the layer I think is most underbuilt in the rest of the agent-framework ecosystem, and it's worth showing the actual math. Evaluation signals are defined as YAML specs with uncertainty bounds, synthetic test cases are auto-generated from production traffic, and every signal is checked for **drift** using two independent statistics that have to agree before an alert fires.

Two-sample Kolmogorov–Smirnov D-statistic, measuring how far apart two empirical distributions have drifted:

python def _kolmogorov_smirnov_statistic(sample_a, sample_b): combined = sorted(set(sample_a + sample_b)) max_diff = 0.0 for val in combined: cdf_a = sum(1 for x in sample_a if x <= val) / len(sample_a) cdf_b = sum(1 for x in sample_b if x <= val) / len(sample_b) max_diff = max(max_diff, abs(cdf_a - cdf_b)) return max_diff # threshold: 0.3

Population Stability Index, measuring binned proportion shift with Laplace smoothing so empty bins don't blow up the log:

python def _population_stability_index(expected, actual, n_bins=10): ... for i in range(n_bins): p_exp = (exp_counts[i] + 0.5) / (n_exp + 0.5 * n_bins) p_act = (act_counts[i] + 0.5) / (n_act + 0.5 * n_bins) psi += (p_act - p_exp) * math.log(p_act / p_exp) return psi # threshold: 0.25

Requiring both KS *and* PSI to cross threshold before flagging drift is a deliberate design choice against false positives — KS is sensitive to shape changes, PSI is sensitive to mass movement between bins, and real capability regressions tend to show up in both. There's also a validation guard that rejects synthetic or degenerate inputs before they can pollute the signal: if the last three scores for an objective are all identical, the new score is rejected outright, since real model output has variance and a suspiciously flat signal is more likely a broken pipeline than a stable one.

4. Tacit Judgment Extractor — mining expertise nobody wrote down

This is the component that answers a question most eval frameworks don't even ask: how do you capture the knowledge an expert *isn't articulating* while they review outputs? The extractor

Sources

DanielKliewer.com blog · source

Related (8)

discusses Intelligence Observatory conf=0.96
discusses Local-First / Sovereignty conf=0.96
discusses Apprenticeship Engine conf=0.8
discusses Tacit Judgment conf=0.7
discusses Signal Router conf=0.6
discusses Evaluation Loop conf=0.6
discusses Model Context Protocol conf=0.5

← all Blog