'The Model Is Not the Product: Residual State, Compiled Agents, and Optimization Loops' [post] deterministic
"Three converging research threads — Apple's Residual Context Diffusion, LMSYS/SGLang agentic execution graphs, and constrained optimization for agent loops — collapse into a single architectural claim: the model is no l
The Model Is Not the Product: Residual State, Compiled Agents, and Optimization Loops
**July 3, 2026**
---
The model is no longer the product. The loop is.
That idea keeps getting reinforced every time I look at new research from Apple, LMSYS, and the recent work on autoresearch and constrained optimization. They're not converging on a better chatbot. They're converging on something closer to a reconfigurable system of computation where "reasoning" is just one phase inside a larger machine.
What's changing isn't just capability. It's where intelligence lives.
It's shifting out of the model and into three places at once: **residual state**, **execution graphs**, and **optimization loops**.
This isn't abstract. Each of these threads has concrete implementations — and when you wire them together, you get something that looks less like a chatbot and more like a continuously recompiled cognitive engine. I've been building toward this architecture across several systems: [Objective05](https://github.com/kliewerdaniel/objective05) (persistent intelligence infrastructure in Rust), [Sovereign Memory Bank](https://github.com/kliewerdaniel/sovereignBank) (7-layer autonomous cognitive memory), [Dynamic Persona MoE RAG](https://github.com/kliewerdaniel/dynamic_persona_moe_rag) (persona-driven mixture-of-experts over local graphs), and [SovereignSpec](https://github.com/kliewerdaniel/sovereignSpec) (spec-driven development with GraphRAG). This post is the synthesis of what those systems are converging on — and what the research confirms.
---
1. From Tokens to Residual State
Apple's Residual Context Diffusion
Apple's **Residual Context Diffusion (RCD)** quietly breaks one of the core assumptions behind most LLM systems: that intermediate uncertainty should be discarded.
**Paper:** *Residual Context Diffusion Language Models* — [arXiv:2601.22954](https://arxiv.org/abs/2601.22954) (Hu et al., 2026) **Code:** [github.com/yuezhouhu/residual-context-diffusion](https://github.com/yuezhouhu/residual-context-diffusion)
In standard generation pipelines, we sample, reject, and move on. Low-confidence paths disappear. Only the final sequence matters.
RCD changes that. Instead of throwing away "failed" intermediate states during diffusion, it feeds them forward as **contextual residuals** — entropy-weighted continuous embedding vectors injected into subsequent denoising steps.
Here's the core mechanism in pseudocode:
```python import torch import torch.nn.functional as F
def residual_diffusion_step( x_t: torch.Tensor, # masked embedding at step t logits: torch.Tensor, # model logits over vocabulary embed_weight: torch.Tensor, # vocabulary embedding matrix residual_buffer: list, # accumulated residuals from prior steps temperature: float = 1.0, entropy_threshold: float = 0.5, ) -> tuple[torch.Tensor, torch.Tensor]: """ One step of RCD decoding.
Instead of hard-committing to argmax tokens and discarding the rest, RCD converts the full predictive distribution into a residual vector and feeds it forward into the next step. """ # Compute token probabilities probs = F.softmax(logits / temperature, dim=-1)
Entropy-weighted residual: sum over vocab weighted by uncertainty # High-entropy (uncertain) positions contribute more residual signal entropy = -(probs * torch.log(probs + 1e-8)).sum(dim=-1, keepdim=True) normalized_entropy = entropy / entropy.max()
Residual = weighted sum of all vocabulary embeddings # NOT just the argmax token — every candidate contributes residual = torch.einsum("b v, v d -> b d", probs, embed_weight) residual = residual * (normalized_entropy > entropy_threshold).float()
Accumulate residual into buffer residual_buffer.append(residual.detach())
Blend: combine original masked embedding with residual history # The mixing weight is itself entropy-dependent blend_weight = torch.sigmoid(2.0 * normalized_entropy - 1.0) x_next = (1 - blend_weight) * x_t + blend_weight * residuals.mean(dim=0)
return x_next, probs ```
That sounds like a small tweak. It isn't.
**Results:** RCD achieves 5–10 point accuracy gains on frontier diffusion LLMs, nearly 2× baseline on AIME, and 4–5× fewer denoising steps at equivalent accuracy — all from converting a standard dLLM with ~300M tokens of additional training. The paper shows this works because the residual buffer captures **discarded hypotheses, low-probability reasoning paths, and partial structures that didn't resolve cleanly** — everything we normally optimize away becomes state for the next iteration.
What This Means for System Architecture
In most LLM systems (including RAG), memory is treated as *retrieval*:
python
def standard_rag(query: str, top_k: int = 5) -> str:
embedding = embedder.embed(query)
results = vector_store.similarity_search(embedding, k=top_k)
return format_context(results)
But RCD suggests a different model:
> Memory is not retrieval. Memory is **residue**.
In my Sovereign Memory Bank architecture ([post](https://www.danielkliewer.com/blog/sovereign-memory-bank-a-deep-dive-into-autonomous-cognitive-memory-for-agent-systems), [repo](https://github.com/kliewerdaniel/sovereignBank)), I implemented exactly this principle through the 7-layer memory hierarchy. Layer 0 (source) and Layer 1 (extracted concepts/claims/entities) are the residual accumulation layer — nothing is discarded, everything feeds forward:
```python # From Sovereign Memory Bank's memory hierarchy: # Every extraction round preserves all intermediate representations # as first-class graph nodes, regardless of "confidence"
class ExtractedClaim(BaseModel): text: str source_chunk_id: str confidence: float # low-confidence claims are NOT filtered — they persist residual_embedding: list[float] # distributional residual, not just argmax extraction_round: int # provenance for evolution tracking status: Literal["candidate", "verified", "contradicted", "superseded"] ```
The principle is structural: **even failure becomes state**. In the context of Dynamic Persona MoE RAG ([post](https://www.danielkliewer.com/blog/dynamic-persona-moe-rag), [repo](https://github.com/kliewerdaniel/dynamic_persona_moe_rag)), this means a persona that produces a low-confidence response doesn't get ignored — its partial output feeds into the next persona's conditioning. The activation_cost and historical_performance fields on each persona schema become the residual signal that shapes future routing decisions.
---
2. From Tool Use to Executable Systems
LMSYS and SGLang Agents
The **LMSYS** work on agent-assisted SGLang development pushes the next abstraction shift: the agent is no longer just a consumer of tools. It becomes part of the system that *defines execution*.
**Paper:** *SGLang: Efficient Execution of Structured Language Model Programs* — [arXiv:2312.07104](https://arxiv.org/abs/2312.07104) (Zheng et al., NeurIPS 2024) **Repo:** [github.com/sgl-project/sglang](https://github.com/sgl-project/sglang) (29.9k+ stars, 400k+ GPUs in production)
Instead of:
python
prompt → model → tool call → result
We start seeing:
python
agent → compiles execution graph → optimizes inference paths → rewrites runtime behavior → executes
SGLang already treats inference as a structured program through its Python-embedded DSL with primitives like gen, select, fork, join, and extend. What the agent layer adds is adaptability at the level of the execution graph itself.
Here's how SGLang represents a multi-step inference as a compilable graph:
```python import sglang as sgl
@sgl.function def multi_step_reasoning(context: str, question: str): """ SGLang compiles this into a computational graph that the runtime can optimize via code motion, instruction selection, and auto-tuning. """ # Step 1: Ana