'The Telemetry Intelligence Engine: A Local-First GraphRAG System for Website Analytics' [post] deterministic
'A spec-driven walkthrough of the Telemetry Intelligence Engine (TIE): a local-first GraphRAG system that turns GA4 telemetry and site content into a behavioral knowledge graph an operator can query in natural language.'
Every analytics dashboard I have ever used answers the same narrow question well: *what happened*. Pageviews, sessions, bounce rate, referral source. What none of them answer is *why it matters* — which pieces of content are actually building toward something, which pathways are quietly leaking high-value visitors, and what I should write next to close the gap between what people are looking for and what I've actually published.
That gap is a reasoning problem, not a reporting problem. And reasoning problems are exactly what local LLMs plus a knowledge graph are good at, provided you're willing to build the plumbing yourself instead of waiting for a SaaS dashboard to grow a brain.
This post is the specification and MVP plan for the **Telemetry Intelligence Engine (TIE)** — a local-first GraphRAG system that treats website analytics as a behavioral knowledge graph rather than a spreadsheet, enriches it with local inference, and lets an operator ask it questions in plain language. It's a direct extension of the Dynamic Persona MoE RAG architecture I've written about previously, retargeted at analytics intelligence instead of general knowledge retrieval.
If you've been following the sovereign AI thread on this site, the pattern will be familiar: the information architecture — the graph, the audit trail, the relationships between entities — is the actual product. The model is just the reasoning engine you point at it.
The core idea
Standard analytics tools store *events*. TIE stores *relationships between events, content, and outcomes*, and lets an LLM walk that graph to answer questions no dashboard was designed to answer:
- "What topics are attracting the highest-value visitors?"
- "What content pathways lead people toward my projects?"
- "What concepts are underrepresented compared to visitor interest?"
- "What should I write next based on observed knowledge gaps?"
- "Why are visitors leaving after reading certain pages?"
These aren't aggregation queries. They require connecting a visitor's session, to the content they touched, to the topics that content covers, to the conversion events (or lack thereof) that followed — and then reasoning over that structure. That's a graph traversal problem wrapped in a retrieval-augmented generation problem, which is precisely the combination GraphRAG architectures are built for.
Architecture overview
At a high level, the system has four moving parts: a telemetry processor that normalizes raw GA4 exports, a behavioral graph that encodes relationships, a vector store that encodes semantic similarity, and a local RAG analyst that reasons over both.
GA4 Export
|
v
Raw Telemetry JSON
|
v
Telemetry Processor
/ \
v v
Behavioral Graph Vector Database
(NetworkX / Neo4j) (ChromaDB)
\ /
v v
Local RAG Analyst
(Ollama / llama.cpp)
|
v
Insights + Recommendations
The split between graph and vector store matters. The graph captures *explicit structural relationships* — this article discusses this topic, this session viewed this page, this page leads to this conversion event. The vector store captures *semantic similarity* — which graph summaries and content chunks are conceptually close to a given question, even when no explicit edge connects them. Query time uses both: semantic retrieval narrows the search space, then graph traversal pulls in the connected neighborhood the LLM actually reasons over.
Data sources
Analytics data
The initial data source is a GA4 export, normalized into a consistent event schema:
json
{
"timestamp": "",
"event_name": "",
"page_path": "",
"session_id": "",
"user_country": "",
"device_category": "",
"traffic_source": "",
"referrer": "",
"engagement_time": "",
"scroll_depth": "",
"events": []
}
This is deliberately the minimum viable schema. Search Console data, GitHub traffic analytics, newsletter open/click metrics, social referral data, server logs, and error telemetry are all planned as future ingestion sources, but the MVP doesn't need them to prove the architecture out. Get one clean pipe of data flowing before adding more.
Content knowledge layer
The site's existing content — blog posts, project pages, essays — becomes graph entities in their own right, not just URLs that telemetry events point at:
/content
|
+-- blog/
+-- projects/
+-- essays/
Each document is parsed into a structured entity:
json
{
"id": "dynamic_persona_rag",
"type": "article",
"title": "Dynamic Persona MoE RAG",
"topics": ["RAG", "agents", "knowledge graphs"],
"entities": ["Ollama", "ChromaDB", "LLMs"]
}
This is the piece most analytics tools skip entirely — they know a URL got 1,200 views, but they have no model of what that URL is actually *about*, or how it relates conceptually to everything else you've published. Without this layer, "what should I write next" isn't answerable at all.
Knowledge graph schema
The schema is organized into three node families, which keeps the graph legible as it grows instead of collapsing into an undifferentiated blob of "things."
**Content nodes:** Article, Project, Page, Repository, Topic, Keyword, Technology
**User behavior nodes:** Visitor Segment, Session, Traffic Source, Device Type, Conversion Event
**Analytical nodes:** Hypothesis, Recommendation, Opportunity, Knowledge Gap, Trend
That third category is the important one and the one most graph-based analytics prototypes leave out. Most systems model content and behavior; few model *the analysis itself* as first-class graph entities. Making hypotheses and recommendations nodes — rather than throwaway text in a report — means the system can later reason about which hypotheses it already tested, which recommendations it already made, and whether outcomes changed after implementation. That's what makes the self-improving loop in Phase 5 possible at all.
Relationships
The edges are what turn a pile of nodes into something queryable:
Visitor --viewed--> Article
Article --discusses--> Topic
Topic --related_to--> Project
Article --leads_to--> Conversion
And behavior paths chain these into traversable sequences:
Google Search
|
v
Ollama Article
|
v
Dynamic Persona RAG
|
v
GitHub Click
A path like this is exactly the kind of thing a traditional funnel report *approximates* with drop-off percentages, but a graph traversal states explicitly: this session entered through this search query, read this article, followed an internal link to this project page, and then clicked out to the repository. Once paths like this are graph-native, you can ask the LLM to generalize across hundreds of them and surface the pattern rather than eyeballing a funnel chart.
LLM enrichment pipeline
Raw telemetry is not semantic. "1,200 views, 240 seconds average time on page" doesn't mean anything on its own — it needs an interpretive layer between the raw numbers and the graph. That's the job of the enrichment pipeline, run locally, in three stages.
**Stage 1 — Event summarization.** Convert raw aggregates into a plain-language characterization:
```json // input {"page": "/projects/rag", "views": 1200, "time": 240}
// output {"meaning": "High-interest technical content attracting AI engineering audience"} ```
**Stage 2 — Entity extraction.** Pull structured topics, audience, and intent out of content and behavior:
Topics:
- AI Agents
- Retrieval Systems
- Local Inference
Audience:
- Developers
- Researchers
Intent:
- Technical exploration
**Stage 3 — Relationship discovery.** Propose new graph edges with a stated rationale, ra