'objective03: My Laptop Eats the News and Talks Back' [post] deterministic
An autonomous news ingestion, claim extraction, contradiction tracking,
<iframe src="https://www.youtube.com/embed/-qL7OtkNQ80" title="objective03 demo" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
<br>
[Free Apple Silicon Download](https://6340588028610.gumroad.com/l/qkxkt)
objective03: A Locally-Run Autonomous News Ingestion and Contradiction Tracking System
Architecture Overview: From RSS Feeds to Audio Broadcasts on Consumer Hardware
---
objective03 is a Python daemon that performs autonomous news ingestion, atomic claim extraction, entity resolution, event clustering, contradiction detection, narrative analysis, and text-to-speech broadcast — all running on local hardware via llama.cpp (Metal GPU backend), KuzuDB (embedded temporal property graph), Qdrant (vector similarity search), and Qwen3-TTS (mlx_audio). The system ingests content from RSS feeds, Reddit subreddits, and YouTube channels, extracts structured factual claims with GBNF-enforced JSON schemas, detects typed contradictions across sources, clusters claims into events and narrative threads, and generates TTS-optimized audio broadcasts with voice cloning and procedural ambient audio. Zero cloud dependencies. Zero API calls.
---
Pipeline Architecture
The system operates as a five-group task scheduler running on independent intervals. Each group contains a sequence of subprocesses with configurable timeouts, failure limits, and circuit-breakers.
Task Group 1: Ingestion (default interval: 60s)
The ingestion module polls three source types:
- **RSS feeds** — HTTP GET with
If-None-Match/ETagsupport for conditional requests. Parsed viafeedparser. Documents normalized (HTML stripped, Unicode NFKC normalized, whitespace collapsed). - **Reddit subreddits** — OAuth2 authenticated API calls. Posts and comments extracted, metadata preserved (author, subreddit, upvotes, timestamps).
- **YouTube channels** —
yt-dlpfor metadata extraction and audio transcription. Channel upload schedules polled on configurable intervals.
All documents undergo SHA-256 deduplication before graph insertion. The normalized document is stored as a Document node in KuzuDB with a FROM_SOURCE edge pointing to the originating Source node.
Task Group 2: Analysis Pipeline (default interval: 120s)
This is the core processing pipeline, executing sequentially:
#### 2a. Claim Extraction
Each document is chunked and passed to a local LLM (llama.cpp, Metal backend). The model extracts atomic factual claims using a GBNF-defined grammar that enforces a strict JSON schema:
json
{
"claim": "string",
"confidence": "float (0.0-1.0)",
"stance": "positive | negative | neutral",
"topic": "string (tag)",
"evidence": "string (verbatim text span)"
}
GBNF grammar enforcement ensures the model's output is structurally valid — no schema drift, no optional fields appearing as required. Each claim node in KuzuDB carries a confidence property and an EXTRACTED_FROM edge to its source document.
#### 2b. Entity Resolution
A second local LLM call extracts named entities (PERSON, ORG, LOC, EVENT) from each document. Extracted entities are resolved against existing graph nodes via:
1. **Exact match** on entity name 2. **Fuzzy matching** using Levenshtein distance with configurable threshold 3. **Alias tracking** — multiple names resolved to the same entity node over time
Resolved entities receive a MENTIONS edge to the source document and an APPEARS_IN edge to any event nodes they participate in.
#### 2c. Event Clustering
Claims are assigned to events based on entity overlap. The algorithm:
1. Extracts all entities from the new claim
2. Queries the graph for existing event nodes connected to any of those entities
3. If matches found, the claim's ABOUT_EVENT edge points to the existing event
4. If no matches, a new Event node is created with emerging status
Events track:
- importance_score — computed from entity frequency, claim count, and temporal recency
- status — emerging -> active -> resolved
- temporal_start / temporal_end — bounded by earliest and latest claim timestamps
#### 2d. Contradiction Detection
New claims are embedded using BGE-Small-EN-v1.5 (384-dimensional) and indexed in Qdrant. For each new claim:
1. **Vector search** — cosine similarity query against the Qdrant collection. Candidates with similarity > 0.75 are returned. 2. **LLM classification** — each candidate pair is passed to the local LLM with a prompt template that classifies the relationship into one of five typed categories:
| Type | Definition | Example |
|------|-----------|---------|
| DIRECT_CONTRADICTION | Same proposition, opposite truth value | "GDP grew 3%" vs "GDP shrank 3%" |
| NUMERICAL_DISCREPANCY | Same proposition, different values | "100 casualties" vs "200 casualties" |
| FRAMING_DIFFERENCE | Same event, different narrative lens | "Tax relief" vs "Tax cut for corporations" |
| TEMPORAL_DISCREPANCY | Same event, different timing | "Signed Monday" vs "Signed Tuesday" |
| COMPATIBLE | No contradiction; semantic overlap warrants review | — |
Contradictions are persisted as CONTRADICTS edges between claim nodes with a type property. Contradictions are **never auto-resolved** — the system preserves the raw disagreement for downstream consumption.
#### 2e. Narrative Analysis
Claims not assigned to events (i.e., no entity overlap with existing event nodes) are grouped into narrative threads via embedding cosine similarity clustering (>0.75 threshold). Each cluster receives an LLM-generated label. Active narratives track:
- drift_score — semantic shift over time within the narrative
- framing_classification — dominant narrative frame (e.g., "economic," "political," "social")
#### 2f. Source Reliability Scoring
Each Source node accumulates a reliability score based on historical claim accuracy — measured by the frequency of that source's claims being contradicted by other sources. Sources that consistently produce contradictory claims see their reliability scores degrade over time.
#### 2g. Graph Update
All extracted nodes and edges are committed to KuzuDB in a single transaction per batch.
Task Group 3: Broadcast Generation (default interval: 90s)
A local LLM queries the KuzuDB graph via Cypher-like queries for:
- Top N events by importance_score
- Unresolved contradictions (all CONTRADICTS edges with no resolution)
- Active narratives (narratives with status == "active")
- System metrics (sources ingested, claims extracted, contradictions detected)
The LLM produces an 800–1200 word broadcast script optimized for TTS. The script uses <think> blocks for internal reasoning before the spoken output. The prompt template includes structural directives: opening summary, top events, contradiction deep-dives, narrative shifts, and closing metrics.
Task Group 4: Audio Production (default interval: 90s)
The broadcast script undergoes preprocessing: 1. **Chunking** — split into ~100-word segments 2. **Abbreviation expansion** — "U.S." -> "United States", "Dr." -> "Doctor" 3. **Number normalization** — "3.5%" -> "three and a half percent", "$500M" -> "five hundred million dollars" 4. **Date formatting** — "Jan 15, 2026" -> "January fifteenth, twenty twenty-six" 5. **Punctuation normalization** — ellipses, em-dashes, and other TTS-sensitive characters
Preprocessed segments are synthesized via Qwen3-TTS using mlx_audio. Voice cloning is supported via reference audio input. Synthesized audio segments are crossfaded at boundaries and queued for playback via afplay on macOS.
Task Group 5: Maintenance (default interval: 24h)
- **Memory consolidation** — low-importance events and old narratives pruned based on
importance_scorethresholds - **Graph evaluation** — sample of claims re-verified against source documents for accuracy metrics
- **E