A local-first, multi-stage AI research agent. You give it a question; it parses the intent, expands it into targeted searches, fetches the open web + Reddit + Hacker News + PDFs, and then runs an LLM relevance gate so the final report cites sources that actually answer the question — instead of everything that happened to return HTTP 200.
Runs entirely on your machine (CLI or a small FastAPI web UI). Extraction and cheap stages can run on local Ollama models; synthesis uses Anthropic (Haiku for the fast path, Opus for the final report).
Most "research agent" demos fetch a pile of URLs and summarize whatever comes back. The failure mode isn't fetching — it's precision. On an early reference run (query: "indian SME preferred AP/AR software and problems no AP/AR is solving") the pipeline fetched 78 pages and "approved" all 78. Only 8 were genuinely relevant — 10% precision. The rest were disambiguation collisions (APAR Industries the cable company, AR/VR, ARPU, the AP exam, "SME" = Subject Matter Expert), SEO-spam subdomains, IIS 404 pages returned as 200s, and the same registration page fetched four times under URL variants.
DeepResearch is built around fixing exactly that. See instructions.md for the full design doc — the P0/P1 work that took precision from 10% toward a 32% floor, with a regression fixture set pinning known-good and known-bad URLs.
query
│
├─ parse intent, subject, subject_domain (disambiguation), polarity, qualifiers
├─ decompose split multi-clause queries ("preferred X AND problems with Y") into sub-questions
├─ expand generate targeted search queries (topic-conditional, not blanket site: injection)
├─ search Tavily
├─ fetch Jina Reader (JS/bot-blocked pages), Reddit JSON, Hacker News, PDF
│ + content validation (reject 404-templates, bot-challenges, title-only stubs)
│ + URL canonicalization & dedupe, domain/subdomain reputation, size caps
├─ relevance per-candidate LLM gate → yes / no / maybe (drop no, second-pass the maybes)
├─ approval human-in-the-loop (or --auto-approve --top N)
├─ per-source signal extraction, source-class weighting (Reddit vs vendor blog vs gov)
├─ coverage check each sub-question is actually answered
├─ synthesis report generation (Opus)
└─ verify critique pass + flag unverified claims
Every LLM call is logged to llm_calls.jsonl; every run persists intermediate state per stage, so a failed run resumes from its last checkpoint instead of re-fetching everything.
# install (uses uv; pip works too)
uv sync # or: pip install -e .
cp .env.example .env # fill in your keys
# CLI
uv run python main.py "what are indian saas founders saying about CRMs"
uv run python main.py --list
uv run python main.py --show 2026-05-06-1430-indian-saas
uv run python main.py --auto-approve --top 10 "your query"
uv run python main.py --retry <RUN_ID> # resume a failed run from checkpoint
# Web UI
uv run uvicorn app:app --reload --port 8000 # then open http://localhost:8000Copy .env.example → .env:
| Var | Purpose |
|---|---|
ANTHROPIC_API_KEY |
extraction (Haiku) + synthesis (Opus) |
TAVILY_API_KEY |
web search |
JINA_API_KEY |
reader/fetcher for JS-heavy & bot-blocked pages |
OLLAMA_BASE_URL / OLLAMA_*_MODEL |
optional local models for cheap stages |
PULSE_*_MODEL |
per-stage model overrides (parse/expand/relevance/synthesis/…) |
Per-domain reputation lives in domains.yaml; pre-fetch noise filters in url_blocklist_patterns.yaml.
main.py CLI
app.py FastAPI web server + SSE run streaming
src/
parse.py constraints.py expand.py query understanding
search.py fetchers/{jina,reddit,hn,pdf}.py retrieval
extract.py relevance.py approval.py filtering & the precision gate
per_source_signal.py source_inventory.py coverage_check.py
synthesis.py verify.py critique.py unverified.py report + self-check
storage.py schemas.py run persistence
domains.yaml url_blocklist_patterns.yaml
instructions.md design doc / engineering log
Working prototype. Local-first by design; run data (data/) and model weights are gitignored and never leave your machine.