Builds a searchable SQLite corpus of Reddit threads and comments about dating & relationship dynamics, then surfaces the threads worth reading by a behavioral taxonomy — ghosting, breadcrumbing, love-bombing, future-faking, hot/cold, pressure tactics, and more.
Retrieval, not detection. This tool ranks and exports threads worth reading; a downstream detector consumes the JSONL exports. The taxonomy keywords here are a retrieval aid (which threads to surface), not a classifier — buckets describe observable behaviors only.
Data source priority: Arctic Shift → PullPush → Reddit live JSON, stored in SQLite.
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python corpus.py init- Preconfigured subreddits — edit
config.yamlundersubreddits:(defaults target arranged-marriage / matrimony / relationship subs) - Discovery by search —
python corpus.py discover --query "arranged marriage", thenpython corpus.py select <name> - Seed URLs — drop Reddit thread URLs into
seeds.txt, runpython corpus.py pull-seeds
python corpus.py init
python corpus.py discover --query "arranged marriage" --limit 25
python corpus.py select arrangedmarriage
python corpus.py list-subs
python corpus.py pull --since 90d --max-posts 500
python corpus.py pull --subreddit arrangedmarriage
python corpus.py pull-thread https://www.reddit.com/r/arrangedmarriage/comments/abc123/title/
python corpus.py pull-seeds
python corpus.py stats
python corpus.py scan "ghosted" --subreddit arrangedmarriage --limit 50
python corpus.py surface --subreddit arrangedmarriage # per-bucket keyword hits + top threads
python corpus.py export --kind comments --subreddit arrangedmarriage --output am_comments.jsonlThe behavioral buckets and their retrieval keywords live in config.yaml under taxonomy_keywords:.
There's also a small web UI under web/ for browsing and cleaning the corpus.
- PullPush coverage currently goes up to ~May 2025. Use
pull-thread(Reddit live) for newer threads. - Don't reproduce verbatim Reddit content in any external artifact — aggregate and paraphrase.
- SQLite, single-writer. To share the corpus, rsync the
.dbfile (the DB itself is gitignored — code only).
See SPEC.md for the storage schema, full CLI contract, and config contract.