Skip to content

Repository files navigation

Reddit Data Analyser

Builds a searchable SQLite corpus of Reddit threads and comments about dating & relationship dynamics, then surfaces the threads worth reading by a behavioral taxonomy — ghosting, breadcrumbing, love-bombing, future-faking, hot/cold, pressure tactics, and more.

Retrieval, not detection. This tool ranks and exports threads worth reading; a downstream detector consumes the JSONL exports. The taxonomy keywords here are a retrieval aid (which threads to surface), not a classifier — buckets describe observable behaviors only.

Data source priority: Arctic Shift → PullPush → Reddit live JSON, stored in SQLite.

Install

python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python corpus.py init

Three ways to define what you pull

  1. Preconfigured subreddits — edit config.yaml under subreddits: (defaults target arranged-marriage / matrimony / relationship subs)
  2. Discovery by search — python corpus.py discover --query "arranged marriage", then python corpus.py select <name>
  3. Seed URLs — drop Reddit thread URLs into seeds.txt, run python corpus.py pull-seeds

Commands

python corpus.py init
python corpus.py discover --query "arranged marriage" --limit 25
python corpus.py select arrangedmarriage
python corpus.py list-subs
python corpus.py pull --since 90d --max-posts 500
python corpus.py pull --subreddit arrangedmarriage
python corpus.py pull-thread https://www.reddit.com/r/arrangedmarriage/comments/abc123/title/
python corpus.py pull-seeds
python corpus.py stats
python corpus.py scan "ghosted" --subreddit arrangedmarriage --limit 50
python corpus.py surface --subreddit arrangedmarriage        # per-bucket keyword hits + top threads
python corpus.py export --kind comments --subreddit arrangedmarriage --output am_comments.jsonl

The behavioral buckets and their retrieval keywords live in config.yaml under taxonomy_keywords:. There's also a small web UI under web/ for browsing and cleaning the corpus.

Caveats

  • PullPush coverage currently goes up to ~May 2025. Use pull-thread (Reddit live) for newer threads.
  • Don't reproduce verbatim Reddit content in any external artifact — aggregate and paraphrase.
  • SQLite, single-writer. To share the corpus, rsync the .db file (the DB itself is gitignored — code only).

See SPEC.md for the storage schema, full CLI contract, and config contract.

About

Reddit data analyser for dating & relationship dynamics — builds a SQLite corpus from Arctic Shift / PullPush / Reddit JSON and surfaces threads by behavioral taxonomy (ghosting, breadcrumbing, love-bombing…).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages