Research-paper RAG assistant built with LangChain + LangGraph.
The project lets you:
- download papers from arXiv,
- ingest PDFs and optional Markdown into Chroma,
- ask one-off questions from the CLI,
- run interactive chat with persistent session memory.
From repository root:
- Create and activate a virtual environment.
PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1Git Bash:
python -m venv .venv
source .venv/Scripts/activate- Install dependencies.
pip install --upgrade pip
pip install -r requirements.txtIf you want local embeddings or Markdown ingestion, also install the optional extras:
pip install -r requirements-optional.txt- Create
.envfrom template and set your OpenAI key.
cp .env.example .envRequired minimum in .env:
OPENAI_API_KEY=sk-...
EMBEDDING_PROVIDER=openai
OPENAI_EMBEDDING_MODEL=text-embedding-3-small- Add PDFs to
data/papers(or download from arXiv):
python scripts/download_arxiv_pdfs.py \
--query "cat:cs.AI AND (all:retrieval OR all:RAG OR all:agents)" \
--max-results 20 \
--out-dir data/papers \
--metadata data/papers/metadata.csv \
--skip-existing- Build the vectorstore and ask a question.
python -m src.main ingest
python -m src.main ask "What is RAG?"- Start interactive chat.
python -m src.main chat --session-id demo- CLI entrypoint with
ingest,ask, andchatcommands. - Configurable embeddings provider:
- OpenAI embeddings (
EMBEDDING_PROVIDER=openai) - Local Hugging Face embeddings (
EMBEDDING_PROVIDER=local)
- OpenAI embeddings (
- Retrieval tool with configurable
RETRIEVAL_KandDOC_PREVIEW_CHARS. - Persistent chat memory via SQLite session history.
- arXiv bulk PDF downloader script with metadata CSV export.
src/main.py: CLI commands (ingest,ask,chat)populate_vectorstore.py: ingestion pipeline entrysrc/ingestion/*: loaders, chunking, embedding model selectionsrc/retrieval/*: vectorstore creation/loading and retrieval helperssrc/agent/*: LangGraph nodes, router, tools, memorysrc/eval/*: eval set generation, retrieval metrics, LLM-as-judge, run reportingscripts/download_arxiv_pdfs.py: arXiv query + PDF downloaderscripts/generate_eval_set.py: generates an eval set from ingested documentsscripts/evaluate.py: runs the eval set against the agent and writes a run reportdata/papers: local PDF corpusdata/eval: eval set and run reportstests/test_memory.py: memory persistence regression test
- Python 3.8+
- pip
- Internet access for OpenAI/Hugging Face/arXiv usage
Windows PowerShell:
python -m venv .venv
.\.venv\Scripts\Activate.ps1Git Bash on Windows:
python -m venv .venv
source .venv/Scripts/activatemacOS/Linux:
python3 -m venv .venv
source .venv/bin/activatepip install --upgrade pip
pip install -r requirements.txtCopy the template and edit values:
cp .env.example .envWindows PowerShell alternative:
Copy-Item .env.example .envAt minimum, set:
OPENAI_API_KEYEMBEDDING_PROVIDER(openaiorlocal)
Use the included script to fetch PDFs and metadata into data/papers.
Example:
python scripts/download_arxiv_pdfs.py \
--query "cat:cs.AI AND (all:retrieval OR all:RAG OR all:agents)" \
--max-results 30 \
--out-dir data/papers \
--metadata data/papers/metadata.csv \
--skip-existingCommon useful flags:
--query(repeatable): add multiple topic pulls--sort-by relevance|lastUpdatedDate|submittedDate--sort-order ascending|descending--delay-seconds 1.5: polite throttling between downloads--cafile <path>: custom CA bundle--insecure: disable SSL verification (troubleshooting only)
After adding PDFs (and optional Markdown), ingest once:
python -m src.main ingestWhenever you add, remove, or change documents afterward, re-run ingestion with
--clean so the vectorstore is rebuilt from scratch instead of appending to what's
already there (without --clean, re-ingesting duplicates every chunk that was
already indexed):
python -m src.main ingest --cleanOptional custom paths:
python -m src.main ingest \
--pdf-directory data/papers \
--md-directory docs \
--persist-directory ./chroma_db \
--cleanpython -m src.main ask "What is RAG?"You will see timing breakdown output:
pre-ask,imports,init,graph,agent,ask-total
python -m src.main chat --session-id my_sessionNotes:
- Type
exitorquitto end chat. - Session memory is stored in
chat_history.db(SQLite).
In .env:
EMBEDDING_PROVIDER=openai
OPENAI_EMBEDDING_MODEL=text-embedding-3-smallIn .env:
EMBEDDING_PROVIDER=local
LOCAL_EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L6-v2
HF_QUIET=trueInstall the optional dependency set first:
pip install -r requirements-optional.txtOptional for higher Hugging Face rate limits:
HF_TOKEN=hf_...Important: if you switch embedding providers/models, rebuild chroma_db.
Adjust in .env:
RETRIEVAL_K=3
DOC_PREVIEW_CHARS=800- Lower values reduce prompt size and latency.
- Higher values may increase context coverage at the cost of speed.
The project includes an eval harness that measures retrieval quality, answer quality, latency, and cost against a known set of questions — a baseline to compare future pipeline changes (e.g. a fine-tuned reranker) against.
- Build a real corpus and ingest it (see "Download Papers from arXiv" and "Build the Vector Store" above).
- Generate an eval set from the ingested documents:
python scripts/generate_eval_set.py \
--persist-directory ./chroma_db \
--output data/eval/eval_set.jsonl \
--questions-per-doc 2This uses an LLM to write a question + reference answer per sampled chunk, grounded
in the actual ingested content, and writes them to data/eval/eval_set.jsonl.
- Run the eval:
python -m src.main evalUseful flags:
--limit N: only run the first N examples (good for a quick sanity check).--skip-judge: skip LLM-as-judge answer scoring, retrieval metrics only (no extra API cost).--k,--search-type: tune the direct-retrieval comparison.
Each run computes:
- Retrieval metrics (hit rate, MRR, precision@k), both for a direct retriever call and for what the agent actually retrieved end-to-end.
- Answer quality, via LLM-as-judge scoring of faithfulness/relevance/correctness.
- Latency and cost per query.
Results are written to data/eval/runs/<timestamp>/ (results.jsonl + summary.json +
summary.csv), with a row appended to data/eval/runs/history.csv so runs can be
diffed over time (e.g. before/after adding a reranker).
pytest -qCurrent test coverage includes persistent memory behavior in tests/test_memory.py.
Alembic scaffolding (alembic.ini, alembic/env.py, alembic/versions/) is set up ahead of
any schema it needs to own — SQLChatMessageHistory (src/agent/memory.py) already manages its
own message_store table automatically, so the first migration is a no-op that just documents
that table's shape (alembic/versions/bf40fc10c929_*.py). This establishes the migrations
workflow now, so the next real schema need (session metadata, document metadata) has it ready
rather than retrofitting it under time pressure. See issue #36.
alembic/env.py builds its connection string via build_connection_string()
(src/config.py, added in #35) — DATABASE_URL, or the discrete
DB_HOST/DB_PORT/DB_NAME/DB_USER/DB_PASSWORD vars — the same helper src/agent/memory.py
is intended to be wired to (a separate, still-pending ticket), so there's exactly one place
connection details are assembled rather than two. alembic.ini's sqlalchemy.url is deliberately
left unset.
Not part of the default pytest run (needs a live Postgres), runnable on demand as a manual
smoke test:
docker run --rm -d --name alembic-smoke-test \
-e POSTGRES_PASSWORD=postgres -p 5433:5432 postgres:16
# Postgres takes a few seconds to accept connections after the container starts
until docker exec alembic-smoke-test pg_isready -U postgres -q; do sleep 1; done
DATABASE_URL=postgresql+psycopg://postgres:postgres@localhost:5433/postgres \
alembic upgrade head
# expect: "Running upgrade -> bf40fc10c929, document sql chat message history table"
# with no errors, and no `message_store` table created (the migration is a no-op).
docker stop alembic-smoke-testRun:
python -m src.main ingestVerify:
OPENAI_API_KEYis valid- billing/quota is available
The first startup can take longer due to initialization. Once prompt appears:
You:
type your question and press Enter.
Try one of:
- provide
--cafile <path> - install/update
certifi - use
--insecureonly for local troubleshooting
- Never commit
.env. - Rotate keys if they were ever exposed.
- Keep API keys scoped to least privilege where possible.
Built iteratively, one capability per commit (ingestion → retriever → LangGraph
agent → persistent memory → FastAPI → Streamlit UI), each followed by a test
before moving on. AI coding assistants (Claude Code) were used throughout for
implementation and review, with every change run through pytest, ruff,
and black via pre-commit before being accepted.
Larger feature work (e.g. the reranker in docs/specs/fine-tuned-reranker.md) is broken into
GitHub issues with explicit acceptance criteria and native blocking dependencies between them.
scripts/ralph/ralph.sh drives Claude Code headlessly and repeatedly — the "Ralph Wiggum" pattern
(one-shot, non-interactive invocations in a loop, checked against a completion sigil) — inside a
disposable Docker container (scripts/ralph/ralph.Dockerfile), so each run:
- Finds the lowest-numbered open, unblocked, unassigned
ready-for-agentissue viagh. - Claims it, implements it, and runs the full test suite/typecheck.
- Stops for human review — it does not commit or close the issue itself.
Authentication uses a long-lived token from claude setup-token, tied to an existing Claude
subscription rather than metered API credits. Running inside a container means a bad iteration's
blast radius is contained to a disposable filesystem rather than the host, aside from the one
bind-mounted checkout it's deliberately allowed to change. Full instructions live in
scripts/ralph/ralph-prompt.md.