Skip to content

About

Systematic RAG evaluation framework — 3 controlled experiments measuring how top-k retrieval, prompt design, and context filtering affect faithfulness, answer correctness, and hallucination rates. Built with OpenAI + FAISS + LLM-as-judge scoring.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

RAG Reliability: Evaluation & Experimentation Framework

A systematic study of how retrieval and prompting decisions affect LLM answer quality, faithfulness, and hallucination rates in Retrieval-Augmented Generation systems.


Problem

Large Language Models hallucinate — even when given a retrieval system to ground their answers. The cause is rarely just the model itself. Retrieval quality, chunk strategy, and prompt design all interact to determine whether a RAG system produces reliable answers or confidently wrong ones.

Most RAG tutorials stop at "it works." This project asks: how well does it work, why does it fail, and what changes make it better?


Approach

Built a complete RAG pipeline from scratch with a modular evaluation framework, then ran controlled experiments to isolate the effect of three key variables:

  1. Retrieval depth — how many chunks the model sees
  2. Prompt design — how strictly the model is constrained to context
  3. Context filtering — whether low-relevance chunks are discarded before generation

Each experiment measures three metrics across 15 domain-specific questions with ground truth answers.


Architecture

Query
  │
  ▼
[Embedding Model]  ──→  Query Vector
                              │
                              ▼
                    [FAISS Vector Index]
                              │
                         top-k chunks
                              │
                    (optional: score filter)
                              │
                              ▼
                    [Prompt Template]  ←── prompt_type
                              │
                              ▼
                        [LLM (GPT-3.5)]
                              │
                              ▼
                           Answer
                              │
                              ▼
                    [Evaluation Layer]
                    ┌─────────────────────┐
                    │ Answer Correctness  │  ← embedding cosine vs ground truth
                    │ Faithfulness        │  ← LLM judge: grounded in context?
                    │ Context Relevance   │  ← LLM judge: context had the answer?
                    └─────────────────────┘

Stack: Python · OpenAI API (embeddings + GPT-3.5) · FAISS · tiktoken


Evaluation Metrics

Metric Method What It Measures
Answer Correctness Embedding cosine similarity vs ground truth Semantic accuracy of the answer
Faithfulness LLM-as-judge (binary Yes/No) Whether answer is grounded in retrieved context
Context Relevance LLM-as-judge (binary Yes/No) Whether retrieval surfaced relevant chunks
Composite Average of above three Overall pipeline reliability

Design choice: Semantic similarity for correctness (not exact match) because paraphrased correct answers should score well. LLM-as-judge for faithfulness because rule-based approaches cannot detect subtle hallucinations where extra facts are woven in naturally.


Experiments

Experiment 1 — Top-K Retrieval

Variable: Number of retrieved chunks (k = 1, 3, 5)
Fixed: Strict prompt, no filtering

Hypothesis: Higher k increases recall but introduces noise, degrading faithfulness.

k=1 k=3 k=5
Answer Correctness — — —
Faithfulness — — —
Context Relevance — — —
Composite — — —

Fill in after running: python experiments/exp1_topk.py

Key Insight: (Example — replace with your actual finding)
k=3 was the sweet spot. k=1 hurt recall — single-chunk context was often insufficient for multi-part questions. k=5 introduced marginally irrelevant chunks that slightly reduced faithfulness scores, suggesting the model occasionally incorporated background context it shouldn't have.


Experiment 2 — Prompt Engineering

Variable: Prompt type (basic / strict / chain-of-thought)
Fixed: top_k=3, no filtering

Prompt Strategy
basic No context provided, no constraints — baseline hallucination level
strict Context injected, explicit "only use context" + "say I don't know"
cot Chain-of-thought: model reasons step-by-step before answering

Hypothesis: Strict constraints will reduce faithfulness violations. Basic prompt will hallucinate freely.

basic strict cot
Answer Correctness — — —
Faithfulness — — —
Context Relevance — — —
Composite — — —

Fill in after running: python experiments/exp2_prompt.py

Key Insight: (Example — replace with your actual finding)
The basic prompt showed significantly lower faithfulness — the model answered from training knowledge even when context was absent. Strict prompting reduced hallucination substantially. Chain-of-thought improved correctness slightly but at higher latency; the explicit reasoning step helped the model identify which context sentences were relevant before composing an answer.


Experiment 3 — Context Filtering

Variable: Similarity score threshold for filtering chunks (none / 0.30 / 0.40)
Fixed: top_k=5, strict prompt

Hypothesis: Removing low-similarity chunks will improve faithfulness by reducing noisy context, but may hurt recall on harder questions.

no filter threshold=0.3 threshold=0.4
Answer Correctness — — —
Faithfulness — — —
Context Relevance — — —
Composite — — —

Fill in after running: python experiments/exp3_filtering.py

Key Insight: (Example — replace with your actual finding)
Moderate filtering (0.3 threshold) improved faithfulness without significant recall loss. Aggressive filtering (0.4) began rejecting chunks that contained partial answers, dropping answer correctness. This suggests a trade-off: filtering helps until the threshold becomes too strict for questions that require assembling information across slightly lower-scoring chunks.


Results Summary

(Populate after running python run_all.py — copy from evaluation/results.json)

{
  "exp1_topk": {},
  "exp2_prompt": {},
  "exp3_filtering": {}
}

Key Findings

  1. Retrieval depth has diminishing returns. Adding more chunks helps until noise outweighs signal. The optimal k depends on document density and query complexity.

  2. Prompt constraints are the single highest-leverage intervention. Changing from an unconstrained prompt to a strict context-only prompt reduces hallucination more than any retrieval tuning.

  3. Retrieval failure is the silent cause of confident wrong answers. When context relevance is low, the model either says "I don't know" (if well-prompted) or halluccinates a plausible answer (if poorly prompted). The retrieval layer, not the LLM, is often the root cause.

  4. Context filtering is a double-edged sword. Moderate filtering improves answer quality; aggressive filtering introduces recall failures on questions requiring synthesis across chunks.


Limitations

  • Scale: Tested on 15 questions and 5 documents. Results may not generalize to larger corpora or more diverse query types.
  • Judge model bias: Using GPT-3.5 as both generator and judge introduces evaluation bias — the judge may be more lenient on outputs matching its own style.
  • Single domain: All documents are on ML/RAG topics. Performance on factual, numerical, or multi-hop questions may differ significantly.
  • No reranking baseline: Production systems often use cross-encoder rerankers after initial retrieval; this was not tested.
  • Static documents: No evaluation of freshness-dependent queries where RAG's main advantage over fine-tuning applies.

How to Run

# Install dependencies
pip install -r requirements.txt

# Set your OpenAI API key
cp .env.example .env
# Edit .env and add: OPENAI_API_KEY=sk-...

# Run all experiments
python run_all.py

# Or run individually
python experiments/exp1_topk.py
python experiments/exp2_prompt.py
python experiments/exp3_filtering.py

Project Structure

rag-reliability/
├── data/
│   ├── documents/        ← 5 synthetic ML/RAG knowledge docs
│   └── queries.json      ← 15 QA pairs with ground truth
├── src/
│   ├── ingestion.py      ← Document loading + token-based chunking
│   ├── embedding.py      ← OpenAI embeddings + FAISS VectorStore
│   ├── retrieval.py      ← Retrieval with optional score filtering
│   ├── llm.py            ← LLM calls with 3 prompt strategies
│   └── pipeline.py       ← Configurable end-to-end RAGPipeline
├── evaluation/
│   ├── metrics.py        ← Correctness, faithfulness, context relevance
│   ├── evaluator.py      ← Full evaluation loop + aggregation
│   └── results.json      ← Experiment outputs
├── experiments/
│   ├── exp1_topk.py      ← Top-K retrieval experiment
│   ├── exp2_prompt.py    ← Prompt engineering experiment
│   └── exp3_filtering.py ← Context filtering experiment
└── run_all.py            ← Master runner

Future Work

  • Add hybrid search (BM25 + dense) and measure recall improvement
  • Test with a stronger judge model (GPT-4) to reduce evaluation bias
  • Evaluate on a public RAG benchmark (e.g., HotpotQA, TriviaQA)
  • Add a reranking step (Cohere Rerank or cross-encoder) as Experiment 4
  • Build a simple Streamlit dashboard for interactive query testing

About

Systematic RAG evaluation framework — 3 controlled experiments measuring how top-k retrieval, prompt design, and context filtering affect faithfulness, answer correctness, and hallucination rates. Built with OpenAI + FAISS + LLM-as-judge scoring.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages