A systematic study of how retrieval and prompting decisions affect LLM answer quality, faithfulness, and hallucination rates in Retrieval-Augmented Generation systems.
Large Language Models hallucinate — even when given a retrieval system to ground their answers. The cause is rarely just the model itself. Retrieval quality, chunk strategy, and prompt design all interact to determine whether a RAG system produces reliable answers or confidently wrong ones.
Most RAG tutorials stop at "it works." This project asks: how well does it work, why does it fail, and what changes make it better?
Built a complete RAG pipeline from scratch with a modular evaluation framework, then ran controlled experiments to isolate the effect of three key variables:
- Retrieval depth — how many chunks the model sees
- Prompt design — how strictly the model is constrained to context
- Context filtering — whether low-relevance chunks are discarded before generation
Each experiment measures three metrics across 15 domain-specific questions with ground truth answers.
Query
│
▼
[Embedding Model] ──→ Query Vector
│
▼
[FAISS Vector Index]
│
top-k chunks
│
(optional: score filter)
│
▼
[Prompt Template] ←── prompt_type
│
▼
[LLM (GPT-3.5)]
│
▼
Answer
│
▼
[Evaluation Layer]
┌─────────────────────┐
│ Answer Correctness │ ← embedding cosine vs ground truth
│ Faithfulness │ ← LLM judge: grounded in context?
│ Context Relevance │ ← LLM judge: context had the answer?
└─────────────────────┘
Stack: Python · OpenAI API (embeddings + GPT-3.5) · FAISS · tiktoken
| Metric | Method | What It Measures |
|---|---|---|
| Answer Correctness | Embedding cosine similarity vs ground truth | Semantic accuracy of the answer |
| Faithfulness | LLM-as-judge (binary Yes/No) | Whether answer is grounded in retrieved context |
| Context Relevance | LLM-as-judge (binary Yes/No) | Whether retrieval surfaced relevant chunks |
| Composite | Average of above three | Overall pipeline reliability |
Design choice: Semantic similarity for correctness (not exact match) because paraphrased correct answers should score well. LLM-as-judge for faithfulness because rule-based approaches cannot detect subtle hallucinations where extra facts are woven in naturally.
Variable: Number of retrieved chunks (k = 1, 3, 5)
Fixed: Strict prompt, no filtering
Hypothesis: Higher k increases recall but introduces noise, degrading faithfulness.
| k=1 | k=3 | k=5 | |
|---|---|---|---|
| Answer Correctness | — | — | — |
| Faithfulness | — | — | — |
| Context Relevance | — | — | — |
| Composite | — | — | — |
Fill in after running:
python experiments/exp1_topk.py
Key Insight: (Example — replace with your actual finding)
k=3 was the sweet spot. k=1 hurt recall — single-chunk context was often insufficient for multi-part questions. k=5 introduced marginally irrelevant chunks that slightly reduced faithfulness scores, suggesting the model occasionally incorporated background context it shouldn't have.
Variable: Prompt type (basic / strict / chain-of-thought)
Fixed: top_k=3, no filtering
| Prompt | Strategy |
|---|---|
basic |
No context provided, no constraints — baseline hallucination level |
strict |
Context injected, explicit "only use context" + "say I don't know" |
cot |
Chain-of-thought: model reasons step-by-step before answering |
Hypothesis: Strict constraints will reduce faithfulness violations. Basic prompt will hallucinate freely.
| basic | strict | cot | |
|---|---|---|---|
| Answer Correctness | — | — | — |
| Faithfulness | — | — | — |
| Context Relevance | — | — | — |
| Composite | — | — | — |
Fill in after running:
python experiments/exp2_prompt.py
Key Insight: (Example — replace with your actual finding)
The basic prompt showed significantly lower faithfulness — the model answered from training knowledge even when context was absent. Strict prompting reduced hallucination substantially. Chain-of-thought improved correctness slightly but at higher latency; the explicit reasoning step helped the model identify which context sentences were relevant before composing an answer.
Variable: Similarity score threshold for filtering chunks (none / 0.30 / 0.40)
Fixed: top_k=5, strict prompt
Hypothesis: Removing low-similarity chunks will improve faithfulness by reducing noisy context, but may hurt recall on harder questions.
| no filter | threshold=0.3 | threshold=0.4 | |
|---|---|---|---|
| Answer Correctness | — | — | — |
| Faithfulness | — | — | — |
| Context Relevance | — | — | — |
| Composite | — | — | — |
Fill in after running:
python experiments/exp3_filtering.py
Key Insight: (Example — replace with your actual finding)
Moderate filtering (0.3 threshold) improved faithfulness without significant recall loss. Aggressive filtering (0.4) began rejecting chunks that contained partial answers, dropping answer correctness. This suggests a trade-off: filtering helps until the threshold becomes too strict for questions that require assembling information across slightly lower-scoring chunks.
(Populate after running python run_all.py — copy from evaluation/results.json)
{
"exp1_topk": {},
"exp2_prompt": {},
"exp3_filtering": {}
}-
Retrieval depth has diminishing returns. Adding more chunks helps until noise outweighs signal. The optimal k depends on document density and query complexity.
-
Prompt constraints are the single highest-leverage intervention. Changing from an unconstrained prompt to a strict context-only prompt reduces hallucination more than any retrieval tuning.
-
Retrieval failure is the silent cause of confident wrong answers. When context relevance is low, the model either says "I don't know" (if well-prompted) or halluccinates a plausible answer (if poorly prompted). The retrieval layer, not the LLM, is often the root cause.
-
Context filtering is a double-edged sword. Moderate filtering improves answer quality; aggressive filtering introduces recall failures on questions requiring synthesis across chunks.
- Scale: Tested on 15 questions and 5 documents. Results may not generalize to larger corpora or more diverse query types.
- Judge model bias: Using GPT-3.5 as both generator and judge introduces evaluation bias — the judge may be more lenient on outputs matching its own style.
- Single domain: All documents are on ML/RAG topics. Performance on factual, numerical, or multi-hop questions may differ significantly.
- No reranking baseline: Production systems often use cross-encoder rerankers after initial retrieval; this was not tested.
- Static documents: No evaluation of freshness-dependent queries where RAG's main advantage over fine-tuning applies.
# Install dependencies
pip install -r requirements.txt
# Set your OpenAI API key
cp .env.example .env
# Edit .env and add: OPENAI_API_KEY=sk-...
# Run all experiments
python run_all.py
# Or run individually
python experiments/exp1_topk.py
python experiments/exp2_prompt.py
python experiments/exp3_filtering.pyrag-reliability/
├── data/
│ ├── documents/ ← 5 synthetic ML/RAG knowledge docs
│ └── queries.json ← 15 QA pairs with ground truth
├── src/
│ ├── ingestion.py ← Document loading + token-based chunking
│ ├── embedding.py ← OpenAI embeddings + FAISS VectorStore
│ ├── retrieval.py ← Retrieval with optional score filtering
│ ├── llm.py ← LLM calls with 3 prompt strategies
│ └── pipeline.py ← Configurable end-to-end RAGPipeline
├── evaluation/
│ ├── metrics.py ← Correctness, faithfulness, context relevance
│ ├── evaluator.py ← Full evaluation loop + aggregation
│ └── results.json ← Experiment outputs
├── experiments/
│ ├── exp1_topk.py ← Top-K retrieval experiment
│ ├── exp2_prompt.py ← Prompt engineering experiment
│ └── exp3_filtering.py ← Context filtering experiment
└── run_all.py ← Master runner
- Add hybrid search (BM25 + dense) and measure recall improvement
- Test with a stronger judge model (GPT-4) to reduce evaluation bias
- Evaluate on a public RAG benchmark (e.g., HotpotQA, TriviaQA)
- Add a reranking step (Cohere Rerank or cross-encoder) as Experiment 4
- Build a simple Streamlit dashboard for interactive query testing