Skip to content

About

Local-first Document Intelligence service with hybrid RAG, evidence tracing, answerability gating and reproducible evaluation.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Document Intelligence — Week 2 AI Engineering Project

This repository is a clean, standalone copy of the Week-2 Document Intelligence service. It is a local-first PDF intelligence system designed to make a wrong answer diagnosable: the trace separates ingestion, retrieval, fusion, reranking, answerability, prompt construction and generation.

No private development PDFs, model weights, Qdrant databases, credentials or parent-repository history are included.

Start here (recommended)

Requires Docker Desktop or Docker Engine with Compose v2. This single command starts the API, worker, Qdrant, Ollama and Demo UI; no second terminal or host Ollama installation is required:

docker compose -f compose.yaml -f compose.ollama.yaml \
  --profile bundled-ollama up --build -d

Compose automatically seeds the six fictional NOVA demo PDFs through the normal POST /v1/documents ingestion API. It waits for those ingestion jobs to finish before starting the Demo UI. Open http://127.0.0.1:8501 after the stack reports the API healthy and demo-seed completed.

Packaging note: this command is for a GitHub clone, where compose.yaml is at the repository root. If you are using the separate Document_Intelligence_2_Hafta_Teslim_Paketi ZIP, run its root-safe command from the extracted document-intelligence-delivery directory instead:

docker compose \
  -f source/compose.yaml \
  -f source/compose.ollama.yaml \
  --project-directory source \
  --profile bundled-ollama \
  up --build -d

To stop the project while preserving downloaded models and indexed data:

docker compose -f compose.yaml -f compose.ollama.yaml \
  --profile bundled-ollama down --remove-orphans

What the project demonstrates

The service accepts arbitrary parseable PDFs, indexes them with a deterministic pipeline, and answers only when the retrieved evidence passes the configured answerability policy. Every answer has application-generated canonical source metadata; source cards are never reconstructed from model text.

The core flow is:

PDF
 → parse → normalize → AUTO chunk selection
 → dense + BM25 → RRF
 → optional reranker → canonical evidence
 → answerability → structured prompt → local LLM
 → answer + canonical sources

The ASK flow runs one retrieval strategy at a time: Dense only, BM25 only, or Hybrid RRF. Hybrid means Dense + BM25 → RRF; BENCHMARKS is the separate place where multiple strategies and reranker states are compared.

The mentor UI is organized as ASK | DOCUMENTS | BENCHMARKS; the ASK result page contains the single Stage Explorer and keeps the full engineering trace under progressive disclosure.

When the Demo UI is opened, its tenant defaults to the isolated final-demo-v1 corpus containing the six fictional NOVA PDFs. Compose's demo-seed init service uploads those PDFs through the normal POST /v1/documents path and waits for successful ingestion before the UI is released. Compose seeds that corpus automatically on first startup and skips files that already have the current active version on later restarts. Set DEMO_SEED_ENABLED=false only when an empty/manual tenant is intentionally required. The historical default tenant is not deleted: it remains available when explicitly entered for benchmark/validation work, but its private or mentor source documents are not part of the normal six-question demo scope.

Measured engineering decisions

The final frozen mentor corpus contains 26 points in snapshot c5e87f7e063769adef368866854d8e45f7b7f9856f905abe9cebe31783262b25. The current final retrieval artifact reports:

Variant Recall@5 MRR@10 nDCG@10
Dense 0.9011 0.8750 0.9296
BM25 0.8178 0.7844 0.8376
Hybrid RRF 0.9233 0.8778 0.9518
Hybrid + reranker 0.9122 0.8333 0.9329

Therefore the demo default is Hybrid RRF with reranker OFF. The reranker remains available for measured ablation; it is not disabled by assumption.

Other validated behaviors:

  • no-answer and security-policy decisions skip the LLM;
  • duplicate ingestion reuses the same document/version when content and the effective pipeline fingerprint match;
  • arbitrary uploads default to AUTO, resolving to a structure-aware strategy only when reliable structure is present and otherwise to bounded generic_v1;
  • the frozen mentor corpus remains an explicit mentor_program_v1 evaluation membership, not a global admission rule;
  • canonical evidence retains document, page, parent/child and rank metadata;
  • installed models and ready models are reported separately.

Quick start

Requirements:

  • Docker Engine/Desktop with Compose v2;
  • Python 3.12 and the development tools for local tests.

Recommended: fully Docker-managed demo

No host Ollama installation or second terminal is required. Compose starts an Ollama container, pulls gemma3:4b on the first run, persists it in a named Docker volume, and does not start the UI until the model is reachable.

docker compose -f compose.yaml -f compose.ollama.yaml \
  --profile bundled-ollama up --build -d

This command works with Docker Engine on Linux and Docker Desktop on macOS or Windows. For a readiness-gated Bash launcher with the same Docker-managed runtime, use:

./scripts/start_demo.sh --bundled-ollama

Open the Demo UI at http://127.0.0.1:8501 after docker compose ps shows the API healthy. The launcher additionally waits for API liveness and readiness, then prints the exact API, health, Qdrant and UI URLs. It exits non-zero and prints the real dependency status if a required check fails.

Stop the standalone stack without deleting its persisted data:

docker compose -f compose.yaml -f compose.ollama.yaml \
  --profile bundled-ollama down --remove-orphans

down keeps the Qdrant and Ollama named volumes. Do not add -v unless you intentionally want to delete indexed data and the downloaded model.

Optional: use an existing host Ollama

The base Compose file still supports an already-managed host Ollama runtime. This is useful when the model is already installed and you do not want a second Ollama container, but it requires that the runtime be reachable from Docker.

./scripts/start_demo.sh --host-ollama

The host mode uses host.docker.internal:11434 and maps API/Qdrant/UI to 8010/6335/8501. Override those host ports or the Ollama URL when needed:

API_HOST_PORT=8011 QDRANT_HOST_PORT=6336 UI_HOST_PORT=8502 \
DIS_OLLAMA_URL=http://host.docker.internal:11434 \
./scripts/start_demo.sh --host-ollama

On Linux, the most portable option is to run Ollama on a separate local port if an existing service already occupies 11434:

# Terminal 1
OLLAMA_HOST=0.0.0.0:11435 ollama serve

# Terminal 2
OLLAMA_HOST=http://127.0.0.1:11435 ollama pull gemma3:4b
DIS_OLLAMA_URL=http://host.docker.internal:11435 \
API_HOST_PORT=8011 QDRANT_HOST_PORT=6336 UI_HOST_PORT=8502 \
./scripts/start_demo.sh --host-ollama

Docker Desktop normally provides host.docker.internal on macOS and Windows. The Compose file adds the same host-gateway mapping on Linux. Keep the Ollama listener restricted to the local machine/network; this project does not require public model exposure.

Readiness is intentionally not equivalent to process liveness. If the model, Ollama or Qdrant check is unavailable, the API reports not-ready and the UI does not pretend that generation is available. The API Compose healthcheck is also readiness-based, so dependent UI startup is blocked until the selected model is actually installed and reachable.

Demo flow

  1. Check live and ready health.
  2. Upload a parseable PDF in the DOCUMENTS tab. Product uploads use AUTO and can fall back to generic_v1; no mentor headings are required.
  3. Select only documents with an active searchable version.
  4. Run a direct fact, a paraphrase and an exact/numeric query.
  5. Inspect Dense, BM25, RRF, evidence, answerability and the canonical source in the single Stage Explorer. Prompt Packing exposes the actual bounded fragments sent to generation, including source/page and omitted-window metadata.
  6. Run the unrelated-question case and verify NO_ANSWER with the LLM skipped.
  7. Run the prompt-injection regressions. A direct, high-confidence injection is blocked with SECURITY_POLICY; the full Atlas indirect-injection case removes unsafe evidence and returns NO_ANSWER · INSUFFICIENT_COVERAGE when no supported answer remains.
  8. Compare retrieval/reranker variants in BENCHMARKS.

For the complete Turkish multi-document mentor corpus, the Compose seed is automatic. To explicitly reset and re-ingest only the demo tenant (for example, after deleting its persistent volumes), use demo/final_demo_pack/README.md. It includes six fictional NOVA PDFs, 14 measured questions, the Best Demo 6 speaking notes, screenshots and safe reset/verification scripts:

./scripts/verify_final_demo.sh

prepare_final_demo.sh remains available as an explicit reset/rebuild tool; it is not required for normal startup.

Unavailable historical ingestion records remain visible for diagnosis but are not selectable and never enter retrieval. Re-upload or retry them through the DOCUMENTS tab so the current AUTO path is used.

Evaluation and reproducibility

The immutable membership manifest is:

data/evaluations/week2_final_corpus_snapshot_v1.json

It records the exact 26 point IDs, document/version membership, collection, pipeline fingerprint and snapshot ID. The final evaluation dataset has 44 cases split into development (19), validation (11) and test (14). The latest live test-split smoke recorded Recall@5 0.9583, MRR@10 1.0000 and nDCG@10 0.9834; its evaluation run ID and raw evidence are preserved in projects/document_intelligence_service/eval/results/.

The repository does not contain a live Qdrant dump or the private mentor PDF. Large candidate reports that would repeat verbatim document text are also excluded; the committed raw CSV/JSONL outputs retain IDs, ranks, decisions and metrics without republishing source-document content. To reconstruct the frozen collection, obtain an approved copy of the source input, ingest it with the explicit mentor_program_v1 evaluation profile, and verify the resulting point manifest before running the benchmark. Normal product ingestion must continue to use AUTO.

The V11 prompt-packing ablation selected a bounded 2,400-character production context. It was the smallest measured configuration that retained both the date and time in the real deadline-style selected evidence; expected answers and trusted gold labels never enter prompt construction. See docs/prompt_packing_ablation_v11.md.

Typical benchmark invocation after the approved corpus has been reconstructed:

DIS_QDRANT_URL=http://127.0.0.1:6335 \
DIS_QDRANT_COLLECTION=document_chunks_week2_final_v1 \
DIS_SECTION_MARKER_PROFILE=mentor_program_v1 \
python -m projects.document_intelligence_service.eval.run_benchmark \
  --mode hybrid --top-k 5 --point-count 26 \
  --output projects/document_intelligence_service/eval/results/hybrid_baseline.json \
  --raw-output-dir projects/document_intelligence_service/eval/results

Local development and tests

python3.12 -m venv .venv
.venv/bin/pip install --index-url https://download.pytorch.org/whl/cpu 'torch==2.13.0+cpu'
.venv/bin/pip install -e 'projects/document_intelligence_service[dev]'
.venv/bin/pytest -q projects/document_intelligence_service/tests
.venv/bin/ruff check projects/document_intelligence_service/app projects/document_intelligence_service/eval projects/document_intelligence_service/tests
.venv/bin/mypy projects/document_intelligence_service/app projects/document_intelligence_service/eval projects/document_intelligence_service/tests
docker compose config --quiet

The standalone Compose smoke validates health, the demo UI, optional PDF ingestion, duplicate-ingestion idempotency and Qdrant restart persistence. It does not publish a PDF in this repository; provide a local input explicitly:

SMOKE_PDF=/path/to/a/parseable.pdf BUNDLED_OLLAMA=true \
  ./scripts/compose_smoke.sh

Without SMOKE_PDF, the smoke still checks service startup, health, UI and Qdrant persistence.

Repository structure

app source:  projects/document_intelligence_service/app/
tests:       projects/document_intelligence_service/tests/
evaluation:  projects/document_intelligence_service/eval/
datasets:    data/evaluations/
UI:          demo_ui/
docs/        architecture, compliance, model compatibility and ADR material
scripts/     standalone Compose smoke

The nested service path is intentional: it preserves the tested Python import package and Docker build layout without adding parent-repository dependencies.

Local models and limitations

INSTALLED does not mean READY. The service exposes model readiness from the runtime probe and keeps Gemma as the measured default. A historical local machine observation (2026-08-10) found qwen3:4b installed, but its controlled think=false, stream=false probe still ended with bounded, incomplete output and no reliable final answer; that observation is not a current-machine guarantee. Thinking text is never treated as a user-facing answer.

Known MVP limits include local CPU generation latency, a small generic-document answerability calibration set, process-local model probe state, local ACL-ready filters rather than a full identity provider, process-local metrics rather than a metrics platform, and MVP-scoped retention/purge. These limitations do not change the frozen benchmark labels or corpus membership.

Architecture and decision records

About

Local-first Document Intelligence service with hybrid RAG, evidence tracing, answerability gating and reproducible evaluation.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages