This repository is a clean, standalone copy of the Week-2 Document Intelligence service. It is a local-first PDF intelligence system designed to make a wrong answer diagnosable: the trace separates ingestion, retrieval, fusion, reranking, answerability, prompt construction and generation.
No private development PDFs, model weights, Qdrant databases, credentials or parent-repository history are included.
Requires Docker Desktop or Docker Engine with Compose v2. This single command starts the API, worker, Qdrant, Ollama and Demo UI; no second terminal or host Ollama installation is required:
docker compose -f compose.yaml -f compose.ollama.yaml \
--profile bundled-ollama up --build -dCompose automatically seeds the six fictional NOVA demo PDFs through the
normal POST /v1/documents ingestion API. It waits for those ingestion jobs
to finish before starting the Demo UI. Open http://127.0.0.1:8501 after the
stack reports the API healthy and demo-seed completed.
Packaging note: this command is for a GitHub clone, where compose.yaml is
at the repository root. If you are using the separate
Document_Intelligence_2_Hafta_Teslim_Paketi ZIP, run its root-safe command
from the extracted document-intelligence-delivery directory instead:
docker compose \
-f source/compose.yaml \
-f source/compose.ollama.yaml \
--project-directory source \
--profile bundled-ollama \
up --build -dTo stop the project while preserving downloaded models and indexed data:
docker compose -f compose.yaml -f compose.ollama.yaml \
--profile bundled-ollama down --remove-orphansThe service accepts arbitrary parseable PDFs, indexes them with a deterministic pipeline, and answers only when the retrieved evidence passes the configured answerability policy. Every answer has application-generated canonical source metadata; source cards are never reconstructed from model text.
The core flow is:
PDF
→ parse → normalize → AUTO chunk selection
→ dense + BM25 → RRF
→ optional reranker → canonical evidence
→ answerability → structured prompt → local LLM
→ answer + canonical sources
The ASK flow runs one retrieval strategy at a time: Dense only, BM25 only, or
Hybrid RRF. Hybrid means Dense + BM25 → RRF; BENCHMARKS is the separate place
where multiple strategies and reranker states are compared.
The mentor UI is organized as ASK | DOCUMENTS | BENCHMARKS; the ASK result
page contains the single Stage Explorer and keeps the full engineering trace
under progressive disclosure.
When the Demo UI is opened, its tenant defaults to the isolated
final-demo-v1 corpus containing the six fictional NOVA PDFs. Compose's
demo-seed init service uploads those PDFs through the normal
POST /v1/documents path and waits for successful ingestion before the UI is
released. Compose seeds
that corpus automatically on first startup and skips files that already have
the current active version on later restarts. Set DEMO_SEED_ENABLED=false
only when an empty/manual tenant is intentionally required. The historical
default tenant is not deleted: it remains available when explicitly entered
for benchmark/validation work, but its private or mentor source documents are
not part of the normal six-question demo scope.
The final frozen mentor corpus contains 26 points in snapshot
c5e87f7e063769adef368866854d8e45f7b7f9856f905abe9cebe31783262b25.
The current final retrieval artifact reports:
| Variant | Recall@5 | MRR@10 | nDCG@10 |
|---|---|---|---|
| Dense | 0.9011 | 0.8750 | 0.9296 |
| BM25 | 0.8178 | 0.7844 | 0.8376 |
| Hybrid RRF | 0.9233 | 0.8778 | 0.9518 |
| Hybrid + reranker | 0.9122 | 0.8333 | 0.9329 |
Therefore the demo default is Hybrid RRF with reranker OFF. The reranker remains available for measured ablation; it is not disabled by assumption.
Other validated behaviors:
- no-answer and security-policy decisions skip the LLM;
- duplicate ingestion reuses the same document/version when content and the effective pipeline fingerprint match;
- arbitrary uploads default to
AUTO, resolving to a structure-aware strategy only when reliable structure is present and otherwise to boundedgeneric_v1; - the frozen mentor corpus remains an explicit
mentor_program_v1evaluation membership, not a global admission rule; - canonical evidence retains document, page, parent/child and rank metadata;
- installed models and ready models are reported separately.
Requirements:
- Docker Engine/Desktop with Compose v2;
- Python 3.12 and the development tools for local tests.
No host Ollama installation or second terminal is required. Compose starts an
Ollama container, pulls gemma3:4b on the first run, persists it in a named
Docker volume, and does not start the UI until the model is reachable.
docker compose -f compose.yaml -f compose.ollama.yaml \
--profile bundled-ollama up --build -dThis command works with Docker Engine on Linux and Docker Desktop on macOS or Windows. For a readiness-gated Bash launcher with the same Docker-managed runtime, use:
./scripts/start_demo.sh --bundled-ollamaOpen the Demo UI at http://127.0.0.1:8501 after docker compose ps shows the
API healthy. The launcher additionally waits for API liveness and readiness,
then prints the exact API, health, Qdrant and UI URLs. It exits non-zero and
prints the real dependency status if a required check fails.
Stop the standalone stack without deleting its persisted data:
docker compose -f compose.yaml -f compose.ollama.yaml \
--profile bundled-ollama down --remove-orphansdown keeps the Qdrant and Ollama named volumes. Do not add -v unless you
intentionally want to delete indexed data and the downloaded model.
The base Compose file still supports an already-managed host Ollama runtime. This is useful when the model is already installed and you do not want a second Ollama container, but it requires that the runtime be reachable from Docker.
./scripts/start_demo.sh --host-ollamaThe host mode uses host.docker.internal:11434 and maps API/Qdrant/UI to
8010/6335/8501. Override those host ports or the Ollama URL when needed:
API_HOST_PORT=8011 QDRANT_HOST_PORT=6336 UI_HOST_PORT=8502 \
DIS_OLLAMA_URL=http://host.docker.internal:11434 \
./scripts/start_demo.sh --host-ollamaOn Linux, the most portable option is to run Ollama on a separate local port
if an existing service already occupies 11434:
# Terminal 1
OLLAMA_HOST=0.0.0.0:11435 ollama serve
# Terminal 2
OLLAMA_HOST=http://127.0.0.1:11435 ollama pull gemma3:4b
DIS_OLLAMA_URL=http://host.docker.internal:11435 \
API_HOST_PORT=8011 QDRANT_HOST_PORT=6336 UI_HOST_PORT=8502 \
./scripts/start_demo.sh --host-ollamaDocker Desktop normally provides host.docker.internal on macOS and Windows.
The Compose file adds the same host-gateway mapping on Linux. Keep the Ollama
listener restricted to the local machine/network; this project does not
require public model exposure.
Readiness is intentionally not equivalent to process liveness. If the model, Ollama or Qdrant check is unavailable, the API reports not-ready and the UI does not pretend that generation is available. The API Compose healthcheck is also readiness-based, so dependent UI startup is blocked until the selected model is actually installed and reachable.
- Check live and ready health.
- Upload a parseable PDF in the DOCUMENTS tab. Product uploads use
AUTOand can fall back togeneric_v1; no mentor headings are required. - Select only documents with an active searchable version.
- Run a direct fact, a paraphrase and an exact/numeric query.
- Inspect Dense, BM25, RRF, evidence, answerability and the canonical source in the single Stage Explorer. Prompt Packing exposes the actual bounded fragments sent to generation, including source/page and omitted-window metadata.
- Run the unrelated-question case and verify
NO_ANSWERwith the LLM skipped. - Run the prompt-injection regressions. A direct, high-confidence injection
is blocked with
SECURITY_POLICY; the full Atlas indirect-injection case removes unsafe evidence and returnsNO_ANSWER · INSUFFICIENT_COVERAGEwhen no supported answer remains. - Compare retrieval/reranker variants in BENCHMARKS.
For the complete Turkish multi-document mentor corpus, the Compose seed is
automatic. To explicitly reset and re-ingest only the demo tenant (for
example, after deleting its persistent volumes), use
demo/final_demo_pack/README.md. It includes
six fictional NOVA PDFs, 14 measured questions, the Best Demo 6 speaking
notes, screenshots and safe reset/verification scripts:
./scripts/verify_final_demo.shprepare_final_demo.sh remains available as an explicit reset/rebuild tool;
it is not required for normal startup.
Unavailable historical ingestion records remain visible for diagnosis but are
not selectable and never enter retrieval. Re-upload or retry them through the
DOCUMENTS tab so the current AUTO path is used.
The immutable membership manifest is:
data/evaluations/week2_final_corpus_snapshot_v1.json
It records the exact 26 point IDs, document/version membership, collection,
pipeline fingerprint and snapshot ID. The final evaluation dataset has 44
cases split into development (19), validation (11) and test (14). The latest
live test-split smoke recorded Recall@5 0.9583, MRR@10 1.0000 and nDCG@10
0.9834; its evaluation run ID and raw evidence are preserved in
projects/document_intelligence_service/eval/results/.
The repository does not contain a live Qdrant dump or the private mentor PDF.
Large candidate reports that would repeat verbatim document text are also
excluded; the committed raw CSV/JSONL outputs retain IDs, ranks, decisions and
metrics without republishing source-document content.
To reconstruct the frozen collection, obtain an approved copy of the source
input, ingest it with the explicit mentor_program_v1 evaluation profile, and
verify the resulting point manifest before running the benchmark. Normal
product ingestion must continue to use AUTO.
The V11 prompt-packing ablation selected a bounded 2,400-character production
context. It was the smallest measured configuration that retained both the
date and time in the real deadline-style selected evidence; expected answers
and trusted gold labels never enter prompt construction. See
docs/prompt_packing_ablation_v11.md.
Typical benchmark invocation after the approved corpus has been reconstructed:
DIS_QDRANT_URL=http://127.0.0.1:6335 \
DIS_QDRANT_COLLECTION=document_chunks_week2_final_v1 \
DIS_SECTION_MARKER_PROFILE=mentor_program_v1 \
python -m projects.document_intelligence_service.eval.run_benchmark \
--mode hybrid --top-k 5 --point-count 26 \
--output projects/document_intelligence_service/eval/results/hybrid_baseline.json \
--raw-output-dir projects/document_intelligence_service/eval/resultspython3.12 -m venv .venv
.venv/bin/pip install --index-url https://download.pytorch.org/whl/cpu 'torch==2.13.0+cpu'
.venv/bin/pip install -e 'projects/document_intelligence_service[dev]'
.venv/bin/pytest -q projects/document_intelligence_service/tests
.venv/bin/ruff check projects/document_intelligence_service/app projects/document_intelligence_service/eval projects/document_intelligence_service/tests
.venv/bin/mypy projects/document_intelligence_service/app projects/document_intelligence_service/eval projects/document_intelligence_service/tests
docker compose config --quietThe standalone Compose smoke validates health, the demo UI, optional PDF ingestion, duplicate-ingestion idempotency and Qdrant restart persistence. It does not publish a PDF in this repository; provide a local input explicitly:
SMOKE_PDF=/path/to/a/parseable.pdf BUNDLED_OLLAMA=true \
./scripts/compose_smoke.shWithout SMOKE_PDF, the smoke still checks service startup, health, UI and
Qdrant persistence.
app source: projects/document_intelligence_service/app/
tests: projects/document_intelligence_service/tests/
evaluation: projects/document_intelligence_service/eval/
datasets: data/evaluations/
UI: demo_ui/
docs/ architecture, compliance, model compatibility and ADR material
scripts/ standalone Compose smoke
The nested service path is intentional: it preserves the tested Python import package and Docker build layout without adding parent-repository dependencies.
INSTALLED does not mean READY. The service exposes model readiness from the
runtime probe and keeps Gemma as the measured default. A historical local
machine observation (2026-08-10) found qwen3:4b installed, but its controlled
think=false, stream=false probe still ended with bounded, incomplete output and
no reliable final answer; that observation is not a current-machine guarantee.
Thinking text is never treated as a user-facing answer.
Known MVP limits include local CPU generation latency, a small generic-document answerability calibration set, process-local model probe state, local ACL-ready filters rather than a full identity provider, process-local metrics rather than a metrics platform, and MVP-scoped retention/purge. These limitations do not change the frozen benchmark labels or corpus membership.