Skip to content

Add retrieval and tool-layer evaluation with measured results - #6

Merged
xfznprojects merged 5 commits into
mainfrom
retrieval-evaluation
Sep 17, 2026
Merged

xfznprojects merged 5 commits into
mainfrom
retrieval-evaluation

Conversation

@xfznprojects

Copy link
Copy Markdown
Owner

The README's central claim is that answers are grounded in the files they cite. Nothing measured whether retrieval actually finds the right files. This adds that measurement.

What it does

evals/ scores two things against known answers, on the library scripts/seed_demo.py generates:

  • Retrieval — 30 hand-labeled questions, each carrying the files that should answer it
  • The tool layer — 7 computable questions, each with the tool call that should answer it

The results

configuration hit@1 hit@3 hit@5 recall@5 MRR p95
lexical + metadata (offline default) 0.733 0.867 0.900 0.871 0.806 1.9 ms
+ vector (ChromaDB) 0.767 0.867 0.900 0.879 0.822 851 ms

Every expected file was retrieved for all 30 questions — the difference is how often the right one ranks first.

Topical questions retrieve almost perfectly. Key, status, tag, format and project-scoped questions all score 1.000 hit@1. Content questions score 0.833.

Superlative questions do not, by design. "Which track is the loudest" is a comparison rather than a similarity search: every audio file contains the words loud, track and bpm about equally often, so text-overlap ranking has no basis to pick a winner (hit@1 0.167). Those questions are routed through the tool layer instead, which answered 7/7 correctly by comparing real numbers.

The vector trade-off is now quantified. It buys 3 points of hit@1 and costs roughly two orders of magnitude of latency. That is the honest reason it stays optional.

Method notes

  • Labels follow what the analyzers detect, not what the synthesizer was asked for. These differ — cue_draft.wav is synthesized on D and detected in A.
  • The corpus is built in a throwaway directory, and the runner refuses to start unless the resolved upload root sits inside it, so it cannot purge a real library.
  • The weight ablation holds component scores fixed and varies only the blend. Before any ablation number is trusted, it re-ranks under the production weights and asserts the order matches search() — passing on all 30 questions. If that ever fails the evaluation exits non-zero rather than reporting numbers from a drifted harness.
  • Results the retriever scored at zero are not counted as retrieved; the recent-file fallback is not a candidate answer.

Limits, stated in evals/README.md

One synthetic corpus, 30 questions, hand-written labels, no generated-answer scoring, and the vector configuration is not perfectly repeatable across runs (MRR moved 0.821–0.822, recall@5 0.879–0.887). The offline configuration reproduced identically every run.

Verification

  • 167 tests pass (143 existing + 24 new), ruff clean
  • Both reports are committed under evals/results/ so the numbers can be read without running anything
  • New CI job runs the evaluation and publishes the report to the job summary

evals/ scores retrieval against 30 hand-labeled questions over the library
scripts/seed_demo.py generates, and scores the tool layer separately against
7 computable questions.

The corpus is built in a throwaway directory and the runner refuses to start
unless the resolved upload root is inside it, so it cannot touch a real
library. Labels follow what the analyzers actually detect rather than what
the synthesizer was asked to produce; those differ.

The weight ablation holds component scores fixed and varies only the blend,
and asserts first that re-ranking under the production weights reproduces
search() exactly, so the ablation cannot drift from the code it describes.
Covers the metric arithmetic, the aggregate-limit rule, and the integrity of
both labeled sets: every expected file must exist in the demo corpus, and every
tool case must name a field the toolbox actually exposes.

The corpus itself is not built here — that needs audio analysis, so the full
evaluation stays a separate command.
The README claimed trustworthy citations with nothing measured behind it. It
now reports what retrieval and the tool layer actually score, on which
questions, at what latency, and what the evaluation does not cover.

Both configurations are checked in so the numbers can be read without running
anything: lexical plus metadata, and the same with the vector layer enabled.
The report lands in the job summary so a regression is visible without
downloading the artifact.
The evaluation now also runs on the Linux CI runner, where hit rates match
Windows exactly but MRR differs by 0.001 through tie ordering.
@xfznprojects
xfznprojects merged commit a775a65 into main Sep 17, 2026
5 checks passed
@xfznprojects
xfznprojects deleted the retrieval-evaluation branch September 17, 2026 17:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant