Add retrieval and tool-layer evaluation with measured results - #6
Merged
Merged
Conversation
evals/ scores retrieval against 30 hand-labeled questions over the library scripts/seed_demo.py generates, and scores the tool layer separately against 7 computable questions. The corpus is built in a throwaway directory and the runner refuses to start unless the resolved upload root is inside it, so it cannot touch a real library. Labels follow what the analyzers actually detect rather than what the synthesizer was asked to produce; those differ. The weight ablation holds component scores fixed and varies only the blend, and asserts first that re-ranking under the production weights reproduces search() exactly, so the ablation cannot drift from the code it describes.
Covers the metric arithmetic, the aggregate-limit rule, and the integrity of both labeled sets: every expected file must exist in the demo corpus, and every tool case must name a field the toolbox actually exposes. The corpus itself is not built here — that needs audio analysis, so the full evaluation stays a separate command.
The README claimed trustworthy citations with nothing measured behind it. It now reports what retrieval and the tool layer actually score, on which questions, at what latency, and what the evaluation does not cover. Both configurations are checked in so the numbers can be read without running anything: lexical plus metadata, and the same with the vector layer enabled.
The report lands in the job summary so a regression is visible without downloading the artifact.
The evaluation now also runs on the Linux CI runner, where hit rates match Windows exactly but MRR differs by 0.001 through tie ordering.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The README's central claim is that answers are grounded in the files they cite. Nothing measured whether retrieval actually finds the right files. This adds that measurement.
What it does
evals/scores two things against known answers, on the libraryscripts/seed_demo.pygenerates:The results
Every expected file was retrieved for all 30 questions — the difference is how often the right one ranks first.
Topical questions retrieve almost perfectly. Key, status, tag, format and project-scoped questions all score 1.000 hit@1. Content questions score 0.833.
Superlative questions do not, by design. "Which track is the loudest" is a comparison rather than a similarity search: every audio file contains the words loud, track and bpm about equally often, so text-overlap ranking has no basis to pick a winner (hit@1 0.167). Those questions are routed through the tool layer instead, which answered 7/7 correctly by comparing real numbers.
The vector trade-off is now quantified. It buys 3 points of hit@1 and costs roughly two orders of magnitude of latency. That is the honest reason it stays optional.
Method notes
cue_draft.wavis synthesized on D and detected in A.search()— passing on all 30 questions. If that ever fails the evaluation exits non-zero rather than reporting numbers from a drifted harness.Limits, stated in
evals/README.mdOne synthetic corpus, 30 questions, hand-written labels, no generated-answer scoring, and the vector configuration is not perfectly repeatable across runs (MRR moved 0.821–0.822, recall@5 0.879–0.887). The offline configuration reproduced identically every run.
Verification
ruffcleanevals/results/so the numbers can be read without running anything