Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
19 changes: 19 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,25 @@ jobs:
- run: pnpm test
- run: pnpm build

eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: actions/setup-python@v7
with:
python-version: "3.12"
cache: pip
cache-dependency-path: pyproject.toml
- run: pip install -e ".[dev]"
- name: Run retrieval evaluation
run: python evals/run.py
- name: Publish report in the job summary
run: cat evals/results/retrieval-report.md >> "$GITHUB_STEP_SUMMARY"
- uses: actions/upload-artifact@v7
with:
name: retrieval-report
path: evals/results/

docker:
runs-on: ubuntu-latest
steps:
Expand Down
23 changes: 23 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -214,6 +214,29 @@ A persistent bottom player continues across views and remembers the last track a
</tr>
</table>

## 📊 Measured results

`evals/` holds a labeled evaluation: 30 questions over a generated 17-file corpus, each with the
files that should answer it, plus 7 computable questions checked against the tool layer.

| configuration | hit@1 | hit@3 | hit@5 | recall@5 | MRR | p95 latency |
| --- | --- | --- | --- | --- | --- | --- |
| lexical + metadata (offline default) | 0.733 | 0.867 | 0.900 | 0.871 | 0.806 | 1.9 ms |
| + vector search (ChromaDB) | 0.767 | 0.867 | 0.900 | 0.879 | 0.822 | 851 ms |

Every expected file was retrieved for all 30 questions. Topical questions — key, status, tags,
project scope, note content — put the right file first essentially every time.

Superlative questions do not, and are not meant to. *"Which track is the loudest"* asks for a
comparison, not a similar file: every audio file matches the words about equally well, so ranking
cannot pick a winner. Those go through the tool layer instead, which answered **7 of 7** correctly
by comparing real numbers.

The trade-off is the honest one: vector search buys 3 points of hit@1 and costs roughly two orders
of magnitude in latency. That is why it is optional.

Methodology, metric definitions and known limits: [`evals/README.md`](evals/README.md).

## 🏗️ Architecture

```mermaid
Expand Down
128 changes: 128 additions & 0 deletions evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,128 @@
# Retrieval evaluation

SessionIQ's claim is that answers are grounded in the files they cite. That claim is only worth
something if retrieval actually finds the right files, so this directory measures whether it does —
and measures it against known answers rather than against a feeling.

```
python evals/run.py # offline: lexical + metadata
python evals/run.py --vector # adds the ChromaDB vector layer
```

Both runs write a Markdown and JSON report to `evals/results/`. The JSON is the machine-readable
form; the Markdown is what the numbers in the main README come from.

## Results

Measured 2026-09-17 over 30 labeled questions on a 17-file corpus.

| configuration | hit@1 | hit@3 | hit@5 | recall@5 | MRR | p50 | p95 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| lexical + metadata | 0.733 | 0.867 | 0.900 | 0.871 | 0.806 | 1.0 ms | 1.9 ms |
| + vector (ChromaDB) | 0.767 | 0.867 | 0.900 | 0.879 | 0.822 | 322 ms | 851 ms |

Every expected file was retrieved somewhere in all 30 questions. What separates the two
configurations is how often the right file comes *first*, and what that costs in latency.

### Where retrieval is strong, and where it isn't

| question type | n | hit@1 | hit@5 | MRR |
| --- | --- | --- | --- | --- |
| key | 3 | 1.000 | 1.000 | 1.000 |
| status | 3 | 1.000 | 1.000 | 1.000 |
| tag | 4 | 1.000 | 1.000 | 1.000 |
| format | 2 | 1.000 | 1.000 | 1.000 |
| project-scoped | 2 | 1.000 | 1.000 | 1.000 |
| content | 6 | 0.833 | 1.000 | 0.917 |
| synonym-expanded | 3 | 0.333 | 1.000 | 0.556 |
| superlative | 6 | 0.167 | 0.500 | 0.336 |

Topical questions — which files are in this key, which are ready, which carry a tag, what a note
says — are answered essentially perfectly. Superlative questions are not, and that is by design
rather than by accident: *"which track is the loudest"* is a comparison, not a similarity search.
Every audio file contains the words *loud*, *track* and *bpm* about equally often, so ranking by
text overlap has no basis on which to pick a winner. It surfaces the right neighbourhood and stops
there.

That is what the tool layer is for, so this harness measures it separately.

## Computable questions (tool layer)

`evals/tool_set.py` lists seven superlative questions and the tool call that should answer each.
The harness executes the call and checks the asset it names.

**7 / 7 returned the correct asset.** Each one is an exact comparison over real numbers — the
loudest file, the slowest tempo, the brightest spectral centroid — rather than a guess about which
file sounds relevant.

Reading the two tables together is the point: retrieval finds candidates, tools compute answers.
Neither replaces the other, and a single end-to-end "accuracy" number would hide that.

## How the corpus works

The evaluation runs against the library that `scripts/seed_demo.py` generates, not against a real
one. That generator synthesizes every file in code, so the corpus is identical on every machine and
in CI, and no licensing question arises from committing it (nothing is committed — it is rebuilt
each run).

Ground truth is labeled against what the analyzers actually report, not against what the
synthesizer was asked for. These differ: `cue_draft.wav` is synthesized on D and detected in A.
Labels follow the analyzers, because the analyzers are what retrieval indexes.

The corpus lives in a temporary directory. `evals/run.py` refuses to start unless the resolved
upload root sits inside the directory it created, so it cannot purge a real library.

## What is measured

**Ranking quality** over the rank-ordered result list, at k = 1, 3, 5:

- **hit@k** — did any expected file appear in the top k
- **recall@k** — what share of the expected files appeared in the top k
- **precision@k** — what share of the top k slots held an expected file, divided by k rather than
by the number of results returned, so configurations that return different counts stay comparable
- **MRR** — 1 / rank of the first expected file

Results the retriever scored at zero are not counted as retrieved; the retriever returns
recently-analyzed files as a fallback when nothing matches, and those are not candidate answers.

**Latency** as p50, p95 and max over 150 calls.

**Blend weights** by holding the component scores fixed and varying only the weights, so each row
isolates what the weighting contributes. Before any ablation number is trusted, the harness re-ranks
the library under the production weights and asserts the order matches what `search()` really
returns. It passes on all 30 questions; if it ever stops passing, the harness has drifted from the
code it describes and the evaluation exits non-zero.

## Limits

- **One corpus, one domain.** Thirty questions over seventeen synthesized audio, MIDI and note
files. It is large enough to separate retrieval configurations and far too small to be a
benchmark.
- **Synthetic audio.** The loops are tonal and rhythmic enough for librosa, but they are not real
music, and real libraries have messier filenames and metadata.
- **Generated answers are not scored.** This measures retrieval and the tool layer. It does not
score the wording of an answer, and it does not measure hallucination. With no model configured
the offline engine answers deterministically, and scoring it would measure that engine rather
than a model.
- **Vector latency is machine-dependent.** The 322 ms p50 above includes embedding the query with
ChromaDB's default local model; repeated runs on one machine varied between roughly 320 ms and
851 ms p95 depending on load. Read it as "hundreds of milliseconds", not as a precise figure.
- **The vector configuration is not perfectly repeatable.** Across runs its MRR moved between 0.821
and 0.822 and its recall@5 between 0.879 and 0.887, while hit@1, hit@3 and hit@5 held steady.
Differences that small are within noise at 30 questions — treat the vector rows as approximate.
- **The offline configuration is stable but not bit-identical across platforms.** Hit rates
reproduced exactly on every run and on both Windows and the Linux CI runner; MRR came out 0.806
on Windows against 0.805 on Linux, which is tie ordering rather than a behavioural difference.
- **Labels are hand-written.** They were read off a real ingestion run, and
`tests/test_evals.py` checks that every label still names a file the corpus creates, but a
mistaken label would quietly depress the score rather than announce itself.

## Reproducing

```
python evals/run.py --keep # leave the generated corpus on disk for inspection
python -m pytest tests/test_evals.py
```

Regenerate the corpus and re-check the labels after changing `seed_demo.py`; the detected tempo,
key and centroid values the labels depend on are recorded at the top of `evals/golden_set.py`.
1 change: 1 addition & 0 deletions evals/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
"""Retrieval evaluation for SessionIQ."""
191 changes: 191 additions & 0 deletions evals/golden_set.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,191 @@
"""Labeled questions over the generated demo library.

Ground truth is keyed on ``<project>/<file name>`` because the demo corpus
contains several files that share a name (four ``cover.png`` artwork files).

Every expected value here was read off a real ingestion run of
``scripts/seed_demo.py`` — the detected tempo, key, loudness and centroid in the
labels are what the analyzers actually produce, not the values the synthesizer
was asked for. Those two differ: ``cue_draft.wav`` is synthesized on D and
detected in A, for example. Labels follow the analyzers, because the analyzers
are what retrieval indexes.

Re-generate the corpus after changing ``seed_demo.py`` and re-check the labels;
``tests/test_evals.py`` asserts that every label still names a real file.
"""

from __future__ import annotations

from dataclasses import dataclass

# Every asset scripts/seed_demo.py produces, project-qualified.
DEMO_CORPUS: tuple[str, ...] = (
"Neon Horizon/Intro/intro.wav",
"Neon Horizon/Intro/cover.png",
"Neon Horizon/Midnight Drive/midnight_drive.wav",
"Neon Horizon/Midnight Drive/mix_notes.txt",
"Neon Horizon/Midnight Drive/cover.png",
"Neon Horizon/Afterglow/afterglow.wav",
"Neon Horizon/Afterglow/afterglow_master.wav",
"Neon Horizon/Afterglow/afterglow_reference.wav",
"Neon Horizon/Afterglow/notes.txt",
"Neon Horizon/Afterglow/cover.png",
"Lo-Fi Study/lofi_sketch.wav",
"Lo-Fi Study/chords.mid",
"Lo-Fi Study/ideas.txt",
"Lo-Fi Study/cover.png",
"Client Cue 03/cue_draft.wav",
"Client Cue 03/reference.wav",
"Client Cue 03/brief.txt",
)

# Detected values, for readers of this file.
# intro.wav 89.1 bpm key G centroid 837.8 -18.8 LUFS
# midnight_drive.wav 117.5 bpm key E centroid 1218.1 -18.3 LUFS
# afterglow.wav 117.5 bpm key E centroid 1202.1 -18.4 LUFS
# afterglow_master.wav 117.5 bpm key E centroid 1208.0 -18.4 LUFS
# afterglow_reference.wav 117.5 bpm key E centroid 1190.4 -25.0 LUFS (quiet bounce)
# lofi_sketch.wav 76.0 bpm key C centroid 813.6 -18.9 LUFS
# cue_draft.wav 99.4 bpm key A centroid 871.0 -19.9 LUFS
# reference.wav 99.4 bpm key A centroid 858.2 -15.5 LUFS (loudest)

FASTEST_TRACKS = (
"Neon Horizon/Midnight Drive/midnight_drive.wav",
"Neon Horizon/Afterglow/afterglow.wav",
"Neon Horizon/Afterglow/afterglow_master.wav",
"Neon Horizon/Afterglow/afterglow_reference.wav",
)

KEY_OF_E = FASTEST_TRACKS

ARTWORK = (
"Neon Horizon/Intro/cover.png",
"Neon Horizon/Midnight Drive/cover.png",
"Neon Horizon/Afterglow/cover.png",
"Lo-Fi Study/cover.png",
)

# Notes whose text literally contains a TODO marker.
TODO_NOTES = (
"Neon Horizon/Midnight Drive/mix_notes.txt",
"Lo-Fi Study/ideas.txt",
"Client Cue 03/brief.txt",
)


@dataclass(frozen=True)
class GoldenCase:
question: str
expected: tuple[str, ...]
category: str
project: str | None = None


GOLDEN_SET: tuple[GoldenCase, ...] = (
# --- acoustic superlatives ------------------------------------------------
GoldenCase("which track is the loudest?", ("Client Cue 03/reference.wav",), "superlative"),
GoldenCase(
"which track is the quietest?",
("Neon Horizon/Afterglow/afterglow_reference.wav",),
"superlative",
),
GoldenCase("which track is the slowest?", ("Lo-Fi Study/lofi_sketch.wav",), "superlative"),
GoldenCase("which tracks are the fastest?", FASTEST_TRACKS, "superlative"),
GoldenCase(
"which track is the brightest?",
("Neon Horizon/Midnight Drive/midnight_drive.wav",),
"superlative",
),
GoldenCase("which track is the darkest?", ("Lo-Fi Study/lofi_sketch.wav",), "superlative"),
# Same intent, natural wording that has to reach the metadata vocabulary
# through the synonym map rather than by literal token overlap.
GoldenCase("which song is the loudest?", ("Client Cue 03/reference.wav",), "synonym"),
GoldenCase("show me the fastest tune", FASTEST_TRACKS, "synonym"),
GoldenCase("which track has the highest volume?", ("Client Cue 03/reference.wav",), "synonym"),
# --- key ------------------------------------------------------------------
GoldenCase("which track is in the key of G?", ("Neon Horizon/Intro/intro.wav",), "key"),
GoldenCase(
"which tracks are in the key of C?",
("Lo-Fi Study/lofi_sketch.wav", "Lo-Fi Study/chords.mid"),
"key",
),
GoldenCase("which audio files are in the key of E?", KEY_OF_E, "key"),
# --- note content ---------------------------------------------------------
GoldenCase("which notes mention a dusty piano loop?", ("Lo-Fi Study/ideas.txt",), "content"),
GoldenCase("which notes still have todo items?", TODO_NOTES, "content"),
GoldenCase(
"which notes mention sidechaining the pads?",
("Neon Horizon/Midnight Drive/mix_notes.txt",),
"content",
),
GoldenCase("what did the client ask for?", ("Client Cue 03/brief.txt",), "content"),
GoldenCase("which notes mention tape saturation?", ("Lo-Fi Study/ideas.txt",), "content"),
GoldenCase(
"where is the note about checking the low end on headphones?",
("Neon Horizon/Afterglow/notes.txt",),
"content",
),
# --- status ---------------------------------------------------------------
GoldenCase("what still needs work?", ("Client Cue 03/cue_draft.wav",), "status"),
GoldenCase(
"which files are ready?",
("Neon Horizon/Afterglow/afterglow_master.wav",),
"status",
),
GoldenCase(
"what is in progress?",
("Neon Horizon/Midnight Drive/midnight_drive.wav",),
"status",
),
# --- tags -----------------------------------------------------------------
GoldenCase(
"which track is tagged synthwave?",
("Neon Horizon/Midnight Drive/midnight_drive.wav",),
"tag",
),
GoldenCase("which track is my favorite?", ("Neon Horizon/Afterglow/afterglow.wav",), "tag"),
GoldenCase("which track is chill?", ("Lo-Fi Study/lofi_sketch.wav",), "tag"),
GoldenCase(
"which track is marked as a single?",
("Neon Horizon/Midnight Drive/midnight_drive.wav",),
"tag",
),
# --- format ---------------------------------------------------------------
GoldenCase("which project has a midi file?", ("Lo-Fi Study/chords.mid",), "format"),
GoldenCase("where is the artwork?", ARTWORK, "format"),
# --- project-scoped -------------------------------------------------------
GoldenCase(
"what files are in this project?",
(
"Lo-Fi Study/lofi_sketch.wav",
"Lo-Fi Study/chords.mid",
"Lo-Fi Study/ideas.txt",
"Lo-Fi Study/cover.png",
),
"project",
project="Lo-Fi Study",
),
GoldenCase(
"which file is the reference?",
("Client Cue 03/reference.wav",),
"project",
project="Client Cue 03",
),
# --- aggregate ------------------------------------------------------------
GoldenCase(
"list all the audio tracks",
(
"Neon Horizon/Intro/intro.wav",
"Neon Horizon/Midnight Drive/midnight_drive.wav",
"Neon Horizon/Afterglow/afterglow.wav",
"Neon Horizon/Afterglow/afterglow_master.wav",
"Neon Horizon/Afterglow/afterglow_reference.wav",
"Lo-Fi Study/lofi_sketch.wav",
"Client Cue 03/cue_draft.wav",
"Client Cue 03/reference.wav",
),
"aggregate",
),
)

CATEGORIES: tuple[str, ...] = tuple(sorted({case.category for case in GOLDEN_SET}))
Loading