From 72bd2bc1177fa7ba6129e0a46663f719f1a93725 Mon Sep 17 00:00:00 2001 From: Julian Payne Date: Tue, 29 Sep 2026 15:45:22 +0200 Subject: [PATCH] docs(guides): add Results & Scoring guide Document how EvalHub calculates evaluation test results so users can understand and reproduce the numbers on the job resource. The existing collections guide covers the config side but not the results shape itself. This adds a dedicated guide explaining: - the two test fields: per-benchmark results.benchmarks[].test and the job-level results.test - how each benchmark's primary score, threshold, and pass/fail are derived, including lower_is_better handling - the weighted-mean job aggregation and its skip conditions - the override precedence for primary metric, threshold, and weight - a dedicated section on non-numeric metrics: only numeric values are scored, and how non-numeric primary metrics cascade to an omitted test block and are excluded from the job average Generated with: Claude Code Co-Authored-By: Claude Opus 4.8 Signed-off-by: Julian Payne --- astro.config.mjs | 1 + .../docs/guides/results-and-scoring.mdx | 323 ++++++++++++++++++ 2 files changed, 324 insertions(+) create mode 100644 src/content/docs/guides/results-and-scoring.mdx diff --git a/astro.config.mjs b/astro.config.mjs index 8e8ada3..6967641 100644 --- a/astro.config.mjs +++ b/astro.config.mjs @@ -104,6 +104,7 @@ export default defineConfig({ { label: 'MLflow', slug: 'guides/mlflow' }, { label: 'Bring Your Own Framework', slug: 'guides/bring-your-own-framework' }, { label: 'Job Lifecycle & States', slug: 'guides/job-lifecycle' }, + { label: 'Results & Scoring', slug: 'guides/results-and-scoring' }, { label: 'Collections', slug: 'guides/collections' }, { label: 'Hardware Profiles', slug: 'guides/hardware-profiles' }, { label: 'OpenTelemetry', slug: 'guides/opentelemetry' }, diff --git a/src/content/docs/guides/results-and-scoring.mdx b/src/content/docs/guides/results-and-scoring.mdx new file mode 100644 index 0000000..e68a988 --- /dev/null +++ b/src/content/docs/guides/results-and-scoring.mdx @@ -0,0 +1,323 @@ +--- +title: "Results & Scoring" +description: "How EvalHub calculates the results.test and results.benchmarks[].test values, and the user overrides that affect them." +--- + +When an evaluation job completes, EvalHub populates a `results` section on the job +resource. This page explains **exactly how the `test` values are calculated** — +both the per-benchmark pass/fail (`results.benchmarks[].test`) and the job-level +aggregate (`results.test`) — and which configuration and request overrides change +the outcome. + +:::note[Where the raw numbers come from] +EvalHub does **not** compute benchmark metrics itself. The metric values (for +example `acc`, `acc_norm`, `exact_match`, `attack_success_rate`) are produced by +the external evaluation adapter/provider harness that runs the benchmark, and are +reported back to EvalHub in the benchmark status event's `metrics` map. EvalHub's +job is to **select** a primary metric, **compare** it against a threshold, and +**aggregate** across benchmarks into a single job score. +::: + +## The two `test` fields + +The `results` section carries two distinct `test` concepts: + +| Field | Type | Meaning | +|-------|------|---------| +| `results.benchmarks[].test` | per-benchmark | Pass/fail for a **single** benchmark, based on its chosen primary metric vs. its threshold. | +| `results.test` | job-level | Overall pass/fail for the **whole job**, a weighted average of every benchmark's primary score vs. the job threshold. | + +## Shape of the `results` section + +```json +{ + "results": { + "test": { + "score": 0.72, + "threshold": 0.45, + "pass": true + }, + "benchmarks": [ + { + "id": "arc_easy", + "provider_id": "lm_evaluation_harness", + "benchmark_index": 0, + "metrics": { "acc": 0.81, "acc_norm": 0.78 }, + "test": { + "primary_score": 0.78, + "primary_score_metric": "acc_norm", + "threshold": 0.25, + "pass": true + } + }, + { + "id": "toxigen", + "provider_id": "lm_evaluation_harness", + "benchmark_index": 1, + "metrics": { "acc": 0.60 }, + "test": { + "primary_score": 0.60, + "primary_score_metric": "acc", + "threshold": 0.70, + "pass": false + } + } + ] + } +} +``` + +- `metrics` — the full map of metric values the adapter reported for that benchmark (passthrough). +- `benchmarks[].test` — the derived pass/fail for that benchmark (see below). +- `test` — the derived job-level result (see below). + +## How each benchmark's `test` is calculated + +`results.benchmarks[].test` is recomputed on every benchmark status event and stored +once the benchmark reaches a terminal state. The steps are: + +1. **Resolve the effective benchmark config.** If the job references a collection, + the benchmark configuration is taken from the collection (with any per-benchmark + job overrides merged in); otherwise it comes from the job's inline `benchmarks`. +2. **Choose the primary metric.** Use `primary_score.metric` from the + job/collection benchmark config. If that is unset or empty, fall back to the + **provider's** default `primary_score.metric` for that benchmark. +3. **Look up the value.** Read `metrics[primary_metric]` from the values the adapter + reported. The value **must be numeric** and is cast to a float + (see [Non-numeric metrics](#non-numeric-metrics) below). + - If the primary metric is **not present** in the reported metrics, **no `test` + block is produced** for that benchmark (the field is omitted). + - If the value is present but **not numeric**, the same thing happens — the `test` + block is omitted. +4. **Choose the threshold.** Use `pass_criteria.threshold` from the job/collection + benchmark config. If unset, fall back to the **provider's** default + `pass_criteria.threshold` for that benchmark. + - If neither defines a threshold, **no `test` block is produced**. +5. **Decide pass/fail.** + + ```text + pass = primary_score >= threshold # default + pass = primary_score <= threshold (if lower_is_better) + ``` + +The resulting object is: + +```json +{ + "primary_score": 0.78, // the reported value of the chosen metric + "primary_score_metric": "acc_norm", + "threshold": 0.25, + "pass": true +} +``` + +:::note +`primary_score` is simply the reported value of the **one** chosen metric — EvalHub +does not average multiple metrics together at the benchmark level. Other metrics +remain visible under `benchmarks[].metrics`. +::: + +## Non-numeric metrics + +Scoring is purely arithmetic: the primary score must be a **number**, and the +job-level aggregate is a weighted mean of numbers. EvalHub has **no notion of +combining categorical, string, or boolean outcomes** into a `test` result. + +### What counts as numeric + +When EvalHub reads the primary metric's value, it casts it to a float and accepts +only these types: + +| Accepted | Rejected | +|----------|----------| +| `float64`, `float32` | strings (**including numeric-looking strings** like `"0.78"`) | +| `int`, `int32`, `int64` | booleans (`true` / `false`) | +| | arrays and objects | +| | `null` / absent | + +:::caution +A numeric **string** such as `"0.78"` is **not** accepted. The value is cast by +type, never parsed from text, so `"0.78"` is treated the same as any other +non-numeric value. The adapter must report the metric as a JSON number +(`0.78`), not a quoted string (`"0.78"`). +::: + +JSON numbers arrive as `float64`, so any adapter that reports a plain numeric +metric works without special handling. + +### What happens when the primary metric is non-numeric + +The effect cascades from the single benchmark up to the whole job: + +1. **Benchmark level.** If the chosen `primary_score.metric` value is non-numeric, + the benchmark's `test` computation fails the cast and produces **no `test` + block**. The failure is logged server-side, but the API response simply omits + `benchmarks[].test` for that benchmark. The raw value is still preserved under + `benchmarks[].metrics`. +2. **Job level.** Any benchmark with no `test` block is **skipped** in the weighted + average — it contributes to neither the weighted sum nor the total weight, so it + does not drag the score up or down; it is simply absent from the calculation. +3. **Whole job.** If *no* benchmark yields a numeric primary score (so no weights + accumulate), the job score cannot be computed and `results.test` is **omitted + entirely**. + +:::note +This is a silent drop from the caller's perspective: a benchmark configured with a +non-numeric primary metric will quietly disappear from both its own pass/fail and +the job aggregate. If you expected a `test` block and don't see one, check that the +adapter reports the chosen metric as a JSON number. +::: + +### Working with non-numeric outputs + +- **Report numeric metrics for scoring.** If an adapter naturally emits a + categorical or textual result, convert it to a numeric metric (for example a + `0`/`1` pass indicator, a rate, or a normalized score) and point + `primary_score.metric` at that numeric field. +- **Keep the rich output as passthrough.** Non-numeric values (labels, category + breakdowns, free-text notes) can still be reported and are retained under + `benchmarks[].metrics` — they are just ignored by scoring, not discarded. + +## How the job-level `test` is calculated + +`results.test` is computed **once**, when the overall job reaches the `completed` +state. It is a **weighted arithmetic mean** of every benchmark's primary score: + +```text + Σ ( weightᵢ × scoreᵢ ) +job.score = ───────────────────────────── + Σ weightᵢ + +pass = job.score >= job.threshold +``` + +Where, for each benchmark `i`: + +- **`scoreᵢ`** is that benchmark's `test.primary_score`. + - If the benchmark's metric is `lower_is_better`, the contribution is inverted to + `1 - primary_score` so that a higher aggregate always means "better". (Because + of this inversion, the job-level comparison is always `>=`.) +- **`weightᵢ`** is the benchmark's configured `weight`. A missing or `0` weight is + treated as `1`. + +Benchmarks that have **no `test` block** (for example, a missing primary metric or +threshold) are **skipped** — they contribute to neither the numerator nor the +denominator. If no weights accumulate at all (`Σ weightᵢ == 0`), the job score is +not computed and `results.test` is omitted. + +### Worked example + +Two benchmarks, one higher-is-better and one lower-is-better: + +| Benchmark | primary_score | lower_is_better | weight | contribution | +|-----------|---------------|-----------------|--------|--------------| +| `arc_easy` | 0.78 | no | 2 | `2 × 0.78 = 1.56` | +| `toxicity` | 0.10 | yes | 1 | `1 × (1 − 0.10) = 0.90` | + +```text +job.score = (1.56 + 0.90) / (2 + 1) = 2.46 / 3 = 0.82 +``` + +With a job threshold of `0.45`, `0.82 >= 0.45` → `pass: true`. + +## User overrides that affect the `test` results + +The scoring inputs — **primary metric**, **threshold**, **weight**, and +**`lower_is_better`** — can each be set in several places. When more than one is +present, the most specific wins. + +### Primary score metric + +Selects which metric becomes `primary_score` for a benchmark. + +| Priority | Source | +|----------|--------| +| 1 (highest) | `primary_score.metric` on the **job/collection benchmark** config | +| 2 (fallback) | The **provider's** default `primary_score.metric` for that benchmark | + +If neither defines a metric, the benchmark gets **no `test` block**. + +### Per-benchmark threshold + +Determines whether a single benchmark passes. + +| Priority | Source | +|----------|--------| +| 1 (highest) | `pass_criteria.threshold` on the **job/collection benchmark** config | +| 2 (fallback) | The **provider's** default `pass_criteria.threshold` for that benchmark | + +If neither defines a threshold, the benchmark gets **no `test` block** (and is +therefore excluded from the job-level average). + +### Benchmark weight + +Controls each benchmark's influence on the job-level average. + +| Priority | Source | +|----------|--------| +| 1 (highest) | `weight` on the **job/collection benchmark** config | +| 2 (fallback) | Default **`1`** (a `0` weight is also treated as `1`) | + +### Job-level threshold + +Determines whether the whole job passes. + +| Priority | Source | Example use case | +|----------|--------|------------------| +| 1 (highest) | `pass_criteria.threshold` on the **job request** | "For just this run, use a stricter bar of 0.9" | +| 2 | `pass_criteria.threshold` on the **collection definition** | The collection's default bar | +| 3 (fallback) | Hard-coded default **`0.5`** | Neither job nor collection defines one | + +### `lower_is_better` + +Set on a benchmark's `primary_score`. It changes two things: + +- **Per benchmark:** the pass comparison flips to `primary_score <= threshold`. +- **Job level:** the benchmark's contribution is inverted to `1 - primary_score` + before it is weighted and averaged. + +Use it for metrics where a smaller number is better (for example +`attack_success_rate` or a toxicity rate). + +### Example: overriding scoring on a job request + +```json +POST /api/v1/evaluations/jobs + +{ + "name": "stricter-safety-run", + "model": { "...": "..." }, + "pass_criteria": { "threshold": 0.9 }, + "collection": { + "id": "safety-and-fairness-v1", + "benchmarks": [ + { + "id": "toxigen", + "provider_id": "lm_evaluation_harness", + "weight": 3, + "primary_score": { "metric": "acc", "lower_is_better": false }, + "pass_criteria": { "threshold": 0.7 } + } + ] + } +} +``` + +Here the job raises the overall bar to `0.9`, and re-weights/re-thresholds the +`toxigen` benchmark just for this run, without changing the shared collection. + +## When the `test` section is missing + +The `test` block is deliberately omitted (rather than showing a misleading `0`) when: + +- the chosen **primary metric is not present** in the adapter's reported metrics, or +- the primary metric value is **non-numeric** (see [Non-numeric metrics](#non-numeric-metrics)), or +- **no threshold** can be resolved (neither the benchmark config nor the provider + defines one), or +- (job level) **no benchmarks contributed** a valid weight/score. + +## Related + +- [Collections](/guides/collections/) — where weights, primary scores, and thresholds are configured, and the list of built-in collections and their thresholds. +- [Evaluation Event Violations](/kubernetes/event-violations/) — how a failing benchmark `test` triggers a threshold-violation notification. +- [Server API](/reference/server-api/) — the full evaluation job and results schema.