You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
feat(criteria): add system_one_judge, a typed-rubric grader on a System One model (#192)
* feat(criteria): add system_one_judge, a typed-rubric grader on a System One model
A System One model (TypeSafe's jev) generates no text: it reads one state and
answers a map of typed questions with calibrated probabilities in a single round
trip. That removes the two things a text judge has to defend against — an
unparseable verdict and a tool call the model may decline to make — so the rubric
the author writes IS the grading schema.
The grade stays on our side: each question resolves to a value in [0,1] via its
own expected/values, and the criterion score is their weighted mean. That buys
determinism, a grade that can be recomputed from an archived transcript, and
findings that print the arithmetic per question instead of arguing it in prose.
A choice question's options may be written as a bare list when the names speak
for themselves; it widens to the option->description map the API wants, with null
descriptions, which the live endpoint accepts. Order is preserved and a repeated
option is a load-time error.
The judge reads the files it is given plus, opt-in per criterion, the agent's own
final message (include_agent_output), its tool-call trajectory
(include_tool_calls) and the simulated dialog (include_dialog) — so a rubric can
grade how the agent worked and whether its summary was honest, not only the
artifact left behind.
tasks/smoke_system_one_judge.yaml exercises all three primitives over all three
of those state slices, and runs in CI's smoke-pass bucket beside smoke_llm_judge.
tests/test_system_one_judge_live.py pins the external wire shape (notably the
STRING level keys in a score answer's probabilities) in the live-tests job, since
the unit tests mock the invoker and would keep agreeing with a stale contract.
Both need the TYPESAFE_API_KEY repo secret. A missing key escalates the row to
ERROR rather than scoring it 0.0, so each job preflights it and fails with a
message naming the key instead of reporting an unexplained tasks_errored.
Deliberately not shared with llm_judge: checker_context.api_route resolves a TEXT
judge model, which is not a substitute, and a transport failure escalates the row
rather than scoring an ungraded row 0.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(criteria): close scoring-integrity and credential-disclosure gaps in system_one_judge
Addresses PR review findings. Each was reproduced before it was fixed.
Scoring failed OPEN. json.loads accepts the NaN/Infinity tokens, so a non-finite
float arrives through an ordinary 200, and base_url may point at any gateway. NaN
is unordered, so max(0.0, min(1.0, nan)) is 1.0: a malformed answer graded as FULL
credit. An infinite question weight did the same through inf/inf in the weighted
mean. Every wire value is now checked with math.isfinite, _clamp maps a non-finite
input to 0.0, and weight is bounded and rejects inf/NaN. An unusable answer earns
no credit and says so in findings.
The detached-grading consent gate did not disclose this judge. embedded_commands()
had no system_one_judge arm, so `evaluate <run_dir>` on an untrusted run record
would read a host env var into a Bearer header and POST it to a recorded URL with
no prompt. Both halves are now disclosed by name.
The reference scrub ran after serialization. The state persists as json.dumps,
which escapes newlines, while scrub_reference matches raw file text — so every
multi-line reference survived into the archived transcript. Confirmed by a new
leak-canary test that failed before the fix. The state's strings are scrubbed
first, then serialized.
Also: a 200 with no usable `answers` escalates instead of scoring an ungraded row
0.0; argmax noul uses a strict > 0.5 so a coin flip no longer passes in both
directions; httpx2.InvalidURL escalates rather than escaping the retry loop to be
downgraded to 0.0; the bare-list choice form rejects non-string options, so YAML's
bare `yes`/`no` cannot become options named 'True'/'False'; the transcript records
the whole response, whose `model` field is the only record of which version a
floating alias actually served.
Parallel paths the change had missed: TYPESAFE_API_KEY joins the container env
allowlist (grading runs in-container, so a docker task ERRORed on every row), and
the $REFERENCE_DIR reference-consumer validator now covers this judge too.
Adds tests/test_judge_system_one.py — the retry/escalation gate had no hermetic
tests at all, since every checker test patches the invoker out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* style(criteria): use deferred annotations in system_one_judge so the TYPE_CHECKING import reads as used
CodeQL flagged `TurnRecord` as an unused import. It is not unused — it is used at
the `turn_records` parameter — but the annotation was a string literal, which
CodeQL does not resolve, so the only reference was invisible to it.
Rather than suppress the alert, adopt agent_judge.py's shape: `from __future__
import annotations` plus unquoted annotations. The names become real AST
references, so the import reads as used without deferring anything at runtime
that was not already deferred.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
**Task success:** all *gating* criteria must score >= their `pass_threshold`. A
721
722
criterion with `weight: 0` is informational — it is still checked, stored, and
@@ -1335,6 +1336,66 @@ Observed label is `"yes"` when either signal is found, else `"no"`. Expected lab
1335
1336
1336
1337
**Typical pattern.** Label each dataset row with its true skill (`expected_skill`, `""` for negatives) and stack one `skill_triggered` criterion per skill against the same dataset — each gets its own confusion matrix from the same agent traces. This is the natural companion to a skill A/B experiment (skill plugin on vs. off); see the [A/B Experiment Guide](AB_EXPERIMENTS.md#recipe-ab-a-skill).
1337
1338
1339
+
### `system_one_judge`
1340
+
1341
+
Grade with a **System One model** ([TypeSafe's `jev`](https://docs.typesafe.ai/concepts/system-one)) instead of a text LLM. A System One model generates no text: it reads one state and answers a map of typed questions with calibrated probabilities, all in one round trip. The rubric you write **is** the grading schema, so there is no prompt to follow, no tool call to force, and no verdict to parse.
1342
+
1343
+
```yaml
1344
+
- type: "system_one_judge"
1345
+
description: "Rubric grade of the refactor"
1346
+
prompt: "The agent was asked to extract the retry loop into a helper."
1347
+
files: ["src/client.py"]
1348
+
questions:
1349
+
helper_extracted:
1350
+
type: noul
1351
+
instructions: "Is the retry loop extracted into a named helper function?"
1352
+
behaviour_preserved:
1353
+
type: noul
1354
+
instructions: "Does the refactor preserve the original retry semantics?"
1355
+
weight: 2.0
1356
+
naming:
1357
+
type: score
1358
+
instructions: "How well does the helper's name describe what it does?"
| `choice` | the top option plus a distribution over all of them | `expected: <option>` scores that one option 1.0; `values: {option: 0.0–1.0}` gives partial credit per option |
1372
+
| `score` | a position on an ordered spectrum, plus a distribution over levels | levels ramp evenly from 0.0 (first) to 1.0 (last) unless `values:` overrides them |
1373
+
1374
+
`choice`needs exactly one of `expected` or `values`. `score` takes 2–10 levels, ordered worst-first. Every question takes a `weight` (default 1.0).
1375
+
1376
+
A `choice` question's `criteria` is a map of option to a description of when it applies, but when the option names speak for themselves you can write a bare list instead — it widens to that map with null descriptions, which the API accepts:
1377
+
1378
+
```yaml
1379
+
exception_handling:
1380
+
type: choice
1381
+
instructions: "How does the function catch failures from requests.get?"
Option order is preserved as written, and a repeated option is a load-time error rather than a silently collapsed map.
1388
+
1389
+
**What the judge reads.** By default the state is `prompt` plus the `files` you list. Three flags widen it, each off by default: `include_agent_output` adds the agent's own final message, `include_tool_calls` adds a summary of its tool-call trajectory, and `include_dialog` adds the multi-turn user/agent exchange (only meaningful in [simulation mode](#simulation)). Turning the first two on is what lets a rubric grade *how* the agent worked and whether its summary was honest, not just the artifact it left behind — see `tasks/smoke_system_one_judge.yaml` for a rubric that does both. `include_reference` (on by default) adds the reference solution. Each section is capped independently by `max_state_chars`.
1390
+
1391
+
**Scoring** — the criterion score is computed by the harness, not the model: each question resolves to a value in [0.0, 1.0] and the score is their weighted mean. `scoring: expected` (default) weights every outcome by its probability, so a half-confident answer lands mid-scale; `scoring: argmax` reads only the top answer and discards the confidence. Either way the reduction is deterministic given the answers, and `findings` records the arithmetic per question so the grade is auditable line by line.
1392
+
1393
+
**Credentials** — the bearer token comes from the env var named by `api_key_env` (default `TYPESAFE_API_KEY`); only the *name* is stored in the task and in run records. `base_url` (default `https://api.typesafe.ai/v1`) points the criterion at a gateway or a recording proxy. Unlike `llm_judge`, this criterion does **not** honour `checker_context.api_route` — a System One model is not interchangeable with a text model, so the eval route's judge model would be the wrong default.
1394
+
1395
+
**When to reach for it over `llm_judge`** — a rubric with many small, repeated questions; a large dataset where a text judge's per-row cost dominates; or a grade you need to be reproducible and inspectable rather than argued in prose. Reach for `llm_judge` instead when the grade genuinely needs open-ended reasoning you cannot enumerate in advance.
1396
+
1397
+
**Failure modes** — a transport failure escalates the row to `ERROR` (it is eval infrastructure, not agent quality) rather than scoring 0.0. A question the API leaves unanswered, or answers with the wrong primitive, scores 0.0 at its full weight and says so in `findings`.
1398
+
1338
1399
## Checker Context
1339
1400
1340
1401
`checker_context` carries task-authored config for the success-checking side, namespaced by reserved key. Currently the only recognized namespace is **`api_route`**:
0 commit comments