Skip to content

Commit f3db499

Browse files
akshayliveclaude
andauthored
feat(criteria): add system_one_judge, a typed-rubric grader on a System One model (#192)
* feat(criteria): add system_one_judge, a typed-rubric grader on a System One model A System One model (TypeSafe's jev) generates no text: it reads one state and answers a map of typed questions with calibrated probabilities in a single round trip. That removes the two things a text judge has to defend against — an unparseable verdict and a tool call the model may decline to make — so the rubric the author writes IS the grading schema. The grade stays on our side: each question resolves to a value in [0,1] via its own expected/values, and the criterion score is their weighted mean. That buys determinism, a grade that can be recomputed from an archived transcript, and findings that print the arithmetic per question instead of arguing it in prose. A choice question's options may be written as a bare list when the names speak for themselves; it widens to the option->description map the API wants, with null descriptions, which the live endpoint accepts. Order is preserved and a repeated option is a load-time error. The judge reads the files it is given plus, opt-in per criterion, the agent's own final message (include_agent_output), its tool-call trajectory (include_tool_calls) and the simulated dialog (include_dialog) — so a rubric can grade how the agent worked and whether its summary was honest, not only the artifact left behind. tasks/smoke_system_one_judge.yaml exercises all three primitives over all three of those state slices, and runs in CI's smoke-pass bucket beside smoke_llm_judge. tests/test_system_one_judge_live.py pins the external wire shape (notably the STRING level keys in a score answer's probabilities) in the live-tests job, since the unit tests mock the invoker and would keep agreeing with a stale contract. Both need the TYPESAFE_API_KEY repo secret. A missing key escalates the row to ERROR rather than scoring it 0.0, so each job preflights it and fails with a message naming the key instead of reporting an unexplained tasks_errored. Deliberately not shared with llm_judge: checker_context.api_route resolves a TEXT judge model, which is not a substitute, and a transport failure escalates the row rather than scoring an ungraded row 0.0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(criteria): close scoring-integrity and credential-disclosure gaps in system_one_judge Addresses PR review findings. Each was reproduced before it was fixed. Scoring failed OPEN. json.loads accepts the NaN/Infinity tokens, so a non-finite float arrives through an ordinary 200, and base_url may point at any gateway. NaN is unordered, so max(0.0, min(1.0, nan)) is 1.0: a malformed answer graded as FULL credit. An infinite question weight did the same through inf/inf in the weighted mean. Every wire value is now checked with math.isfinite, _clamp maps a non-finite input to 0.0, and weight is bounded and rejects inf/NaN. An unusable answer earns no credit and says so in findings. The detached-grading consent gate did not disclose this judge. embedded_commands() had no system_one_judge arm, so `evaluate <run_dir>` on an untrusted run record would read a host env var into a Bearer header and POST it to a recorded URL with no prompt. Both halves are now disclosed by name. The reference scrub ran after serialization. The state persists as json.dumps, which escapes newlines, while scrub_reference matches raw file text — so every multi-line reference survived into the archived transcript. Confirmed by a new leak-canary test that failed before the fix. The state's strings are scrubbed first, then serialized. Also: a 200 with no usable `answers` escalates instead of scoring an ungraded row 0.0; argmax noul uses a strict > 0.5 so a coin flip no longer passes in both directions; httpx2.InvalidURL escalates rather than escaping the retry loop to be downgraded to 0.0; the bare-list choice form rejects non-string options, so YAML's bare `yes`/`no` cannot become options named 'True'/'False'; the transcript records the whole response, whose `model` field is the only record of which version a floating alias actually served. Parallel paths the change had missed: TYPESAFE_API_KEY joins the container env allowlist (grading runs in-container, so a docker task ERRORed on every row), and the $REFERENCE_DIR reference-consumer validator now covers this judge too. Adds tests/test_judge_system_one.py — the retry/escalation gate had no hermetic tests at all, since every checker test patches the invoker out. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * style(criteria): use deferred annotations in system_one_judge so the TYPE_CHECKING import reads as used CodeQL flagged `TurnRecord` as an unused import. It is not unused — it is used at the `turn_records` parameter — but the annotation was a string literal, which CodeQL does not resolve, so the only reference was invisible to it. Rather than suppress the alert, adopt agent_judge.py's shape: `from __future__ import annotations` plus unquoted annotations. The names become real AST references, so the import reads as used without deferring anything at runtime that was not already deferred. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent b5d8ba8 commit f3db499

31 files changed

Lines changed: 2012 additions & 24 deletions

‎.claude/notes/contracts.md‎

Lines changed: 66 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -391,3 +391,69 @@ with a plain synchronous call, because offloading a single fast syscall only wid
391391
cancellation window. The cleanup itself is deliberately synchronous: a bare `await` inside
392392
`finally` is cancellable, and cancelling as that line is reached would skip cleanup and leak
393393
the copy with no reaper.
394+
395+
## System One rubric scoring
396+
397+
A System One model (TypeSafe's `jev`) does not generate text. It reads one state and answers
398+
a map of typed questions — `noul` (P(yes)), `choice` (a distribution over options), `score` (a
399+
distribution over ordered levels) — with calibrated probabilities. That inverts what a judge
400+
criterion has to defend against. There is no prose to parse, so the whole `submit_verdict`
401+
tool channel has no counterpart here: the response schema is fixed by the questions that were
402+
asked, and a model that answers off-schema is a wire fault, not a grading fault. It also means
403+
there is no system prompt to hold an identity, so the transcript's system-prompt slot carries
404+
the rubric as sent — that, not a persona, is what a reviewer needs to replay a grade.
405+
406+
### The score is ours, not the model's
407+
408+
The model reports a distribution; the grade is arithmetic we do on it. Each question resolves
409+
to a value in [0.0, 1.0] through the author's own `expected` / `values`, and the criterion
410+
score is their weighted mean. Keeping that reduction on our side buys three things a text
411+
judge cannot have: the same answers always produce the same grade, the arithmetic is printable
412+
(`findings` carries one line per question, with the value and the weight), and a rubric can be
413+
re-scored from an archived transcript without another call.
414+
415+
`scoring: expected` weights every outcome by its probability, so a model that is genuinely
416+
torn lands mid-scale instead of being rounded into a confident-looking verdict — the point of
417+
a calibrated model. `argmax` exists for the case where partial credit is misleading rather
418+
than informative: a gate. The default is `expected` because discarding the confidence is the
419+
lossy choice and should be the one you ask for.
420+
421+
A distribution that is absent, non-numeric or sums to zero falls back to the point answer
422+
(`choice` / `score` / `noul`) rather than grading as 0.0 — a broken `probabilities` block is a
423+
provider fault, and the point answer is still a real answer. A question that is *unanswered*,
424+
or answered with the wrong primitive, is the opposite case: it scores 0.0 at its full weight
425+
and names itself in `findings`, because the rubric asked something the grade depends on and
426+
dropping it would quietly inflate the mean.
427+
428+
### What it does not share with `llm_judge`
429+
430+
It does not read `checker_context.api_route`. The eval route resolves a TEXT judge model, and
431+
substituting one for a System One model is not a fallback, it is a different API. The
432+
credential comes from the env var *named* by `api_key_env`, so only the name is ever stored on
433+
the criterion or persisted into a run record. A transport failure raises
434+
`JudgeInfrastructureError` and escalates the row, rather than following `llm_judge`'s
435+
unconfigured-transport arm into a scored 0.0 — an ungraded row must not read as a failed one.
436+
437+
A 200 response whose body carries no usable `answers` map is the same class of fault. Reducing
438+
it would score every question "no answer returned" and finalize the row as a graded 0.0, which
439+
reads as a failure the agent earned; it escalates instead.
440+
441+
### Non-finite values fail CLOSED
442+
443+
`json.loads` accepts the `NaN` and `Infinity` tokens, so a non-finite float arrives through an
444+
ordinary 200 — and `base_url` may point at any gateway. NaN is unordered, so the obvious clamp
445+
`max(0.0, min(1.0, x))` returns **1.0** for it: the naive reading of a malformed answer is FULL
446+
credit. Every wire value is therefore checked with `math.isfinite` before it is used (`_is_number`),
447+
`_clamp` maps a non-finite input to 0.0, and a question's `weight` is bounded and rejects inf/NaN
448+
so the weighted mean cannot become `inf/inf`. The rule is that an unusable answer never earns
449+
credit; it scores 0.0 and says so in `findings`.
450+
451+
`argmax` uses a strict `> 0.5` for `noul`. At exactly 0.5 the agreement is 0.5 whichever way
452+
`expected` points, so `>=` passed a maximally uncertain answer in *both* directions.
453+
454+
### The reference scrub runs before serialization
455+
456+
`system_one_judge` persists its state as `json.dumps(state)`, which escapes newlines and tabs.
457+
`scrub_reference` matches raw file text, so scrubbing the *rendered* JSON silently misses every
458+
multi-line reference. The state's strings are scrubbed first, then serialized — a leak shipped
459+
exactly this way, so the leak-canary test uses a multi-line, quoted reference on purpose.

‎.env.example‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -31,6 +31,14 @@ LOG_TO_FILE=false # Set to true to enable file logging
3131
# BEDROCK_MODEL="eu.anthropic.claude-sonnet-4-5-20250929-v1:0"
3232
# BEDROCK_SMALL_MODEL="eu.anthropic.claude-haiku-4-5"
3333

34+
# TypeSafe System One (the system_one_judge criterion, model `jev`). NOT routed
35+
# through API_BACKEND — a System One model answers typed questions with
36+
# probabilities rather than text, so it is its own endpoint and its own key.
37+
# The criterion stores only this variable's NAME (api_key_env), never the token.
38+
# Unlike llm_judge, a missing key escalates the row to ERROR instead of scoring
39+
# it 0.0: an ungraded row must not read as a failed one.
40+
# TYPESAFE_API_KEY="<your_typesafe_api_key_here>"
41+
3442
# Codex agent settings (requires the [codex] extra: `uv sync --extra codex`).
3543
# Only CODEX_API_KEY is read for auth (sent as a Bearer token). CODEX_BASE_URL
3644
# routes to a custom OpenAI-/responses-compatible endpoint; unset = standard

‎.github/workflows/pr-checks.yml‎

Lines changed: 42 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -488,16 +488,21 @@ jobs:
488488
AWS_BEARER_TOKEN_BEDROCK: ${{ secrets.AWS_BEARER_TOKEN_BEDROCK }}
489489
AWS_REGION: ${{ secrets.AWS_REGION }}
490490
BEDROCK_MODEL: ${{ secrets.BEDROCK_MODEL }}
491-
# tasks_run for --tags smoke-pass. 8 task files (hello_date, dataset_example,
492-
# smoke_llm_judge, smoke_agent_judge, byod_smoke_test, agentless_smoke_test,
493-
# anti_cheat_reference, record_cli_responses); dataset_example fans out to 2
494-
# inline rows, so 9 sub-tasks. If you add/remove a smoke-pass task or change
495-
# the dataset row count, bump these.
491+
# Grades smoke_system_one_judge. Not a Bedrock credential: the System One
492+
# judge calls TypeSafe directly and ignores the run's API backend. A
493+
# missing key escalates that task to ERROR rather than failing a criterion,
494+
# so the preflight step below fails loudly instead.
495+
TYPESAFE_API_KEY: ${{ secrets.TYPESAFE_API_KEY }}
496+
# tasks_run for --tags smoke-pass. 9 task files (hello_date, dataset_example,
497+
# smoke_llm_judge, smoke_agent_judge, smoke_system_one_judge, byod_smoke_test,
498+
# agentless_smoke_test, anti_cheat_reference, record_cli_responses);
499+
# dataset_example fans out to 2 inline rows, so 10 sub-tasks. If you
500+
# add/remove a smoke-pass task or change the dataset row count, bump these.
496501
#
497502
# anti_cheat_reference lives in a SUBDIRECTORY, which `tasks/*.yaml` does not
498503
# match — the smoke-pass step names its path explicitly. Keep that in sync.
499-
EXPECTED_SMOKE_PASS_RUN: "9"
500-
EXPECTED_SMOKE_PASS_SUCCEEDED: "9"
504+
EXPECTED_SMOKE_PASS_RUN: "10"
505+
EXPECTED_SMOKE_PASS_SUCCEEDED: "10"
501506
# smoke-fail bucket: three tasks expected to fail.
502507
# 1. smoke_negative_path: file_contains criterion is unsatisfiable
503508
# (sentinel-string regression detection for success-checker).
@@ -571,6 +576,15 @@ jobs:
571576
# record_cli_responses is the record_cli per-invocation-response probe and
572577
# is also driver: docker, so it needs that same image; it is flat in
573578
# tasks/, so the glob already matches it.
579+
- name: Verify smoke secrets present
580+
# smoke_system_one_judge grades through TypeSafe. Without the key the
581+
# criterion raises JudgeInfrastructureError and the task lands in
582+
# tasks_errored, which reads as "the harness broke" rather than "the
583+
# secret is missing". Fail here, where the message says which.
584+
run: |
585+
: "${TYPESAFE_API_KEY:?TYPESAFE_API_KEY missing — needed by smoke_system_one_judge}"
586+
echo "All smoke secrets present."
587+
574588
- name: Run smoke-pass bucket (expect all to succeed)
575589
run: |
576590
.venv/bin/coder-eval run tasks/*.yaml tasks/anti_cheat_reference/*.yaml \
@@ -697,6 +711,8 @@ jobs:
697711
AWS_BEARER_TOKEN_BEDROCK: ${{ secrets.AWS_BEARER_TOKEN_BEDROCK }}
698712
AWS_REGION: ${{ secrets.AWS_REGION }}
699713
BEDROCK_MODEL: ${{ secrets.BEDROCK_MODEL }}
714+
# The System One judge's own endpoint — unrelated to either route above.
715+
TYPESAFE_API_KEY: ${{ secrets.TYPESAFE_API_KEY }}
700716

701717
steps:
702718
- name: Checkout code
@@ -746,6 +762,7 @@ jobs:
746762
: "${AWS_BEARER_TOKEN_BEDROCK:?AWS_BEARER_TOKEN_BEDROCK missing}"
747763
: "${AWS_REGION:?AWS_REGION missing}"
748764
: "${BEDROCK_MODEL:?BEDROCK_MODEL missing}"
765+
: "${TYPESAFE_API_KEY:?TYPESAFE_API_KEY missing}"
749766
echo "All live-test secrets present."
750767
751768
- name: Run claude-settings enforcement live tests (DirectRoute)
@@ -773,6 +790,17 @@ jobs:
773790
-m live -v --tb=short --strict-markers -ra -n 4 \
774791
--junit-xml=tmp/junit-settings-bedrock.xml
775792
793+
- name: Run System One judge wire-contract live tests
794+
# The only thing in CI that talks to TypeSafe. The unit tests mock the
795+
# invoker, so nothing else catches a change to the answer shape the
796+
# reduction assumes — notably the STRING level keys ("0", "1", ...) in a
797+
# score answer's probabilities. `-n0`: three questions in one round trip,
798+
# so there is nothing to parallelize.
799+
run: |
800+
.venv/bin/pytest tests/test_system_one_judge_live.py \
801+
-m live -n0 -v --tb=short --strict-markers -ra \
802+
--junit-xml=tmp/junit-system-one.xml
803+
776804
- name: Assert live tests actually ran (not silently skipped)
777805
# Parse JUnit XML for *passed* count, not collected count. Pytest collects
778806
# @pytest.mark.skipif-marked tests even when the predicate is True, so a
@@ -798,11 +826,17 @@ jobs:
798826
return total - skipped - errors - failures
799827
p_settings_direct = passed("tmp/junit-settings.xml")
800828
p_settings_bedrock = passed("tmp/junit-settings-bedrock.xml")
801-
print(f"Passed: settings(direct)={p_settings_direct}, settings(bedrock)={p_settings_bedrock}")
829+
p_system_one = passed("tmp/junit-system-one.xml")
830+
print(
831+
f"Passed: settings(direct)={p_settings_direct}, "
832+
f"settings(bedrock)={p_settings_bedrock}, system_one={p_system_one}"
833+
)
802834
if p_settings_direct < 1:
803835
sys.exit("test_claude_settings_enforcement_live.py (DirectRoute) reported zero PASSED tests")
804836
if p_settings_bedrock < 1:
805837
sys.exit("test_claude_settings_enforcement_live.py (BedrockRoute) reported zero PASSED tests")
838+
if p_system_one < 1:
839+
sys.exit("test_system_one_judge_live.py reported zero PASSED tests")
806840
PY
807841
808842
- name: Run cost-budget smoke (max_usd → COST_BUDGET_EXCEEDED via DirectRoute)

‎docs/REPORT_SCHEMA.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -184,7 +184,7 @@ fields so subclass keys round-trip.
184184
`transcript_path` (a sibling `judge-N.yaml`, or `post-failure-judge-N.yaml` for
185185
diagnostic records). The full `transcript` is **stripped
186186
from `task.json`** — read it from the referenced file. Emitted by `llm_judge`,
187-
`agent_judge`.
187+
`agent_judge`, `system_one_judge`.
188188

189189
### Post-failure criterion evidence
190190

‎docs/TASK_DEFINITION_GUIDE.md‎

Lines changed: 62 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -35,6 +35,7 @@ Complete reference for defining evaluation tasks in Coder Eval.
3535
- [llm_judge](#llm_judge)
3636
- [agent_judge](#agent_judge)
3737
- [skill_triggered](#skill_triggered)
38+
- [system_one_judge](#system_one_judge)
3839
- [Checker Context](#checker-context)
3940
- [Reference Solutions](#reference-solutions)
4041
- [Pre-Run Commands](#pre-run-commands)
@@ -715,7 +716,7 @@ All criteria share these fields:
715716
**Scoring types:**
716717
- **Binary** (1.0 or 0.0): `file_exists`, `run_command`, `file_matches_regex`, `cli_called`, `classification_match`, `skill_triggered`
717718
- **Fractional** (0.0–1.0): `file_contains`, `file_check`, `json_check`, `command_executed`, `uipath_eval`
718-
- **Continuous** (0.0–1.0): `reference_comparison`, `commands_efficiency`, `llm_judge`, `agent_judge`
719+
- **Continuous** (0.0–1.0): `reference_comparison`, `commands_efficiency`, `llm_judge`, `agent_judge`, `system_one_judge`
719720

720721
**Task success:** all *gating* criteria must score >= their `pass_threshold`. A
721722
criterion with `weight: 0` is informational — it is still checked, stored, and
@@ -1335,6 +1336,66 @@ Observed label is `"yes"` when either signal is found, else `"no"`. Expected lab
13351336

13361337
**Typical pattern.** Label each dataset row with its true skill (`expected_skill`, `""` for negatives) and stack one `skill_triggered` criterion per skill against the same dataset — each gets its own confusion matrix from the same agent traces. This is the natural companion to a skill A/B experiment (skill plugin on vs. off); see the [A/B Experiment Guide](AB_EXPERIMENTS.md#recipe-ab-a-skill).
13371338

1339+
### `system_one_judge`
1340+
1341+
Grade with a **System One model** ([TypeSafe's `jev`](https://docs.typesafe.ai/concepts/system-one)) instead of a text LLM. A System One model generates no text: it reads one state and answers a map of typed questions with calibrated probabilities, all in one round trip. The rubric you write **is** the grading schema, so there is no prompt to follow, no tool call to force, and no verdict to parse.
1342+
1343+
```yaml
1344+
- type: "system_one_judge"
1345+
description: "Rubric grade of the refactor"
1346+
prompt: "The agent was asked to extract the retry loop into a helper."
1347+
files: ["src/client.py"]
1348+
questions:
1349+
helper_extracted:
1350+
type: noul
1351+
instructions: "Is the retry loop extracted into a named helper function?"
1352+
behaviour_preserved:
1353+
type: noul
1354+
instructions: "Does the refactor preserve the original retry semantics?"
1355+
weight: 2.0
1356+
naming:
1357+
type: score
1358+
instructions: "How well does the helper's name describe what it does?"
1359+
criteria: ["opaque", "workable", "self-explanatory"]
1360+
leftovers:
1361+
type: noul
1362+
instructions: "Is any dead code left behind?"
1363+
expected: false
1364+
```
1365+
1366+
**Question types**
1367+
1368+
| Type | What the model returns | How the rubric turns it into 0.0–1.0 |
1369+
| --- | --- | --- |
1370+
| `noul` | P(yes) | `expected: true` (default) scores P(yes); `expected: false` scores 1 − P(yes) |
1371+
| `choice` | the top option plus a distribution over all of them | `expected: <option>` scores that one option 1.0; `values: {option: 0.0–1.0}` gives partial credit per option |
1372+
| `score` | a position on an ordered spectrum, plus a distribution over levels | levels ramp evenly from 0.0 (first) to 1.0 (last) unless `values:` overrides them |
1373+
1374+
`choice` needs exactly one of `expected` or `values`. `score` takes 2–10 levels, ordered worst-first. Every question takes a `weight` (default 1.0).
1375+
1376+
A `choice` question's `criteria` is a map of option to a description of when it applies, but when the option names speak for themselves you can write a bare list instead — it widens to that map with null descriptions, which the API accepts:
1377+
1378+
```yaml
1379+
exception_handling:
1380+
type: choice
1381+
instructions: "How does the function catch failures from requests.get?"
1382+
criteria: [none, bare_except, broad_exception, specific_timeout]
1383+
expected: specific_timeout
1384+
weight: 2.0
1385+
```
1386+
1387+
Option order is preserved as written, and a repeated option is a load-time error rather than a silently collapsed map.
1388+
1389+
**What the judge reads.** By default the state is `prompt` plus the `files` you list. Three flags widen it, each off by default: `include_agent_output` adds the agent's own final message, `include_tool_calls` adds a summary of its tool-call trajectory, and `include_dialog` adds the multi-turn user/agent exchange (only meaningful in [simulation mode](#simulation)). Turning the first two on is what lets a rubric grade *how* the agent worked and whether its summary was honest, not just the artifact it left behind — see `tasks/smoke_system_one_judge.yaml` for a rubric that does both. `include_reference` (on by default) adds the reference solution. Each section is capped independently by `max_state_chars`.
1390+
1391+
**Scoring** — the criterion score is computed by the harness, not the model: each question resolves to a value in [0.0, 1.0] and the score is their weighted mean. `scoring: expected` (default) weights every outcome by its probability, so a half-confident answer lands mid-scale; `scoring: argmax` reads only the top answer and discards the confidence. Either way the reduction is deterministic given the answers, and `findings` records the arithmetic per question so the grade is auditable line by line.
1392+
1393+
**Credentials** — the bearer token comes from the env var named by `api_key_env` (default `TYPESAFE_API_KEY`); only the *name* is stored in the task and in run records. `base_url` (default `https://api.typesafe.ai/v1`) points the criterion at a gateway or a recording proxy. Unlike `llm_judge`, this criterion does **not** honour `checker_context.api_route` — a System One model is not interchangeable with a text model, so the eval route's judge model would be the wrong default.
1394+
1395+
**When to reach for it over `llm_judge`** — a rubric with many small, repeated questions; a large dataset where a text judge's per-row cost dominates; or a grade you need to be reproducible and inspectable rather than argued in prose. Reach for `llm_judge` instead when the grade genuinely needs open-ended reasoning you cannot enumerate in advance.
1396+
1397+
**Failure modes** — a transport failure escalates the row to `ERROR` (it is eval infrastructure, not agent quality) rather than scoring 0.0. A question the API leaves unanswered, or answers with the wrong primitive, scores 0.0 at its full weight and says so in `findings`.
1398+
13381399
## Checker Context
13391400

13401401
`checker_context` carries task-authored config for the success-checking side, namespaced by reserved key. Currently the only recognized namespace is **`api_route`**:

‎evalboard/lib/pricing.generated.ts‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -69,6 +69,8 @@ export const PRICING: Record<string, Pricing> = {
6969
"gpt-5.6-luna": { inputPerMTok: 0.2, outputPerMTok: 1.2, cacheWritePerMTok: 0.2, cacheReadPerMTok: 0.02 },
7070
"gpt-5.6-sol": { inputPerMTok: 4.0, outputPerMTok: 20.0, cacheWritePerMTok: 4.0, cacheReadPerMTok: 0.4 },
7171
"gpt-5.6-terra": { inputPerMTok: 2.0, outputPerMTok: 12.0, cacheWritePerMTok: 2.0, cacheReadPerMTok: 0.2 },
72+
"jev-1.13.0": { inputPerMTok: 0.042, outputPerMTok: 0.0, cacheWritePerMTok: 0.042, cacheReadPerMTok: 0.0 },
73+
"jev-latest": { inputPerMTok: 0.042, outputPerMTok: 0.0, cacheWritePerMTok: 0.042, cacheReadPerMTok: 0.0 },
7274
"kimi-k2-7-code": { inputPerMTok: 0.95, outputPerMTok: 4.0, cacheWritePerMTok: 0.0, cacheReadPerMTok: 0.19 },
7375
"moonshotai.kimi-k2.5": { inputPerMTok: 0.72, outputPerMTok: 3.6, cacheWritePerMTok: 0.72, cacheReadPerMTok: 0.0 },
7476
"virtuoso-1-5": { inputPerMTok: 0.95, outputPerMTok: 4.0, cacheWritePerMTok: 0.0, cacheReadPerMTok: 0.16 },

0 commit comments

Comments
 (0)