Skip to content

Commit a6b828d

Browse files
uipreligaclaude
andcommitted
feat(plugin): add optimize-skill and the --split row filter it measures with
Adds the `/coder-eval:optimize-skill` skill — prose only, no Python — plus the one product change it needs to measure anything: `Dataset.split_field` and a `--split` row filter in `task_loader.expand_dataset`. Both are inert when unused: a dataset with no `split_field` and a run with no `--split` behave exactly as before. Stage B here is a manual reading of two runs. There is no gate, no statistics and no `coder_eval.optimize` package in this PR. One published skill changes behaviour: `analyze`'s frontmatter `description` is replaced by the measured winner of the round that tutorial 08 documents. Also fixes a dependency defect this PR's CI surfaced, unrelated to the split: `evaluation/judge_bedrock.py` imports `httpx`, which `pyproject.toml` never declared. It arrived transitively via `anthropic` until `anthropic` 1.0.0 (released 2026-08-20) moved to `httpx2`. A FRESH resolve — what `uv tool install` and `pip install coder-eval` do, since a wheel carries no lockfile — then stopped installing `httpx`, so `criteria/llm_judge.py` raised on import, the discovery loop swallowed it, and `llm_judge` vanished from the registry. `uv.lock` hid this from every `uv sync --frozen` job. `httpx` is now declared, and `tests/test_declared_dependencies.py` asserts the invariant on the DECLARATION rather than the installed set — the only form that fails in the locked jobs, where this bug was invisible. Imports guarded by `try/except ImportError` are derived as optional and exempt, so the soft dependencies (`google.antigravity`, `openai`) need no allowlist. Squashed from feat/plugin-optimize-skill: 7376062 feat(dataset): 1/3 — add Dataset.split_field and the --split row filter 240d66c feat(plugin): 2/3 — add the optimize-skill skill and split-label the activation template 8920410 docs: 3/3 — add the skill-optimization tutorial, and fix the reachability guidance it disproved 2c30397 style: apply ruff format to the reachability lint assertion ae3c39b docs(harness): record the all-skipped-run-exits-0 gap found while adding --split b53c7d4 fix: code review fixes for the split-field / optimize-skill plan d7d56f1 feat(plugin): promote a measured `analyze` description, and close the two open findings 844348d docs(tutorial): Stage C completed — the analyze promotion is confirmed on holdout e340b58 docs: record the bare-name collision hazard, and mark the plan complete Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 3d09f0f commit a6b828d

37 files changed

Lines changed: 1894 additions & 56 deletions

‎.claude/harness-candidates.md‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -455,3 +455,18 @@ with the two `action.yml` items above — one considered change to the action's
455455
`final_status`, which does not exist in `run.json` and would have made a new
456456
assertion dead on arrival. Guard: assert the key set that non-Python consumers
457457
depend on, mirroring how CE030 pins doc/schema parity.
458+
## From the split-field / optimize-skill plan (2026-08-12)
459+
460+
- [ ] **A run whose every task is skipped exits 0 — a green run of zero tasks.** When
461+
`resolve_all_tasks` demotes every task to `skipped_tasks` (a load failure, `skip: true`,
462+
or now a `--split` selector matching no labelled row), the run reports success: nothing
463+
failed, so the exit gate in `cli/run_command.py` — which keys only on failed/errored tasks
464+
and suite gates — passes. Verified directly: `coder-eval run <suite> --split holdou`
465+
prints one yellow "1 task file(s) skipped" line and exits 0. This is pre-existing, but
466+
`--split` makes it reachable by a one-character CLI typo rather than a broken file, and
467+
the whole point of a holdout confirmation is that you trust its verdict. Not guarded, and
468+
not a five-minute fix: making an all-skipped run non-green changes exit semantics for
469+
every skipped-task path (including deliberate `skip: true` suites and tag filters that
470+
match nothing), so it needs a decision about which of those should be fatal, plus tests
471+
per case. A narrower option is to fail only when a CLI *selector* (`--split`, `--tags`)
472+
eliminated everything, since that is unambiguously a user error rather than repo state.

‎.claude/shared/run-layout.md‎

Lines changed: 47 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,53 @@ runs/<run_id>/<variant_id>/<task_id>/<NN>/{task.json, task.log, artifacts/}
1414
- `task.json.malformed` — present only on the docker degrade path: when an existing `task.json` fails to parse (schema skew from a stale `:latest` image, or a truncated/torn write), the docker runner moves the unparseable original aside to this sidecar and writes a synthetic `final_status=ERROR` `task.json` in its place. Diagnostic-only; `rglob("task.json")` consumers do not match it.
1515
- `task.log` — the human-readable task log; `artifacts/` — files the agent produced.
1616

17+
## Suite rollups (dataset-backed tasks only)
18+
19+
A task carrying `dataset:` fans out into one row-task per row and additionally writes a
20+
per-suite rollup:
21+
22+
```
23+
runs/<run_id>/<variant_id>/<suite_id>/{suite.json, suite.md}
24+
```
25+
26+
`<suite_id>` is the original (pre-fan-out) `task_id`. Nothing is written for a task
27+
without `dataset:`.
28+
29+
`suite.json` carries the suite's pass counts plus `criterion_aggregates[]` — one entry per
30+
criterion that opted into across-row aggregation, each with:
31+
32+
- `criterion_type`, and `description` (set when a task stacks several criteria of the same
33+
type, e.g. one `skill_triggered` per skill — that is what distinguishes them);
34+
- `rows_total` and `rows_excluded` — the denominator and what was dropped from it. A row
35+
that errored before criteria ran (a timeout, say) is **excluded** rather than scored, so
36+
metrics are computed over `rows_total - rows_excluded`;
37+
- `metrics` — a **flat** name → float map. Classification-style criteria emit
38+
`accuracy`, `macro_f1`, and per-label `precision.<label>` / `recall.<label>` /
39+
`f1.<label>` (for `skill_triggered` the labels are `yes` / `no`). It also carries
40+
`completion_rate` — the surviving fraction of the denominator above — which is an
41+
ordinary metric, so `suite_thresholds: {completion_rate: 1.0}` gates a run whose sample
42+
eroded;
43+
- `details` — `labels`, `per_label`, `confusion`, `total_pairs`;
44+
- `threshold_checks` + `passed`, from the criterion's `suite_thresholds`.
45+
46+
Alongside the aggregates, `failed_samples[]` lists failed/errored rows **by `row_id`** with
47+
their failure reasons, capped at a fixed limit. It is the only place in `suite.json` that
48+
carries row identity — `metrics` and `details` are counts only — so a consumer that needs
49+
to know *which* rows failed reads `failed_samples`, then falls back to the per-row
50+
`<variant_id>/<suite_id>/<row_id>/<NN>/task.json` for anything past the cap.
51+
52+
**Read `metrics`; never recompute a metric from `details.confusion`.** The criterion layer
53+
owns that arithmetic, including its division-by-zero convention, and a consumer that
54+
re-derives F1 will disagree with the gate the run already applied.
55+
56+
**Rollups pool replicates.** The grouping key is `(variant_id, suite_id)` — the replicate
57+
index is *not* part of it. So `--repeats N` over an M-row suite writes **one** `suite.json`
58+
per variant, with `rows_total: N × M` and a single pooled confusion matrix. There is no
59+
per-replicate metric in that file. A consumer that needs replicate-to-replicate spread must
60+
invoke the run N times and read N run directories; `--repeats` cannot serve that purpose.
61+
(Contrast `experiment.md`'s `## Paired Comparison` block, which averages replicates per row
62+
before pairing — there `--repeats` is exactly the right tool.)
63+
1764
**Scope-marker files** (used to detect what a given path represents):
1865

1966
- `run.json` at the run root → **run scope**. If `experiment.json` (+ `experiment.md`) is also present → multi-variant experiment.

‎CLAUDE.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -124,7 +124,7 @@ tests/ # Test suite
124124
docs/ # Documentation
125125
templates/ # Sandbox template directories
126126
.claude-plugin/marketplace.json # Makes this repo a Claude Code plugin marketplace (`/plugin marketplace add UiPath/coder_eval`); lists the one plugin below.
127-
plugins/coder-eval/ # The published Claude Code plugin: `.claude-plugin/plugin.json` (its `version` is a derived pin of pyproject's, bumped by release.yml, guarded by tests/test_action_version_pin.py), `skills/<name>/SKILL.md` × 6 (`/coder-eval:init`, `/coder-eval:check-skill`, `/coder-eval:task`, `/coder-eval:lint-tasks`, `/coder-eval:analyze`, `/coder-eval:ci`), and `reference/` — everything a skill reads must live here, since an installed plugin is copied to ~/.claude/plugins/cache/ WITHOUT its parent dirs (address it via `${CLAUDE_PLUGIN_ROOT}`). `reference/criteria.md` is generated (`make plugin-reference`, CE033); `reference/run-layout.md` is a verbatim mirror of `.claude/shared/run-layout.md`; `reference/task-rubric.md` is the shared task-quality rubric that `task` and `lint-tasks` both read (plugin-only — no repo-side twin); `reference/repo-layout.md` is the eval-tree DISCOVERY policy every skill reads (`SKILL_NEEDS_EVAL_ROOT_DISCOVERY`, which a new skill must declare a stance in) — glob for `task_id:` files and `run.json`, never assume `tasks/`/`runs/latest` — as distinct from `run-layout.md`, which describes what is inside a run directory. Every skill must appear in all four surfaces in `SKILL_DOC_SURFACES` (derived test), and their combined frontmatter `description` length is capped (`SKILL_LISTING_BUDGET_CHARS`) because the skill listing's budget is shared with every skill the user has installed. **Skill naming is verb-first imperative** — a skill is a command you issue (`/coder-eval:<name>`) and every one of them takes an action, so name it for the action: a bare verb where that is unambiguous (`init`, `analyze` — the object comes from the argument), otherwise `<verb>-<object>` (`lint-tasks`, `check-skill`). Never `<object>-<verb>`: `skill-check` was renamed to `check-skill` precisely because it read backwards next to `lint-tasks`. `task` and `ci` predate the rule and stay — renaming a published skill breaks every user's muscle memory for no functional gain, since activation keys on the `description`, never the name. Distinct from `.claude/commands/`, which stays repo-local contributor tooling.
127+
plugins/coder-eval/ # The published Claude Code plugin: `.claude-plugin/plugin.json` (its `version` is a derived pin of pyproject's, bumped by release.yml, guarded by tests/test_action_version_pin.py), `skills/<name>/SKILL.md` × 7 (`/coder-eval:init`, `/coder-eval:check-skill`, `/coder-eval:optimize-skill`, `/coder-eval:task`, `/coder-eval:lint-tasks`, `/coder-eval:analyze`, `/coder-eval:ci`), and `reference/` — everything a skill reads must live here, since an installed plugin is copied to ~/.claude/plugins/cache/ WITHOUT its parent dirs (address it via `${CLAUDE_PLUGIN_ROOT}`). `reference/criteria.md` is generated (`make plugin-reference`, CE033); `reference/run-layout.md` is a verbatim mirror of `.claude/shared/run-layout.md`; `reference/task-rubric.md` is the shared task-quality rubric that `task` and `lint-tasks` both read (plugin-only — no repo-side twin); `reference/repo-layout.md` is the eval-tree DISCOVERY policy every skill reads (`SKILL_NEEDS_EVAL_ROOT_DISCOVERY`, which a new skill must declare a stance in) — glob for `task_id:` files and `run.json`, never assume `tasks/`/`runs/latest` — as distinct from `run-layout.md`, which describes what is inside a run directory. Every skill must appear in all four surfaces in `SKILL_DOC_SURFACES` (derived test), and their combined frontmatter `description` length is capped (`SKILL_LISTING_BUDGET_CHARS`) because the skill listing's budget is shared with every skill the user has installed. **Skill naming is verb-first imperative** — a skill is a command you issue (`/coder-eval:<name>`) and every one of them takes an action, so name it for the action: a bare verb where that is unambiguous (`init`, `analyze` — the object comes from the argument), otherwise `<verb>-<object>` (`lint-tasks`, `check-skill`, `optimize-skill`). Never `<object>-<verb>`: `skill-check` was renamed to `check-skill` precisely because it read backwards next to `lint-tasks`, and `optimize-skill` was named that way rather than `skill-optimize` for the same reason — the rule is applied at authoring time now, not repaired later. `task` and `ci` predate the rule and stay — renaming a published skill breaks every user's muscle memory for no functional gain, since activation keys on the `description`, never the name. Distinct from `.claude/commands/`, which stays repo-local contributor tooling.
128128
action.yml # Published composite GitHub Action (coder-eval as a CI gate). release.yml's `release` job maintains its `version:` default; its `promote` job (gated on publish-pypi) moves the `v<major>` tag + cuts the Release, so nothing consumer-visible moves before the wheel is on PyPI. verify-published-action.yml then verifies the published composite (tag/pin/PyPI/Marketplace parity, plus a real consumer run) after each Release and nightly. Runbook: CONTRIBUTING.md § Releasing.
129129
```
130130

@@ -139,7 +139,7 @@ action.yml # Published composite GitHub Action (coder-ev
139139
- **Single declarative merge resolver**: All five config layers merge through ONE engine (`orchestration/config_merge.py::resolve_root`) for the three `-D`-reachable roots (`agent`/`run_limits`/`sandbox`). Each field declares *how it merges* once, on the model, via `MergeField(strategy="deep"|"append"|"replace")` (or a type-aware default: nested `BaseModel`/free-form `dict` → `deep`; `list`/scalar → `replace`). `resolve_task_for_variant` (layers 1–4) and `apply_overrides` (layer 5) build `Layer` lists and call the same `resolve_root`, so a field merges identically regardless of which layer supplied it (the unification invariant, enforced by `tests/test_merge_unification.py`). Lint rule CE014 forces every list field to declare its strategy explicitly.
140140
- **Generic CLI overrides (`-D`/`--set`)**: Layer 5 is a thin wrapper (`orchestration/overrides.py`) over the resolver above. `coder-eval run -D agent.model=opus -D run_limits.max_turns=30` overrides any field on the resolved `TaskDefinition` (`agent`/`run_limits`/`sandbox` roots), schema-validated with did-you-mean. Only `--model` (→ `agent.model`) and `--driver` (→ `sandbox.driver`) survive as active thin aliases that emit the equivalent `-D` entry; an alias and `-D` targeting the same path is a hard error. `--type` (→ `agent.type`) is a separate, lighter alias that does NOT route through that collision check — `--type` and `-D agent.type=…` last-win rather than hard-error (the `-D` value wins). Tools, plugins, and SDK options are `-D`-only.
141141
- **All core models importable from `coder_eval.models`** regardless of submodule
142-
- **Dataset fan-out**: `TaskDefinition.dataset` (inline rows or JSONL path) expands a single task into N row-tasks with `${row.<field>}` substitution in `initial_prompt` and `success_criteria` string fields. Expansion runs in `task_loader.expand_dataset` **before** variant resolution, so variants cannot override the dataset. Row sampling: CLI `--sample N` (fixed-seed uniform-random N over the whole dataset) overrides `--sample-per-stratum N` / `dataset.sample_per_stratum` (stratified random N-per-stratum, keyed on `stratify_field`, default `expected_skill` — for classification suites like activation). Stratified sampling (whether the N-per-stratum count comes from the **CLI** `--sample-per-stratum` flag or **YAML** `dataset.sample_per_stratum`) is **nondeterministic** by default — it re-draws each run (so the nightly activation suite broadens coverage over time). Set `dataset.sample_seed` to pin a reproducible sample; an explicit seed always wins. (Only `--sample N` uses a fixed seed, since a smoke test wants the same N rows each run.)
142+
- **Dataset fan-out**: `TaskDefinition.dataset` (inline rows or JSONL path) expands a single task into N row-tasks with `${row.<field>}` substitution in `initial_prompt` and `success_criteria` string fields. Expansion runs in `task_loader.expand_dataset` **before** variant resolution, so variants cannot override the dataset. Row selection is filter-then-sample: CLI `--split <name>` (keep only rows whose `dataset.split_field` value matches, default field `split`) runs **first** and is orthogonal to the sampler win-order — a row is unlabelled when the field is absent/`null`/`""` and a task whose rows are all unlabelled passes through unfiltered; partial labelling drops the unlabelled rows. A *labelled* task with no matching row raises a `ValueError` listing the splits that exist — which `resolve_all_tasks` catches into `skipped_tasks` like any load failure, so a mistyped selector yields a zero-task run that still exits 0. Filtering before sampling is a correctness requirement: sampling first would leave an unpredictable (possibly zero) number of rows per split, destroying the tune/holdout comparison. Then sampling: CLI `--sample N` (fixed-seed uniform-random N over the whole dataset) overrides `--sample-per-stratum N` / `dataset.sample_per_stratum` (stratified random N-per-stratum, keyed on `stratify_field`, default `expected_skill` — for classification suites like activation). Stratified sampling (whether the N-per-stratum count comes from the **CLI** `--sample-per-stratum` flag or **YAML** `dataset.sample_per_stratum`) is **nondeterministic** by default — it re-draws each run (so the nightly activation suite broadens coverage over time). Set `dataset.sample_seed` to pin a reproducible sample; an explicit seed always wins. (Only `--sample N` uses a fixed seed, since a smoke test wants the same N rows each run.)
143143
- **Per-criterion aggregation**: Each `BaseCriterion` subclass exposes `aggregate(criterion, per_row_results) -> CriterionAggregate | None`. Default emits `count / mean / median / std / min / max` so every criterion is suite-thresholdable for free. Classification-style criteria return `ClassificationCriterionResult` (subclass of `CriterionResult`) and layer accuracy / P/R/F1 / confusion via the shared `overlay_classification_metrics` utility. `BaseSuccessCriterion.suite_thresholds` gates the suite on those metrics; CLI exits non-zero on any gate failure.
144144
- **Sub-agent token accounting**: There is NO separate per-sub-agent field. Every sub-agent generation is captured as a `parent_tool_use_id`-tagged `AssistantMessage` in the turn transcript, so per-sub-agent usage is derived by grouping those messages on that id (the evalboard's `aggregateSubAgentUsage` does exactly this). Claude bubbles its sub-agent's intermediate generations into the parent stream natively, and the **terminal** generation (delivered as the Agent tool result, never streamed) is synthesized into one via `_synthesize_subagent_terminal_message` from `tool_use_result.usage`. Codex reconstructs all child generations from the child rollout (`_recover_subagent_tool_calls`). The turn total already includes sub-agent cost — Claude via the SDK's cumulative `model_usage`; Codex via `_fold_subagent_tokens`, which folds the child messages (their real per-generation tokens) into the parent total. `CommandTelemetry.result_summary` is stored **untruncated** (no 200-char cap) so sub-agent returns are preserved whole. Set `CODER_EVAL_RAW_SDK_LOG=1` to dump every raw SDK event to the task log for inspection.
145145
- **Reconciliation message (stream self-reconciles to the turn total)**: The per-message stream consistently under-reports the authoritative turn total — a fixed prompt slice (~512 input tokens on Claude) is billed on no SDK-emitted message, and sub-agent input/cache only partially bubbles up. So `EventCollector.build_turn_record` appends one synthetic `ReconciliationMessage` (`role="reconciliation"`, in the `TranscriptMessage` union) per turn, carrying the per-bucket residual = `token_usage` − Σ(assistant message buckets). The invariant: **summing the four token buckets across `TurnRecord.messages` (assistant + reconciliation) equals `token_usage` exactly**, for both Claude and Codex (Codex's stream is already complete after `_recover_subagent_tool_calls`, so its residual is usually 0 and no entry is emitted). This is what lets the evalboard SUM the message stream as the source of truth instead of reading a separate aggregate ("agent tokens"): `selectTokenTotals` returns the stream sum whenever a reconciliation entry is present, and the timeline renders it as its own row. It is agent-agnostic (booked at the single `EventCollector` seam), carries no cost (cost stays on `token_usage`), and is excluded from generation/turn counts and the cost simulator. The LiteLLM open-weight actual-cost join (`litellm_cost.apply_actual_cost`) deliberately writes cost at the TURN level only (`token_usage.total_cost_usd` = the real OpenRouter bill) plus the per-call `TurnRecord.provider_call_costs` audit record; it does NOT touch the message token buckets, so `EventCollector` stays the single writer and this invariant holds on every backend. The Python `token_usage`/`total_token_usage` aggregate is unchanged and still authoritative for budget/judges/reports.

‎README.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -107,9 +107,9 @@ results — runs inside the agent:
107107
/plugin install coder-eval@coder-eval
108108
```
109109

110-
That adds six slash commands: `/coder-eval:init`, `/coder-eval:check-skill`,
111-
`/coder-eval:task`, `/coder-eval:lint-tasks`, `/coder-eval:analyze` and
112-
`/coder-eval:ci`. They drive the `coder-eval` CLI, so install it too
110+
That adds seven slash commands: `/coder-eval:init`, `/coder-eval:check-skill`,
111+
`/coder-eval:optimize-skill`, `/coder-eval:task`, `/coder-eval:lint-tasks`,
112+
`/coder-eval:analyze` and `/coder-eval:ci`. They drive the `coder-eval` CLI, so install it too
113113
(`uv tool install coder-eval`). See [Claude Code Plugin](docs/PLUGIN.md).
114114

115115
## Use as a GitHub Action

0 commit comments

Comments
 (0)