Skip to content

Commit af5e07c

Browse files
uipreligaclaude
andcommitted
docs: slim CLAUDE.md, and drop a directory that never existed
CLAUDE.md is loaded into every session, so a copy kept there costs context on every turn and goes stale with nothing to sense it. Three sections were copies. - **Directory Structure** was wrong. It listed `optimize/`, which exists on neither this branch nor main, and omitted `errors/` and `plugins.py`, both present since the initial public release. The tree was never the value — the annotations were. It becomes a list of what a filename cannot tell you (the CE048 twins, the CE063 seam, the CE057 sidecar, CE053's filename ownership, CE056's env var, the pricing.ts mirror), opening with "run `ls`". 454 -> 176w. - **Success Criteria** was a third copy of one registry, behind the CE033-generated plugin reference and the Task Definition Guide, with no parity test — the shape CE030's own CLAUDE.md test exists for. Accurate today; stale eventually. Now the shared-field paragraph plus three pointers. 259 -> 76w. - **Extension Points** restated docs/EXTENDING.md's checklists step for step while the section 190 lines above it says "Each entry is a pointer". Keeps the three orientation facts — pkgutil discovery, the plugin SPI with no closed enum, pricing on the same register hook — and the obligations (CE036, CE047, parity). 48 -> 21 lines. Also reverts the run-limit parity sentence claiming an adapter rejects an unsupported field at load time. It does not: opencode_agent.py:785 and pi_agent.py:738 both warn and continue. Neither lint-pinned surface is touched — the CE030 model sentence and the six skill names are unchanged — and nothing links into the cut sections. CLAUDE.md: 2,802 -> 2,028 words. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent cd11bb2 commit af5e07c

1 file changed

Lines changed: 49 additions & 132 deletions

File tree

‎CLAUDE.md‎

Lines changed: 49 additions & 132 deletions
Original file line numberDiff line numberDiff line change
@@ -26,69 +26,28 @@ data-driven analysis.
2626

2727
## Directory Structure
2828

29-
```
30-
coder_eval/
31-
├── agent.py # Agent ABC (start, communicate, stop, get_state)
32-
├── config.py # Settings via pydantic-settings (.env loading)
33-
├── sandbox.py # Sandbox manager (tempdir, venv, templates, adopt)
34-
├── orchestrator.py # Main evaluation loop
35-
├── reports.py # Markdown/JSON run reports + per-suite rollups
36-
├── reports_experiment.py # Cross-variant experiment reports
37-
├── reports_junit.py # JUnit XML from a finalized run dir (CI ingestion)
38-
├── reports_html.py # Single-file HTML report (the evalboard's static twin)
39-
├── reports_stats.py # Shared report statistics + ungraded rendering helpers
40-
├── formatting.py # Number/duration formatting shared by the renderers
41-
├── analysis.py # Command statistics aggregation
42-
├── logging_config.py # Structured logging setup
43-
├── path_utils.py # Run IDs, path utilities, atomic writes, tree digests
44-
├── fs_permissions.py # set_permissions: stacked chmod window
45-
├── pricing.py # Model pricing (mirrored by evalboard/lib/pricing.ts)
46-
├── litellm_cost.py # Join proxy-captured actual per-call cost onto turns
47-
├── timing.py # TurnClock + turn decomposition (single subtraction seam)
48-
├── invocation_log.py # record_cli recording shim + JSON Lines reader
49-
├── argv_match.py # Structured argv matcher (STDLIB-ONLY sidecar — CE057)
50-
├── telemetry.py # App Insights / OpenTelemetry emission
51-
├── isolation/ # driver: docker — one container per task
52-
├── harbor/ # Harbor export + coder-eval as a Harbor agent
53-
├── optimize/ # Prompt/config optimization helpers
54-
├── utils.py # Version info helpers
55-
│
56-
├── agents/ # Agent implementations (claude_code, codex, antigravity,
57-
│ # opencode, pi, noop) + registry, watchdog
58-
│
59-
├── models/ # Pure Pydantic data models (see __init__ for exports)
60-
│ ├── enums.py # AgentKind, AgentState, FinalStatus, ApiBackend
61-
│ ├── criteria.py # 15 success criterion types + base + union
62-
│ ├── experiment.py # ExperimentDefinition, ExperimentVariant, ResolvedTask
63-
│ ├── cli_match.py # FlagMatch + CliMatch (cycle-free leaf)
64-
│ ├── container_paths.py # IN_CONTAINER_ENV + container path constants (CE056)
65-
│ ├── mutations.py # PromptMutation variants
66-
│ ├── results.py # CriterionResult, TurnRecord, EvaluationResult, rollups
67-
│ ├── routing.py # ApiRoute (DirectRoute/BedrockRoute)
68-
│ ├── sandbox.py # SandboxConfig, ResourceLimits, RecordedCli, CliResponse
69-
│ ├── tasks.py # TaskDefinition, AgentConfig, Dataset, RunLimits
70-
│ ├── telemetry.py # CommandTelemetry, TokenUsage, TranscriptMessage
71-
│ └── templates.py # RepoSource, TemplateDirSource, StarterFilesSource
72-
│
73-
├── criteria/ # Criterion checker plugins (one file per type)
74-
│ ├── __init__.py # CriterionRegistry with auto-discovery
75-
│ └── base.py # BaseCriterion + @handle_criterion_errors
76-
│
77-
├── evaluation/ # checker.py (SuccessChecker), judge_context, judge_verdict,
78-
│ # sub_agent, summaries
79-
├── orchestration/ # batch, config, config_merge, early_stop, evaluation,
80-
│ # regrade, experiment, overrides, task_loader
81-
├── cli/ # Typer commands; each has a plain-Python twin (CE048)
82-
├── scoring/ # AST / token / signature / complexity / quality similarity
83-
├── streaming/ # Event protocol, EventCollector, renderers
84-
├── simulation/ # Multi-turn user simulation (dialog mode)
85-
└── resources/ # Package resources
86-
87-
experiments/ tasks/ tests/ docs/ templates/ evalboard/
88-
plugins/coder-eval/ # Published Claude Code plugin (six skills)
89-
action.yml # Published composite GitHub Action
90-
.claude-plugin/marketplace.json
91-
```
29+
`src/coder_eval/` — run `ls` for the current layout. What a filename does not tell you:
30+
31+
- **`models/`** is the pure-Pydantic layer: the dependency arrow runs `agents` →
32+
`models`, so it may reach `agents` / `plugins` only lazily (CE017). All core models
33+
import from `coder_eval.models`, never from its submodules.
34+
- **`criteria/`** auto-discovers one checker per type via `pkgutil`.
35+
- **`cli/`** holds Typer commands; each has a plain-Python twin (CE048).
36+
- **`timing.py`** owns the single subtraction seam (CE063).
37+
- **`argv_match.py`** is a STDLIB-ONLY sidecar copied beside the recorder (CE057).
38+
- **`fs_permissions.py`** is `set_permissions`, the stacked chmod window.
39+
- **`path_utils.py`** owns run ids, atomic writes and tree digests — and every run-record
40+
filename literal (CE053).
41+
- **`models/container_paths.py`** owns `IN_CONTAINER_ENV` (CE056).
42+
- **`pricing.py`** is hand-mirrored by `evalboard/lib/pricing.ts`; a parity test fails on
43+
drift either way.
44+
- **`reports_html.py`** is the evalboard's static twin.
45+
- **`isolation/`** is `driver: docker`, one container per task.
46+
- **`streaming/`** is the event protocol and `EventCollector`.
47+
48+
Outside the package: `tasks/`, `experiments/`, `templates/`, `tests/`, `docs/`,
49+
`evalboard/`, `plugins/coder-eval/` (the published plugin), `action.yml` (the published
50+
composite Action).
9251

9352
## Key Architectural Patterns
9453

@@ -120,8 +79,7 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
12079
Defense-in-depth, not a boundary — the known gaps are documented in the notes.
12180
Authoring reference: [Reference Solutions](docs/TASK_DEFINITION_GUIDE.md#reference-solutions).
12281
- **Harness run-limit parity**: a shared config field must mean the same thing on every
123-
backend, or the adapter rejects it at load time. A silently ignored field is a defect,
124-
not a table row. Table:
82+
backend, or the divergence is documented. Table:
12583
[Run-Limit Parity](docs/agents/HARNESS_PARITY.md). Caps are authored under
12684
[Run Limits](docs/TASK_DEFINITION_GUIDE.md#run-limits).
12785
- **Execute vs. run**: `execute` is `run` with grading off — rows finalize as
@@ -145,37 +103,20 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
145103
- **Dialog mode**: `simulation/` drives a multi-turn LLM user — see
146104
[Dialog Mode](docs/DIALOG_MODE.md).
147105

148-
## Success Criteria (15 types)
149-
150-
| Type | Scoring | Description |
151-
|------|---------|-------------|
152-
| [`file_exists`](docs/TASK_DEFINITION_GUIDE.md#file_exists) | Binary | File must exist |
153-
| [`file_contains`](docs/TASK_DEFINITION_GUIDE.md#file_contains) | Fractional | String presence/absence |
154-
| [`file_check`](docs/TASK_DEFINITION_GUIDE.md#file_check) | Fractional | Unified file existence + content + regex check |
155-
| [`json_check`](docs/TASK_DEFINITION_GUIDE.md#json_check) | Fractional | JSON validation + JSON Schema + JMESPath assertions |
156-
| [`run_command`](docs/TASK_DEFINITION_GUIDE.md#run_command) | Binary / Continuous | Exit code + optional stdout matching or float scoring |
157-
| [`file_matches_regex`](docs/TASK_DEFINITION_GUIDE.md#file_matches_regex) | Binary | Regex match on file |
158-
| [`reference_comparison`](docs/TASK_DEFINITION_GUIDE.md#reference_comparison) | Continuous | AST/token/complexity similarity |
159-
| [`command_executed`](docs/TASK_DEFINITION_GUIDE.md#command_executed) | Fractional | Agent tool usage verification |
160-
| [`cli_called`](docs/TASK_DEFINITION_GUIDE.md#cli_called) | Binary | Structured match over the `record_cli` invocation log |
161-
| [`commands_efficiency`](docs/TASK_DEFINITION_GUIDE.md#commands_efficiency) | Continuous | Tool-call efficiency against an expected budget |
162-
| [`uipath_eval`](docs/TASK_DEFINITION_GUIDE.md#uipath_eval) | Fractional | UiPath agent evaluation results |
163-
| [`classification_match`](docs/TASK_DEFINITION_GUIDE.md#classification_match) | Binary | File-based label match; emits suite-level P/R/F1 |
164-
| [`skill_triggered`](docs/TASK_DEFINITION_GUIDE.md#skill_triggered) | Binary | Did the agent engage the target skill? Agent-agnostic |
165-
| [`llm_judge`](docs/TASK_DEFINITION_GUIDE.md#llm_judge) | Continuous | LLM grades artifacts + optional trajectory/reference |
166-
| [`agent_judge`](docs/TASK_DEFINITION_GUIDE.md#agent_judge) | Continuous | Sandboxed SDK agent investigates with tools. Expensive |
106+
## Success Criteria
167107

168-
All criteria support `weight` (default 1.0) and `pass_threshold` (default 0.9). Live
169-
criteria also accept `stop_early:`. Dataset-backed tasks may set `suite_thresholds:`
170-
(see [Suite-level scoring](docs/DATASETS.md#suite-level-scoring)).
171-
172-
Each type above links to its own section in the
173-
[Task Definition Guide](docs/TASK_DEFINITION_GUIDE.md#success-criteria), which is the
174-
authoritative per-field reference. Path and env-var resolution inside a checker is
108+
Every criterion type registers in `criteria/` and is listed by
109+
`CriterionRegistry.list_types()`. The authoritative per-field reference is
110+
[Task Definition Guide § Success criteria](docs/TASK_DEFINITION_GUIDE.md#success-criteria);
111+
path and env-var resolution inside a checker is
175112
[Checker Context](docs/TASK_DEFINITION_GUIDE.md#checker-context). The plugin ships a
176113
generated copy at `plugins/coder-eval/reference/criteria.md` — regenerate it with
177114
`make plugin-reference`; never hand-edit it (CE033).
178115

116+
All criteria support `weight` (default 1.0) and `pass_threshold` (default 0.9). Live
117+
criteria also accept `stop_early:`. Dataset-backed tasks may set `suite_thresholds:`
118+
(see [Suite-level scoring](docs/DATASETS.md#suite-level-scoring)).
119+
179120
## Evaluation Flow
180121

181122
```
@@ -284,50 +225,25 @@ documentation: [Claude Code plugin](docs/PLUGIN.md).
284225

285226
## Extension Points
286227

287-
### Adding a New Criterion
288-
289-
1. Define the model in `models/criteria.py` inheriting `BaseSuccessCriterion`.
290-
2. Add it to the `SuccessCriterion` union.
291-
3. Create the checker in `criteria/` inheriting `BaseCriterion`, decorated with
292-
`@register_criterion` — auto-discovered at runtime.
293-
4. If it is a live criterion, add `ContractCase`s (CE036) and run `make plugin-reference`.
228+
**A new criterion**: model in `models/criteria.py` → the `SuccessCriterion` union →
229+
checker in `criteria/` decorated `@register_criterion`, auto-discovered via `pkgutil`.
230+
A live criterion also needs `ContractCase`s (CE036) and `make plugin-reference`.
294231

295-
Worked example: [Custom success criteria](docs/EXTENDING.md). Document the new type in
296-
the [Task Definition Guide](docs/TASK_DEFINITION_GUIDE.md#success-criteria).
297-
298-
### Adding a New Agent
299-
300-
Agents register through the plugin SPI (entry-point group `coder_eval.plugins`) — there
301-
is no closed enum or dispatch to edit. In-tree and third-party agents use the same path.
302-
Full walkthrough: [Extending Coder Eval](docs/EXTENDING.md). Per-agent setup and
303-
credentials: [Claude Code](docs/agents/CLAUDE_CODE.md), [Codex](docs/agents/CODEX.md),
304-
[Antigravity](docs/agents/ANTIGRAVITY.md), [OpenCode](docs/agents/OPENCODE.md),
305-
[Pi](docs/agents/PI.md). A new agent must also be added to every onboarding surface
306-
CE047 tracks, and its run-limit behaviour recorded in
232+
**A new agent**: agents register through the plugin SPI (entry-point group
233+
`coder_eval.plugins`) — there is no closed enum or dispatch to edit, and in-tree and
234+
third-party agents take the same path. A new agent must be named on every onboarding
235+
surface CE047 tracks, and its run-limit behaviour recorded in
307236
[Run-Limit Parity](docs/agents/HARNESS_PARITY.md).
308237

309-
1. Define a `BaseAgentConfig` subclass (its own `type: Literal["your-kind"]`) and
310-
implement the `Agent` ABC.
311-
2. Bind them with `registry.register("your-kind", YourConfig)(YourAgent)` inside a
312-
`register(registry)` hook exposed via a `coder_eval.plugins` entry point.
313-
3. Use the shared turn lifecycle on the base class — `self._begin_turn()`,
314-
`self._end_turn_ok()`, `self._mark_stopped()`. Do not reimplement it.
315-
4. Before raising on a mid-turn failure, set `self.pending_turn` to a `crashed=True`
316-
`TurnRecord`, then raise `AgentCrashError` or `TurnTimeoutError` bare.
317-
5. Emit the standardized event protocol and fan it through an internal `EventCollector`
318-
plus the caller's `stream_callback`. One `AgentStartEvent` and one matching
319-
`AgentEndEvent` on *every* exit path, from `finally`.
320-
6. If the agent shells out or holds OS resources, implement real `stop()` / `kill()` /
321-
`kill_sync()`. `kill_sync()` runs on a non-asyncio thread and must not await.
322-
323-
### Registering Model Pricing (plugins)
324-
325-
Call `register_pricing(YOUR_RATES)` from the same `register(registry)` hook — there is
326-
no separate entry-point group. Keys are bare model ids; vendor/Bedrock prefixes are
327-
normalized off at lookup. Registration is idempotent for identical rates and raises on a
328-
conflicting rate for an existing key, so plugin load order can never silently reprice a
329-
model. `coder_eval_uipath/pricing.py` is the worked example; see also
330-
[Model pricing](docs/EXTENDING.md).
238+
**Model pricing**: `register_pricing(YOUR_RATES)` from the same `register(registry)`
239+
hook — no separate entry-point group.
240+
241+
Steps, checklists and worked examples: [Extending Coder Eval](docs/EXTENDING.md).
242+
Document a new criterion type in the
243+
[Task Definition Guide](docs/TASK_DEFINITION_GUIDE.md#success-criteria). Per-agent setup
244+
and credentials: [Claude Code](docs/agents/CLAUDE_CODE.md) · [Codex](docs/agents/CODEX.md)
245+
· [Antigravity](docs/agents/ANTIGRAVITY.md) · [OpenCode](docs/agents/OPENCODE.md) ·
246+
[Pi](docs/agents/PI.md).
331247

332248
## Task Definition
333249

@@ -362,6 +278,7 @@ bandit, pre-commit, mcp
362278
- **YAGNI** — don't add complexity until actually needed
363279
- **KISS** — keep it simple
364280
- **Clean code** — no dead code, all imports used, all tests passing
281+
- **Greenfield project** — no backward-compatibility burden
365282
- **Delete before you guard** — before you add a lint rule, doc paragraph, criterion
366283
type or config field, try to delete the pattern that needs it. A new type that
367284
subsumes an old one removes the old one in the same change (no back-compat burden)

0 commit comments

Comments
 (0)