@@ -26,69 +26,28 @@ data-driven analysis.
2626
2727## Directory Structure
2828
29- ```
30- coder_eval/
31- ├── agent.py # Agent ABC (start, communicate, stop, get_state)
32- ├── config.py # Settings via pydantic-settings (.env loading)
33- ├── sandbox.py # Sandbox manager (tempdir, venv, templates, adopt)
34- ├── orchestrator.py # Main evaluation loop
35- ├── reports.py # Markdown/JSON run reports + per-suite rollups
36- ├── reports_experiment.py # Cross-variant experiment reports
37- ├── reports_junit.py # JUnit XML from a finalized run dir (CI ingestion)
38- ├── reports_html.py # Single-file HTML report (the evalboard's static twin)
39- ├── reports_stats.py # Shared report statistics + ungraded rendering helpers
40- ├── formatting.py # Number/duration formatting shared by the renderers
41- ├── analysis.py # Command statistics aggregation
42- ├── logging_config.py # Structured logging setup
43- ├── path_utils.py # Run IDs, path utilities, atomic writes, tree digests
44- ├── fs_permissions.py # set_permissions: stacked chmod window
45- ├── pricing.py # Model pricing (mirrored by evalboard/lib/pricing.ts)
46- ├── litellm_cost.py # Join proxy-captured actual per-call cost onto turns
47- ├── timing.py # TurnClock + turn decomposition (single subtraction seam)
48- ├── invocation_log.py # record_cli recording shim + JSON Lines reader
49- ├── argv_match.py # Structured argv matcher (STDLIB-ONLY sidecar — CE057)
50- ├── telemetry.py # App Insights / OpenTelemetry emission
51- ├── isolation/ # driver: docker — one container per task
52- ├── harbor/ # Harbor export + coder-eval as a Harbor agent
53- ├── optimize/ # Prompt/config optimization helpers
54- ├── utils.py # Version info helpers
55- │
56- ├── agents/ # Agent implementations (claude_code, codex, antigravity,
57- │ # opencode, pi, noop) + registry, watchdog
58- │
59- ├── models/ # Pure Pydantic data models (see __init__ for exports)
60- │ ├── enums.py # AgentKind, AgentState, FinalStatus, ApiBackend
61- │ ├── criteria.py # 15 success criterion types + base + union
62- │ ├── experiment.py # ExperimentDefinition, ExperimentVariant, ResolvedTask
63- │ ├── cli_match.py # FlagMatch + CliMatch (cycle-free leaf)
64- │ ├── container_paths.py # IN_CONTAINER_ENV + container path constants (CE056)
65- │ ├── mutations.py # PromptMutation variants
66- │ ├── results.py # CriterionResult, TurnRecord, EvaluationResult, rollups
67- │ ├── routing.py # ApiRoute (DirectRoute/BedrockRoute)
68- │ ├── sandbox.py # SandboxConfig, ResourceLimits, RecordedCli, CliResponse
69- │ ├── tasks.py # TaskDefinition, AgentConfig, Dataset, RunLimits
70- │ ├── telemetry.py # CommandTelemetry, TokenUsage, TranscriptMessage
71- │ └── templates.py # RepoSource, TemplateDirSource, StarterFilesSource
72- │
73- ├── criteria/ # Criterion checker plugins (one file per type)
74- │ ├── __init__.py # CriterionRegistry with auto-discovery
75- │ └── base.py # BaseCriterion + @handle_criterion_errors
76- │
77- ├── evaluation/ # checker.py (SuccessChecker), judge_context, judge_verdict,
78- │ # sub_agent, summaries
79- ├── orchestration/ # batch, config, config_merge, early_stop, evaluation,
80- │ # regrade, experiment, overrides, task_loader
81- ├── cli/ # Typer commands; each has a plain-Python twin (CE048)
82- ├── scoring/ # AST / token / signature / complexity / quality similarity
83- ├── streaming/ # Event protocol, EventCollector, renderers
84- ├── simulation/ # Multi-turn user simulation (dialog mode)
85- └── resources/ # Package resources
86-
87- experiments/ tasks/ tests/ docs/ templates/ evalboard/
88- plugins/coder-eval/ # Published Claude Code plugin (six skills)
89- action.yml # Published composite GitHub Action
90- .claude-plugin/marketplace.json
91- ```
29+ ` src/coder_eval/ ` — run ` ls ` for the current layout. What a filename does not tell you:
30+
31+ - ** ` models/ ` ** is the pure-Pydantic layer: the dependency arrow runs ` agents ` →
32+ ` models ` , so it may reach ` agents ` / ` plugins ` only lazily (CE017). All core models
33+ import from ` coder_eval.models ` , never from its submodules.
34+ - ** ` criteria/ ` ** auto-discovers one checker per type via ` pkgutil ` .
35+ - ** ` cli/ ` ** holds Typer commands; each has a plain-Python twin (CE048).
36+ - ** ` timing.py ` ** owns the single subtraction seam (CE063).
37+ - ** ` argv_match.py ` ** is a STDLIB-ONLY sidecar copied beside the recorder (CE057).
38+ - ** ` fs_permissions.py ` ** is ` set_permissions ` , the stacked chmod window.
39+ - ** ` path_utils.py ` ** owns run ids, atomic writes and tree digests — and every run-record
40+ filename literal (CE053).
41+ - ** ` models/container_paths.py ` ** owns ` IN_CONTAINER_ENV ` (CE056).
42+ - ** ` pricing.py ` ** is hand-mirrored by ` evalboard/lib/pricing.ts ` ; a parity test fails on
43+ drift either way.
44+ - ** ` reports_html.py ` ** is the evalboard's static twin.
45+ - ** ` isolation/ ` ** is ` driver: docker ` , one container per task.
46+ - ** ` streaming/ ` ** is the event protocol and ` EventCollector ` .
47+
48+ Outside the package: ` tasks/ ` , ` experiments/ ` , ` templates/ ` , ` tests/ ` , ` docs/ ` ,
49+ ` evalboard/ ` , ` plugins/coder-eval/ ` (the published plugin), ` action.yml ` (the published
50+ composite Action).
9251
9352## Key Architectural Patterns
9453
@@ -120,8 +79,7 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
12079 Defense-in-depth, not a boundary — the known gaps are documented in the notes.
12180 Authoring reference: [ Reference Solutions] ( docs/TASK_DEFINITION_GUIDE.md#reference-solutions ) .
12281- ** Harness run-limit parity** : a shared config field must mean the same thing on every
123- backend, or the adapter rejects it at load time. A silently ignored field is a defect,
124- not a table row. Table:
82+ backend, or the divergence is documented. Table:
12583 [ Run-Limit Parity] ( docs/agents/HARNESS_PARITY.md ) . Caps are authored under
12684 [ Run Limits] ( docs/TASK_DEFINITION_GUIDE.md#run-limits ) .
12785- ** Execute vs. run** : ` execute ` is ` run ` with grading off — rows finalize as
@@ -145,37 +103,20 @@ Each entry is a pointer. Full rationale: `.claude/notes/` (index: `.claude/notes
145103- ** Dialog mode** : ` simulation/ ` drives a multi-turn LLM user — see
146104 [ Dialog Mode] ( docs/DIALOG_MODE.md ) .
147105
148- ## Success Criteria (15 types)
149-
150- | Type | Scoring | Description |
151- | ------| ---------| -------------|
152- | [ ` file_exists ` ] ( docs/TASK_DEFINITION_GUIDE.md#file_exists ) | Binary | File must exist |
153- | [ ` file_contains ` ] ( docs/TASK_DEFINITION_GUIDE.md#file_contains ) | Fractional | String presence/absence |
154- | [ ` file_check ` ] ( docs/TASK_DEFINITION_GUIDE.md#file_check ) | Fractional | Unified file existence + content + regex check |
155- | [ ` json_check ` ] ( docs/TASK_DEFINITION_GUIDE.md#json_check ) | Fractional | JSON validation + JSON Schema + JMESPath assertions |
156- | [ ` run_command ` ] ( docs/TASK_DEFINITION_GUIDE.md#run_command ) | Binary / Continuous | Exit code + optional stdout matching or float scoring |
157- | [ ` file_matches_regex ` ] ( docs/TASK_DEFINITION_GUIDE.md#file_matches_regex ) | Binary | Regex match on file |
158- | [ ` reference_comparison ` ] ( docs/TASK_DEFINITION_GUIDE.md#reference_comparison ) | Continuous | AST/token/complexity similarity |
159- | [ ` command_executed ` ] ( docs/TASK_DEFINITION_GUIDE.md#command_executed ) | Fractional | Agent tool usage verification |
160- | [ ` cli_called ` ] ( docs/TASK_DEFINITION_GUIDE.md#cli_called ) | Binary | Structured match over the ` record_cli ` invocation log |
161- | [ ` commands_efficiency ` ] ( docs/TASK_DEFINITION_GUIDE.md#commands_efficiency ) | Continuous | Tool-call efficiency against an expected budget |
162- | [ ` uipath_eval ` ] ( docs/TASK_DEFINITION_GUIDE.md#uipath_eval ) | Fractional | UiPath agent evaluation results |
163- | [ ` classification_match ` ] ( docs/TASK_DEFINITION_GUIDE.md#classification_match ) | Binary | File-based label match; emits suite-level P/R/F1 |
164- | [ ` skill_triggered ` ] ( docs/TASK_DEFINITION_GUIDE.md#skill_triggered ) | Binary | Did the agent engage the target skill? Agent-agnostic |
165- | [ ` llm_judge ` ] ( docs/TASK_DEFINITION_GUIDE.md#llm_judge ) | Continuous | LLM grades artifacts + optional trajectory/reference |
166- | [ ` agent_judge ` ] ( docs/TASK_DEFINITION_GUIDE.md#agent_judge ) | Continuous | Sandboxed SDK agent investigates with tools. Expensive |
106+ ## Success Criteria
167107
168- All criteria support ` weight ` (default 1.0) and ` pass_threshold ` (default 0.9). Live
169- criteria also accept ` stop_early: ` . Dataset-backed tasks may set ` suite_thresholds: `
170- (see [ Suite-level scoring] ( docs/DATASETS.md#suite-level-scoring ) ).
171-
172- Each type above links to its own section in the
173- [ Task Definition Guide] ( docs/TASK_DEFINITION_GUIDE.md#success-criteria ) , which is the
174- authoritative per-field reference. Path and env-var resolution inside a checker is
108+ Every criterion type registers in ` criteria/ ` and is listed by
109+ ` CriterionRegistry.list_types() ` . The authoritative per-field reference is
110+ [ Task Definition Guide § Success criteria] ( docs/TASK_DEFINITION_GUIDE.md#success-criteria ) ;
111+ path and env-var resolution inside a checker is
175112[ Checker Context] ( docs/TASK_DEFINITION_GUIDE.md#checker-context ) . The plugin ships a
176113generated copy at ` plugins/coder-eval/reference/criteria.md ` — regenerate it with
177114` make plugin-reference ` ; never hand-edit it (CE033).
178115
116+ All criteria support ` weight ` (default 1.0) and ` pass_threshold ` (default 0.9). Live
117+ criteria also accept ` stop_early: ` . Dataset-backed tasks may set ` suite_thresholds: `
118+ (see [ Suite-level scoring] ( docs/DATASETS.md#suite-level-scoring ) ).
119+
179120## Evaluation Flow
180121
181122```
@@ -284,50 +225,25 @@ documentation: [Claude Code plugin](docs/PLUGIN.md).
284225
285226## Extension Points
286227
287- ### Adding a New Criterion
288-
289- 1 . Define the model in ` models/criteria.py ` inheriting ` BaseSuccessCriterion ` .
290- 2 . Add it to the ` SuccessCriterion ` union.
291- 3 . Create the checker in ` criteria/ ` inheriting ` BaseCriterion ` , decorated with
292- ` @register_criterion ` — auto-discovered at runtime.
293- 4 . If it is a live criterion, add ` ContractCase ` s (CE036) and run ` make plugin-reference ` .
228+ ** A new criterion** : model in ` models/criteria.py ` → the ` SuccessCriterion ` union →
229+ checker in ` criteria/ ` decorated ` @register_criterion ` , auto-discovered via ` pkgutil ` .
230+ A live criterion also needs ` ContractCase ` s (CE036) and ` make plugin-reference ` .
294231
295- Worked example: [ Custom success criteria] ( docs/EXTENDING.md ) . Document the new type in
296- the [ Task Definition Guide] ( docs/TASK_DEFINITION_GUIDE.md#success-criteria ) .
297-
298- ### Adding a New Agent
299-
300- Agents register through the plugin SPI (entry-point group ` coder_eval.plugins ` ) — there
301- is no closed enum or dispatch to edit. In-tree and third-party agents use the same path.
302- Full walkthrough: [ Extending Coder Eval] ( docs/EXTENDING.md ) . Per-agent setup and
303- credentials: [ Claude Code] ( docs/agents/CLAUDE_CODE.md ) , [ Codex] ( docs/agents/CODEX.md ) ,
304- [ Antigravity] ( docs/agents/ANTIGRAVITY.md ) , [ OpenCode] ( docs/agents/OPENCODE.md ) ,
305- [ Pi] ( docs/agents/PI.md ) . A new agent must also be added to every onboarding surface
306- CE047 tracks, and its run-limit behaviour recorded in
232+ ** A new agent** : agents register through the plugin SPI (entry-point group
233+ ` coder_eval.plugins ` ) — there is no closed enum or dispatch to edit, and in-tree and
234+ third-party agents take the same path. A new agent must be named on every onboarding
235+ surface CE047 tracks, and its run-limit behaviour recorded in
307236[ Run-Limit Parity] ( docs/agents/HARNESS_PARITY.md ) .
308237
309- 1 . Define a ` BaseAgentConfig ` subclass (its own ` type: Literal["your-kind"] ` ) and
310- implement the ` Agent ` ABC.
311- 2 . Bind them with ` registry.register("your-kind", YourConfig)(YourAgent) ` inside a
312- ` register(registry) ` hook exposed via a ` coder_eval.plugins ` entry point.
313- 3 . Use the shared turn lifecycle on the base class — ` self._begin_turn() ` ,
314- ` self._end_turn_ok() ` , ` self._mark_stopped() ` . Do not reimplement it.
315- 4 . Before raising on a mid-turn failure, set ` self.pending_turn ` to a ` crashed=True `
316- ` TurnRecord ` , then raise ` AgentCrashError ` or ` TurnTimeoutError ` bare.
317- 5 . Emit the standardized event protocol and fan it through an internal ` EventCollector `
318- plus the caller's ` stream_callback ` . One ` AgentStartEvent ` and one matching
319- ` AgentEndEvent ` on * every* exit path, from ` finally ` .
320- 6 . If the agent shells out or holds OS resources, implement real ` stop() ` / ` kill() ` /
321- ` kill_sync() ` . ` kill_sync() ` runs on a non-asyncio thread and must not await.
322-
323- ### Registering Model Pricing (plugins)
324-
325- Call ` register_pricing(YOUR_RATES) ` from the same ` register(registry) ` hook — there is
326- no separate entry-point group. Keys are bare model ids; vendor/Bedrock prefixes are
327- normalized off at lookup. Registration is idempotent for identical rates and raises on a
328- conflicting rate for an existing key, so plugin load order can never silently reprice a
329- model. ` coder_eval_uipath/pricing.py ` is the worked example; see also
330- [ Model pricing] ( docs/EXTENDING.md ) .
238+ ** Model pricing** : ` register_pricing(YOUR_RATES) ` from the same ` register(registry) `
239+ hook — no separate entry-point group.
240+
241+ Steps, checklists and worked examples: [ Extending Coder Eval] ( docs/EXTENDING.md ) .
242+ Document a new criterion type in the
243+ [ Task Definition Guide] ( docs/TASK_DEFINITION_GUIDE.md#success-criteria ) . Per-agent setup
244+ and credentials: [ Claude Code] ( docs/agents/CLAUDE_CODE.md ) · [ Codex] ( docs/agents/CODEX.md )
245+ · [ Antigravity] ( docs/agents/ANTIGRAVITY.md ) · [ OpenCode] ( docs/agents/OPENCODE.md ) ·
246+ [ Pi] ( docs/agents/PI.md ) .
331247
332248## Task Definition
333249
@@ -362,6 +278,7 @@ bandit, pre-commit, mcp
362278- ** YAGNI** — don't add complexity until actually needed
363279- ** KISS** — keep it simple
364280- ** Clean code** — no dead code, all imports used, all tests passing
281+ - ** Greenfield project** — no backward-compatibility burden
365282- ** Delete before you guard** — before you add a lint rule, doc paragraph, criterion
366283 type or config field, try to delete the pattern that needs it. A new type that
367284 subsumes an old one removes the old one in the same change (no back-compat burden)
0 commit comments