Skip to content

Commit c46c241

Browse files
akshayliveclaude
andcommitted
fix(routing): reinstate litellm+agent_judge rejection guard
Removing the old _reject_litellm_eval_route_if_unsupported() guard wholesale (rather than narrowing it) correctly fixed the simulation.enabled false-positive but also silently reopened the agent_judge half of the same check: a task combining checker_context.api_route.route: litellm with an enabled agent_judge criterion now runs to completion instead of failing loudly, misrouting the judge sub-agent onto the harness's own ambient LiteLLM proxy credentials. Restore a narrow, agent_judge-only guard and update the tests/docs that had locked in the unguarded behavior. Also collapses the redundant duplicate resolve_route(settings) call for simulator_route and fixes two stale doc/comment spots left over from the decoupling. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
1 parent a9ae665 commit c46c241

3 files changed

Lines changed: 55 additions & 20 deletions

File tree

‎docs/TASK_DEFINITION_GUIDE.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -64,7 +64,7 @@ reference: { ... } # Optional reference solution (a directory
6464
pre_run: [ ... ] # Optional pre-run commands (before agent starts)
6565
post_run: [ ... ] # Optional post-run commands
6666
dataset: { ... } # Optional dataset fan-out (one task -> N row-tasks)
67-
checker_context: { ... } # Optional: backend/model for the evaluation side (judge, simulator)
67+
checker_context: { ... } # Optional: backend/model for the judge side (llm_judge/agent_judge)
6868
```
6969
7070
### `dataset`
@@ -1312,7 +1312,7 @@ checker_context:
13121312
model: claude-haiku-4-5 # model override for that route
13131313
```
13141314

1315-
- `route` selects which backend `llm_judge` and `agent_judge` call — they share one resolved eval route per run for `direct`/`bedrock`/`litellm` (this isn't per-criterion) — **decoupled from the agent's own route** (a Claude agent can be graded by a differently-backed judge, or vice versa). The simulator is NOT part of this sharing: it always resolves its own route the same way the agent under test does (subprocess-safe by construction), independent of `checker_context.api_route` entirely — see the `route: litellm` note below. This is a backend *name* (`direct` / `bedrock` / `litellm`), not a route object. For `direct`/`bedrock` credentials are never read from the task, always from the matching environment variables (`ANTHROPIC_API_KEY` for `direct`, `AWS_BEARER_TOKEN_BEDROCK`/`AWS_REGION` for `bedrock`). An unconfigured or unknown backend name raises at dispatch rather than silently falling back. **`route: litellm` dispatches `llm_judge` ONLY** through the `litellm` library (the `coder-eval[litellm]` extra, `litellm.acompletion`) rather than assuming one wire protocol — a gateway-routed judge model (e.g. an Azure AI `/openai/v1` deployment) rarely speaks Anthropic Messages, so this lets `model` carry its own provider hint (e.g. `azure/gpt-5.6-luna`) and get that provider's actual request/response shape handled by the library. `agent_judge` spawns a real Claude Code CLI subprocess that only speaks Anthropic Messages — combining `route: litellm` with an enabled `agent_judge` criterion is not yet guarded against (tracked as a follow-up); the simulator is unaffected since it never reads this override. Unlike the other two backends, `route: litellm` has NO implicit env-var fallback — see `params`/`env_params` below, which is how it's configured. `model` is required for `route: litellm` (there is no default open-weight/gateway model).
1315+
- `route` selects which backend `llm_judge` and `agent_judge` call — they share one resolved eval route per run for `direct`/`bedrock`/`litellm` (this isn't per-criterion) — **decoupled from the agent's own route** (a Claude agent can be graded by a differently-backed judge, or vice versa). The simulator is NOT part of this sharing: it always resolves its own route the same way the agent under test does (subprocess-safe by construction), independent of `checker_context.api_route` entirely — see the `route: litellm` note below. This is a backend *name* (`direct` / `bedrock` / `litellm`), not a route object. For `direct`/`bedrock` credentials are never read from the task, always from the matching environment variables (`ANTHROPIC_API_KEY` for `direct`, `AWS_BEARER_TOKEN_BEDROCK`/`AWS_REGION` for `bedrock`). An unconfigured or unknown backend name raises at dispatch rather than silently falling back. **`route: litellm` dispatches `llm_judge` ONLY** through the `litellm` library (the `coder-eval[litellm]` extra, `litellm.acompletion`) rather than assuming one wire protocol — a gateway-routed judge model (e.g. an Azure AI `/openai/v1` deployment) rarely speaks Anthropic Messages, so this lets `model` carry its own provider hint (e.g. `azure/gpt-5.6-luna`) and get that provider's actual request/response shape handled by the library. `agent_judge` spawns a real Claude Code CLI subprocess that only speaks Anthropic Messages — combining `route: litellm` with an enabled `agent_judge` criterion is rejected at resolution time (see below); the simulator is unaffected since it never reads this override. Unlike the other two backends, `route: litellm` has NO implicit env-var fallback — see `params`/`env_params` below, which is how it's configured. `model` is required for `route: litellm` (there is no default open-weight/gateway model).
13161316
- `model` overrides the model that resolved route uses for **`llm_judge` only** — when the criterion itself leaves `model:` unset (precedence: an explicit per-criterion `model:` always wins; below that, `checker_context.api_route.model`; below that, the built-in `DEFAULT_JUDGE_MODEL`). This floor is deliberate and never the agent's own model — an unpinned judge must grade identically regardless of which model the agent under test is using, so `resolve_evaluation_route` never lets the agent's env-configured model (e.g. `BEDROCK_MODEL`) leak into `route.model` on its own; `route.model` is set only when this override was actually given. This works because every `ApiRoute` (`DirectRoute`/`BedrockRoute`/`LiteLLMRoute`) carries its own `model` field; the orchestrator bakes the override into the resolved route's `model` before any criterion runs, so `llm_judge` just reads `context.route.model` — it never reads `checker_context` directly. **`agent_judge` and the simulator do not honor this override** — `agent_judge`'s sub-agent model comes from the criterion's own `agent:` block (defaulted to a fixed judge model), and the simulator's model is pinned by `SimulationConfig.model` (see [Simulation](#simulation) below) — both independent of `checker_context.api_route.model` by design, for the same "measuring instrument stays fixed" reason.
13171317
- `params`/`env_params` (**`route: litellm` only**) are how the call is actually configured — there is no fallback to the agent's own `LITELLM_BASE_URL`/`LITELLM_AUTH_TOKEN` env vars, since a gateway-routed judge model rarely reuses the agent's own LiteLLM proxy/credential. They also cover any of the dozens of other provider-specific kwargs `litellm.acompletion` accepts (`aws_access_key_id`, `vertex_project`, `api_version`, ...), which have no dedicated field on `LiteLLMRoute`:
13181318
```yaml
@@ -1327,7 +1327,7 @@ checker_context:
13271327
api_key: LITELLM_AUTH_TOKEN
13281328
```
13291329
`params` is merged straight into the `litellm.acompletion(**kwargs)` call — litellm validates the param names itself, so there's no allowlist to keep in sync here. `env_params` maps a kwarg name to the *name* of an environment variable; the value is resolved right before the call, so no secret is ever written into task/experiment YAML — this is how an arbitrary provider's config, including secrets (IAM keys, an Azure AD token, a service-account path, ...), is representable without a dedicated field per provider. `env_params` is resolved after `params`, so it always wins for the same key. Rejected at task-load time if given without `route: litellm`. For an Azure deployment, pin `api_version` via `params` to whatever API version the agent side is actually configured for (e.g. Codex's `CODEX_API_VERSION`) — the judge has no way to inherit it, and a mismatched version can hit a different shape of the same endpoint.
1330-
**`route: litellm` has no bearing on the simulator** — the simulator is a real Claude Code CLI subprocess that speaks the Anthropic Messages protocol, so pointing it at an arbitrary litellm-fronted gateway (which may speak an entirely different wire protocol) isn't representable. Rather than rejecting a task that combines `route: litellm` with `simulation.enabled: true`, the simulator always resolves its own route independently of `checker_context.api_route` — the same way the agent under test resolves its route — so `llm_judge` grades through the gateway named here while the simulator behaves exactly as if this override weren't set. A task can freely combine `route: litellm` with `simulation.enabled: true`; there is nothing to disable or remove. **`agent_judge` is not yet decoupled the same way** — it still shares the `llm_judge` eval route, so combining `route: litellm` with an enabled `agent_judge` criterion is currently unsupported (use `route: bedrock`/`direct` instead if a task needs both).
1330+
**`route: litellm` has no bearing on the simulator** — the simulator is a real Claude Code CLI subprocess that speaks the Anthropic Messages protocol, so pointing it at an arbitrary litellm-fronted gateway (which may speak an entirely different wire protocol) isn't representable. Rather than rejecting a task that combines `route: litellm` with `simulation.enabled: true`, the simulator always resolves its own route independently of `checker_context.api_route` — the same way the agent under test resolves its route — so `llm_judge` grades through the gateway named here while the simulator behaves exactly as if this override weren't set. A task can freely combine `route: litellm` with `simulation.enabled: true`; there is nothing to disable or remove. **`agent_judge` is NOT decoupled the same way** — it still shares the `llm_judge` eval route, and `agent_judge`'s Claude Code CLI subprocess has no way to honor a `LiteLLMRoute` safely, so the orchestrator rejects `route: litellm` combined with an enabled `agent_judge` criterion at resolution time (a clear error, not a silent misroute) — use `route: bedrock`/`direct` instead if a task needs both.
13311331

13321332
`checker_context` merges shallow-per-namespace across `default_experiment.defaults.checker_context` → `experiment.defaults.checker_context` → `task.checker_context` → `variant.checker_context` (same 4-layer precedence as `agent`/`simulation`). So a judge-model A/B, or a judge-backend A/B, is a variant-level config change, not an edit to every task YAML.
13331333

@@ -1613,7 +1613,7 @@ simulation:
16131613
| `check_criteria` | `end_of_dialog` | `end_of_dialog`, `every_turn`, or `both`. |
16141614
| `model` | `anthropic.claude-sonnet-4-6` | Model that plays the simulated user. Auto-translated to the run's backend (Bedrock inference profile / bare Anthropic alias), the same way [`llm_judge`](#llm_judge)'s `model` is. |
16151615

1616-
The simulator runs as a tools-disabled Claude Code agent on the run's resolved *evaluation* `ApiRoute` (the coding agent's own route unless [`checker_context.api_route.route`](#checker-context) overrides it), so temperature and sampling are resolved at the route level (same `-b` flag as the coding agent by default) and are not configured on this block. The **model is not**: it is pinned by `model` above. Inheriting it from the route meant `BEDROCK_MODEL` decided who the simulated user was, so an A/B varying the subject model silently varied its interlocutor too. Hold `model` fixed across variants for the same reason you hold a judge model fixed — the simulator is part of the measuring instrument, not the thing being measured.
1616+
The simulator runs as a tools-disabled Claude Code agent on its own resolved `ApiRoute` (the same resolution as the coding agent's own route — [`checker_context.api_route.route`](#checker-context) has no bearing on it), so temperature and sampling are resolved at the route level (same `-b` flag as the coding agent by default) and are not configured on this block. The **model is not**: it is pinned by `model` above. Inheriting it from the route meant `BEDROCK_MODEL` decided who the simulated user was, so an A/B varying the subject model silently varied its interlocutor too. Hold `model` fixed across variants for the same reason you hold a judge model fixed — the simulator is part of the measuring instrument, not the thing being measured.
16171617

16181618
**Semantics:**
16191619

‎src/coder_eval/orchestrator.py‎

Lines changed: 38 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -35,6 +35,7 @@
3535
CONTAINER_REFERENCE_DIR,
3636
DEFAULT_STOP_EARLY_GATE_THRESHOLD,
3737
ROUTE_NAMES,
38+
AgentJudgeCriterion,
3839
AgentKind,
3940
ApiRoute,
4041
BedrockRoute,
@@ -1500,7 +1501,10 @@ def _resolve_routes(self) -> None:
15001501
assert self.sandbox is not None
15011502
self.route = resolve_route(settings)
15021503
overrides = self._eval_route_overrides()
1503-
self.simulator_route = resolve_route(settings)
1504+
# Same resolution as self.route (subprocess-safe by construction) — not
1505+
# yet independently overridable, so this is an alias rather than a
1506+
# second call, until simulator_route grows its own override mechanism.
1507+
self.simulator_route = self.route
15041508
self.eval_route = resolve_evaluation_route(
15051509
settings,
15061510
self.route,
@@ -1509,9 +1513,37 @@ def _resolve_routes(self) -> None:
15091513
params_override=overrides.params,
15101514
env_params_override=overrides.env_params,
15111515
)
1516+
self._reject_litellm_agent_judge_if_unsupported()
15121517
logger.info("API routing: %s", _format_routing(self.route, self.task.agent.model if self.task.agent else None))
15131518
self.success_checker = SuccessChecker(self.sandbox, route=self.eval_route)
15141519

1520+
def _reject_litellm_agent_judge_if_unsupported(self) -> None:
1521+
"""``checker_context.api_route.route: litellm`` dispatches ``llm_judge``
1522+
through the ``litellm`` library (protocol-agnostic — see
1523+
``invoke_litellm_judge_async``'s module docstring), but ``agent_judge``
1524+
runs as a real Claude Code CLI subprocess that speaks the Anthropic
1525+
Messages protocol only. Handing it a ``LiteLLMRoute`` built from
1526+
arbitrary ``params``/``env_params`` would silently misroute it onto the
1527+
AGENT's own unrelated ambient LiteLLM proxy settings (``ClaudeCodeAgent``
1528+
ignores ``LiteLLMRoute.params``/``env_params`` entirely and falls back
1529+
to ``settings.litellm_base_url``/``litellm_auth_token``) rather than
1530+
raising — the exact misrouting ``LiteLLMRoute``'s own docstring says
1531+
must never happen. Reject the combination loudly at resolution time
1532+
instead. The simulator doesn't need this check: ``simulator_route``
1533+
never derives from ``checker_context.api_route`` in the first place.
1534+
"""
1535+
if not isinstance(self.eval_route, LiteLLMRoute):
1536+
return
1537+
if any(isinstance(c, AgentJudgeCriterion) and c.enabled for c in self.task.success_criteria):
1538+
msg = (
1539+
"checker_context.api_route.route: litellm is llm_judge-only (it dispatches through the "
1540+
"litellm library in-process, not a real Claude Code subprocess), but this task also has "
1541+
"an enabled agent_judge criterion, which runs as a Claude Code sub-agent requiring an "
1542+
"Anthropic-compatible endpoint. Use route: bedrock/direct instead, or remove/disable the "
1543+
"agent_judge criterion."
1544+
)
1545+
raise ValueError(msg)
1546+
15151547
def _record_route_environment_info(self) -> None:
15161548
"""Persist resolved route + judge transport into ``result.environment_info``.
15171549
@@ -1522,10 +1554,11 @@ def _record_route_environment_info(self) -> None:
15221554
assert self.result is not None
15231555
assert self.route is not None
15241556
self.result.environment_info["api_routing"] = ROUTE_NAMES[type(self.route)]
1525-
# The evaluation side (llm_judge / agent_judge / simulated user) may run on
1526-
# a different, constant backend — pinned to Claude when the agent is on
1527-
# LiteLLM — so record it: a run then shows what actually graded/simulated
1528-
# it, distinct from the agent's api_routing.
1557+
# The evaluation side (llm_judge / agent_judge) may run on a different,
1558+
# constant backend — pinned to Claude when the agent is on LiteLLM — so
1559+
# record it: a run then shows what actually graded it, distinct from the
1560+
# agent's api_routing. The simulator is NOT part of this: simulator_route
1561+
# always mirrors self.route, so it never diverges from api_routing above.
15291562
if self.eval_route is not None:
15301563
self.result.environment_info["eval_routing"] = ROUTE_NAMES[type(self.eval_route)]
15311564
# bedrock_model/litellm_model below are sourced from self.route (the

‎tests/test_orchestrator.py‎

Lines changed: 13 additions & 11 deletions
Original file line numberDiff line numberDiff line change
@@ -120,9 +120,9 @@ class TestSimulatorRouteDecoupledFromCheckerContext:
120120
simulation.enabled, so an experiment-wide litellm default broke all of
121121
them at setup.
122122
123-
agent_judge is intentionally NOT covered here — it still shares
124-
eval_route like before (out of scope for this pass; see the follow-up
125-
discussion on PR #2864)."""
123+
agent_judge is NOT decoupled the same way — it still shares eval_route,
124+
so the litellm+agent_judge combination is still rejected outright, exactly
125+
like before this pass (see test_agent_judge_rejects_litellm_eval_route)."""
126126

127127
@staticmethod
128128
def _orchestrator(tmp_path: Path, success_criteria, simulation=None, api_route=None) -> Orchestrator:
@@ -175,20 +175,22 @@ def test_enabled_simulation_never_gets_litellm(self, tmp_path, monkeypatch):
175175
assert isinstance(orchestrator.eval_route, LiteLLMRoute)
176176
assert not isinstance(orchestrator.simulator_route, LiteLLMRoute)
177177

178-
def test_agent_judge_still_shares_eval_route(self, tmp_path, monkeypatch):
179-
"""Documents current scope: agent_judge is NOT decoupled in this pass —
180-
it still reads success_checker.route (== eval_route), which can be a
181-
LiteLLMRoute if checker_context.api_route.route: litellm is set."""
178+
def test_agent_judge_rejects_litellm_eval_route(self, tmp_path, monkeypatch):
179+
"""agent_judge is NOT decoupled from eval_route the way the simulator
180+
is — it still reads success_checker.route (== eval_route), which would
181+
be an unsupported LiteLLMRoute if checker_context.api_route.route:
182+
litellm is set. Rather than silently sharing it (the previous gap),
183+
_resolve_routes rejects the combination loudly, same as before this
184+
pass — this closes the gap the simulator-decoupling PR temporarily left
185+
open (see _reject_litellm_agent_judge_if_unsupported)."""
182186
from coder_eval.models import AgentJudgeCriterion
183187

184188
orchestrator = self._orchestrator(
185189
tmp_path, [AgentJudgeCriterion(description="x", prompt="grade")], api_route=self._litellm_api_route()
186190
)
187191
monkeypatch.setattr(orchestrator_module, "resolve_route", lambda _s: DirectRoute(judge_transport="anthropic"))
188-
orchestrator._resolve_routes()
189-
assert isinstance(orchestrator.eval_route, LiteLLMRoute)
190-
assert orchestrator.success_checker is not None
191-
assert orchestrator.success_checker.route is orchestrator.eval_route
192+
with pytest.raises(ValueError, match="agent_judge"):
193+
orchestrator._resolve_routes()
192194

193195
def test_non_litellm_override_leaves_simulator_route_unaffected(self, tmp_path, monkeypatch):
194196
"""simulator_route always equals the agent's own resolution, independent

0 commit comments

Comments
 (0)