What
The rocket-launch scenario's timer-countdown token baseline (5k I+O, ceiling 7k) no longer reflects main, so PRs that don't touch the scenario are reporting token regressions against it.
Two samples from today, both on branches with current main merged, neither touching evals/ or anything timer-countdown reads:
| PR |
Change under test |
timer-countdown I+O |
Check |
| #36 |
camunda-process-test reference wording |
6k (+22% vs 5k) |
passed, under ceiling |
| #32 |
46 lines in camunda-connectors-development/references/registration-and-hosting.md |
8k (+61% vs 5k) |
scenario:rocket-launch [with_skill] failed |
Outcome scoring passed in both runs (✅ 1/1) — the only breach is the token ceiling.
Why it happens
rocket-launch carries skills: "all" in its METADATA, so every skill in the collection is planted for the run. Any skill that grows anywhere on main raises this scenario's token floor, whether or not the change has anything to do with launching a rocket. The 5k baseline predates a large amount of skill growth on main, so the measurement now sits close enough to the 7k ceiling that ordinary run-to-run variance in an agentic eval pushes individual runs over it.
The failure mode is the one that costs the most review attention: a red check on a PR whose diff cannot plausibly have caused it, which invites either ignoring eval signal generally or chasing a non-existent regression.
Suggested fix
Regenerate the baseline against current main (evals:regenerate-baselines), so the ceiling is derived from where the scenario actually sits today.
Worth considering alongside it, since regenerating only resets the clock:
- Widen the ceiling margin for
skills: "all" targets. They are structurally exposed to growth in every skill, so a percentage headroom that suits a single-skill eval is tighter than it should be here.
- Or scope
rocket-launch to the skills it actually exercises, so unrelated growth stops moving it. This is the more durable fix, but it changes what the scenario is meant to prove — cross-skill routing under a full plant — so it's a judgement call rather than an obvious win.
Notes
Not urgent for merges: the eval comment marks this signal non-blocking ("doesn't block merge"). It is, however, actively misleading on open PRs right now.
Surfaced while converging #32; deliberately not fixed there, since a baseline regeneration is a maintainer-gated action and shouldn't ride along in a docs PR.
What
The
rocket-launchscenario'stimer-countdowntoken baseline (5kI+O, ceiling7k) no longer reflectsmain, so PRs that don't touch the scenario are reporting token regressions against it.Two samples from today, both on branches with current
mainmerged, neither touchingevals/or anythingtimer-countdownreads:camunda-process-testreference wording6k(+22% vs5k)camunda-connectors-development/references/registration-and-hosting.md8k(+61% vs5k)scenario:rocket-launch [with_skill]failedOutcome scoring passed in both runs (
✅ 1/1) — the only breach is the token ceiling.Why it happens
rocket-launchcarriesskills: "all"in itsMETADATA, so every skill in the collection is planted for the run. Any skill that grows anywhere onmainraises this scenario's token floor, whether or not the change has anything to do with launching a rocket. The5kbaseline predates a large amount of skill growth onmain, so the measurement now sits close enough to the7kceiling that ordinary run-to-run variance in an agentic eval pushes individual runs over it.The failure mode is the one that costs the most review attention: a red check on a PR whose diff cannot plausibly have caused it, which invites either ignoring eval signal generally or chasing a non-existent regression.
Suggested fix
Regenerate the baseline against current
main(evals:regenerate-baselines), so the ceiling is derived from where the scenario actually sits today.Worth considering alongside it, since regenerating only resets the clock:
skills: "all"targets. They are structurally exposed to growth in every skill, so a percentage headroom that suits a single-skill eval is tighter than it should be here.rocket-launchto the skills it actually exercises, so unrelated growth stops moving it. This is the more durable fix, but it changes what the scenario is meant to prove — cross-skill routing under a full plant — so it's a judgement call rather than an obvious win.Notes
Not urgent for merges: the eval comment marks this signal non-blocking ("doesn't block merge"). It is, however, actively misleading on open PRs right now.
Surfaced while converging #32; deliberately not fixed there, since a baseline regeneration is a maintainer-gated action and shouldn't ride along in a docs PR.