Skip to content

rocket-launch token baseline is stale against main — unrelated PRs breach the ceiling #91

Description

@HanselIdes

What

The rocket-launch scenario's timer-countdown token baseline (5k I+O, ceiling 7k) no longer reflects main, so PRs that don't touch the scenario are reporting token regressions against it.

Two samples from today, both on branches with current main merged, neither touching evals/ or anything timer-countdown reads:

PR Change under test timer-countdown I+O Check
#36 camunda-process-test reference wording 6k (+22% vs 5k) passed, under ceiling
#32 46 lines in camunda-connectors-development/references/registration-and-hosting.md 8k (+61% vs 5k) scenario:rocket-launch [with_skill] failed

Outcome scoring passed in both runs (✅ 1/1) — the only breach is the token ceiling.

Why it happens

rocket-launch carries skills: "all" in its METADATA, so every skill in the collection is planted for the run. Any skill that grows anywhere on main raises this scenario's token floor, whether or not the change has anything to do with launching a rocket. The 5k baseline predates a large amount of skill growth on main, so the measurement now sits close enough to the 7k ceiling that ordinary run-to-run variance in an agentic eval pushes individual runs over it.

The failure mode is the one that costs the most review attention: a red check on a PR whose diff cannot plausibly have caused it, which invites either ignoring eval signal generally or chasing a non-existent regression.

Suggested fix

Regenerate the baseline against current main (evals:regenerate-baselines), so the ceiling is derived from where the scenario actually sits today.

Worth considering alongside it, since regenerating only resets the clock:

  • Widen the ceiling margin for skills: "all" targets. They are structurally exposed to growth in every skill, so a percentage headroom that suits a single-skill eval is tighter than it should be here.
  • Or scope rocket-launch to the skills it actually exercises, so unrelated growth stops moving it. This is the more durable fix, but it changes what the scenario is meant to prove — cross-skill routing under a full plant — so it's a judgement call rather than an obvious win.

Notes

Not urgent for merges: the eval comment marks this signal non-blocking ("doesn't block merge"). It is, however, actively misleading on open PRs right now.

Surfaced while converging #32; deliberately not fixed there, since a baseline regeneration is a maintainer-gated action and shouldn't ride along in a docs PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions