fix(ci): the mutation gate failed its own baseline, because the report rounded - #69
Conversation
…t rounded The first real run of the job added in agenticsorg#66 went red, correctly, against the suite its floor was measured on. 131/257 is **50.97%**. The report printed it with `toFixed(1)` as "51.0%". I read that printed value back and set the floor to 51. Then `score < threshold` is `50.97 < 51`, which is true, so the gate rejected the exact baseline it was calibrated from. The instrument rounded, and its own output was treated as ground truth. ## Fix `fmtScore` truncates toward zero at two decimals instead of rounding, so the printed number is always a **lower bound** on the real score. A floor copied from the job output can then never exceed the score that produced it. Two decimals because one does not separate 50.97 from 51. Floor corrected to `50.97`, with the reason recorded next to it in the workflow so the next person to move it knows why it is not a round number. ## Tests Four new, suite 189/189. Red-first: restoring `toFixed(1)` turns all four red. The fourth is the property rather than an example: for a range of killed/total pairs, a floor copied from `fmtScore` must never fail the score that produced it. That is the invariant the gate actually depends on, and stating it as a property is what stops the next off-by-one at a different denominator. ## Why this is worth its own commit The gate did its job on its first outing, and what it caught was a defect in the gate. Two things follow that are worth writing down. The failure was **real, not a false alarm**. A floor that rejects its own baseline would have failed every nightly run until someone edited the number, and the obvious edit under time pressure is to lower the floor, which quietly removes the ratchet the floor exists to provide. And it is the same class of error as the one the job was built to find. agenticsorg#66's own commit message says a high mutation score is not evidence of correctness. Here a *rounded* score was not evidence of the floor being safe. In both cases the measurement was fine and the reading of it was wrong.
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The fix is small, self-contained, and well-tested, the floor-safety invariant was verified numerically, and the only finding is an optional style nit.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
This PR fixes a self-inflicted bug in the nightly mutation-testing gate added in #66. The gate's floor (MUTATION_FLOOR) was copied from a report that printed the mutation score with toFixed(1), which rounded 50.97% up to "51.0%". With the floor set to 51, the raw score 50.97 < 51 tripped the gate against the exact baseline it was calibrated from. The fix introduces a fmtScore helper that truncates toward zero at two decimals, guaranteeing the printed value is a lower bound on the real score, and corrects the floor to 50.97.
Changes:
- Add
fmtScore(truncate toward zero, 2 decimals) and route all three score printouts through it instead oftoFixed(1). - Correct
MUTATION_FLOORfrom51to50.97, with a comment explaining the truncation guarantee. - Add four tests, including a property test that a floor copied from
fmtScorenever fails the score that produced it.
| File | Description |
|---|---|
scripts/mutate.mjs |
Adds fmtScore truncating helper and replaces the three toFixed(1) score renders (console, gate error, report). |
test/mutate.test.js |
Adds four tests for truncation behavior and the floor-safety invariant. |
.github/workflows/mutation.yml |
Lowers the floor to 50.97 and documents why the number is truncated, not rounded. |
I verified the truncation logic, the gate comparison (raw score < threshold with threshold = Number('50.97')), all four test cases, and the property invariant numerically — all hold. All prior toFixed(1) score sites were updated with no stragglers. The only finding is a trivial brace-spacing inconsistency.
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| * so a floor copied from this output can never exceed the score that produced | ||
| * it. Two decimals, because one does not separate 50.97 from 51. | ||
| */ | ||
| export function fmtScore(score){ |

fix(ci): the mutation gate failed its own baseline, because the report rounded
The first real run of the job added in #66 went red, correctly, against the
suite its floor was measured on.
131/257 is 50.97%. The report printed it with
toFixed(1)as "51.0%". I readthat printed value back and set the floor to 51. Then
score < thresholdis50.97 < 51, which is true, so the gate rejected the exact baseline it wascalibrated from.
The instrument rounded, and its own output was treated as ground truth.
Fix
fmtScoretruncates toward zero at two decimals instead of rounding, so theprinted number is always a lower bound on the real score. A floor copied from
the job output can then never exceed the score that produced it. Two decimals
because one does not separate 50.97 from 51.
Floor corrected to
50.97, with the reason recorded next to it in the workflowso the next person to move it knows why it is not a round number.
Tests
Four new, suite 189/189. Red-first: restoring
toFixed(1)turns all four red.The fourth is the property rather than an example: for a range of killed/total
pairs, a floor copied from
fmtScoremust never fail the score that produced it.That is the invariant the gate actually depends on, and stating it as a property
is what stops the next off-by-one at a different denominator.
Why this is worth its own commit
The gate did its job on its first outing, and what it caught was a defect in the
gate. Two things follow that are worth writing down.
The failure was real, not a false alarm. A floor that rejects its own
baseline would have failed every nightly run until someone edited the number,
and the obvious edit under time pressure is to lower the floor, which quietly
removes the ratchet the floor exists to provide.
And it is the same class of error as the one the job was built to find. #66's own
commit message says a high mutation score is not evidence of correctness. Here a
rounded score was not evidence of the floor being safe. In both cases the
measurement was fine and the reading of it was wrong.