Skip to content

fix(ci): the mutation gate failed its own baseline, because the report rounded - #69

Merged
michaeloboyle merged 1 commit into
agenticsorg:mainfrom
michaeloboyle:fix/mutation-floor-precision
Oct 1, 2026
Merged

michaeloboyle merged 1 commit into
agenticsorg:mainfrom
michaeloboyle:fix/mutation-floor-precision

Conversation

@michaeloboyle

Copy link
Copy Markdown
Collaborator

fix(ci): the mutation gate failed its own baseline, because the report rounded

The first real run of the job added in #66 went red, correctly, against the
suite its floor was measured on.

131/257 is 50.97%. The report printed it with toFixed(1) as "51.0%". I read
that printed value back and set the floor to 51. Then score < threshold is
50.97 < 51, which is true, so the gate rejected the exact baseline it was
calibrated from.

The instrument rounded, and its own output was treated as ground truth.

Fix

fmtScore truncates toward zero at two decimals instead of rounding, so the
printed number is always a lower bound on the real score. A floor copied from
the job output can then never exceed the score that produced it. Two decimals
because one does not separate 50.97 from 51.

Floor corrected to 50.97, with the reason recorded next to it in the workflow
so the next person to move it knows why it is not a round number.

Tests

Four new, suite 189/189. Red-first: restoring toFixed(1) turns all four red.

The fourth is the property rather than an example: for a range of killed/total
pairs, a floor copied from fmtScore must never fail the score that produced it.
That is the invariant the gate actually depends on, and stating it as a property
is what stops the next off-by-one at a different denominator.

Why this is worth its own commit

The gate did its job on its first outing, and what it caught was a defect in the
gate. Two things follow that are worth writing down.

The failure was real, not a false alarm. A floor that rejects its own
baseline would have failed every nightly run until someone edited the number,
and the obvious edit under time pressure is to lower the floor, which quietly
removes the ratchet the floor exists to provide.

And it is the same class of error as the one the job was built to find. #66's own
commit message says a high mutation score is not evidence of correctness. Here a
rounded score was not evidence of the floor being safe. In both cases the
measurement was fine and the reading of it was wrong.

…t rounded

The first real run of the job added in agenticsorg#66 went red, correctly, against the
suite its floor was measured on.

131/257 is **50.97%**. The report printed it with `toFixed(1)` as "51.0%". I read
that printed value back and set the floor to 51. Then `score < threshold` is
`50.97 < 51`, which is true, so the gate rejected the exact baseline it was
calibrated from.

The instrument rounded, and its own output was treated as ground truth.

## Fix

`fmtScore` truncates toward zero at two decimals instead of rounding, so the
printed number is always a **lower bound** on the real score. A floor copied from
the job output can then never exceed the score that produced it. Two decimals
because one does not separate 50.97 from 51.

Floor corrected to `50.97`, with the reason recorded next to it in the workflow
so the next person to move it knows why it is not a round number.

## Tests

Four new, suite 189/189. Red-first: restoring `toFixed(1)` turns all four red.

The fourth is the property rather than an example: for a range of killed/total
pairs, a floor copied from `fmtScore` must never fail the score that produced it.
That is the invariant the gate actually depends on, and stating it as a property
is what stops the next off-by-one at a different denominator.

## Why this is worth its own commit

The gate did its job on its first outing, and what it caught was a defect in the
gate. Two things follow that are worth writing down.

The failure was **real, not a false alarm**. A floor that rejects its own
baseline would have failed every nightly run until someone edited the number,
and the obvious edit under time pressure is to lower the floor, which quietly
removes the ratchet the floor exists to provide.

And it is the same class of error as the one the job was built to find. agenticsorg#66's own
commit message says a high mutation score is not evidence of correctness. Here a
*rounded* score was not evidence of the floor being safe. In both cases the
measurement was fine and the reading of it was wrong.
Copilot AI balanced review requested due to automatic review settings October 1, 2026 13:50
@michaeloboyle
michaeloboyle merged commit 3bbb4ea into agenticsorg:main Oct 1, 2026
1 check passed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The fix is small, self-contained, and well-tested, the floor-safety invariant was verified numerically, and the only finding is an optional style nit.

Review effort: Balanced
Findings: 1 Low severity

Open (1)
What changed in this PR

This PR fixes a self-inflicted bug in the nightly mutation-testing gate added in #66. The gate's floor (MUTATION_FLOOR) was copied from a report that printed the mutation score with toFixed(1), which rounded 50.97% up to "51.0%". With the floor set to 51, the raw score 50.97 < 51 tripped the gate against the exact baseline it was calibrated from. The fix introduces a fmtScore helper that truncates toward zero at two decimals, guaranteeing the printed value is a lower bound on the real score, and corrects the floor to 50.97.

Changes:

  • Add fmtScore (truncate toward zero, 2 decimals) and route all three score printouts through it instead of toFixed(1).
  • Correct MUTATION_FLOOR from 51 to 50.97, with a comment explaining the truncation guarantee.
  • Add four tests, including a property test that a floor copied from fmtScore never fails the score that produced it.
File Description
scripts/​mutate.mjs Adds fmtScore truncating helper and replaces the three toFixed(1) score renders (console, gate error, report).
test/​mutate.test.js Adds four tests for truncation behavior and the floor-safety invariant.
.github/​workflows/​mutation.yml Lowers the floor to 50.97 and documents why the number is truncated, not rounded.

I verified the truncation logic, the gate comparison (raw score < threshold with threshold = Number('50.97')), all four test cases, and the property invariant numerically — all hold. All prior toFixed(1) score sites were updated with no stragglers. The only finding is a trivial brace-spacing inconsistency.


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/mutate.mjs
* so a floor copied from this output can never exceed the score that produced
* it. Two decimals, because one does not separate 50.97 from 51.
*/
export function fmtScore(score){
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants