Skip to content

fix: close an area's stranded reports once a live one is open - #13

Merged
kim-em merged 3 commits into
mainfrom
fix/supersede-stranded-reports
Sep 17, 2026
Merged

kim-em merged 3 commits into
mainfrom
fix/supersede-stranded-reports

Conversation

@kim-em

@kim-em kim-em commented Sep 16, 2026

Copy link
Copy Markdown
Contributor

This PR closes the reports an area can no longer publish, once a live one is open for it. An area can hold more than one open report and nothing reconciled the ones that are not the current window, so the roadmap repository grew one unmergeable report per area per round and never shed them: 13 of the 15 open progress pull requests were stranded this way when this was written.

Two things strand a report, and neither closes itself.

A sibling merging advances the cursor, which orphans every other open report for that area instantly: the merge gate requires a byte-exact append at the cursor recorded in PROGRESS.md, so a report whose from_sha is no longer that cursor is dead by construction rather than merely behind. This happens in ordinary healthy operation, and accounted for 12 of the 13.

Separately, a repository-wide breakage fails every report's build. The planner's staleness expiry then releases the area after STALE_PR_HOURS, the next round recomputes a wider window, and because the branch name is a pure function of the window that is a new branch and a new pull request, which also cannot merge. TauCetiRoadmap's main went red on 2026-09-14 and the ModularForms reports opened at 20:40, 04:56, 14:16 and 22:24 -- every gap the 8-hour expiry.

The sweep runs only after the new report is open, so a failure costs a tidy-up and never a publication, and the only open report for an area is never closed. Ordering by creation date rather than comparing to_sha keeps the decision computable from the branch name and makes the newest report win deterministically, so two workers racing cannot each decide the other is superseded. Only reports this operator could have opened are closed, for the reason the rest of the module is careful about ownership: branch names are a pure function of the window, so a stranger can create one, and closing theirs would be a silent veto of someone else's contribution.

Note that this does not stop a red main generating stranded reports; it stops them accumulating. The generator-side guard is a separate change.

🤖 Prepared with Claude Code

Kim Morrison and others added 3 commits September 17, 2026 09:51
An area can hold more than one open report, and nothing reconciled the ones that are not the
current window. Two things strand a report and neither closes itself: a sibling merging advances
the cursor, which orphans every other open report for that area instantly, since the merge gate
requires a byte-exact append at the cursor; and any repository-wide breakage fails every report's
`build`, so the planner's staleness expiry releases the area each round and the next round opens a
fresh report that also cannot merge. The repository grew one unmergeable report per area per round
and never shed them.

After the new report is open, close the area's reports that can no longer merge: those whose window
does not start at the cursor, and those that start at the cursor but predate ours, whose window ours
therefore covers. Ordering by creation date keeps the newest report the winner deterministically, so
two workers racing cannot each close the other. Only reports we could have opened are closed, for
the reason the rest of the module is careful about ownership: branch names are a pure function of
the window, so a stranger can create one, and closing theirs would be a silent veto. A failure to
close is reported and never fatal -- by then the report is already published.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JedNDof6uYKixz6MvtZzCX
The first version described creation-order arbitration in its docstring and then never read a
timestamp, so every report at the same cursor was closed unconditionally and two workers racing
would each close the other -- the one outcome the ordering existed to prevent.

Order by pull request number instead, which GitHub allocates monotonically per repository. It is a
total order with no clock in it, so neither a shared timestamp nor a clock running backwards can
invert it, and exactly one direction of a racing pair ever fires. A report numbered above ours
supersedes us and is left for its own sweep. When our own number cannot be read, no same-cursor
report is closed at all: asserting an order we cannot establish risks discarding a live report,
which is far worse than leaving a duplicate for the next round. Orphans are unaffected either way,
since being dead is a property of the report alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JedNDof6uYKixz6MvtZzCX
A review found three ways the first version could destroy work or quietly stop working.

A worker whose plan predates a sibling's merge swept on ITS cursor, not the live one, so it read the
one report that could still merge as an orphan and closed it, keeping its own dead report. The live
cursor is now read from the roadmap repository at sweep time, and a run whose planned cursor no
longer matches it sweeps nothing: it is the stale one.

Ordering by creation, or by pull request number, never established what it claimed. A worker that
plans early and stalls opens a NARROW report late, so "newest wins" closes a wider report and loses
coverage already written. Neither a timestamp nor a number says which window contains which. Close a
same-cursor report only on proof: identical `from_sha`, and its window's pull requests a proper
subset of ours, both read from the metadata a report already carries in its body. Containment is
antisymmetric, so racing workers still cannot close each other, and an incomparable or identical
window is kept. The body is mutable, so it is only ever used to prove a report covers LESS; a forged
wider claim merely spares it.

A failed sweep was never retried. The next run for the same window stops at the in-flight check long
before reaching the sweep, so a cleanup outage left the pile-up forever while every run reported
success. The sweep now runs on that path too, and only `gh` failures are caught, so a parsing or
programming error surfaces instead of reading as a clean sweep.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JedNDof6uYKixz6MvtZzCX
@kim-em
kim-em merged commit dbb229d into main Sep 17, 2026
1 check passed
@kim-em
kim-em deleted the fix/supersede-stranded-reports branch September 17, 2026 05:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant