fix(upstream-watch): retry and pace routine dispatch - #280
Merged
Merged
Conversation
The first scheduled dispatch fired 7 released fixes back to back and all got HTTP 503 without a session, failing the run. The error also dropped the response body and request id. - retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried) - include request-id and the response error message in the failure - fire at most 3 fixes per run, 15s apart; stop at the first failed call - defer the rest to the next run without recording an error
Surface the dispatcher's deferred list in watch.json, the console output and the workflow summary so a backlog is visible, not silent. Pause between firings uses the script's injectable sleep.
This was referenced Sep 29, 2026
pacphi
added a commit
that referenced
this pull request
Sep 29, 2026
) * fix(upstream-watch): retry and pace routine dispatch (#280) * fix(upstream-watch): retry and pace routine dispatch The first scheduled dispatch fired 7 released fixes back to back and all got HTTP 503 without a session, failing the run. The error also dropped the response body and request id. - retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried) - include request-id and the response error message in the failure - fire at most 3 fixes per run, 15s apart; stop at the first failed call - defer the rest to the next run without recording an error * feat(upstream-watch): report fixes deferred to the next run Surface the dispatcher's deferred list in watch.json, the console output and the workflow summary so a backlog is visible, not silent. Pause between firings uses the script's injectable sleep. * fix(upstream-watch): bound trigger retries and redact error metadata * fix(upstream-watch): preserve deferred backlog in previews and blind results * docs(upstream-watch): explain bounded dispatch and deferred evidence * fix(upstream-watch): stop retries on response body transport failures * docs(upstream-watch): archive reviewed main reconciliation
pacphi
added a commit
that referenced
this pull request
Sep 29, 2026
* chore: establish develop remediation integration (#272) * docs(plan): confirm develop remediation execution * ci: validate develop branch pushes * fix(paths): ignore relative XDG locations consistently (#273) * docs(plan): map V4 follow ups v2 branch * fix(paths): ignore a relative XDG_* value, as the XDG Base Directory spec requires * docs(plan): specify V4 follow-up mappings and decisions * fix(footprint): align deep runtime log root with state base * fix(test-runner): finish guarded execution and environment hygiene (#274) * docs(plan): map runner hygiene tasks and evidence gates * docs(research): how test temp folders leak and how a run root is proven abandoned * docs(research): complete test creator lifecycle census * docs(research): account for parent cleanup in credential census * test(about): render the Ruflo install-edit line on the About card * test(tripwire): list the live Ruflo session's .claude-flow folder and proven-config files as concurrent writers * fix(test-runner): guard owner roots and keep sibling cleanup list-only * fix(test-runner): fail hygiene on own-root cleanup errors * feat(test-runner): guard focused runs and prove interrupted retention * test(test-runner): observe orphan exit before retention check * test(runner): preserve tool selector propagation * fix(ui): isolate Chrome launch environment * test(status): await owned spawn guard child before cleanup * test(runner): verify owned fork exits before sandbox cleanup * test(runner): retain own root for unresolved child holds * test(runner): await owned process cleanup on cancellation * docs(archive): record completed runner hygiene work * test(runner): keep cancellation fixture alive on Node 22 * test(runner): use owned native termination for signal retention * test(runner): retry fixture removal after confirmed child exit * feat(dashboard): deliver explicit refresh and live-session corrections (#276) * docs(plan): map dashboard refresh delivery * feat(dashboard): one server refresh operation behind POST /api/refresh * docs(plan): complete dashboard refresh carry-ins * docs: repair dashboard usage citation * feat(dashboard): one Refresh control with the CLI three strengths; Reload re-reads the view * fix(dashboard): show Unknown for unassessed Claude Code configuration * fix(dashboard): meet icon contrast in the host header * fix(dashboard): honor changed Maintenance URL state * test(dashboard): prove refresh request and write boundaries * docs(dashboard): record refresh client gate status * test(dashboard): restore Maintenance write safety journeys * fix(dashboard): reread active System views on Reload * fix(dashboard): retain write guard until refresh status reconciles * fix(dashboard): reconcile superseded refresh operations * fix(dashboard): start scans with POST requests * docs: align dashboard docs with the one Refresh control * docs: correct refresh labels and machine stage order * docs(adr): reconcile dashboard refresh supersessions * test(dashboard): cover ruflo component cwd forwarding * fix(live): resume displaced native transcript readers * docs(live): state re-entry and structured-source limits * docs(dashboard): reconcile final V3 evidence * fix(activity): show paused scan record time * fix(activity): reject impossible scan dates * docs(archive): record dashboard refresh implementation * fix(dashboard): honor current project-tree refresh selection * test(maintenance): align evidence method with Refresh vocabulary * test: trace native learning dependencies on macOS (#277) * docs(trace): plan upstream native resolution evidence * feat(trace): record resolved Transformers and ORT package roots * test(nightly): retain macOS learning resolution traces * test(trace): preload hook with portable file URL * docs(archive): record native learning trace evidence * fix(watch): complete bounded polling and watcher follow-ups (#278) * fix(upstream-watch): skip retries for deterministic fetch failures * fix(upstream-watch): bound dispatch pull request observation * test(upstream-watch): cover ledger git and spawn failures * test(upstream-watch): reject invalid ledger branch before git * perf(upstream-watch): measure notice fit with prefix lengths * test(upstream-watch): cover singular notice and latest firing * fix(upstream-watch): report invalid record registry once * test(upstream-watch): pin invalid registry precedence over future since * refactor(upstream-watch): share ledger event vocabulary * fix(upstream-watch): fail ledger on invalid registry * docs(upstream-watch): remove stale Codex reprobe instruction * fix(upstream-watch): format exhausted firing sessions as a list * fix(watch): require a commit before sending a notice * fix(watch): explain blind workflow summaries * docs(watch): archive completed follow-up plan * fix: remediate CLI, memory and upstream integration follow-ups (#275) * fix(cli): keep command usage failures machine-readable * fix(cli): validate usage and adapter revocation arguments * fix(host): keep dry-run JSON previews structured * fix(host): avoid status evidence writes in dry runs * fix(versions): throttle offline self retries by channel * fix(versions): scope self cache freshness to checked tags * test(maintain): guard injected refresh construction * fix(setup): isolate memory probe native mirror * test(evidence): prove repair commands re-record fresh facts * test(evidence): exercise setup machine host lifecycle wiring * test(aqe): verify live-lock fallback on installed artifact * test(aqe): guard live-lock probe cancellation * fix(memory): explain unsuitable locations and strict temp nesting * fix(memory): bind coexistence status to routing evidence * test(status): align offline self cache with checked tags * fix(memory): preserve cli proof across route checks * fix(memory): discover stray stores in ordinary dot directories * fix(status): report AQE home store separately * fix(status): replace broken AQE Codex setup hint * fix(memory): keep dependency markers and metadata errors out of complete scans * fix(status): give one ruflo component restart instruction * fix(daemon): preserve YAML configuration precedence * fix(sync): preview version-triggered daemon convergence * fix(daemon): respect explicit config and active restart source * fix(discovery): restore paused coverage from durable summary * fix(exec): abort owned process trees * fix(exec): bound uncertain Windows cleanup and byte caps * test(setup): isolate host rerecord project fixture * fix(test): hold C1 run root until owned children close * fix(test): accept Windows bootstrap env casing * test(ci): seed both requested self-drift tags * test(ci): probe pinned upstream conformance in isolated homes * docs(host-support): align upstream risks with verified releases * docs(upstream): register aqe init settings churn report * docs(host-support): clarify Ruflo source caveat * test(exec): retain a ready descendant after Windows parent exit * test(exec): launch a script through the native PowerShell fixture * test(ruflo): diagnose native Windows MCP transport boundaries * test(ruflo): isolate diagnostic launches and retain uncertain cleanup * test(ci): capture native Windows MCP transport diagnostics * fix(aqe): retire FsyncFailed live-lock exception * test(exec): hardcode the PowerShell fixture entry point * fix(exec): launch recognized npm Windows shims through their public bins * fix(identity): preserve exact persisted file IDs * test(ruflo): await natural closure for EOF diagnostics * test(exec): preserve native extensions in PowerShell ownership fixture * test(exec): keep reported parent out of cleanup authority * fix(live-checks): report skipped deja-vu and clean proof temp dirs * test(identity): correct Windows persisted identity fixtures * test(ci): retire completed native integration proof job * fix(test-runner): compare exact file identities before cleanup * docs(adr): scope self-version retry evidence to checked channels * docs(v4): record bounded C6 checks and accepted work * docs(remediation): archive verified V4 follow-ups plan * fix(usage): complete session attribution and accounting remediation (#282) * docs(usage): plan usage accuracy capture units * feat(usage): add shared session surface vocabulary * fix(usage): preserve fixed initiators and bound raw evidence * docs(usage): map capture units to source and tests * fix(usage): classify local managed Claude statusline settings * fix(usage): keep ambiguous managed statusline values unknown * fix(usage): classify statusline shell wrappers as custom * fix(usage): reject option-shaped statusline targets * feat(usage): persist session surface classification in cache * fix(usage): retain bounded unfamiliar origin metadata * fix(usage): classify Codex child threads and unpriced reviews * fix(usage): preserve rejected source and imported origin * fix(footprint): count Claude sessions by declared identity * fix(footprint): keep recovered project evidence out of session counts * fix(runtime): distinguish desktop apps from hosted CLI sessions * fix(runtime): recognize Codex service after global options * fix(runtime): preserve quoted Codex config boundaries * fix(system): preserve project census count basis in management API * fix(system): label desktop applications in runtime views * fix(usage): bind Claude provider detail to session model evidence * fix(usage): tighten Claude provider evidence validation * fix(usage): exclude imported Codex turns individually * fix(usage): reject incomplete mixed-turn ownership * fix(usage): count Codex component usage without responses * fix(usage): attribute OpenCode totals by response provider * fix(usage): retain provider bucket session metrics * feat(usage): capture Codex effort timing and compaction * feat(usage): preserve session surface presentation evidence across project DTOs * feat(dashboard): render session surfaces and import census disclosures * fix(usage): bound Codex compactions and gate unprovable replay * fix(dashboard): preserve legacy surface filters and refresh selections * feat(usage): reconcile Claude cost-state checkpoints * fix(usage): qualify Claude cost-state comparison scope * test(ui): include session surfaces in default suite * fix(usage): mark OpenCode reported zero as unpriced when untrusted * fix(usage): deduplicate copied Claude messages across files * fix(usage): scope Claude dedup to display windows * fix(usage): elect one Claude message owner across windows * fix(usage): bind Claude ownership to source eligibility * fix(usage): gate OpenCode cache reuse on parse semantics * fix(usage): invalidate local buckets when timezone changes * fix(usage): count unknown Claude transcript records * fix(usage): classify valid Claude JSON record shapes * fix(usage): require unambiguous OpenCode database selection * feat(usage): detect unsupported OpenCode storage presence * fix(usage): reject incomplete OpenCode source discovery * fix(usage): disclose unsupported OpenCode storage coverage * fix(usage): capture OpenCode compactions and reconcile counters * fix(usage): refresh OpenCode evidence and retain uncertain bounds * fix(usage): exclude OpenCode children from prompt fingerprints * docs(usage): accept bounded session and accounting contracts * docs(adr): record accepted session surface contracts * docs(usage): correct window provider and delegation metrics * fix(usage): preserve legacy filters and integration contracts * test(dashboard): align served provider helper assertion * docs(usage): archive completed V6 plans and evidence * test(maintenance): use a native census project fixture * test(usage): model unavailable runtime timezone portably * test(hooks): isolate CLI probes and assert no drift launches * test(host): frame lifecycle logs as test diagnostics * fix(upstream-watch): reconcile main pacing with develop safeguards (#283) * fix(upstream-watch): retry and pace routine dispatch (#280) * fix(upstream-watch): retry and pace routine dispatch The first scheduled dispatch fired 7 released fixes back to back and all got HTTP 503 without a session, failing the run. The error also dropped the response body and request id. - retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried) - include request-id and the response error message in the failure - fire at most 3 fixes per run, 15s apart; stop at the first failed call - defer the rest to the next run without recording an error * feat(upstream-watch): report fixes deferred to the next run Surface the dispatcher's deferred list in watch.json, the console output and the workflow summary so a backlog is visible, not silent. Pause between firings uses the script's injectable sleep. * fix(upstream-watch): bound trigger retries and redact error metadata * fix(upstream-watch): preserve deferred backlog in previews and blind results * docs(upstream-watch): explain bounded dispatch and deferred evidence * fix(upstream-watch): stop retries on response body transport failures * docs(upstream-watch): archive reviewed main reconciliation * chore(closeout): reconcile remediation v2 evidence and enforce comment hygiene (#284) * docs(plan): define final remediation closeout units * test(comments): guard durable references and remove transient labels * docs(closeout): record v2 scope and qualified execution evidence * fix(tests): distinguish test declarations from ordinary member calls * fix(tests): exclude table data from test context inference * docs(closeout): reconcile integration status and upstream evidence * test(quality): isolate parser guard from unit matrix * ci(quality): require parser guard after dependency installation * docs(closeout): archive validated V7 execution plan
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Scheduled run 36583034986 was the first to dispatch. It fired 7 released AQE fixes (#735, #753-755, #757-759) back to back and every call got
HTTP 503 without a session, so the Verdict step failed. Nofiredrecords were written, so no state was corrupted.The cause of the 503 is not established: the error discarded the response body and request id, so an outage, a capacity or concurrency limit, and a burst limit cannot be told apart. A 401/403 (token) or 404 (routine) would have shown as such.
Change
fire()retries 5xx answers up to 3 attempts (2s, then 4s backoff); 4xx answers are not retried.request-idand the first 200 characters of the response error message (never the token).deferred, not errors, and the next run picks them up.deferredis reported inwatch.json, the console output and the workflow run summary.Not done
MAX_FIRES_PER_RUN/FIRE_SPACING_MS.deferred; the cap and defer logic is covered by the dispatcher tests.Verification
node --test tests/kit/upstream-watch-*.test.mjs: 146 pass (new tests written first)pnpm run typecheck,pnpm run lint:md: cleanpnpm testnot runTo re-run the workflow after merge:
gh workflow run upstream-watch. The 7 fixes then go out 3 per run.🤖 Generated with Claude Code