Skip to content

fix(upstream-watch): retry and pace routine dispatch - #280

Merged
pacphi merged 2 commits into
mainfrom
fix/upstream-watch-dispatch-503
Sep 29, 2026
Merged

pacphi merged 2 commits into
mainfrom
fix/upstream-watch-dispatch-503

Conversation

@pacphi

@pacphi pacphi commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Problem

Scheduled run 36583034986 was the first to dispatch. It fired 7 released AQE fixes (#735, #753-755, #757-759) back to back and every call got HTTP 503 without a session, so the Verdict step failed. No fired records were written, so no state was corrupted.

The cause of the 503 is not established: the error discarded the response body and request id, so an outage, a capacity or concurrency limit, and a burst limit cannot be told apart. A 401/403 (token) or 404 (routine) would have shown as such.

Change

  • fire() retries 5xx answers up to 3 attempts (2s, then 4s backoff); 4xx answers are not retried.
  • The failure now carries the request-id and the first 200 characters of the response error message (never the token).
  • A run fires at most 3 fixes, 15s apart, and stops at the first failed call.
  • Fixes not fired are deferred, not errors, and the next run picks them up. deferred is reported in watch.json, the console output and the workflow run summary.

Not done

  • The limits (3 per run, 15s) are estimates; the routine's real capacity is unknown. If 503s continue, the new error text should show why; then adjust MAX_FIRES_PER_RUN / FIRE_SPACING_MS.
  • No script-level test covers a non-empty deferred; the cap and defer logic is covered by the dispatcher tests.

Verification

  • node --test tests/kit/upstream-watch-*.test.mjs: 146 pass (new tests written first)
  • eslint on the changed files, pnpm run typecheck, pnpm run lint:md: clean
  • pnpm test not run

To re-run the workflow after merge: gh workflow run upstream-watch. The 7 fixes then go out 3 per run.

🤖 Generated with Claude Code

The first scheduled dispatch fired 7 released fixes back to back and all
got HTTP 503 without a session, failing the run. The error also dropped
the response body and request id.

- retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried)
- include request-id and the response error message in the failure
- fire at most 3 fixes per run, 15s apart; stop at the first failed call
- defer the rest to the next run without recording an error
Surface the dispatcher's deferred list in watch.json, the console output
and the workflow summary so a backlog is visible, not silent. Pause between
firings uses the script's injectable sleep.
@pacphi
pacphi merged commit cb54f42 into main Sep 29, 2026
15 checks passed
@pacphi
pacphi deleted the fix/upstream-watch-dispatch-503 branch September 29, 2026 17:24
pacphi added a commit that referenced this pull request Sep 29, 2026
)

* fix(upstream-watch): retry and pace routine dispatch (#280)

* fix(upstream-watch): retry and pace routine dispatch

The first scheduled dispatch fired 7 released fixes back to back and all
got HTTP 503 without a session, failing the run. The error also dropped
the response body and request id.

- retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried)
- include request-id and the response error message in the failure
- fire at most 3 fixes per run, 15s apart; stop at the first failed call
- defer the rest to the next run without recording an error

* feat(upstream-watch): report fixes deferred to the next run

Surface the dispatcher's deferred list in watch.json, the console output
and the workflow summary so a backlog is visible, not silent. Pause between
firings uses the script's injectable sleep.

* fix(upstream-watch): bound trigger retries and redact error metadata

* fix(upstream-watch): preserve deferred backlog in previews and blind results

* docs(upstream-watch): explain bounded dispatch and deferred evidence

* fix(upstream-watch): stop retries on response body transport failures

* docs(upstream-watch): archive reviewed main reconciliation
pacphi added a commit that referenced this pull request Sep 29, 2026
* chore: establish develop remediation integration (#272)

* docs(plan): confirm develop remediation execution

* ci: validate develop branch pushes

* fix(paths): ignore relative XDG locations consistently (#273)

* docs(plan): map V4 follow ups v2 branch

* fix(paths): ignore a relative XDG_* value, as the XDG Base Directory spec requires

* docs(plan): specify V4 follow-up mappings and decisions

* fix(footprint): align deep runtime log root with state base

* fix(test-runner): finish guarded execution and environment hygiene (#274)

* docs(plan): map runner hygiene tasks and evidence gates

* docs(research): how test temp folders leak and how a run root is proven abandoned

* docs(research): complete test creator lifecycle census

* docs(research): account for parent cleanup in credential census

* test(about): render the Ruflo install-edit line on the About card

* test(tripwire): list the live Ruflo session's .claude-flow folder and proven-config files as concurrent writers

* fix(test-runner): guard owner roots and keep sibling cleanup list-only

* fix(test-runner): fail hygiene on own-root cleanup errors

* feat(test-runner): guard focused runs and prove interrupted retention

* test(test-runner): observe orphan exit before retention check

* test(runner): preserve tool selector propagation

* fix(ui): isolate Chrome launch environment

* test(status): await owned spawn guard child before cleanup

* test(runner): verify owned fork exits before sandbox cleanup

* test(runner): retain own root for unresolved child holds

* test(runner): await owned process cleanup on cancellation

* docs(archive): record completed runner hygiene work

* test(runner): keep cancellation fixture alive on Node 22

* test(runner): use owned native termination for signal retention

* test(runner): retry fixture removal after confirmed child exit

* feat(dashboard): deliver explicit refresh and live-session corrections (#276)

* docs(plan): map dashboard refresh delivery

* feat(dashboard): one server refresh operation behind POST /api/refresh

* docs(plan): complete dashboard refresh carry-ins

* docs: repair dashboard usage citation

* feat(dashboard): one Refresh control with the CLI three strengths; Reload re-reads the view

* fix(dashboard): show Unknown for unassessed Claude Code configuration

* fix(dashboard): meet icon contrast in the host header

* fix(dashboard): honor changed Maintenance URL state

* test(dashboard): prove refresh request and write boundaries

* docs(dashboard): record refresh client gate status

* test(dashboard): restore Maintenance write safety journeys

* fix(dashboard): reread active System views on Reload

* fix(dashboard): retain write guard until refresh status reconciles

* fix(dashboard): reconcile superseded refresh operations

* fix(dashboard): start scans with POST requests

* docs: align dashboard docs with the one Refresh control

* docs: correct refresh labels and machine stage order

* docs(adr): reconcile dashboard refresh supersessions

* test(dashboard): cover ruflo component cwd forwarding

* fix(live): resume displaced native transcript readers

* docs(live): state re-entry and structured-source limits

* docs(dashboard): reconcile final V3 evidence

* fix(activity): show paused scan record time

* fix(activity): reject impossible scan dates

* docs(archive): record dashboard refresh implementation

* fix(dashboard): honor current project-tree refresh selection

* test(maintenance): align evidence method with Refresh vocabulary

* test: trace native learning dependencies on macOS (#277)

* docs(trace): plan upstream native resolution evidence

* feat(trace): record resolved Transformers and ORT package roots

* test(nightly): retain macOS learning resolution traces

* test(trace): preload hook with portable file URL

* docs(archive): record native learning trace evidence

* fix(watch): complete bounded polling and watcher follow-ups (#278)

* fix(upstream-watch): skip retries for deterministic fetch failures

* fix(upstream-watch): bound dispatch pull request observation

* test(upstream-watch): cover ledger git and spawn failures

* test(upstream-watch): reject invalid ledger branch before git

* perf(upstream-watch): measure notice fit with prefix lengths

* test(upstream-watch): cover singular notice and latest firing

* fix(upstream-watch): report invalid record registry once

* test(upstream-watch): pin invalid registry precedence over future since

* refactor(upstream-watch): share ledger event vocabulary

* fix(upstream-watch): fail ledger on invalid registry

* docs(upstream-watch): remove stale Codex reprobe instruction

* fix(upstream-watch): format exhausted firing sessions as a list

* fix(watch): require a commit before sending a notice

* fix(watch): explain blind workflow summaries

* docs(watch): archive completed follow-up plan

* fix: remediate CLI, memory and upstream integration follow-ups (#275)

* fix(cli): keep command usage failures machine-readable

* fix(cli): validate usage and adapter revocation arguments

* fix(host): keep dry-run JSON previews structured

* fix(host): avoid status evidence writes in dry runs

* fix(versions): throttle offline self retries by channel

* fix(versions): scope self cache freshness to checked tags

* test(maintain): guard injected refresh construction

* fix(setup): isolate memory probe native mirror

* test(evidence): prove repair commands re-record fresh facts

* test(evidence): exercise setup machine host lifecycle wiring

* test(aqe): verify live-lock fallback on installed artifact

* test(aqe): guard live-lock probe cancellation

* fix(memory): explain unsuitable locations and strict temp nesting

* fix(memory): bind coexistence status to routing evidence

* test(status): align offline self cache with checked tags

* fix(memory): preserve cli proof across route checks

* fix(memory): discover stray stores in ordinary dot directories

* fix(status): report AQE home store separately

* fix(status): replace broken AQE Codex setup hint

* fix(memory): keep dependency markers and metadata errors out of complete scans

* fix(status): give one ruflo component restart instruction

* fix(daemon): preserve YAML configuration precedence

* fix(sync): preview version-triggered daemon convergence

* fix(daemon): respect explicit config and active restart source

* fix(discovery): restore paused coverage from durable summary

* fix(exec): abort owned process trees

* fix(exec): bound uncertain Windows cleanup and byte caps

* test(setup): isolate host rerecord project fixture

* fix(test): hold C1 run root until owned children close

* fix(test): accept Windows bootstrap env casing

* test(ci): seed both requested self-drift tags

* test(ci): probe pinned upstream conformance in isolated homes

* docs(host-support): align upstream risks with verified releases

* docs(upstream): register aqe init settings churn report

* docs(host-support): clarify Ruflo source caveat

* test(exec): retain a ready descendant after Windows parent exit

* test(exec): launch a script through the native PowerShell fixture

* test(ruflo): diagnose native Windows MCP transport boundaries

* test(ruflo): isolate diagnostic launches and retain uncertain cleanup

* test(ci): capture native Windows MCP transport diagnostics

* fix(aqe): retire FsyncFailed live-lock exception

* test(exec): hardcode the PowerShell fixture entry point

* fix(exec): launch recognized npm Windows shims through their public bins

* fix(identity): preserve exact persisted file IDs

* test(ruflo): await natural closure for EOF diagnostics

* test(exec): preserve native extensions in PowerShell ownership fixture

* test(exec): keep reported parent out of cleanup authority

* fix(live-checks): report skipped deja-vu and clean proof temp dirs

* test(identity): correct Windows persisted identity fixtures

* test(ci): retire completed native integration proof job

* fix(test-runner): compare exact file identities before cleanup

* docs(adr): scope self-version retry evidence to checked channels

* docs(v4): record bounded C6 checks and accepted work

* docs(remediation): archive verified V4 follow-ups plan

* fix(usage): complete session attribution and accounting remediation (#282)

* docs(usage): plan usage accuracy capture units

* feat(usage): add shared session surface vocabulary

* fix(usage): preserve fixed initiators and bound raw evidence

* docs(usage): map capture units to source and tests

* fix(usage): classify local managed Claude statusline settings

* fix(usage): keep ambiguous managed statusline values unknown

* fix(usage): classify statusline shell wrappers as custom

* fix(usage): reject option-shaped statusline targets

* feat(usage): persist session surface classification in cache

* fix(usage): retain bounded unfamiliar origin metadata

* fix(usage): classify Codex child threads and unpriced reviews

* fix(usage): preserve rejected source and imported origin

* fix(footprint): count Claude sessions by declared identity

* fix(footprint): keep recovered project evidence out of session counts

* fix(runtime): distinguish desktop apps from hosted CLI sessions

* fix(runtime): recognize Codex service after global options

* fix(runtime): preserve quoted Codex config boundaries

* fix(system): preserve project census count basis in management API

* fix(system): label desktop applications in runtime views

* fix(usage): bind Claude provider detail to session model evidence

* fix(usage): tighten Claude provider evidence validation

* fix(usage): exclude imported Codex turns individually

* fix(usage): reject incomplete mixed-turn ownership

* fix(usage): count Codex component usage without responses

* fix(usage): attribute OpenCode totals by response provider

* fix(usage): retain provider bucket session metrics

* feat(usage): capture Codex effort timing and compaction

* feat(usage): preserve session surface presentation evidence across project DTOs

* feat(dashboard): render session surfaces and import census disclosures

* fix(usage): bound Codex compactions and gate unprovable replay

* fix(dashboard): preserve legacy surface filters and refresh selections

* feat(usage): reconcile Claude cost-state checkpoints

* fix(usage): qualify Claude cost-state comparison scope

* test(ui): include session surfaces in default suite

* fix(usage): mark OpenCode reported zero as unpriced when untrusted

* fix(usage): deduplicate copied Claude messages across files

* fix(usage): scope Claude dedup to display windows

* fix(usage): elect one Claude message owner across windows

* fix(usage): bind Claude ownership to source eligibility

* fix(usage): gate OpenCode cache reuse on parse semantics

* fix(usage): invalidate local buckets when timezone changes

* fix(usage): count unknown Claude transcript records

* fix(usage): classify valid Claude JSON record shapes

* fix(usage): require unambiguous OpenCode database selection

* feat(usage): detect unsupported OpenCode storage presence

* fix(usage): reject incomplete OpenCode source discovery

* fix(usage): disclose unsupported OpenCode storage coverage

* fix(usage): capture OpenCode compactions and reconcile counters

* fix(usage): refresh OpenCode evidence and retain uncertain bounds

* fix(usage): exclude OpenCode children from prompt fingerprints

* docs(usage): accept bounded session and accounting contracts

* docs(adr): record accepted session surface contracts

* docs(usage): correct window provider and delegation metrics

* fix(usage): preserve legacy filters and integration contracts

* test(dashboard): align served provider helper assertion

* docs(usage): archive completed V6 plans and evidence

* test(maintenance): use a native census project fixture

* test(usage): model unavailable runtime timezone portably

* test(hooks): isolate CLI probes and assert no drift launches

* test(host): frame lifecycle logs as test diagnostics

* fix(upstream-watch): reconcile main pacing with develop safeguards (#283)

* fix(upstream-watch): retry and pace routine dispatch (#280)

* fix(upstream-watch): retry and pace routine dispatch

The first scheduled dispatch fired 7 released fixes back to back and all
got HTTP 503 without a session, failing the run. The error also dropped
the response body and request id.

- retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried)
- include request-id and the response error message in the failure
- fire at most 3 fixes per run, 15s apart; stop at the first failed call
- defer the rest to the next run without recording an error

* feat(upstream-watch): report fixes deferred to the next run

Surface the dispatcher's deferred list in watch.json, the console output
and the workflow summary so a backlog is visible, not silent. Pause between
firings uses the script's injectable sleep.

* fix(upstream-watch): bound trigger retries and redact error metadata

* fix(upstream-watch): preserve deferred backlog in previews and blind results

* docs(upstream-watch): explain bounded dispatch and deferred evidence

* fix(upstream-watch): stop retries on response body transport failures

* docs(upstream-watch): archive reviewed main reconciliation

* chore(closeout): reconcile remediation v2 evidence and enforce comment hygiene (#284)

* docs(plan): define final remediation closeout units

* test(comments): guard durable references and remove transient labels

* docs(closeout): record v2 scope and qualified execution evidence

* fix(tests): distinguish test declarations from ordinary member calls

* fix(tests): exclude table data from test context inference

* docs(closeout): reconcile integration status and upstream evidence

* test(quality): isolate parser guard from unit matrix

* ci(quality): require parser guard after dependency installation

* docs(closeout): archive validated V7 execution plan
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant