Skip to content

feat: integrate remediation v2 for human review - #285

Merged
pacphi merged 12 commits into
mainfrom
develop
Sep 29, 2026
Merged

pacphi merged 12 commits into
mainfrom
develop

Conversation

@pacphi

@pacphi pacphi commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Summary

Remediation v2 corrects dashboard refresh behavior and usage accounting, improves
integration diagnostics, and records the evidence and remaining limits for all
189 Appendix A plus 57 Appendix B scope entries. This PR is for human review;
the final merge into main remains human-owned.

  • Dashboard refresh has one explicit control and POST operation; reading a GET
    route does not start a scan. Bounded live-session re-entry preserves observed
    totals, and the native plain-folder binding case is documented with its limits.
  • Usage views distinguish session surfaces and imported records, avoid duplicate
    Claude charges within the documented identity pool, and require explicit
    OpenCode database selection when multiple sources are ambiguous. Cowork's
    missing reader is disclosed. Legacy and incomplete data retain honest labels.
  • CLI and integration fixes improve retries, discovery and failure reporting.
    Ordinary AQE lock contention remains busy; unrelated storage errors fail closed.
    Upstream watching retains bounded documented retries and visible failed reads.
  • Windows fixture work preserves the test matrix and assertions. The fixed
    ten-before/ten-after observations show lower medians, with failed and cancelled
    observations retained. Documentation preserves the original dispositions,
    qualified native proofs and 18 ordered execution rulings.

Source and integration

Inherited on main: V1 #267 and V2 #264/#266/#269. This PR accumulates bootstrap
#272, V3 #276, V4 #273/#275/#277/#278, V5 #274, V6 #282, watcher reconciliation
#283 and V7 #284. Source-bound feature/squash mappings are in the integration receipt.

  • Reviewed final feature: ebce0758b3845bd3cd63087b16ef974c27d9928a.
  • Final develop head: 04e3de3c336fce47ce85a6684b39608d916036cc.
  • Current main base: cb54f42e11fb6ff152ec278eb2dc1e119ee05ac5, an ancestor of develop.
  • Feature and squash trees are identical: 373babefb2874b9ac43a6498e5e127141e5d0470.
  • V7 CI 36625700956
    and all 16 applicable PR checks passed; two intentional jobs were skipped.
  • Develop CI 36626621714: all 13 checks passed.

Independent whole-branch, scoped correction and final archival reviews passed.
V7 corrected its new guard's Git environment and runs all 12 ESLint-dependent guard
tests in the mandatory installed-dependency quality job. All nine runtime matrix
legs still run the full suite without installing development dependencies.
The content-identical ancestry merge records the already-reviewed main watcher
input after #283's squash; its tree and parents were verified. The main watcher changes were incorporated through reviewed #283.

The parent plans and dated archive receipts retain their capture-time status; this
PR provides the subsequent integration, CI and issue-action record.

Validation and evidence

All nine local gate steps passed at implementation head
795a0f7266de0bf6a04358359b75a1ad1b609e1d: 6,440 unit tests passed with
7 platform skips; all legacy suites passed; browser checks passed 514 assertions
plus 16 further tests; the new guard passed all 12 tests. Coverage measured
94.15% lines / 83.83% branches / 93.49% functions, above the enforced
70/70/70 floors and the 80% line target. Types, lint, complexity, Markdown, build
and offline links passed. The documentation-only archival commit then passed
17 focused tests, 240-file Markdown and offline links.

The state tripwire guarded the suites. Local ESLint excluded ignored immutable
scratch probes; fresh-checkout CI used normal lint. The Node 22 runner workaround
is test-local and does not claim to fix the upstream runtime defect.

Actual executed tests and native probes driven by Codex provide the verification
evidence. AQE tool calls returned unusable placeholder or empty scenario output;
no AQE fleet execution, generated-scenario coverage or predicted coverage is claimed.

The historical Windows all-job medians are 649/705.5/875 seconds before and
221.5/225/210 after for Node 22/24/26. These are observational comparisons.
The preserved 19:20 snapshot failed on a 358-second Windows leg. The newer
20:27 capture of the latest three PR runs also fails: #284 Windows Node 24
took 311 seconds (Node 22/26: 201/230); #283 and #282 qualify. No rerun
replaced the slow observation. The full table and current gate
are posted on #262, which remains open.

Issue disposition and remaining gates

This PR performs no release, global installation, deletion, real-store merge or
paid provider run. Those operations retain separate authorization and evidence
gates. Real routine execution and installed-release conformance are not inferred
from source fixes. AQE 3.14.5 setup churn remains observed; later-release conformance,
Windows AQE, old-fork repair and signed-entry validation remain unproved here.
D-5's four readiness items remain deferred to v5; future #239 W1–W8/D-20+ work is
outside this approval. The memory loop stays OFF until main approval. Plans remain
active while human review and operational gates remain; branches and worktrees are
retained. Attended time was not instrumented.

* docs(plan): confirm develop remediation execution

* ci: validate develop branch pushes
* docs(plan): map V4 follow ups v2 branch

* fix(paths): ignore a relative XDG_* value, as the XDG Base Directory spec requires

* docs(plan): specify V4 follow-up mappings and decisions

* fix(footprint): align deep runtime log root with state base
)

* docs(plan): map runner hygiene tasks and evidence gates

* docs(research): how test temp folders leak and how a run root is proven abandoned

* docs(research): complete test creator lifecycle census

* docs(research): account for parent cleanup in credential census

* test(about): render the Ruflo install-edit line on the About card

* test(tripwire): list the live Ruflo session's .claude-flow folder and proven-config files as concurrent writers

* fix(test-runner): guard owner roots and keep sibling cleanup list-only

* fix(test-runner): fail hygiene on own-root cleanup errors

* feat(test-runner): guard focused runs and prove interrupted retention

* test(test-runner): observe orphan exit before retention check

* test(runner): preserve tool selector propagation

* fix(ui): isolate Chrome launch environment

* test(status): await owned spawn guard child before cleanup

* test(runner): verify owned fork exits before sandbox cleanup

* test(runner): retain own root for unresolved child holds

* test(runner): await owned process cleanup on cancellation

* docs(archive): record completed runner hygiene work

* test(runner): keep cancellation fixture alive on Node 22

* test(runner): use owned native termination for signal retention

* test(runner): retry fixture removal after confirmed child exit
#276)

* docs(plan): map dashboard refresh delivery

* feat(dashboard): one server refresh operation behind POST /api/refresh

* docs(plan): complete dashboard refresh carry-ins

* docs: repair dashboard usage citation

* feat(dashboard): one Refresh control with the CLI three strengths; Reload re-reads the view

* fix(dashboard): show Unknown for unassessed Claude Code configuration

* fix(dashboard): meet icon contrast in the host header

* fix(dashboard): honor changed Maintenance URL state

* test(dashboard): prove refresh request and write boundaries

* docs(dashboard): record refresh client gate status

* test(dashboard): restore Maintenance write safety journeys

* fix(dashboard): reread active System views on Reload

* fix(dashboard): retain write guard until refresh status reconciles

* fix(dashboard): reconcile superseded refresh operations

* fix(dashboard): start scans with POST requests

* docs: align dashboard docs with the one Refresh control

* docs: correct refresh labels and machine stage order

* docs(adr): reconcile dashboard refresh supersessions

* test(dashboard): cover ruflo component cwd forwarding

* fix(live): resume displaced native transcript readers

* docs(live): state re-entry and structured-source limits

* docs(dashboard): reconcile final V3 evidence

* fix(activity): show paused scan record time

* fix(activity): reject impossible scan dates

* docs(archive): record dashboard refresh implementation

* fix(dashboard): honor current project-tree refresh selection

* test(maintenance): align evidence method with Refresh vocabulary
* docs(trace): plan upstream native resolution evidence

* feat(trace): record resolved Transformers and ORT package roots

* test(nightly): retain macOS learning resolution traces

* test(trace): preload hook with portable file URL

* docs(archive): record native learning trace evidence
* fix(upstream-watch): skip retries for deterministic fetch failures

* fix(upstream-watch): bound dispatch pull request observation

* test(upstream-watch): cover ledger git and spawn failures

* test(upstream-watch): reject invalid ledger branch before git

* perf(upstream-watch): measure notice fit with prefix lengths

* test(upstream-watch): cover singular notice and latest firing

* fix(upstream-watch): report invalid record registry once

* test(upstream-watch): pin invalid registry precedence over future since

* refactor(upstream-watch): share ledger event vocabulary

* fix(upstream-watch): fail ledger on invalid registry

* docs(upstream-watch): remove stale Codex reprobe instruction

* fix(upstream-watch): format exhausted firing sessions as a list

* fix(watch): require a commit before sending a notice

* fix(watch): explain blind workflow summaries

* docs(watch): archive completed follow-up plan
* fix(cli): keep command usage failures machine-readable

* fix(cli): validate usage and adapter revocation arguments

* fix(host): keep dry-run JSON previews structured

* fix(host): avoid status evidence writes in dry runs

* fix(versions): throttle offline self retries by channel

* fix(versions): scope self cache freshness to checked tags

* test(maintain): guard injected refresh construction

* fix(setup): isolate memory probe native mirror

* test(evidence): prove repair commands re-record fresh facts

* test(evidence): exercise setup machine host lifecycle wiring

* test(aqe): verify live-lock fallback on installed artifact

* test(aqe): guard live-lock probe cancellation

* fix(memory): explain unsuitable locations and strict temp nesting

* fix(memory): bind coexistence status to routing evidence

* test(status): align offline self cache with checked tags

* fix(memory): preserve cli proof across route checks

* fix(memory): discover stray stores in ordinary dot directories

* fix(status): report AQE home store separately

* fix(status): replace broken AQE Codex setup hint

* fix(memory): keep dependency markers and metadata errors out of complete scans

* fix(status): give one ruflo component restart instruction

* fix(daemon): preserve YAML configuration precedence

* fix(sync): preview version-triggered daemon convergence

* fix(daemon): respect explicit config and active restart source

* fix(discovery): restore paused coverage from durable summary

* fix(exec): abort owned process trees

* fix(exec): bound uncertain Windows cleanup and byte caps

* test(setup): isolate host rerecord project fixture

* fix(test): hold C1 run root until owned children close

* fix(test): accept Windows bootstrap env casing

* test(ci): seed both requested self-drift tags

* test(ci): probe pinned upstream conformance in isolated homes

* docs(host-support): align upstream risks with verified releases

* docs(upstream): register aqe init settings churn report

* docs(host-support): clarify Ruflo source caveat

* test(exec): retain a ready descendant after Windows parent exit

* test(exec): launch a script through the native PowerShell fixture

* test(ruflo): diagnose native Windows MCP transport boundaries

* test(ruflo): isolate diagnostic launches and retain uncertain cleanup

* test(ci): capture native Windows MCP transport diagnostics

* fix(aqe): retire FsyncFailed live-lock exception

* test(exec): hardcode the PowerShell fixture entry point

* fix(exec): launch recognized npm Windows shims through their public bins

* fix(identity): preserve exact persisted file IDs

* test(ruflo): await natural closure for EOF diagnostics

* test(exec): preserve native extensions in PowerShell ownership fixture

* test(exec): keep reported parent out of cleanup authority

* fix(live-checks): report skipped deja-vu and clean proof temp dirs

* test(identity): correct Windows persisted identity fixtures

* test(ci): retire completed native integration proof job

* fix(test-runner): compare exact file identities before cleanup

* docs(adr): scope self-version retry evidence to checked channels

* docs(v4): record bounded C6 checks and accepted work

* docs(remediation): archive verified V4 follow-ups plan
…282)

* docs(usage): plan usage accuracy capture units

* feat(usage): add shared session surface vocabulary

* fix(usage): preserve fixed initiators and bound raw evidence

* docs(usage): map capture units to source and tests

* fix(usage): classify local managed Claude statusline settings

* fix(usage): keep ambiguous managed statusline values unknown

* fix(usage): classify statusline shell wrappers as custom

* fix(usage): reject option-shaped statusline targets

* feat(usage): persist session surface classification in cache

* fix(usage): retain bounded unfamiliar origin metadata

* fix(usage): classify Codex child threads and unpriced reviews

* fix(usage): preserve rejected source and imported origin

* fix(footprint): count Claude sessions by declared identity

* fix(footprint): keep recovered project evidence out of session counts

* fix(runtime): distinguish desktop apps from hosted CLI sessions

* fix(runtime): recognize Codex service after global options

* fix(runtime): preserve quoted Codex config boundaries

* fix(system): preserve project census count basis in management API

* fix(system): label desktop applications in runtime views

* fix(usage): bind Claude provider detail to session model evidence

* fix(usage): tighten Claude provider evidence validation

* fix(usage): exclude imported Codex turns individually

* fix(usage): reject incomplete mixed-turn ownership

* fix(usage): count Codex component usage without responses

* fix(usage): attribute OpenCode totals by response provider

* fix(usage): retain provider bucket session metrics

* feat(usage): capture Codex effort timing and compaction

* feat(usage): preserve session surface presentation evidence across project DTOs

* feat(dashboard): render session surfaces and import census disclosures

* fix(usage): bound Codex compactions and gate unprovable replay

* fix(dashboard): preserve legacy surface filters and refresh selections

* feat(usage): reconcile Claude cost-state checkpoints

* fix(usage): qualify Claude cost-state comparison scope

* test(ui): include session surfaces in default suite

* fix(usage): mark OpenCode reported zero as unpriced when untrusted

* fix(usage): deduplicate copied Claude messages across files

* fix(usage): scope Claude dedup to display windows

* fix(usage): elect one Claude message owner across windows

* fix(usage): bind Claude ownership to source eligibility

* fix(usage): gate OpenCode cache reuse on parse semantics

* fix(usage): invalidate local buckets when timezone changes

* fix(usage): count unknown Claude transcript records

* fix(usage): classify valid Claude JSON record shapes

* fix(usage): require unambiguous OpenCode database selection

* feat(usage): detect unsupported OpenCode storage presence

* fix(usage): reject incomplete OpenCode source discovery

* fix(usage): disclose unsupported OpenCode storage coverage

* fix(usage): capture OpenCode compactions and reconcile counters

* fix(usage): refresh OpenCode evidence and retain uncertain bounds

* fix(usage): exclude OpenCode children from prompt fingerprints

* docs(usage): accept bounded session and accounting contracts

* docs(adr): record accepted session surface contracts

* docs(usage): correct window provider and delegation metrics

* fix(usage): preserve legacy filters and integration contracts

* test(dashboard): align served provider helper assertion

* docs(usage): archive completed V6 plans and evidence

* test(maintenance): use a native census project fixture

* test(usage): model unavailable runtime timezone portably

* test(hooks): isolate CLI probes and assert no drift launches

* test(host): frame lifecycle logs as test diagnostics
)

* fix(upstream-watch): retry and pace routine dispatch (#280)

* fix(upstream-watch): retry and pace routine dispatch

The first scheduled dispatch fired 7 released fixes back to back and all
got HTTP 503 without a session, failing the run. The error also dropped
the response body and request id.

- retry 5xx trigger answers up to 3 attempts with backoff (4xx not retried)
- include request-id and the response error message in the failure
- fire at most 3 fixes per run, 15s apart; stop at the first failed call
- defer the rest to the next run without recording an error

* feat(upstream-watch): report fixes deferred to the next run

Surface the dispatcher's deferred list in watch.json, the console output
and the workflow summary so a backlog is visible, not silent. Pause between
firings uses the script's injectable sleep.

* fix(upstream-watch): bound trigger retries and redact error metadata

* fix(upstream-watch): preserve deferred backlog in previews and blind results

* docs(upstream-watch): explain bounded dispatch and deferred evidence

* fix(upstream-watch): stop retries on response body transport failures

* docs(upstream-watch): archive reviewed main reconciliation
Main cb54f42 was incorporated and independently reviewed in PR #283. Its squash 9dc018b has exactly the reviewed feature tree. This content-identical merge records that existing main ancestry after the required feature squash; it introduces no source changes.
…t hygiene (#284)

* docs(plan): define final remediation closeout units

* test(comments): guard durable references and remove transient labels

* docs(closeout): record v2 scope and qualified execution evidence

* fix(tests): distinguish test declarations from ordinary member calls

* fix(tests): exclude table data from test context inference

* docs(closeout): reconcile integration status and upstream evidence

* test(quality): isolate parser guard from unit matrix

* ci(quality): require parser guard after dependency installation

* docs(closeout): archive validated V7 execution plan
@pacphi
pacphi merged commit 0511d57 into main Sep 29, 2026
31 checks passed
@pacphi
pacphi deleted the develop branch September 29, 2026 22:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: live view re-entering files, plain-folder bind, and structured live-events need real-machine evidence

1 participant