Skip to content

test: fail the suite when tests touch real user state - #245

Merged
pacphi merged 28 commits into
mainfrom
fix/test-hermeticity
Sep 27, 2026
Merged

pacphi merged 28 commits into
mainfrom
fix/test-hermeticity

Conversation

@pacphi

@pacphi pacphi commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

Tests wrote real user state four times during the #237/#238 work: the statusline version and environment pins, the maintenance state, and a heal receipt. This branch makes that a failing run instead of a silent write.

  • Real-state tripwire. pnpm test and pnpm run test:ui now run through scripts/run-tests.mjs. It hashes ak's real config and state folders (POSIX and %APPDATA%/%LOCALAPPDATA%), the files ak manages in other tools' homes, and the repository's .claude, .swarm, .agentic-qe, .claude-flow and .harness before and after the run. It exits 3 and names each changed path. Locally, files a live Claude Code session writes on its own (statusline tee, ~/.claude.json, Ruflo/AQE hook stores) are listed, not failed. CI and AK_TRIPWIRE_STRICT=1 fail on any change.
  • Suite temp root. The suite runs in its own temporary folder, refuses a temp root inside a git repository, and fails if a test leaves a folder behind. 59 test files that leaked temp folders were fixed.
  • Dashboard. An injected test server with any collector but no control root now refuses the default maintenance and management services (logged, 503). This closes the hermeticity gap recorded in the audit.
  • Spawned processes. spawnEnv, sandboxHome and redirectToolState keep spawned processes and in-process tests away from the developer's XDG_*, OpenCode and Codex state. A static guard fails on any child_process call that inherits process.env implicitly.
  • Live test. The memory-routing live test runs in a disposable home and project.
  • UI. The stale polyglot-badge test now follows DDD-09. maintenance-focus.mjs and maintenance-guidance.mjs join test:ui. Two flakes are fixed in the product:
    • a closing health dialog no longer pulls focus off the next badge (host-readiness);
    • the Maintenance results list is marked busy while an inventory query runs (MNT-UX-007).
  • Docs. AGENTS.md, MAINTAINER.md and the audit record's Branch 2 implementation status.

Verification

  • Guarded runner, Node 26.4: 5,151 pass / 0 fail. Two consecutive full passes before the rebase, 5,140 / 0 each, with the fingerprint unchanged across both.
  • Node 22.22.3: exit 0.
  • Seven .cjs suites pass. UI: 495/0 and 11/11.
  • Coverage: 93.82 / 82.45 / 93.03. tsc, eslint (0 errors), complexity ceiling, markdownlint, build-check, doc-citations and ga-surface-guard: all clean.
  • An injected write to a watched path fails the run with exit 3 and names the path.
  • Rebased on fix(discovery): keep imported Codex copies out of project origins #244. The guarded run is clean, including fix(discovery): keep imported Codex copies out of project origins #244's new tests.
  • Windows roots and the case-insensitive root merge have unit tests only. This CI run is their first real Windows check.

Refs #239 · remediation program Branch 2

🤖 Generated with Claude Code

Adds scripts/real-state-tripwire.mjs: a builtins-only library and CLI that
snapshots ~/.config/agentic-kit, ~/.local/state/agentic-kit, the
APPDATA/LOCALAPPDATA equivalents, the repo's .claude/.swarm/.agentic-qe/
.claude-flow/.harness and the ak-managed guidance files, then names each
created, removed or rewritten path. Live-session writers are reported as
concurrent (not failing) outside CI/AK_TRIPWIRE_STRICT.
pnpm test and pnpm run test:ui now run through scripts/run-tests.mjs, which
snapshots the real-state roots before and after the same commands the old
&& chain ran, keeps the first failing command's exit code, and exits 3 when
a test wrote real user state. AGENTS.md documents the tripwire, concurrent
writers and AK_TRIPWIRE_STRICT=1.
… injected test servers

A server given any injected collector (fetchStatus, usage, limits, hooks,
live, transcripts, intelWatch, discoverProjects, machineWideIntel, models,
system, hostReadiness) and neither an injected maintenance service nor
maintenanceOptions.controlRoot now refuses the default maintenance service
(logged once, 503 on its routes) and the default management facade, so a
deep System refresh can no longer write the real control root. The real
ak dashboard command injects none of them and is unaffected. No existing
dashboard caller relied on the real control root.

The lookbackDays citation in docs/USAGE-SCORECARD-METRICS.md moves with the
25 added lines (1824 -> 1849).
Adds spawnEnv/sandboxEnvFor/INHERITED_STATE_KEYS/redirectToolState to
tests/kit/helpers/home-sandbox.mjs (sandboxHome now builds on sandboxEnvFor,
leaving the temp keys alone) and a static guard, spawn-env-guard.test.mjs,
that fails on any `...process.env` / `env: process.env` in tests/ without a
`spawn-env: inherits (<reason>)` marker.

The guard found 30 sites (the plan estimated 22; none are in tests/*.cjs).
24 now build their child env with spawnEnv(home, extra); 6 carry the marker:
the hostile-env sandbox tripwire, the npm install in the stock-OpenCode test
(needs ~/.npmrc; cache pinned), two paid live AQE release-proof sites and the
two live disposable-memory-project spawns (the caller sandboxes process.env).

adapter-aqe-provider's hidden-transport test overrode only HOME, so its child
read the developer's XDG_CONFIG_HOME kit.json; it still reports "not active"
inside the sandbox. opencode-stock-ruflo-gateway's `opencode --version` probe
passed no env at all (a form the guard cannot see) and created OpenCode's
state/data/cache/temp folders under the inherited bases; it now runs in the
test's sandbox home too.
integration-command-facts, opencode and provider-credentials call code that
spawns the real `opencode` with process.env; each now runs under
redirectToolState(), so OpenCode's state/data/cache/temp folders land in a
per-file temp base that is removed afterwards. The three direct-spawn files
(opencode-stock-ruflo-gateway, provider-cli, provider-refresh-cli) were fixed
by the spawnEnv conversion in the previous commit; a per-file probe with the
XDG bases and TMPDIR redirected shows no leftovers for any of the six.

opencode-state-hermeticity.test.mjs runs the three in-process files in a
nested `node --test` and requires the watched bases to stay empty. The
nested run must not inherit NODE_TEST_CONTEXT: with it, the child runs no
test and the check passes vacuously (observed), so the variable is dropped
and the child's TAP output must report at least one pass.
Adds tests/kit/helpers/temp-dir.mjs: tempDir(prefix, t?) makes a real-path
folder under os.tmpdir() and removes it when the test (or, without t, the
file or the calling test) ends, with maxRetries for Windows EBUSY.

A per-file probe (each tests/kit file alone with its own TMPDIR and XDG
bases) found 53 leaking files at this branch's head (the plan's table had 59;
the six OpenCode files were fixed by the previous two commits). 52 test files
plus tests/kit/helpers/codex-rollout.mjs change:
- inline fs.mkdtempSync(path.join(os.tmpdir(), ...)) sites become tempDir();
  conformance-tiers imports it as makeTempDir (its locals are named tempDir);
- every sandboxHome()/sandboxProject() folder without a removal gets
  after(() => rmrf(...));
- codex-plugins nests its root one level down: its traversal case resolves
  cache/../../outside-market, which landed in os.tmpdir() itself;
- statusline runs under redirectToolState('ak-sl'), so the rendered footer's
  30 s caches ruflo-daemon-count.json and ruvnet-brain-kb-size.json stop
  landing in the os.tmpdir() the developer's live footer reads.

After: the same probe over those files reports no leftovers and no failures.
…eftovers

runGuarded now makes an ak-suite-* folder under os.tmpdir(), passes it to
every command as TMPDIR/TEMP/TMP, lists and fails (exit 4) on anything left in
it except node-compile-cache, and always removes it. It refuses to start
(exit 2) when that folder sits inside a git repository, because the "outside
a git repository" tests would then write into the enclosing one.

Running the seven .cjs suites through `run-tests.mjs exec` found leftovers the
tests/kit probe could not see (statusline-segments 46, statusline-brain 21,
statusline-window-ledger 12, dashboard 2, including the footer's
ruflo-daemon-count.json / ruvnet-brain-kb-size.json caches). Those plain-script
suites now call usePrivateTmpdir() from tests/kit/helpers/private-tmpdir.cjs,
which points os.tmpdir() at a private folder removed on exit. The UI suite
gets the same in dashboard-ui.mjs, and maintenance-host-alignment.mjs writes
its screenshot to .ui-artifacts/ (as dashboard-project-context does) instead
of the temp folder. After: all seven .cjs suites and `run-tests.mjs ui` exit 0
with no leftovers.

AGENTS.md (Running Tests) describes the temp root, the refusal, tempDir() and
spawnEnv().
Probe: every tests/kit file run alone with TMPDIR inside a fake enclosing git
repository. Two files wrote into it: aqe-embedding-projection (the two
"outside a git repository" cases: .claude/settings.local.json and its
ownership sidecar) and status-aqe-drift (the no-project case: .claude and
.agentic-qe/llm-config.json). Both cases now skip with the enclosing
repository named when repoRoot() of their fixture is not null; with the
default TMPDIR they run as before (28/28 pass, 0 skipped).

opencode-stock-ruflo-gateway gives the workspace OpenCode runs in its own
.git, so OpenCode can never climb to an enclosing package.json/node_modules.

Eleven other files fail (without writing) when TMPDIR is inside a
repository; the suite runner refuses to start in that layout, so they are
recorded in the slice report rather than changed here.
tests/live/ruflo-memory-routing.test.mjs now calls sandboxHome() before any
kit module loads (restoring PATH so the real ruflo, git and npm are found),
imports the kit modules dynamically, and runs the probe in
createDisposableMemoryProject()'s git-initialised, daemon-off project, removed
by its cleanup(). A new first assertion requires paths.home to lie under
os.tmpdir(); it failed against the old module head ("not /Users/...").

The test also drops inherited RUFLO_* and CLAUDE_FLOW_* variables: a shell or
Claude Code session in an ak-managed project exports CLAUDE_FLOW_DB_PATH at
the real store and RUFLO_MCP_ENFORCE_POLICY=1, which made Ruflo 3.46.1 fail
closed ("RUFLO_MCP_ENFORCE_POLICY is set but .harness/mcp-policy.json is
missing") in the disposable project.

Run on macOS with Ruflo 3.46.1 (XDG_* unset): 1/1 pass, observed routing
cliToMcp=visible, mcpToCli=not-visible, mcpStore=agentdb-memory.db; the
brief's fingerprint and `ls -la ~/.claude-flow ~/.swarm` identical before
and after; no leftover temp folder or process. The paid qe-court live test
stays manual. installedRoutingVersion() still finds the installed Ruflo
(npm's prefix comes from the Node install, not HOME).
…both Maintenance specs in test:ui

The polyglot check failed "0 !== 3" because its harness did not load
maintenance-cards: with a pageerror listener the render throws
"mntProjectKindBadge is not defined" and no card appears. It also still
asserted the retired three-icon cap and "+2 more languages" disclosure
(DDD-09, docs/LANGUAGE-LOGOS.md). The check now loads maintenance-cards
whole (no further stubs needed), fails on any page error, and asserts all
five icons inline, no disclosure, and no horizontal overflow at 390px.

Screenshots from both specs go to .ui-artifacts/ instead of /tmp.
maintenance-focus.mjs and maintenance-guidance.mjs join SUITES.ui; the
runner test pins the full list.
…he next badge

The dialog's `close` event is a queued task. Its handler focused the badge
that opened the dialog unconditionally, so when focus had already moved on
(the UI spec's Escape, then focus the next badge, then Enter), a late
close event pulled focus back and Enter reopened the previous host. That
is the host-readiness flake: under 8-way parallel load the unchanged spec
failed 3 of 32 runs with the previous host's title, e.g.
'Claude Code health' !== 'Codex health'.

The handler now returns focus only when focus is still in the dialog or
on the body. A new check closes the dialog and focuses another badge in
one task, then awaits the close event itself (no timer): it failed
'claude' !== 'codex' before the fix and passes after. The per-host loop
now waits for the open dialog's title and for the closed dialog with the
badge focused, as conditions. Screenshots go to .ui-artifacts/.

After: 20/20 serial runs and 24/24 runs under 8-way parallel load pass.
… runs

MNT-UX-007 failed once on CI ("arrow keys did not move the roving tab
stop"): after Clear all, the check focused a row and pressed ArrowDown
while the Clear all query could still be in flight, so its response could
re-render the rows between the key press and the tabindex read.

#mnt-results now carries aria-busy from mntInventoryBusy, set in
renderMntResults (the element is never replaced, and every busy
transition renders). The UI spec holds the next inventory response,
checks aria-busy="true" while held (MNT-UX-007a), releases it, waits for
aria-busy="false", checks Clear all issued exactly one query
(MNT-UX-007b), and waits for the second row to take focus before reading
tab indexes.

Before: "MNT-UX-007a ... aria-busy read null", then a 30 s wait timeout.
After: 495 passed, 0 failed; 10/10 serial runs pass.
Mark the spawn-test XDG_* inheritance (Task 2.1) and the review-deferred temp-folder and enclosing-repository writes (Tasks 2.3-2.5) as fixed on fix/test-hermeticity.
…g base

integration-command-facts, opencode and provider-credentials still started the
real `opencode` with the developer's XDG_CONFIG_HOME (APPDATA on Windows), and
each run created <config>/opencode there. redirectToolState() now moves the
config base into its temporary folder along with state, data, cache and temp.

opencode-state-hermeticity.test.mjs watches a config folder as well, and a new
deterministic check asserts that redirectToolState() places every tool base,
config included, inside its base and restores them all, so the fix holds on a
machine without opencode installed.
MAINTAINER.md section 4 still described test isolation as two halves. It now
covers scripts/run-tests.mjs (real-state fingerprint, the ak-suite temp root,
exit codes 2/3/4, concurrent writers and AK_TRIPWIRE_STRICT, the exec mode) and
the helpers tests use: sandboxHome(), spawnEnv(), redirectToolState(), tempDir()
and isolateProject(). AGENTS.md keeps the exact list of watched paths.
The comment above each redirectToolState() call listed state/data/cache/temp only; the helper now also moves XDG_CONFIG_HOME and APPDATA.
…hat stays unwatched

The tripwire now fingerprints ~/.claude/settings.json, ~/.claude.json,
~/.codex/config.toml and the repository's root CLAUDE.md, AGENTS.md and
.mcp.json alongside the guidance files it already covered. ~/.claude.json is
a developer-mode concurrent writer because Claude Code rewrites it during
every session; strict mode still fails on it. The header, AGENTS.md and
MAINTAINER.md now list the watched paths and name the tool paths left to
spawnEnv().
… implicitly

The guard only matched lines that spread process.env, so a spawn with no env
option passed it while inheriting everything. It now scans each child_process
call (code only, strings and comments blanked) and fails when the call passes
no env option and its first line carries no spawn-env marker. The thirteen calls
that ran kit code (the CLI binary, the statusline, the diagnostic script and
module imports) now get spawnEnv(home); git in throwaway repositories, inert
node sleepers, PATH probes and the opt-in live tests carry a stated reason.
MAINTAINER.md and the audit record describe what the guard checks and what it
cannot see.
The audit record carried Branch 2 only as inline notes under Open items. A
new section records the guarded runner and tripwire, the watched paths, the
isolation helpers, the plan-item-5 replacement, the commits, the gate-1,
adversarial and post-fix runner results with their fingerprint evidence, and
what stays unproven.
…rocess helpers

ak writes ~/.codex/AGENTS.md and config.toml whatever CODEX_HOME says
(src/lib/paths.mjs:58-60), so the tripwire watched the wrong files when a
developer exported CODEX_HOME. It now watches both homes; the dedupe keeps
one when they match. The header, AGENTS.md and the audit record also name
sandboxHome() and redirectToolState(), which keep in-process tests away from
the unwatched tool paths.
…ertion

The report names native absolute paths, so on Windows the changed file reads
...\maintenance\latest-scan.json and /maintenance\// never matched.
A copy of process.env keeps Windows spellings (Path, SystemRoot, sometimes
Temp), so spawnEnv's exact-name delete and override could leave the parent
value beside a second spelling, and env.PATH read undefined on the copy.
spawnEnv now folds names on win32, keeps the parent spelling of a replaced
key, and takes an injectable env and platform; envValue() reads a name the
way the platform does.
…tay exact

NTFS file IDs put a sequence number in the top 16 bits and exceed 2^53 once it
reaches 32. As Numbers, two files one record apart can read as the same inode,
so the replaced-file check could pass a swapped file; the test's stat.ino++ mock
is a no-op for the same reason, the only mechanism consistent with the Windows
Node 26 failure.
System Chrome leaves com.google.Chrome.chrome_chrome_url_fetcher_.* folders in
its temp dir. On Linux that is TMPDIR, the runner's suite temp root, so the
leftover check failed the ui job on 14 of them. launchChrome() points TMPDIR,
TEMP, TMP and MAC_CHROMIUM_TMPDIR (macOS Chrome ignores TMPDIR) at a folder it
removes on browser.close(); every UI test launches through it, and a unit test
keeps it that way. The runner's leftover check stays strict.
@pacphi
pacphi force-pushed the fix/test-hermeticity branch from fdcaf17 to 43cb0c2 Compare September 27, 2026 18:43
Windows file IDs can exceed 2^53, where two different files read as the
same Number inode. CI run 36341703517 (Windows Node 22) let a source
replaced between inspection and open pass readBoundedFile for this reason.
The hook-audit, hook-remediation receipt and target, adapter hook-file and
maintenance inventory readers now stat with { bigint: true } on both sides
of each in-process identity check and convert size only after the bounds
check. inspectHookTarget keeps its Number fstat for the persisted image and
adds a BigInt fstat for the identity compare; the adapter post-read check
compares mtimeNs instead of whole-millisecond mtimeMs. Recorded identities
(receipt parent dev/ino, discovery partitions) are unchanged.
@pacphi
pacphi merged commit a0b7549 into main Sep 27, 2026
16 checks passed
@pacphi
pacphi deleted the fix/test-hermeticity branch September 27, 2026 19:01
pacphi added a commit that referenced this pull request Sep 27, 2026
…kups (#247)

* docs(plan): Branch 3 Ruflo support window code-level plan

* feat(upstream): record the Ruflo support window on its dependency policy

The Ruflo dependency policy now carries supportWindow (newest 6 minors,
never fewer than those first published in the last 30 days). The loader
validates the optional field: positive integer newestMinors and minDays
and a string basis; anything else invalidates the policy.

* feat(versions): support a rolling window of Ruflo minors (n-5, at least 30 days)

ak status adds a Ruflo support-window row computed from release dates
remembered in kit.json (versionCheck.rufloMinors): inside the window is
info, below it is a fail row that sync repairs by upgrading, and with no
remembered dates the row says the window is not yet known and names
ak sync. A plain read never calls the network. A non-dry, upgrading sync
records each minor's first stable npm publish right after its forced
drift lookup; a failed lookup keeps the old dates. The versions section
gains drift/loadConfig/now seams and sync.run a releaseDatesRunner seam.

* feat(upstream): hold workaround removals until the oldest supported Ruflo has the fix

The watch computes the Ruflo support-window floor from the npm release
dates it reads (reusing a gate's fetch) and the registry's supportWindow.
A released Ruflo entry whose fix version is above the floor moves from
the act-now groups to 'Released, waiting for the support window' with no
dispatch; its 'released' ledger line carries no branch. With no floor
(offline, npm unreadable, no window) nothing is held. The upstream-status
skill names the new group.

* refactor(statusline): remove the retired CVE-counter overlay

Ruflo fixed its fabricated CVE count (ruvnet/ruflo#2694) in 3.32.2, below
the support window's floor. ak no longer injects the getStatuslineData
wrapper, the footer drops rufloLocalSecurity and rufloHonestInsight, and
status and sync drop the statusline/cve subsystem. SEC_WRAP_STRIP stays
for one release so a statusline an older ak patched is cleaned on the
next sync. The registry marks ruvnet/ruflo#2694 adopted with its removal
proof.

* docs(status): drop the stale "#2986 pending" note

Ruflo shipped `migrate fix --agents` in 3.38.2, below the support window.
The scaffold-agents info row now says the installed Ruflo predates it and
points to ak sync. The section gains gaps/fixAvailable seams for a
hermetic test. The registry marks ruvnet/ruflo#2986 and ruvnet/ruflo#2985
(the same ak change) adopted.

* docs(adr): record the Ruflo support window (ADR-0041 §7)

ADR-0041 §7 records the rule, where it lives (supportWindow on the Ruflo
dependency policy), the remembered evidence (kit.json
versionCheck.rufloMinors, written only by an upgrading ak sync) and the
watch's hold. The hook-assurance DDD context gains SupportWindow.
UPGRADING, UPSTREAM-WATCH and TROUBLESHOOTING describe the window, the
waiting-for-window group and the retired statusline overlay; ak sync
--help says when release dates are read.

* feat(status): show whether Ruflo's backup and distillation are running

A live daemon that deferred backup or distillation (CPU load or low free
memory, read from .claude-flow/logs/daemon.log after the last daemon start)
now gets a daemons warning until the job's own metrics are newer. On macOS
a low-memory deferral names ruvnet/ruflo#2935 and is a sync repair;
elsewhere it names the flat threshold key. With no daemon and start-on-use
on, the row says Ruflo starts it on the next ruflo command.

* feat(ruflo-daemon): enable auto-start with Ruflo's supported daemon settings

ak setup no longer turns claudeFlow.daemon.autoStart to false. It writes
flat keys in .claude-flow/config.json before starting the daemon:
daemon.idleSecs 0 below 3.46.0 (ruvnet/ruflo#3194) and, on macOS,
daemon.resourceThresholds.minFreeMemoryPercent 0 (ruvnet/ruflo#2935), never
through ruflo config set (ruvnet/ruflo#3449). It turns init's autoStart
false to true unless kit.json rufloDaemon.autoStart is false, with receipts
in kit.json that ak uninstall uses to restore. A malformed file or a user
value is left alone. ak status reports drift in a Ruflo repository and ak
sync applies it, removing keys a newer Ruflo no longer needs and restarting
a live daemon so it reads them. Setup discloses both writes.

* docs(ruflo-daemon): document managed daemon settings and start-on-use

SETUP and TROUBLESHOOTING describe what ak writes in .claude-flow/config.json
and .claude/settings.json, the new daemons rows, and the kit.json opt-out.
ADR-0016 and the ubiquitous language record the managed daemon settings;
the audit record's Addendum 3 Item 1 gets the 3.46.1 implementation note
and the disposable-project proof. ak setup --help and ak sync --help
mention the daemon step.

* fix(status): offer the daemon repair only where sync performs it

The macOS deferral row promised a sync repair in any project with a live
daemon, but sync manages daemon settings only in a Ruflo repository, so
elsewhere the planned fix could never converge. Outside one it is now the
manual step. Sync also no longer restarts a daemon whose deferred job has
run since: status and sync share one pendingDeferral check.

* fix(security): stop reporting defend as non-functional when Ruflo ships the built-in engine

Ruflo 3.32.2+ falls back to security/builtin-aidefence.js when
@claude-flow/aidefence is missing (ruvnet/ruflo#2670), so a missing
aidefence now costs only adaptive learning and the aidefence_* MCP tools.

- natives: rufloBuiltinDefence() probes the CLI's built-in engine.
- status: warn (sync repair kept) instead of fail when the engine exists.
- verify: run defend with -o json and read the verdict
  (parseDefendVerdict); the text-mode post-detection crash
  (ruvnet/ruflo#3473, still in 3.46.1) is reported as a crash.
- footer: a fourth, silent "builtin" state; no alarm.
- about and healAidefence detail describe what aidefence adds.
- registry: #2670 adopted; #3473 re-checked on 3.46.1.

* feat(setup): drop the init opt-out suppression on Ruflo 3.46.0 and newer

Ruflo 3.46.0 honours --no-codex-detect and --no-skills-sh (ruvnet/ruflo#3167,
PR #3434). rufloProjectInitInvocation(version) passes the flags alone from
3.46.0; below it (and for an unknown version) it keeps --format json and
RUFLO_NO_SKILLS_SH=1, since 3.39 to 3.45 are inside the support window.

The registry's affected ranges now accept <major>.<minor>.x, the
suppression constraint covers 3.38.21 to 3.45.x, and #3167 is adopted
with the removal proof run on 3.46.1.

* fix(ruflo-components): governance is enforced on stdio from Ruflo 3.46.0

Ruflo 3.46.0 wires the MCP policy enforcer into both stdio entry points
(ruvnet/ruflo#3415, PR #3423); 3.45.0 has no evaluateToolCall there. The
not-wired explanation now applies below 3.46.0 (it stopped at 3.44.0),
and the catalogue no longer calls the policy inert.

Live proof on 3.46.1 with ak's policy capped at 2: two memory_stats calls
allowed, the third refused, all three audited. From a subfolder with no
.harness/ every call is refused; that exposure stays open until the
launcher's Claude mode (plan Task 4.1).

* feat(ruflo-components): keep ak's MCP policy file out of git

A committed .harness/mcp-policy.json has no receipt on a teammate's
machine, so reconcilePolicy reports it user-managed there and stops
managing it. Other tools (Agentic QE) commit their own policy file on
purpose, so ak excludes only its own file, never .harness/ as a whole,
and never touches .gitignore.

excludeFromGit adds one anchored line under a "# agentic-kit" comment to
the repository's info/exclude, resolved without spawning git (a linked
worktree's .git file and commondir lead to the main repository's file,
the only one git reads). It runs whenever ak's file is written or already
converged; removing the policy removes only that line and its comment.
Sync reports the change and the trust manifest discloses it.

* chore(upstream): retire ruvnet/ruflo#3166 (doctor reports agent-browser since 3.46.0)

PR #3441 shipped in Ruflo 3.46.0: ruflo doctor reports the agent-browser
CLI (ADR-122) against a 0.27.0 minimum, confirmed in the @claude-flow/cli
3.46.0 tarball and the installed 3.46.1. ak never carried a workaround.

* chore(upstream): record the partial #3193 fix shipped in Ruflo 3.46.0

PR #3420 (3.46.0) makes the daemon parse .claude-flow/config.yaml, but the
memory root still reads JSON only (memory-initializer.js, rechecked on
3.46.1), so ak keeps its memory pin and the entry stays watching. The
entry now names the ak files that cite it.

* docs(ci): correct the nightly note for ruvnet/ruflo#2885

The note said "open as of 2026-08-13". It now cites the Aug 31 triage and
the 2026-09-27 hosted probe on Ruflo 3.46.1 (run 36333572972): 10/10
aborts by default and 10/10 with single-threaded ONNX Runtime sessions.
The step stays continue-on-error. The registry records the comment.

* docs(adr): record governance on stdio and the security wording (ADR-0058)

ADR-0058 Updated 2026-09-27: Ruflo 3.46.0 enforces the policy on stdio
(#3415); ak's boundary corrected from 3.44.0 to below 3.46.0; ak excludes
its policy file from git (section 5, the section 7 table row, Consequences
and Implementation status).

User-facing docs describe the current state: the security rows in
TROUBLESHOOTING (built-in engine, the #3473 crash message), the version-
gated init flags in SETUP and HOST-SUPPORT, the 3.46.0 governance boundary
and git exclusion in MANAGED-TOOLS and DASHBOARD, and the footer's alarm
in README. The audit record gets the slice 3 implementation note.

* test(ruflo-components): keep the git-exclude tests portable to Windows

GIT_CONFIG_GLOBAL=/dev/null has no Windows equivalent path, so the tests
set it only off Windows; -c user.name/user.email already give git what the
tests need.

* docs(upgrading): ak keeps its MCP policy file out of git

Sync and project setup now add a line to the repository's
.git/info/exclude where MCP governance is on; the upgrade note says what
is written, what is never touched, and how to untrack an already
committed copy.

* feat(ruflo-mcp): a Claude mode that pins only the memory location

ak x ruflo-mcp --host claude starts Ruflo at the repository root (else the
folder, else the user-level store) like Codex's mode, but keeps Claude
Code's settings env: no componentEnv and no governance deletion
(ADR-0058 §3). ak's agent-browser config stays. An unknown --host exits 2.
Started from a subfolder, Ruflo now reads <repo>/.harness/mcp-policy.json,
which closes the 3.46.x stdio-governance exposure for Claude sessions
opened in a subfolder (B3-D1).

* feat(mcp): route Claude Code's Ruflo MCP through ak x ruflo-mcp

register() now adds claude-flow at user scope as
`ak x ruflo-mcp --host claude` (with ak's AGENT_BROWSER_CONFIG). ak's
earlier `ruflo mcp start` entry and the launcher entry are both ak's to
replace; any other command, scope or env key stays the user's. When
`ak` is not on PATH, register() changes nothing and returns
reason 'ak-not-on-path', which setup, sync and ak x mcp pick report.
Status warns on ak's old entry (a sync repair when registration is
managed) and, for the launcher, names the store it picks from the
current folder. Hermetic Claude seats use the Claude mode too, and the
memory rows now say the launcher serves Claude Code and Codex (B3-D1).

* fix(memory): Claude-side harvest and setup use the user-level store outside projects

Harvest takes rufloMemoryLocation(cwd) directly: from a repository it
still runs at <repo> with <repo>/.swarm/memory.db; from the home folder,
a temporary root or a tool's own folder it runs in the flat user-level
store with CLAUDE_FLOW_DB_PATH and CLAUDE_FLOW_MEMORY_PATH pinned and
distils that store. An explicit root (ak x verify harvest) is unchanged.
Project setup from such a folder now refuses before spawning anything
("this folder is <reason>; run ak setup from a project folder").
memoryProjectRoot is unchanged for the daemons row and projectMemoryEnv
(B3-D1).

* feat(sync): remove ak's old setup probe rows once, with a backup and receipt

Older ak setup runs left `_setup/verify-<pid>-<ms>` rows (content
setup-verify, namespace _setup) in project and user-level stores; Ruflo
mirrors them into agentdb-memory.db and its own delete leaves the mirror
(ruvnet/ruflo#3450, re-checked on 3.46.1). Status now warns with the
count per store folder, and a new memory-probe-cleanup sync step backs
each affected file up with VACUUM INTO, deletes exactly the matched ids
from both files, and writes a receipt under the state folder. kit.json
cleanups.setupProbeRows records each cleaned file so it is cleaned at
most once. Scope: the current project's store and the user-level store
(B3-D2).

* docs(audit): record Branch 3 decisions and the Claude launcher route

The audit record gains the four Branch 3 decisions (B3-D1 to B3-D4) in
its decision format, with the implementing commits and the disposable
proofs. ADR-0058 §3 and ADR-0016 record that Claude Code reaches Ruflo
through ak x ruflo-mcp --host claude. SETUP, HOST-SUPPORT, UPGRADING,
TROUBLESHOOTING and the DDD glossary describe the launcher for both
hosts, the ak-on-PATH requirement and the one-time probe-row cleanup.

* docs(adr): ADR-0017 no longer calls ruflo mcp start ak's Claude and Codex registration

* fix(ruflo-mcp): start ruflo through the resolved Windows shim

The launcher that Claude Code and Codex register spawned a bare `ruflo`.
On Windows that is an npm .cmd shim CreateProcess cannot start, so every
launch failed with ENOENT while readiness (have/run, which resolve the
shim) passed. It now resolves the command through exec.mjs resolveShim
against the launch env's PATH and refuses with a message when no safe
invocation exists.

* fix(status): report a daemon config.json ak leaves to the user

A malformed .claude-flow/config.json, or one holding the user's own
value for a key ak wants, is left untouched and never counts as drift,
so ak status said nothing about it. The dry-run reconcile now names the
held keys (reconcileRufloDaemon `held`, daemonConfigHeld), and status
adds a manual row that names the file, the user's value and the flat
key to set.

* fix(daemons): offer the macOS threshold fix only when ak manages it

When .claude-flow/config.json was malformed or held the user's own
minFreeMemoryPercent, status still offered the macOS deferral as a sync
repair. Sync then changed nothing and only restarted the daemon, which
hid the row until the next deferral and repeated on every sync. Status
now makes that deferral a manual step naming the key, sync skips the
restart when the floor is held by the user, and its warning says the
daemon was not restarted for that key.

* fix(status): make the mcp fixes manual while ak is not on PATH

register() refuses with 'ak-not-on-path' when it would write a
registration that starts `ak`, so the mcp sync step failed on every run
while status kept offering it as a sync repair. When ak manages the
registration, status now looks ak up and, when it does not resolve,
makes the three rows whose fix goes through register() manual: the
registration that does not start through the launcher, the missing
agent-browser config and the missing registration. Each names "put ak
on PATH, then run ak sync". The legacy-migration row stays a sync fix:
it does not need a new registration when claude-flow is already right.

The offline status fixture has an empty PATH, so the golden snapshot
now records the manual row. The two sync plan tests that expect the mcp
heal run with a fake ak on PATH, and a new test pins that the heal is
not planned without it.

* docs: status names a held daemon key and the ak-on-PATH step

SETUP, TROUBLESHOOTING and ADR-0016 say a .claude-flow/config.json that
is unreadable or holds the user's own value is left alone, ak status
names the key to set, and sync does not restart the daemon for it.
UPGRADING says status lists "put ak on PATH, then run ak sync" when the
launcher registration cannot be written.

* test: keep Branch 3's tests inside the suite's sandbox rules

After the rebase onto #245, the guarded unit run failed on two rules
that the branch's own tests broke. The git-exclude test built git's
environment from process.env, and the spawn-env guard now refuses that.
It now runs git through spawnEnv with its own throwaway home. Three test
files left their sandbox homes behind, and the leftover check reported
ak-about-security-home, ak-probe-cleanup-home and
ak-security-status-home. Each file now removes its home when it
finishes.

* fix(mcp): register the Claude launcher only when the PATH ak supports it

Claude Code starts whatever `ak` it finds on PATH. That can be an older
install than the kit writing the registration. ak 4.0.0-alpha.56 rejects
`--host` (exit 2, "Unknown option"). So a sync run by a newer kit
re-registered claude-flow to a command that cannot start, and every
Claude session lost Ruflo.

register() now asks the PATH `ak` for `ak x ruflo-mcp --help` before it
writes anything. `--help` is answered before the command runs. When the
help does not name `--host`, it returns the new reason
'ak-launcher-outdated' and leaves the working registration in place.
The status mcp section asks the same question, but only when sync would
re-register. It then makes those rows manual, with the fix "update the
`ak` on PATH ..., then run `ak sync`". Setup, sync and `ak x mcp pick`
name the new reason. `ak` joins exec.mjs's Windows shim list, so the
check resolves its npm shim.

The hermetic Claude seat of `ak run` and qe-court now starts Ruflo
through the running kit itself (node plus this kit's
bin/agentic-kit.mjs), never a PATH `ak`.

* test(mcp): convergence expects Claude Code's launcher registration

"fresh dual-host Codex provisioning and repeated refresh retain one
canonical connection" went stale when claude-flow moved onto
`ak x ruflo-mcp --host claude`. It had two faults. It injected no
launcher check, so on the sandbox's empty PATH register() refused
before the runner ran (false !== true). It also still expected
`ruflo mcp start`. It now injects a ready launcher and expects the
launcher arguments with ak's AGENT_BROWSER_CONFIG. The fake
~/.claude.json gets the launcher entry with that env, so later
refreshes find it already in place and claude-flow is added once.

* fix(codex-mcp): disable Codex's imported copy of Claude Code's launcher entry

Codex's Claude config import copies Claude Code's MCP servers by name.
Since B3-D1, Claude Code's claude-flow is `ak x ruflo-mcp --host claude`.
On a machine without the disabled placeholder, Codex therefore gained a
second Ruflo transport. ak recognised only `ruflo mcp start` under that
name, so both tables had no repair kind, the repair plan was empty, and
every sync reported an unrepairable duplicate and exited 1.

A Codex claude-flow table in ak's Claude launcher form now counts as the
same alias. It needs exactly command and args, with at most ak's
AGENT_BROWSER_CONFIG child. Sync disables it in place with the existing
placeholder, after the usual backup and confirmation. Consent is
remembered for that form as well. The Codex-mode launcher under the
claude-flow name stays the user's. The repair reason names the imported
copy instead of calling it a deprecated transport.

* fix(ruflo-daemon): manage daemon settings only in a durable Ruflo project

rufloProjectRoot accepts any Git repository root with a .claude-flow/
folder. Ruflo 3.46.1 does not count a bare folder as a project
(services/daemon-autostart.js:90-123, isRufloProject), because its
startup migration can create one on its own (ruvnet/ruflo#2852). It does
count .claude-flow/config.json. On macOS ak always wants the free-memory
floor, so sync wrote that file in such a repository and made it a Ruflo
project. The next ruflo command there then started a detached daemon.

applyRufloDaemon and the status drift, held and deferral rows now use
rufloDaemonProjectRoot. It requires the same gate plus one of Ruflo's
durable markers: .claude-flow/config.{yaml,yml,json},
claude-flow.config.json, .swarm/memory.db, a claudeFlow block in
.claude/settings.json, or a ruflo/claude-flow server in .mcp.json.
rufloProjectRoot keeps its ADR-0058 role for the policy file.

* docs(audit): mark Branch 3's resolved open items and record fix round 2

The Open items still listed three things this branch resolved. Each is
now marked resolved with its commits:
- Claude-side memory outside a project (B3-D1): 518b4be, 8c91658 and
  8966a18.
- The one-time _setup/verify-* cleanup (B3-D2): ec6c936.
- setup turning daemon start-on-use off (Addendum 3 Item 1): a9d59cd.

The Branch 3 decision notes now cite the SHAs after the rebase onto
main. B3-D1 records the launcher check and the disabled Codex import
copy. The daemon note records the durable-marker gate.

* docs(upgrading): the PATH ak needs a launcher that takes --host

The launcher note said the PATH `ak` "must be this version or newer".
But the released 4.0.0-alpha.56 has the same version number as this
build and still lacks `--host`. The note now names the capability ak
checks (`ak x ruflo-mcp --help` must list `--host`), as SETUP.md does.

* chore(upstream-watch): register the four threads filed on 2026-09-27

* feat(upstream-watch): hold AgentDB fixes until the oldest supported Ruflo bundles them

Decision B3-D5: an AgentDB fix delivered through Ruflo (doneWhen.release.bundledBy)
counts as released for ak only when the oldest Ruflo inside the support window
bundles a fixed agentdb, the same rule Ruflo's own fixes follow.

- fetch: bundled() takes an optional carrier version, so the floor Ruflo's
  dependency chain resolves exactly like the newest one's.
- watch: once the floor is known, resolve what the floor Ruflo bundles for the
  entries the newest carrier was checked for and for bundled entries recorded
  as released with a first fixed version.
- classify: such an entry waits for the window until the floor bundles the
  fix; a floor bundle that could not be resolved is "Could not check". The
  held ledger line carries no branch and a newer Ruflo never changes it.
- render: the held line names what the floor Ruflo bundles.

* docs(schema): document supportWindow and the <major>.<minor>.x affected range

The registry schema now describes a dependency policy's optional supportWindow
(newestMinors, minDays, basis, unsupported) and the affected-range grammar
upstream.mjs applies: an exact version, <major>.x or <major>.<minor>.x. A
package-qualified affected item is noted as reader-only. bundledBy's
description names the support-window floor (decision B3-D5).

* docs(audit): record decision B3-D5

Branch 3's fifth decision: an AgentDB fix delivered through Ruflo counts as
released for ak only when the oldest supported Ruflo bundles a fixed agentdb
(choice A, oldest supported Ruflo; B was the newest Ruflo). The audit records
it with its limit (npm resolves each range to its highest match, so the answer
is a fresh install of the floor Ruflo). UPSTREAM-WATCH.md describes the floor
check and the held group; ADR-0041 section 7 and its Updated line note it.

* test(memory): close the WAL holder before the fixture folder is removed

* test(docs): read ADR headers with LF line endings in the living-plan guard
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant