Skip to content

CI probe: full-suite job on Linux (do not merge) - #357

Closed
stuinfla wants to merge 71 commits into
mainfrom
ci-probe/full-suite-4.4
Closed

stuinfla wants to merge 71 commits into
mainfrom
ci-probe/full-suite-4.4

Conversation

@stuinfla

@stuinfla stuinfla commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Throwaway PR to measure the new required full-suite job on a GitHub ubuntu runner before the 4.4.0 release PR. Will be closed without merging.

🤖 Generated with Claude Code

Fixture added 29 commits September 30, 2026 23:53
…eal only while an update is proven

kit.json ruvnetBrain:true made autoUpdateOptOut stand down forever, but agentic-kit schedules
nothing on the owner's Mac, so the knowledge base aged by hand only. Ownership is now honoured
only while a successful refresh or CURRENT verdict is proven within 36h; otherwise SessionStart
self-heals and the knowledge line names the Brain's own update, never --enable-nightly.
Also de-flakes the 35-day currency test that compared a real-clock age to a fixed NOW.
…t puts the gold store

planSourceRoute is the routing block of searchAllPrimary moved verbatim into an exported
function (git diff -w shows only the move). scripts/route-gold-rank.mjs runs it with no
model and reports gold-in-1/3/5, declines and stores opened with Wilson intervals.
…ole, tolerate a self-updating host CLI

- consoleRestartState prunes a receipt whose pid is dead (ESRCH) and whose port does not answer
  /api/runtime with the same identity (scripts/console-instances.mjs); --doctor re-reads live receipts
  instead of trusting the recorded pending-console-restart snapshot.
- syncHostsAfterUpdate replaces an owned Console still serving the previous runtime by running the
  activated runtime's own launcher (--serve, never --open) in that scope and port, waits for the
  receipt to name the new runtime and the old process to exit, and records exactly why when it cannot.
- claude/codex invocations go through scripts/host-cli.mjs runHostCli: a CLI seen this run (or whose
  PATH link dangles) that is missing is retried 3 times over ~20s, then reported as one line; a CLI
  that was never installed still fails at once.
…the update path

From the customer-state-matrix agent's measured patches (2026-09-30), unchanged except where noted:
1 update-apply: a machine with no host and no active spine has nothing to converge (exit 0, not 1).
2 refresh-run: physicalPath(); lock inheritance compares directories, not spellings (symlinked cache).
3 update-storage-transaction: a kill mid candidate-build quarantines the unsealed candidate instead of
  RECOVERY_REQUIRED forever; the quarantine is released one cycle later only if bytes are unchanged.
6 install --update: recovers an interrupted storage transaction before looking for the updater; no fresh
  reinstall fallback on retained-copy refusal or a transient manifest failure (403/408/429/5xx/offline).
7 corpus-canary: opt-in private-overlay and older-runtime cases (default stays clean) + the
  customer-state-matrix harness (registered in wired-check as a human-run harness).
Also repairs corpus-canary's BREAK-IT test, red on origin/main: its literal no longer matched
carryLiveNodeModules()' guard line.
…data per KB build

metadataSourceRoute broke a top-overlap tie by repository name and kept one store. It now keeps
up to METADATA_ROUTE_TIES (3) stores tied at the best overlap, ordered by how many of their
entries reach it. Route-only on the 206-need set (kb 4.3.37): gold in top 3 42/206 -> 83/206.

The metadata scan reads an in-process inverted index keyed by the KB build identity (manifest
stat + generated stamp) and the store file's dev/inode/mtime/size, so an update never serves the
previous build's index even when a rewritten store keeps its mtime and size.
…s into the brain

Matrix patch 6 (revised): run recoverIncompleteStorageTransactions whenever the transaction ledger
exists, not only when forge-update.mjs is missing. The preflight re-stamps RUNTIME-IDENTITY.json, after
which recovery can never prove live equals the identity sealed at LOCKED, so a pre-activation kill wedged
every later update at RECOVERY_REQUIRED.
…e gist store alone

An unhyphenated "rUv" was read as provenance intent, so "which rUv library runs vector search
in the browser?" searched ruv-gists only (5 of 6 product probes on the 4.3.37 corpus). It now
signals provenance only next to something rUv authored or said (publish/write/post/announce,
tutorial, spec, post ...). After: those 5 route to agentdb, ruvector, ruv-fann, ruview,
agentic-flow; all 7 provenance probes still route to ruv-gists.
…with store identity stripped

One described need per store (180+ stores) so a routing change is validated on far more than the
three repositories the novice need set covers.
…rain home; test --doctor's live re-read

Matrix patch 4 (kb/brain-profile.mjs): published names with dots (dspy.ts, ruv.io) are valid store
names (still no leading dot, '..', separator or .big/.rvf suffix), and stores in the release's
PUBLIC-RVF-GENERATIONS.json (concepts, ruv-gists) are release-managed unless fenced private or opted out.
Matrix patch 5 (kb/lifecycle-evidence-retention.mjs): the brain home (the evidence roots' parent) may be
a symlink to another disk; references are compared by physical path. An evidence ROOT that is itself a
link is still refused — new TEETH test through a real and a linked brain home.
--doctor: the live re-read of Console receipts is now withLiveConsoleState(), exported and tested.
…STER #26)

canonical-qa gains a full-suite job: scripts/full-suite-gate.mjs runs every file
vitest.config.mjs includes and fails on any red not in the reviewed
tests/known-red.json (class, evidence, owner), on a quarantined test that now
passes, on file-level errors, and on any test file vitest silently never runs.
Before this, 552 of 589 test files ran in no automatic gate.

Also: ruflo-daemon-autostart reads the version of the binary it executes, not a
hard-coded $HOME path; drop a dangling QE lane path and fix a wrong
executed-by comment in corpus-seed.yml.
…uth green)

install-release-fallback: the rate-limit reset time is built from a Date and asserted with toContain.
customer-state-matrix: N-k expectations are read from the fixture release list, not typed versions.
no-restated-truth was red on main for the first and on this branch for the second (carried with patch 7).
…me each known-red owner

lesson-migrate-agentdb asserts the developer's own AgentDB stores through the
real ruflo CLI; it was red in every environment but the owner's. It now lives
with the other diagnostics under tests/diagnostics/vitest.config.mjs and is an
explicit excludedFiles entry. tests/known-red.json is transitional: every entry
names its owning branch, and a release ships with it empty.
…--update

Matrix D8 / patch 8: an installed updater run directly (cd kb && node forge-update.mjs --apply) can
predate the current candidate layout and fail the guard (measured on 4.3.28 vs 4.3.39); the npx door
places the current updater first. Changed: the SessionStart BEHIND hint (behavioural test through
heartbeat()), kb/forge-update.mjs --check's 'A newer build exists' line (asserted in the real --check
run in forge-update-apply-rollback), kb/forge-currency.mjs's brain line, docs/ARCHITECTURE-MAP.md.
…orts FLAKY

Measured: a full run at load 100-150 turned four budget/latency tests red that
pass alone (restore-semantics 3745ms vs 1000ms, decision-gate 2000ms budget,
ux-qe probe, issue-83). A test red in the full run but green in one isolated
retry is listed as FLAKY in the verdict, never hidden; red in both runs, or
absent from the retry, stays RED. Sabotage-tested.
…s final)

Local paths are replaced by placeholders (<kb>, <scratch>, <worktree>, ...). The final run's
off-topic and held-out outputs were byte-identical to the baseline's; one copy of each is kept.
…f 4.3.40

Reapplies, without conflicts, the hook commit reverted out of release/4.3.40
(cd385f3, saved as .claude/worktrees/hook-fixes.patch) and the adversarial
review's work-in-progress corrections (hook-review-fixes-wip.patch):

- grounding-stamp.sh: success markers and refusals read from tool_response
  only (a model-written query can no longer forge a stamp); refusals checked
  before any marker; fast-lane card and host-replaced oversize results
  recognised; oversize file read only under $HOME/.claude/projects/*/tool-results/
- grounding-turn-evidence.mjs: brainAnswered() recognises the same shapes
- set -u: initialise the read variable in design-wall, ground-before-write,
  grounding-stamp, protect-brain-state, route-dispatch
- ground-ruvnet.sh: 2>/dev/null before the file redirect (no shell error text)
- codex-hooks.json: SessionEnd launcher 2500ms / 3s (the host's hard cap)
- scripts/hook-qualify*.mjs + fixtures + tests: the hook qualification matrix
  (its path-leak guard assembles the owner's username instead of carrying it)

Follow-up commit finishes the review corrections and the Stop false alarm.
…on a topic; finish review corrections

THE FALSE ALARM. Gate 1 arms on any prompt naming the rUv stack, and in this
repository nearly every prompt does, so the Stop gate said "no successful
search_ruvnet call is in this turn's transcript" on release-status, git/CI,
disk/backup and memory-write answers. grounding-turn-gate now demands a search
only when the final answer asserts what a rUv product does/can/cannot/requires
/says (ruvCapabilityClaims, deterministic, no model). Measured on hand-labelled
real Stop points (kept outside the repository) through the real decide(),
origin/main vs this commit:
  183 real deliveries of the correction: false positives 158/176 -> 0/177
  held-out set (70 Stop points, never tuned on): FP 68/68 -> 0/68,
  FN 0/2 -> 2/2 (both implicit shapes)
tests/fixtures/grounding-turn-stop-points.json is a MINIMAL public-safe
fixture: 45 deciding excerpts (7 must-fire claims, 6 pinned known misses,
3 grounded claims, 29 hard negatives across release/git-CI/disk/memory), no
paths, usernames or sensitive strings, guarded by a test that refuses them.

Review corrections finished:
- grounding-turn-gate: a long turn whose tail cannot see the turn start falls
  back to the stamp evidence (turn = null) instead of returning null (silent pass)
- brainAnswered(): oversize file read only under $HOME/.claude/projects/*/
  tool-results/, never a symlink, never with a refusal before the answer
- stamp forgery tests per case: symlink, `..` path, case-forged markers
- Codex SessionEnd: wrapper budget 2200ms (inside the 2500ms launcher, 3s host
  cap) handed down; session-snapshot plans inside it, captures the NEW snapshot
  first and defers outbox replay under 4000ms (deferred, never dropped)
- set -u: learn-capture.sh and kling-preflight.sh (red on macOS bash 3.2);
  the regression test discovers every set -u + timed-read hook and spawns
  bash through resolveBash(), never a literal path
- grounding-turn-replay.mjs reports gate1WouldFire under the new rule plus
  gate1ArmedUnsearched (the old count); CONTRIBUTING and hooks.json describe it
- wired-check: register scripts/hook-qualify.mjs (human-run matrix)
…n update

identifierScan keyed its warm-process cache by directory path, so after an update swapped kb/
under the same path the MCP worker kept answering with the previous build's stores. It now keys
on the same KB build identity the router metadata index uses (extracted to
kb/kb-build-identity.mjs) and drops the old build's scans.
# Conflicts:
#	data/convergence-manifest.json
#	scripts/wired-check.mjs
# Conflicts:
#	data/convergence-manifest.json
…back, not export

The real-AgentDB progression test ended with a `ruflo memory export` probe that
returned count 0. Root cause is the probe, not the product: export has no --path
and on ruflo 3.49.0 ignores CLAUDE_FLOW_DB_PATH/MEMORY_PATH, reading a
cwd-derived store (0 rows from the project root, 1 from <root>/.swarm). The same
two product-written rows read back byte-identical through `memory list --path`
and exact `memory retrieve --path` on the canonical store (the --path surfaces
ruflo wires per scripts/smoke-memory-db-path.mjs, #2105).

The probe now reads every row back through the CLI with --path and compares
canonical digests. Red when the product writes outside the documented namespace
(capture/restore stay self-consistent; the CLI readback fails), green on the
product as shipped. Drops its tests/known-red.json entry.
# Conflicts:
#	data/convergence-manifest.json
#	tests/known-red.json
….swarm store is created

ruflo creates <cwd>/.swarm on every invocation even when --path names the store. The progression
store ran it from inside <root>/.swarm, leaving an unused nested store in every project (measured
on ruflo 3.49.0). When the store is the default <root>/.swarm/memory.db, run from the project root
so that directory is the store's own. Proven: the real-AgentDB test fails on the old cwd (nested
store present) and passes now; a unit assertion pins the cwd.
…ling

Review finding S4: any test that failed in the full run and passed alone was
waved through, unbounded and unrecorded, so a value-assertion red caused by
real shared state could hide. Now a test is FLAKY only when its full-run
failure message is a timing/budget failure (vitest test/hook timeout, or an
explicit budget/ceiling/latency/'took Nms'/'exceeded Nms' assertion) AND it
passed the isolated retry. Every FLAKY is recorded with its failure in the
JSON verdict and the GitHub job summary; more than 3 in one run fails the gate.
Each rule is sabotage-tested. full-suite job cap 75 -> 120 min.
… the supplied KB; restore aliases without a public repo-aliases.json

S5: the private-overlay canary test compared the verdict with all-checks-ok, a tautology; it now names
every check green and private-store-preserved byte-identical. Making it real exposed a product defect it
had hidden: restorePrivateOverlayState read the candidate's repo-aliases.json unconditionally, so a public
bundle without one failed every private-overlay update (ENOENT). It now starts from an empty map.
S6: with --installed-kb, extra cases reused and mutated the supplied KB. suppliedKbInstaller gives each
extra case its own copy of the pre-clean snapshot, and runCanary refuses an extra case handed the clean
case's KB.
# Conflicts:
#	data/convergence-manifest.json
@vercel

vercel Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
explainer Ready Ready Preview Oct 1, 2026 10:44am UTC
ruvnet-brain Ready Ready Preview Oct 1, 2026 10:44am UTC

Request Review

# Conflicts:
#	data/convergence-manifest.json
Fixture and others added 20 commits October 1, 2026 06:02
# Conflicts:
#	data/convergence-manifest.json
… Brain home

Re-review blocker: ruflo writes snapshot content (hnsw.metadata.json) into <cwd>/.swarm and loads any
metadata it finds there. One shared /tmp/ruvnet-brain-ruflo-cwd-<uid> pooled every project's text and,
on a shared /tmp, could be pre-created or symlinked by another user. rufloCwdFor() now returns
<RUVNET_BRAIN_HOME or ~/.cache/ruvnet-brain>/ruflo-cwd/<sha256(store path)>, and every directory it
creates is verified by ensurePrivateDir(): a real directory (a symlink is refused), owned by this uid
(else refused), mode 0700 (repaired). The test suite points the root at a private temp dir
(RUVNET_RUFLO_CWD_ROOT in vitest.config.mjs) so no test touches the real Brain home.
…ps the user's keys and proxy

S-B: a live pid whose port is silent (a busy Console) or answers another identity (pid reuse) is now a
reported failure with the receipt kept — it is never deleted, so a live stale Console cannot be
orphaned; only dead-pid + silent-port receipts are pruned (readConsoleReceipts). Still reported at once,
with no launch and no 20s wait.
S-C: consoleEnv is a denylist of this run's flags (RUVNET_NIGHTLY*, test/import switches, the refresh
token, NODE_OPTIONS, npm_*/VITEST*), so provider API keys, HTTP(S)_PROXY/NO_PROXY and
NODE_EXTRA_CA_CERTS the Console reads reach it.
…Y timing class

- pre-commit-convergence-heal ran 'git config user.name/email' inside a linked
  worktree of the real repo; linked worktrees share the main .git/config, so
  every run rewrote the owner's identity to Fixture (143 commits since
  2026-09-27). Identity is now passed per command with -c. A guard test runs the
  heal test in a separate process and asserts the real repo's --local
  user.name/user.email are unchanged (it restores them before failing): red
  with the old lines reinstated, green now. No other test configures a
  worktree of the real repo; the rest configure throwaway repos.
- full-suite-gate: FLAKY timing class is now only a vitest test/hook timeout
  or a numeric bound comparison carrying a duration cue. Budget/ceiling/latency
  words in a value assertion stay RED. Proven with real failure messages; red
  when the old loose rule is restored.
…ink clone, else full copy)

NIT 7: copies already use COPYFILE_FICLONE; the doc now says what that costs without reflink (about 3x
the KB at peak) and why hardlinks are not used (an in-place rewrite would reach the caller's brain).
…detector, degraded, node-path, windowsHide nits

4.4.0 re-review.

S-A. Ordering was not guaranteed: a boundary counted only outbox rows as older
work, so with an older boundary queued and a worker holding the lock, a new
boundary at 1900 ms produced+captured inline and at 30000 ms replayed and
captured inline, and the worker later committed the OLDER snapshot after the
NEWER one. The replay lock is now the right to commit in order: every boundary
takes it first; if a live worker holds it, or captures are queued (any
budget), or a short budget has outbox debt, the boundary queues itself behind
that work and hands the lock to a detached worker; otherwise it works inline
while holding it. The lock carries an owner token: only the owner refreshes
(heartbeat between every step) or releases it; a worker whose lock was taken
over stops at once; a lock not refreshed for 120 s is taken over (rename-aside,
then exclusive create). Any boundary therefore drains a queue a dead worker
stranded. Per-step budget 45 s (< half the stale time).

NIT 3. "will not" counts only with a capability verb, "now" only with a
capability verb, and a same-sentence pronoun only after "X is a/an/the ..." or
"X has/provides/ships ...": the reviewer's four false positives are silent,
the S2 recalls still fire. Real labelled sets unchanged: 183 firings FP 0/177
FN 0/6; held-out 94 FP 0/89 FN 2/5; held-out 70 FP 0/68 FN 2/2.
NIT 4. The degraded paragraph is matched to its fixed closing sentence, so a
repo error spanning lines no longer fails a real answer closed.
NIT 5. hook-shim passes its own node as RUVNET_NODE_BIN; grounding-stamp.sh
prefers it over PATH (and no longer needs dirname).
NIT 6. The detached worker is spawned with windowsHide.
# Conflicts:
#	data/convergence-manifest.json
# Conflicts:
#	data/convergence-manifest.json
… authorship words

Review found product questions that still reached ruv-gists alone through words like share,
threads, specs, writes, writing and post. ruvAuthorshipIntent now requires rUv as the subject of
an act of saying or publishing, or rUv's written artifact. Six review questions are negative
tests (red under the previous rule); eight provenance controls reach the gist store only through
this rule (red without it).
…ver a literal path

entrypoint-guard-safety refuses a literal /bin/bash spawn; the new hook-shim
node-path test's PATH precondition used one.
…scan cache is bounded

The private overlay, on-demand ingest and forge-refresh rewrite .passages.jsonl without touching
manifest.json, so a build-identity key alone served stale scans after them. The key now adds a
fingerprint of every passage sidecar (name, dev, inode, mtime, size); the build identity still
covers a rewrite that keeps all four. Scans are LRU-bounded at SCAN_CACHE_MAX (64).
# Conflicts:
#	data/convergence-manifest.json
Linux probe: readSourceIdentity digested untracked files from git ls-files --others --exclude-standard,
which includes .swarm/memory.db (and -wal/-shm, outbox, queue) whenever the customer's project does not
gitignore .swarm. Every capture wrote memory.db, so every later boundary looked changed, the no-op path
was never taken, and memory.db grew at every boundary. It passed on the owner's Mac only because the
global gitignore lists .swarm/. The tracked, untracked and dirty digests now exclude .swarm,
.claude-flow and ruvector.db by pathspec, independent of any gitignore. Tests run git with an empty
GIT_CONFIG_GLOBAL and GIT_CONFIG_NOSYSTEM=1.
…ncy and memory harnesses

One token dictionary per KB directory build and three typed arrays per store (CSR) replace a
Map<token, Uint32Array> per store. scripts/route-index-memory.mjs on the 4.3.37 corpus (199
stores, --expose-gc, two runs each): retained 162.5 MB -> 43.9 MB; cold build 2.2-2.5 s both;
warm call 2-3 ms both. Indexes for at most two KB directories are kept (LRU), so alternating
directories does not rebuild.

scripts/route-latency-warm.mjs is the paired warm latency harness behind the routing 4.4
numbers; --summarize over the committed rows reproduces latency-warm-abc/summary-paired.json.
# Conflicts:
#	data/convergence-manifest.json
# Conflicts:
#	data/convergence-manifest.json
# Conflicts:
#	data/convergence-manifest.json
@stuinfla

stuinfla commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Probe complete: full-suite green on Linux (runs 4 and 5). Closing without merge; 4.4.0 ships from release/4.4.0.

@stuinfla stuinfla closed this Oct 1, 2026

This branch was successfully deployed

2 active deployments
Preview – explainer — 66222aaf Deployed Oct 1, 2026 by vercel[bot]
Preview – ruvnet-brain — 66222aaf Deployed Oct 1, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant