Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
38 commits
Select commit Hold shift + click to select a range
bfd0185
recursive(run-106): lock Phase 0 requirements and worktree isolation
try-works Oct 4, 2026
147e4a8
recursive(run-106): lock Phase 1 AS-IS analysis
try-works Oct 4, 2026
45899a9
recursive(run-106): lock Phase 1.5 root cause
try-works Oct 4, 2026
1878f31
recursive(run-106): lock Phase 2 TO-BE plan
try-works Oct 4, 2026
b720d5e
recursive(run-106): lock Phase 2 TO-BE plan after traceability audit
try-works Oct 4, 2026
2c040dc
recursive(run-106): SP1 effort-policy normalization (RED-GREEN)
try-works Oct 4, 2026
3e63af4
recursive(run-106): SP2 reasoning-effort arm expansion (RED-GREEN)
try-works Oct 4, 2026
1f12993
recursive(run-106): SP3 borrowed quality prior (RED-GREEN)
try-works Oct 4, 2026
e156deb
recursive(run-106): SP4 effort-policy resolution (RED-GREEN)
try-works Oct 4, 2026
525d975
recursive(run-106): SP5 turn-aware hard shortcut (RED-GREEN)
try-works Oct 4, 2026
5e5c9c3
recursive(run-106): SP6 non-inferiority preference (RED-GREEN)
try-works Oct 4, 2026
a5dc829
recursive(run-106): SP7 effort union/intersection (RED-GREEN)
try-works Oct 4, 2026
80ca62a
recursive(run-106): lock Phase 3 implementation summary (SP1-SP7)
try-works Oct 4, 2026
1131aa1
recursive(run-106): SP3b wire borrowed quality prior into getQualityM…
try-works Oct 4, 2026
e0be271
recursive(run-106): SP4b wire effort-policy resolution into applyReas…
try-works Oct 4, 2026
d511a62
recursive(run-106): reflect SP3b/SP4b wiring in Phase 3 summary
try-works Oct 4, 2026
9547319
recursive(run-106): apply Phase 3.5 review repairs (HIGH-1 + MEDIUM-3…
try-works Oct 4, 2026
a478d00
recursive(run-106): re-mark R2/R4/R8/R9 deferred per Phase 3.5 review
try-works Oct 4, 2026
8567ae8
recursive(run-106): lock Phase 3.5 code review
try-works Oct 4, 2026
b1d9344
recursive(run-106): lock Phase 4 test summary
try-works Oct 4, 2026
771a361
recursive(run-106): lock Phase 5 manual QA (isolated-port strict rout…
try-works Oct 4, 2026
dfc4a94
recursive(run-106): lock Phases 6-8 closeout
try-works Oct 4, 2026
f1bf390
recursive(run-106): fix RCS diff accounting and lock hashes for clean…
try-works Oct 4, 2026
1d5c2ea
recursive(run-106): regenerate closeout lock receipts
try-works Oct 4, 2026
b2c4ca7
recursive(run-106): wire R2/R3/R5/R8/R9 into production (batch 1) + r…
try-works Oct 4, 2026
2b9b704
recursive(run-106): implement R4/R7/R11 + reconcile validate-vendors …
try-works Oct 4, 2026
8f511b0
recursive(run-106): implement R6 arm-aware lifecycle + R10 canonical …
try-works Oct 4, 2026
26d07cb
recursive(run-106): reconcile run91 to unsupported_fallback semantics
try-works Oct 4, 2026
d2ce0a4
recursive(run-106): correct R2 arm-expansion semantics + reconcile in…
try-works Oct 4, 2026
4127cfb
recursive(run-106): fix router-policy effort preservation expectation
try-works Oct 4, 2026
45709dd
recursive(run-106): Phase 3.5 review repairs (H-1/H-2/M-1..M-5/L-1..L…
try-works Oct 4, 2026
d3c1965
recursive(run-106): remove committed runtime-state residue (qa-state-…
try-works Oct 4, 2026
9260a10
recursive(run-106): fix M-3 occurrence effortSource regression (binar…
try-works Oct 4, 2026
57f627c
recursive(run-106): re-mark + lock Phases 3, 3.5, 4, 5 (R1-R15 verified)
try-works Oct 4, 2026
936b821
recursive(run-106): close out Phases 6, 7, 8 (decisions/state/memory …
try-works Oct 4, 2026
f255b4d
recursive(106): merge dev into run 106 and make the branch lint-clean
try-works Oct 6, 2026
953e6e7
fix(run-106): project legacy effort_source to named in analytics + dr…
try-works Oct 6, 2026
4919c02
fix(run-106): format the effortSource analytics projection
try-works Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -18,3 +18,4 @@ role-model-router/apps/runtime-ui/test-results/
/.recursive/config/recursive-router-discovered.json
# CI paired-checkout directory (checkout of try-works/role-model-internal nested by the track-b lane)
_private/
/.recursive/run/106-client-neutral-model-effort-routing/qa-state-sea//
42 changes: 42 additions & 0 deletions .recursive/DECISIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,9 +14,51 @@
- `86-runtime-ui-rm3-design-system-frontend` - RM v3 design system + `@role-model/ui` kit migration for router runtime-ui; Paper `4-0`/`5-0`/`6-0`/`7-0` IA; FD#15 config→strategy; SP8 floor green; hybrid Phase 5 QA on rebuilt `:3470` with human Paper sign-off; operator polish P1–P8 (Phases 0-8). Folder: `.recursive/run/86-runtime-ui-rm3-design-system-frontend/`. Soft-closes run `60-runtime-ui-paper-linear-review-alignment` as live styling authority for migrated surfaces (Linear/Paper-Linear historical).
- `92-configured-model-pool-benchmark-convergence` - Endpoint-variant-exact membership revision token stamped at persist/read/portfolio/decision time; honest null candidate-space (no synthetic 0/0%); membership-revision + stale benchmark quarantine; destructive-confirm final-controller eject; decision revision; transactional benchmark clear (Phases 0-8, strict TDD, agent-operated QA on rebuilt `:3501`). Folder: `.recursive/run/92-configured-model-pool-benchmark-convergence/`. Soft-closes run 76's membership-authority contract with a revision-token convergence wave.

- `106-client-neutral-model-effort-routing` - Client-neutral model-effort routing: reasoning effort becomes a first-class routing dimension (`strict`/`preferred`/`router` policy, executable model-endpoint-effort arms, effort-scoped evidence with borrowed/related priors, four-state effort-source vocabulary, canonical decision/telemetry/trace provenance, UI truthfulness, packaged SEA + real-Pi Phase 5 QA on `:3462`) (Phases 0-8). Folder: `.recursive/run/106-client-neutral-model-effort-routing/`.
- `103-agent-strategy-and-scoring-strategy` - Routing posture split (`routing.mode` + `routing.scoring_strategy` + `pin_weights` + `weights`), real scoring strategies, agent-strategy and workload postures with `<name>.<scope>` aliases, strategy/latency/alias receipts on every decision, measured-latency effective metric, three runtime-ui surfaces; strict TDD, two delegated review rounds, agent-operated rebuilt-runtime pi-CLI QA on `:3458` (Phases 0-8). Folder: `.recursive/run/103-agent-strategy-and-scoring-strategy/`.
- `105-route-learning-matching-scope-activation` - Stage-3 matching-scope activation: per-(role,task) ranked endpoint ladder (aggregation, store, advisory walk, dispatch/activation/rollback, Packs UI) plus the Phase 3.5 classification-alignment repair (capture/advisory reads the authoritative taxonomy identity, not runtime-policy ids); strict TDD, delegated review + re-review, agent-operated rebuilt-runtime pi-CLI QA on `:3458` (Phases 0-8). Folder: `.recursive/run/105-route-learning-matching-scope-activation/`.

## Run: `106-client-neutral-model-effort-routing`

Date: `2026-10-04`

### What changed

- Reasoning effort becomes a first-class, client-neutral routing dimension: the normalized input separates `requested_effort` from `effort_policy` (`strict | preferred | router`); omitted effort -> router, legacy scalar -> preferred, explicit policy -> authoritative (R1).
- The router selects over executable model-endpoint-effort arms (`expandReasoningEffortArms` + adapter execution mapping) rather than base endpoints (R2).
- `strict` considers exact-effort arms only (else `reasoning_effort_unavailable`); `preferred` falls back through named hard-eligibility/provider-attempt receipts; `preferred` with zero exact arms records `unsupported_fallback` and performs router-managed selection (R3).
- Four lossless effort states (`named | disabled | provider-default | no-client-preference`) are preserved through core types, serialization, SQLite, discovery, telemetry, trace, and UI; the occurrence boundary keeps a binary `effort_source` vocabulary while the decision keeps the four-state (R4, M-3 fix `9260a10b`).
- Benchmark/operational evidence is keyed by the effort arm; cross-effort evidence acts only as a labeled, discounted prior (`resolveBorrowedQualityPrior` + `getQualityMetric` + `benchmark-summary` + `profile-aggregator`) (R5).
- Arm-aware strategy/controller/cache/fallback (R6), turn-aware difficulty (R7), non-inferiority ranking (R8), arm-level discovery union/intersection (R9), and canonical decision/telemetry/trace/error provenance (R10).
- UI truthfulness: model + effective effort + exact/borrowed evidence on operator surfaces (R11).
- Packaged SEA (`role-model-dev.exe`, sha256 `e181c6011a50a7e9681d6f16314ae7cefcbe7ca26bb21eeac34343b3befc7364`) + real Pi CLI on `:3462` (requestId `req-9d68e76b`, decision `decision-req-9d68e76b` -> deepseek-v4-pro; strict+max -> v4-flash-max, router+max -> v4-flash) (R12/R15).
- Strict TDD (SP1-SP7 RED/GREEN) + delegated audits + controller verification (R13/R14).

### Why

- High-effort traffic mostly selected V4 Pro although Flash was faster, cheaper, and strongly benchmarked; the arithmetic was consistent but the arm/evidence identity and its presentation were not.

### How

- Strict TDD (SP1-SP7) + SP3b/SP4b and batch 1/2/3 wiring; Phase 3.5 delegated re-review (12 findings addressed); Phase 4 test audit (Tier A 23 files green, Tier B 2063/2068); Phase 5 agent-operated packaged-SEA + real-Pi QA on `:3462`.

### What was not done (OOS / residuals)

- Provider-specific effort ordering/auto-mapping without published equivalence (OOS1); re-benchmarking every provider/model/effort (OOS2); stage/main promotion (OOS4).

### Follow-ups

- Merge feature branch `recursive/106-client-neutral-model-effort-routing` (rebase onto post-105 dev) and open PR.
- Non-blocking LOW documentation follow-ups (materiality, threshold documentation, strict-with-no-effort docs).

### Artifact references

- `.recursive/run/106-client-neutral-model-effort-routing/00-requirements.md`
- `.recursive/run/106-client-neutral-model-effort-routing/03-implementation-summary.md`
- `.recursive/run/106-client-neutral-model-effort-routing/03.5-code-review.md`
- `.recursive/run/106-client-neutral-model-effort-routing/04-test-summary.md`
- `.recursive/run/106-client-neutral-model-effort-routing/05-manual-qa.md`

## Run: `92-configured-model-pool-benchmark-convergence`

Date: `2026-08-21`
Expand Down
9 changes: 9 additions & 0 deletions .recursive/STATE.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

## Current State

Run `106-client-neutral-model-effort-routing` is the current increment: reasoning effort is now a first-class, client-neutral routing dimension. The normalized input separates `requested_effort` from `effort_policy` (`strict` | `preferred` | `router`); the router selects over executable model-endpoint-effort arms; `strict` considers exact-effort arms only (else `reasoning_effort_unavailable`), `preferred` falls back through named receipts, and `preferred` with zero exact arms records `unsupported_fallback`. Four lossless effort states (`named` | `disabled` | `provider-default` | `no-client-preference`) are preserved end to end (decision four-state, occurrence binary). Benchmark/operational evidence is keyed by the effort arm with borrowed/related-effort priors; arm-aware lifecycle, turn-aware difficulty, non-inferiority ranking, and arm-level discovery are wired; decision/telemetry/trace provenance is canonical; and operator surfaces co-display model + effective effort + exact/borrowed evidence. Phases 0-8 are locked. Phase 5 ran the packaged SEA (`role-model-dev.exe`, sha256 `e181c6011a50a7e9681d6f16314ae7cefcbe7ca26bb21eeac34343b3befc7364`) on its own isolated port `:3462` driven by a real Pi CLI (requestId `req-9d68e76b`, routingDecisionId `decision-req-9d68e76b` -> deepseek-v4-pro; strict+max -> v4-flash-max, router+max -> v4-flash). Promotion remains a separate release operation.

Run `105-route-learning-matching-scope-activation` is the current increment: stage-3 matching-scope activation is shipped — a pack is a per-(role,task) ranked endpoint ladder (exact match only), materialized and aggregated pairwise, walked advisory-only, with derived floor-based activation and per-task rollback, and the Packs page renders the ladder index. The Phase 3.5 review found and the controller repaired a R1/R8 classification divergence: buildRequestClassificationForPlan read runtime-policy ids instead of the authoritative taxonomy identity, so capture/advisory classification was null while telemetry showed coder|coder.edit. Fixed at public d797a185 (tree 5da40073) + private da40a115; full host suite 2250/5-skip; live pi request classifies coder|coder.edit with lastUnclassifiedCaptures 0 on :3458 (dev channel only). Promotion remains a separate release operation.

**Run 105 closeout — bug 3 verified, four operational caveats carried forward** (not defects in the shipped fix). Verified live on `:3458`: both `coder/coder.edit` and `coder/coder.config` built **4/4 ladders** through counterfactual replays (all configured endpoints admitted, including `kimi-k3` — the configured controller and therefore the judge — admitted as a *challenger*, with `evaluation_judge_switches` recording `kimi-k3 -> deepseek-flash` from `dedupeJudgeAgainstPair`), and the route advisory reached live routing (`selection=advisory_applied`; `deepseek-flash-max` x3, `deepseek-flash` x1). Queue at closeout: 28 dispositions (24 `replayed`), 36/39 replay jobs `complete`, 0 deferred-pending, 0 capture-WAL rows, 0 stuck ledger reservations. Four caveats for the next replay-lane increment:
Expand Down Expand Up @@ -70,6 +72,9 @@ Run `92-configured-model-pool-benchmark-convergence` is the **current closed-out

### Product truths

- **Client-neutral model-effort routing (run 106):** `requested_effort` + `effort_policy` (`strict`|`preferred`|`router`); executable model-endpoint-effort arms; `unsupported_fallback` for zero-exact-arm preferred; four lossless effort states (decision four-state, occurrence binary); effort-scoped evidence with borrowed/related priors; arm-aware lifecycle + turn-aware difficulty + non-inferiority + arm-level discovery; canonical decision/telemetry/trace provenance; UI model+effort+evidence co-display.
- **Run 106 worktree:** `D:\DEV\role-model\.worktrees\106-client-neutral-model-effort-routing` on branch `recursive/106-client-neutral-model-effort-routing` (diff basis `701b8b8fc0b0eeebdfe818b757f5702f50021488`; HEAD `9260a10b`).
- **Run 106 verification floor:** Tier A 23 files green · Tier B 2063/2068 (3 conditional skips) · schemas:validate 37+30 · conformance 53/53 · core 113 · packaged SEA sha256 `e181c6011a50a7e9681d6f16314ae7cefcbe7ca26bb21eeac34343b3befc7364` · Pi QA `:3462`.
- **Configured model pool (run 92):** `computeConfiguredMembershipRevision` (order-stable SHA-256 over endpoint-variant-exact tuples) stamped on router candidates, routing decisions, benchmark manifests/samples, and clear receipts; runtime-ui `fetchRuntimeModels` no longer falls back to `/v1/models`; candidate-space scorers nullable with `—`/`n/a` presentation; `readLatestBenchmarkProfilesByEndpointIds` skips membership-mismatch and `completion_state: "stale"` samples; controller eject is destructive-confirmed; benchmark clear is transactional.
- **Run 92 worktree:** `D:\DEV\role-model\.worktrees\92-configured-model-pool-benchmark-convergence` on branch `recursive/92-configured-model-pool-benchmark-convergence` (diff basis `d59f07b91e7b23c25e7297860a0f9c967b342b7a`; HEAD `01537fb8b402c6808e7a6b69c3a03227acceb17c`).
- **Run 92 verification floor:** host-bridge 756 passed/3 skipped · runtime-ui 454 passed · sqlite-memory 67 passed · profile-aggregator 8 passed · builds green · agent-operated QA on `:3501`.
Expand All @@ -90,6 +95,8 @@ Run `92-configured-model-pool-benchmark-convergence` is the **current closed-out

### Known limitations

- Run 106 residual: non-blocking LOW documentation follow-ups (materiality, threshold documentation, strict-with-no-effort docs).

- Feature-branch merge to origin `dev` remains operator-requested (runs 92, 89, 86, 85, …).
- Run 92 residual: `profileRevision` is membership-keyed (diagnostic-only) until a distinct profile receipt is warranted; no decision-time membership snapshot is persisted (the field reflects current membership at read time).
- Run 89 residual: land `.agents/plugins/marketplace.json` on published `dev` for GitHub marketplace one-liner; optional Desktop UI glance; optional Codex Stop-hook auto-continue (not adapter regex).
Expand All @@ -103,6 +110,8 @@ Run `92-configured-model-pool-benchmark-convergence` is the **current closed-out

### Operational notes

- Prefer run-106 evidence under `.recursive/run/106-client-neutral-model-effort-routing/evidence/` for client-neutral effort routing, four-state effort-source vocabulary, borrowed/related-effort priors, and Phase 5 packaged-SEA + Pi QA.

- Prefer run-92 evidence under `.recursive/run/92-configured-model-pool-benchmark-convergence/evidence/` for membership-revision convergence, honest null candidate-space, benchmark quarantine, controller eject, and Phase 5 `:3501` QA.
- Prefer run-89 evidence under `.recursive/run/89-codex-role-model-package/evidence/` for Codex adapter, tool-bridge, npm/marketplace, and Phase 5 live routing proofs.
- Codex outsider install: `npx --yes @try-works/codex-role-model@latest setup|start`; marketplace via personal npm catalog or (after merge) `codex plugin marketplace add try-works/role-model --ref dev`.
Expand Down
2 changes: 1 addition & 1 deletion .recursive/memory/MEMORY.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ Control-plane docs are not memory docs:
## Registry

- `domains/remote-effort-instance-identity.md` - effort-variant identity,
admission, eligibility, telemetry and packaged Track B boundary (Run 93).
admission, eligibility, telemetry and packaged Track B boundary (Runs 93, 106).

- `domains/release-artifact-provenance.md` - exact CI commit identity across
Stage/production package assembly, runtime startup, and promotion (Run 94).
Expand Down
18 changes: 14 additions & 4 deletions .recursive/memory/domains/remote-effort-instance-identity.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,10 +4,10 @@ Status: CURRENT
Scope: Remote endpoint effort-instance identity, admission, eligibility, and UI projection.
Owns-Paths: role-model-router/apps/runtime-host-bridge/src/; role-model-router/apps/runtime-ui/app/lib/
Watch-Paths: role-model-router/apps/runtime-ui/app/routes/; role-model-router/packages/provider-*/
Source-Runs: 91-reasoning-effort-instance-identity; 92-configured-model-pool-benchmark-convergence; 93-variant-admission-model-pool-integrity
Validated-At-Commit: working-tree Run 93 Phase 5 rebuild
Last-Validated: 2026-08-22
Tags: effort, endpoint, admission, telemetry, benchmark, track-b
Source-Runs: 91-reasoning-effort-instance-identity; 92-configured-model-pool-benchmark-convergence; 93-variant-admission-model-pool-integrity; 106-client-neutral-model-effort-routing
Validated-At-Commit: working-tree Run 106 closeout (HEAD 9260a10b)
Last-Validated: 2026-10-04
Tags: effort, endpoint, admission, telemetry, benchmark, track-b, effort-policy, effort-source
---

# Remote effort-instance identity
Expand All @@ -21,3 +21,13 @@ Managed adapter inventory (for example LiteLLM) is not a user-configurable
provider connection. The paired Track B distribution is mandatory at packaged
runtime startup; extension actions remain accurately labelled when shadow or
gated.
Run 106 makes reasoning effort a first-class, client-neutral routing dimension:
the normalized input separates `requested_effort` from `effort_policy`
(`strict` | `preferred` | `router`), the router selects over executable
model-endpoint-effort arms, and the effort state is preserved losslessly end to
end. The decision surface keeps a four-state vocabulary
(`named` | `disabled` | `provider-default` | `no-client-preference`) while the
occurrence boundary keeps a binary `effort_source` vocabulary; the two must not
be conflated (M-3 regression). Benchmark/operational evidence is keyed by the
effort arm, and cross-effort evidence only ever acts as a labeled, discounted
borrowed/related prior — never as exact evidence.
Original file line number Diff line number Diff line change
Expand Up @@ -68,8 +68,9 @@ Source-Runs:
- `71-runtime-startup-lifecycle-and-health-truth-reconciliation`
- `72-standalone-runtime-config-authority-and-alias-rematerialization`
- `74-kimi-k3-kimi-code-oauth-support`
Validated-At-Commit: `working-tree`
Last-Validated: `2026-07-17`
- `106-client-neutral-model-effort-routing`
Validated-At-Commit: `working-tree Run 106 closeout (HEAD 9260a10b)`
Last-Validated: `2026-10-04`
Tags:
- `runtime`
- `routing`
Expand Down
Loading
Loading