Repository navigation
fix(replay): stop the replay tick starving on an unclassified pending capture - #321
Merged
Merged
Conversation
… capture The tick builds its capture set from the sidecar's cheap pending view and then adds the rich corpus read only for captures the pending view does not already contain. Because only the corpus read carried the route classification, deduping by capture ref discarded the classified row and kept the unclassified one, so the ownership filter matched nothing and the tick threw NoReplayableRequest on every tick for ever. Whenever the focus task's captures all fit inside the pending scan window - a task that has just started being routed, or any small state root - no replay job is ever created, no disposition is written and the pending row is never retired. On a mature root the corpus read usually supplies captures outside the window, which is why the defect stayed hidden on stage and in production. A capture that is already present is now enriched from the corpus read instead of dropped, so an unclassified row can never mask the identity the corpus read proved it has; a capture that is already classified keeps the cheap row because then the two views agree. Adds the regression test that models the real interaction - the same capture ref in both views, with identity only in the corpus read - which the run-105 suite never exercised because it stubbed the two readers independently and almost always left the pending view empty. Also names every gate that rejects a capture in readRouteReplayableCaptures and logs the failing gates under ROLE_MODEL_FOCUS_DIAG, and logs the cause when the corpus read is unavailable instead of collapsing to null. That corpus read had been blamed for the starvation because a rejected capture surfaced only as an unattributable NoReplayableRequest; the transcript was never the problem, so the trial artifact readback added for it is removed.
The supervised-replay comparability read the task family straight off the capture record, while the role and the taxonomy revision next to it have always been read through a fallback to classification. A capture that records the family only where the capture contract puts it therefore produced a comparison whose comparability block named the role and not the task. The knowledge worker scopes a comparison by (comparability.roleId, comparability.taskTypeId) and excludes an incomplete scope fail-closed as incomplete_scope. So the group was finalized, the admission floor never counted it, and no RouteLadderPackV1 could ever be written for the task - the learner saw a task that never accumulated evidence while its comparisons piled up durably. Measured on the dev replay queue: two of eight finalized comparison groups carried taskTypeId null while their source captures' durable classification named writer.summarize. The fixtures that covered the reader always declared a flat taskTypeId, so the gap was invisible to the suite; the new cases pin the shape production actually has.
…fort The focus-dispatch walk returns the first configured endpoint that is not admitted, not an unavailable rung, not already a rung and not excluded - effort is not an input to it at all. The downstream same-model repair (preferEffortMatchedReplayArms) cannot rescue the result when the arm's model IS the source's model, because the only configured endpoint satisfying 'this model at the source effort' is then the source endpoint itself, which the run-105 bug-3 distinct-source clause excludes. The arm keeps the effort it was handed, the comparison is finalized with validityIssues ['arm_effort_mismatch'], and the learner's admission floor discards it fail-closed. Measured on the dev replay queue: 3 of 8 finalized comparison groups, and a whole source endpoint - the max-effort variant of the source's model - could never contribute a single eligible comparison however much traffic it served. The walk now prefers a survivor that runs at the source capture's effort, and the tick resolves that effort from the endpoint registry it already has (the capture names the endpoint it came from; the registry names the effort that endpoint runs at), so no extra capture plumbing is needed. The comparison is what the run-104 comparability dimension already grades, and effort is compared trimmed and case-insensitively exactly as that grading does, so the chosen arm and the record that grades it agree. It never refuses: with no effort-matched survivor, or with no effort view supplied at all, the first survivor is returned exactly as before, so nothing that used to dispatch stops dispatching and the run-104 same-model repair contract is left untouched.
The loop's endpoint descriptor provider read modelId and reasoningEffort off the endpoint candidate's TOP LEVEL, where they do not exist: an EndpointCandidate carries them at identity.model_id and identity.reasoning_effort. Every descriptor therefore answered an empty model and a null effort, which made the focus-dispatch effort view empty and made the queue handoff's arm comparability always arm_effort_unspecified. Read the identity fields, keeping the top-level shape as a fallback.
…n it Operator decision (2026-10-07): reasoning effort is not a requirement for comparing two endpoints. Any configured endpoint is a legitimate counterfactual for any other one no matter what effort each runs at, so an effort difference must not cost a comparison its place in the learner's evidence. Run 104 R9 recorded the dimension but also named a mismatch as a validity issue. validityIssues means 'this comparison is not valid evidence at all' - a judge grading its own arm, or two arms that produced one byte-identical outcome - and every consumer treats a non-empty list as disqualifying: the knowledge worker's projection drops the group out of the admission floor, the ranking and the pack; the learning pass counts it as incomparable:arm_effort_mismatch and skips it; and this repository's finalized-route-challenge read refused to accept it as ladder progress. On a configuration whose source model has no sibling at the source effort, that removed every comparison the endpoint could ever produce, and 24 of 30 finalized groups were excluded on this name alone while the eligible count sat at 4 against a floor of 5. The dimension itself is unchanged and still published in the group's comparability block (matched / mismatched / unspecified on either side), so a receipt can still answer 'was this comparison effort-confounded?' and the learner can still report it. Only the disqualification is removed: the finalized-challenge read now accepts a mismatched arm and still requires the dimension to be present and well formed, and every other validity issue keeps its fail-closed meaning. Groups finalized before the decision still carry the name durably, so the knowledge worker and the learning pass filter that one name out rather than trusting the record to be clean.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Two defects on the path from a live request to a learner pack. Both were invisible because the tests stubbed a shape production does not have.
1. The replay tick starved on an unclassified pending capture
The tick builds its capture set from the sidecar's cheap pending view and then adds the rich corpus read only for captures the pending view does not already contain. Only the corpus read carried the route classification, so deduping by capture ref discarded the classified row and kept the unclassified one: the ownership filter matched nothing, the tick threw
NoReplayableRequeston every tick for ever, no replay job was created, no disposition was written and the pending row was never retired.Whenever the focus task's captures all fit inside the pending scan window - a task that has just started being routed, or any small state root - nothing can ever be dispatched. On a mature root the corpus read usually supplies captures outside the window, which is why the defect stayed hidden on stage and in production.
A capture that is already present is now enriched from the corpus read instead of dropped, so an unclassified row can never mask the identity the corpus read proved it has. A capture that is already classified keeps the cheap row, because then the two views agree.
2. The comparison never named the task family
The supervised-replay comparability read
capture.taskTypeIdstraight off the capture record, whileroleIdandtaxonomyVersionnext to it have always been read through a fallback toclassification. A comparison built from a capture that records the family only where the capture contract puts it was therefore finalized withtaskTypeId: null.The knowledge worker scopes a comparison by
(comparability.roleId, comparability.taskTypeId)and excludes an incomplete scope fail-closed asincomplete_scope. Those groups were durably recorded and then invisible to both the admission floor and the ranking - so noRouteLadderPackV1could ever be written for the task, and the learner reported a task that never accumulated evidence instead of reporting an error.readCaptureTaskTypeIdreads the family exactly like its two siblings. Measured: 2 of 8 finalized groups on the dev root carriedtaskTypeId: nullwhile their source captures decrypted towriter.summarize.Diagnostics
readRouteReplayableCapturesexpressed its 23 admission conditions as one boolean chain behind a barecontinue, so a rejected capture surfaced only as an unattributableNoReplayableRequest. Each condition is now named once in an ordered gate table, andROLE_MODEL_FOCUS_DIAGprints the gates that dropped each capture; the corpus-unavailable path logs its own cause instead of collapsing tonull. This is what made the second defect findable.The trial artifact readback added earlier for a misdiagnosis is removed: the transcript was never the problem. The capsule read already resolves it out of the artifact store, so the refs-only design holds without a second reader.
Tests
run105-review-dispatch.test.ts: a cold-start regression test that models the real interaction - the same capture ref in both views with identity only in the corpus read. The suite previously stubbed the two readers independently and almost always left the pending view empty, so it could not express production's state.run104-r22-learning-scope-travel.test.ts: three cases pinning the family reader, whose fixtures previously always declared a flattaskTypeId.tsc --noEmitandbiome checkclean.Verification
Rebuilt paired distribution (
b5a11efc, sidecar881d393888adbe1f) launched on:3458: the first tick owned and dispatched a capture, 8 captures reached the durablereplayeddisposition, the evaluation core finalized comparison groups with full comparability identity, and the admission floor reached 5 eligible effort-matched comparisons at mean confidence 0.849 for the pair(deepseek-v4-pro-max, deepseek-flash-max).Docs
docs/operations/route-capture-and-replay-queues.md(private) is the canonical map;.recursive/STATE.mdand.recursive/DECISIONS.mdrecord and reference it.Out of scope
Recorded as open findings in the doc, not fixed here: the replay arm is chosen without consulting reasoning effort (so on a configuration with a single low-effort variant every comparison sourced from a max-effort endpoint is voided
arm_effort_mismatch), the learner sweep that materializes a pack is skipped while a long replay is in flight, and a capture whose replay keeps failing keeps re-entering the queue. No stage, main or release promotion.