fix(state): count Todo headings the way the region writer classifies them - #4643
huangruiteng merged 5 commits into
Conversation
…them
`loopx/state_projection.py` carried its own copy of the Todo header marker
tuples, and `loopx/control_plane/goals/active_state_metadata.py` carried
another. Both classify a heading from the same active-state document into the
same `user` / `agent` / `None` roles, so this is one contract with two
implementations -- and they had drifted. Measured over 23 realistic headings,
the two disagree on 15.
Three of those disagreements are defects, not preferences:
- `Agent Todo Archive` counted as **live** agent Todos. The writer already
guards archives through `TODO_ARCHIVE_HEADER_MARKERS`; the counter had no
such guard.
- `Codex Todo` counted as nothing, although the region writer creates exactly
that heading, so a whole class of agent Todos was invisible to the count.
- A bare `owner` marker made prose sections match: a section named `Ownership`
or `Downstream Owner Notes` counted its bullets as user Todos.
The consequence is not a wrong number on a screen. The counts feed
`state_projection_gap_warning`, which reports that a Next Action is executable
while no agent Todo stands behind it. An agent count inflated by archived work
suppresses that warning, so the failure mode is silence. On a state document
with two live Todos, two archived ones, one `Codex Todo` and one `Ownership`
prose section, the counter returned `{'user': 1, 'agent': 5}` where the truth
is `{'user': 0, 'agent': 3}`.
Delete both copies and call `todo_role_for_heading`, so the counter and the
region writer agree by construction. The markers this drops -- `agent backlog`,
`agent action`, `user action`, `用户`, `人工`, bare `owner` -- appear in no
shipped template or fixture as a heading, and the writer never creates them, so
counting them reported work the machinery could not see or act on.
Two undeclared multi-value forks disappear with the copies:
`multi_value_forks` 4 -> 2, `multi_value_forks_semantic` 3 -> 1,
`multi_value_fork_definitions` 10 -> 6, lowered in this diff on both the
registry and `BUDGET_ANCHOR` sides as the equality check requires.
Refs loopx-project#4447
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
exact-head 复核(
|
| 标题 | state_projection |
active_state_metadata |
后果 |
|---|---|---|---|
Agent Todo Archive |
agent |
None |
归档被当成在办 |
Codex Todo |
None |
agent |
写入方创建的标题,计数看不见 |
Ownership |
user |
None |
裸 owner 标记命中散文段落 |
后果是沉默,不是一个错数字
计数喂给 state_projection_gap_warning,后者报告「Next Action 可执行但背后没有 agent todo」。归档把 agent 计数抬高 → agent_open > 0 → 该告警永不触发。
实测(2 条在办 + 2 条归档 + 1 个 Codex Todo + 1 段 Ownership 散文):
修改前: {'user': 1, 'agent': 5}
修改后: {'user': 0, 'agent': 3} <- 真值
修法
删掉两份副本,改调 todo_role_for_heading,让计数与写入方由构造保证一致,而不是靠两份列表保持同步。
被丢弃的标记(agent backlog、agent action、user action、用户、人工、裸 owner)在任何已交付模板或夹具中都不作为标题出现,写入方也从不创建它们。计入它们等于报告一个机制既看不见、也无法据以行动的工作量。
棘轮
两份副本消失后,两个未声明分叉随之消除,在同一 diff 内按相等检查的要求同时下调 registry 与 BUDGET_ANCHOR 两侧:
multi_value_forks 4 -> 2
multi_value_forks_semantic 3 -> 1
multi_value_fork_definitions 10 -> 6
RAW_MATERIAL_KEY_HINTS(issue 列的第三个分叉)本 PR 不碰:它的两份定义值集确实不同(body/chat/credential/dm vs credential/local_path/log/raw),需要单独分类。
验证(当前 exact head)
semantic-vocabulary-drift-smoke: ok # multi_value_forks=2/2 semantic=1/1 definitions=6/6
pytest tests/control_plane/test_todo_machine_region.py -> 20 passed
回归测试钉住三处缺陷,并做了变异验证:只回退 state_projection.py,测试失败并给出 {'user': 1} != {'user': 0}。
…odo-header-markers loopx-project#4608 merged while this branch was open and both sides edit the same anchored ratchet block. Resolved by taking the tighter side of each, as budgets_only_decrease requires: - same_runtime_forks_semantic 13 -> 11 from main, which loopx-project#4608 earned by importing the Todo task-class vocabulary from its owner. - multi_value_forks 4 -> 2, multi_value_forks_semantic 3 -> 1 and multi_value_fork_definitions 10 -> 6 from this branch, which are the two undeclared forks the marker copies were. Registry and BUDGET_ANCHOR carry the same values. Measured after the merge: every counter is at budget, with multi_value_forks=2/2, multi_value_forks_semantic=1/1 and multi_value_fork_definitions=6/6. Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…ted one PR loopx-project#4643 retired the AGENT/USER_TODO_HEADER_MARKERS forks, which four drift tests from loopx-project#4614 had rented as samples for the rename-laundering limit. Track A retires real forks one by one, so renting them makes every retirement a test breakage -- this PR included. The pins now inject their own synthetic fork (one name, two modules, two disagreeing value sets) and measure against the live baseline instead of hardcoded counts, so they pin the machinery -- the RFC Section 9 known limit -- independently of the debt population: - the laundering test measures base / base+1 / base instead of 3/2; - the declaration test proves exclusion by removing SOURCE_SURFACES from the registry and watching the count rise by exactly one, instead of asserting today's undeclared count; - the three-case rename boundary runs its undeclared cases on the synthetic fork; Case 1 (declared rename refused) still uses SOURCE_SURFACES, which is still declared and still 4-way divergent; - the advisory test keeps one real assertion: RAW_MATERIAL_KEY_HINTS, the only undeclared fork surviving this PR, must stay listed, with a comment that the assertion retires with the fork. Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Self-review on exact head
|
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: 0538faa3c70fbe7985ad8d150eda3fd0c9bbfc9e (codex/todo-heading-role-single-source).
动机
#4447 里的第一个未声明 fork:loopx/state_projection.py 自带一份 Todo 标题标记元组,loopx/control_plane/goals/active_state_metadata.py 另有一份。两者都在把同一份 active-state 文档的标题分成 user / agent / None——前者喂 summarize_state_todo_open_counts,后者管区域写入——一份契约两套实现,而且已经漂移出三个真实缺陷:Agent Todo Archive 被当成在办的 agent 待办(没有归档守卫)、Codex Todo(区域写入器自己创建的标题)完全看不见、以及裸 owner 标记把 Ownership 这类散文小节算成 user 待办。
后果不是数字错,而是沉默:计数喂 state_projection_gap_warning,而它只在 agent_open == 0 时才会报出「Next Action 可执行、背后却没有 agent 待办」。
改动思路
入口是 summarize_state_todo_open_counts -> _role_for_heading,现在直接调用写入器的 todo_role_for_heading(先归档守卫、再 user、再 agent),两处副本元组删除。这是「删掉第二份实现」而不是「把两份手工同步」,正是 fork ratchet 想收敛的东西;machine_region.py:137 本来就用这个函数做角色回退,所以计数与区域写入器是按构造一致。registry 与 BUDGET_ANCHOR 同时下调三个 ratchet(multi_value_forks 4→2、multi_value_forks_semantic 3→1、multi_value_fork_definitions 10→6),等值 anchor 因此仍然成立。
具体改动
5 个文件、+160/-68:state_projection.py 删掉两份元组并改为委托(净减生产代码)、registry/anchor 的 ratchet、回归测试,以及 drift 测试相应更新;RAW_MATERIAL_KEY_HINTS(issue 里的第三个 fork)刻意不动,因为它的两份定义值集确实不同,需要单独分类。
我独立跑过前后对比(同一份文档:2 条在办 Agent Todo、2 条归档、1 条 Codex Todo、1 个 Ownership 散文小节):merge-base(origin/main 897e9ae){'user': 1, 'agent': 4} → head {'user': 0, 'agent': 3};pytest tests/control_plane/test_todo_machine_region.py -q → 20 passed;pytest tests/architecture/test_semantic_vocabulary_drift.py -q → 61 passed;smoke → ok,multi_value_forks=2/2 multi_value_forks_semantic=1/1 multi_value_fork_definitions=6/6。我也确认全仓已无对已删元组的引用(只剩 owner 模块自己)。
关键代码讲解
state_projection._role_for_heading(342):改为return todo_role_for_heading(heading),docstring 把三个漂移与「计数被灌水会压掉警告」写清楚——把规则留在被强制执行的地方。active_state_metadata.todo_role_for_heading(89):归档标记优先,其次 user(user todo/owner review reading queue/owner reading queue),其次 agent(agent todo/codex todo/project agent todo);这是唯一实现。state_projection_gap_warning(755):if agent_open == 0 and executable:才产出next_action_executable_without_agent_todo证据——我在 head 上读了这段,确认 PR 的「沉默」因果成立。test_state_counts_classify_headings_exactly_as_the_region_writer:既断言{'user': 0, 'agent': 3},又逐条断言五个标题的角色(含三个缺陷用例)。inventory_ratchets:三个 fork 预算与 anchor 同步下调并锁死(smoke 输出 2/2、1/1、6/6)。
对主干的风险
最强回归是「旧状态文件里用了被删掉的拼法」:裸 owner、用户、人工、agent action、agent backlog、项目 agent、agent 待办 这些标题此前的角色判定会变成 None,计数下降。方向是 fail-closed(可能多报一次 gap 警告),且 PR 说明这些拼法不作为标题出现在任何已交付模板或 fixture 中、写入器也从不创建它们;我没有逐份历史 goal state 复算,因此把它记在 residual 里。第二个潜在风险是 substring 匹配本身较松,但这是写入器既有的规则,本次是把计数对齐到它、而不是新造规则。
一条 P3(F1,非阻塞):正文里的 before 数字在 merge-base 上复现不出来。正文写「同一个文档:before {'user': 1, 'agent': 5}」,但我用 merge-base 的真实函数量到的是 {'user': 1, 'agent': 4}——旧分类器的 agent 标记里没有 codex todo,所以那条 Codex Todo 贡献的是 0 而不是被算成 live;归档小节贡献 2。head 的 {'user': 0, 'agent': 3} 与三个缺陷(归档算在办、Codex Todo 不可见、Ownership 算 user)我都复核为真,只有 before 少算/多算一格。建议把正文数字改成实测值。
我的整体评价
结论 APPROVE。这是一次正确的去重:删掉第二份分类实现、让计数与区域写入器共用 owner,顺带把三处真实漂移修掉并下调对应 ratchet;生产代码净减少,回退成本是一个 commit。我独立复现了 merge-base 与 head 的计数差异、读了警告的抑制分支、跑了 20 + 61 个测试与 smoke,并确认删除的常量没有残留引用。
唯一 P3 是正文示例数字未复现,非阻塞;RAW_MATERIAL_KEY_HINTS 留作后续分类是恰当的边界。
English verdict: APPROVE - exact head 0538faa; state_projection now delegates to the writer's todo_role_for_heading instead of keeping a drifted copy, which removes two of the three undeclared forks (ratchets 4/3/10 -> 2/1/6, locked, smoke 2/2 1/1 6/6). I reproduced the behaviour on both trees for the described document (merge-base {'user': 1, 'agent': 4} -> head {'user': 0, 'agent': 3}), read the agent_open == 0 suppression branch in state_projection_gap_warning, confirmed no references to the deleted tuples remain, and ran 20 + 61 tests. One non-blocking P3: the PR body's illustrative before-count ({'user': 1, 'agent': 5}) does not reproduce - the merge-base measurement is agent 4, because the old classifier never matched 'codex todo'.
No conflict; the branch was only behind main. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
Head moved: merged No conflict — the branch was only behind Revalidated: Per the exact-head rule in |
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: d8eacce197ee63f8ba5c123c97c2d71c2b770cfd (codex/single-source-todo-header-markers, re-review after the earlier head was superseded by a main merge).
动机
loopx/state_projection.py 和 loopx/control_plane/goals/active_state_metadata.py 各存了一份 Todo header marker 表,而它们分类的是同一份 active-state 文档、产出同样的 user/agent/None 角色:一个用于 summarize_state_todo_open_counts 计数,一个用于 find_todo_regions 定位写入区域。两份实现已经漂移:计数那份把裸 owner 当 user 标记(于是 Ownership 这种散文段落被计成 user todo)、没有 archive 守卫(Agent Todo Archive 被算成进行中)、还漏了 writer 会创建的 codex todo。后果是沉默而非数字错——agent 计数虚高使 state_projection_gap_warning(「Next Action 可执行但背后没有 agent Todo」)永远不触发。动机成立。
改动思路
不是把两份表对齐,而是删掉一份:计数改为调用 writer 的 todo_role_for_heading,让计数与写入区域按构造一致。删除副本带来的两个多值 fork 同步下调(registry 与 BUDGET_ANCHOR 两侧一起动,保持等值 anchor)。
具体改动
5 个文件、+160/-68:
loopx/state_projection.py:删除USER_TODO_HEADER_MARKERS/AGENT_TODO_HEADER_MARKERS,_role_for_heading改为委托todo_role_for_heading(净减 34 行)。loopx/semantics/vocabulary_v0.json+examples/semantic-vocabulary-drift-smoke.py:multi_value_forks 4→2、multi_value_forks_semantic 3→1、multi_value_fork_definitions 10→6。tests/architecture/test_semantic_vocabulary_drift.py(+128)、tests/control_plane/test_todo_machine_region.py(+54):钉住「计数与 writer 同规则」与精确计数。
我复核的关键点(都在这个 head 上自己跑过):
- before/after 复现:同一份文档(2 条进行中 agent todo + 2 条 archive + 1 条
Codex Todo+ 1 段Ownership)在 basea96c9aa91c上得{'user': 1, 'agent': 4},在 head 上得{'user': 0, 'agent': 3}——archive 不再算进行中、Codex Todo被计入、Ownership不再算 user。 - 分类器探针:
Agent Todo Archive/Ownership/Downstream Owner Notes/用户/人工/Agent Backlog→ None;Agent Todo/Codex Todo→ agent;Owner Reading Queue→ user。 pytest -q tests/control_plane/test_todo_machine_region.py→ 20 passed;三个依赖 state_projection 的 control-plane 套件 → 51 passed;架构套件 → 67 passed;drift smoke → exit 0 且三个 ratchet 打印2/2、1/1、6/6(等于实测)。- 真实数据回归扫描:对 5 份真实
ACTIVE_GOAL_STATE.md的全部##标题(67 个)用 base 与 head 两套分类器分别判定,差异 0——说明这次修的是潜在漂移,当前真实文档行为不变。
遗留问题(非阻塞,P3)
被删掉的标记里包含中文 用户/人工(archive 表里仍保留中文)。也就是说,手写 ## 用户待办 这类标题在 base 上会算作 user todo,在 head 上返回 None。我全仓与真实状态文档扫描确认:没有已发布模板/fixture 使用这些标题(唯一的 用户 出现在 README 正文),所以今天不改变任何行为。但如果非英文状态标题是受支持的用法,最小修法是把中文 user/agent 标记补进共享表;否则应在模块 docstring 与状态模板文档里写明「只有列出的英文标题会被识别」。
对主干的风险
最强回归是计数继续把归档/散文当工作进行中,或漏掉 writer 创建的标题,从而压掉 gap 警告。这次改动把两侧收敛到一个函数,方向正确;剩下的机制性风险是依旧靠有序子串匹配:这次通过把 owner 换成 owner reading queue 这类更长短语消除了 Ownership 误判,但将来若加入短标记,同类误判可以回来。未验证维度:中文/非英文标题(见 P3)。不涉及写路径、quota 或持久化状态;回退成本一个 commit。本次 head 只多了一次 main 合并,我对当前 main 重比了内容增量(5 文件、+160/-68)并重跑了全部证据。
我的整体评价
结论 APPROVE。这是「删掉重复知识而不是同步两份表」的正确做法:两侧按构造一致,三处实际缺陷(归档、Codex Todo、Ownership 散文)都被修正,且我用真实文档验证了当前行为不变。唯一 P3 是非英文标题的识别收窄,属文档/标记补充问题,不阻塞。
English verdict: APPROVE - exact head d8eacce (re-review after a main merge; content delta versus main unchanged at 5 files, +160/-68). The projection now calls the writer's todo_role_for_heading instead of keeping a second marker table: I reproduced the counts on base a96c9aa vs head ({'user': 1, 'agent': 4} -> {'user': 0, 'agent': 3}), the classifier rejects archived/prose headings and keeps Agent Todo, Codex Todo and Owner Reading Queue, 20 + 51 + 67 tests pass, the drift smoke exits 0 with multi_value_forks=2/2, multi_value_forks_semantic=1/1 and multi_value_fork_definitions=6/6, and all 67 headings across five live ACTIVE_GOAL_STATE.md files classify identically at base and head. One non-blocking P3: the removed Chinese user markers mean a hand-written heading such as 用户待办 no longer counts, which no shipped template or fixture exercises today.
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: d8eacce197ee63f8ba5c123c97c2d71c2b770cfd.
动机
按 #4447 当前 Track A 的修正结论,这两份 Todo marker 分类的是同一份文档、同一种角色,应合并到已有 owner。早期“只补 scope declaration”的建议不足以修复这里已经存在的行为分歧。此 PR 完成一个独立有效的增量:计数不再把归档当在办、不再漏掉 Codex Todo,也不再把 Ownership 散文算作 User Todo;剩余 RAW_MATERIAL_KEY_HINTS 分类仍归原 tracker,不宣称整个语义收敛工作完成。
独立实测同一合成文档:base a96c9aa91c9d92454efcd7937483bac76e35fd4a 的 {'user': 1, 'agent': 4} → 当前 head 的 {'user': 0, 'agent': 3}。原 PR 正文的 before agent=5 应为 4。更重要的是:只有归档任务、但 Next Action 可执行时,真实 refresh_state_run 从无告警变为返回完整的缺失 Agent Todo 告警。
改动思路
计数复用 active_state_metadata.todo_role_for_heading,删除第二份分类规则。只同步两份 tuple 会保留未来漂移的来源;增加新分类框架也没有必要。调用方向仍是外层 projection → 既有 control-plane metadata owner,没有新增 capability、provider、状态表或权限。
本次保留已有私有 delegate:它没有独立决策,内联可以进一步省掉一跳,但当前形态只承载兼容性说明,无额外状态或维护分支,不构成阻塞。最有价值的简化已经由删除重复知识实现。
具体改动
完整范围为 5 文件、+160/-68:运行时改为共享分类器;registry 与 smoke anchor 同步收紧三个 fork 预算;Todo 回归测试保护三个实际缺陷;架构测试改用注入的双模块分歧样本,避免依赖正在消除的真实债务。
关键代码讲解
state_projection._role_for_heading(342 行):只调用 canonical classifier;归档优先级、user/agent marker 不再本地重述。summarize_state_todo_open_counts(439 行):遍历完整文档、按标题切换角色、排除已完成 checkbox,结果交给 gap warning。state_projection_gap_warning(755 行):缺省时采用上述计数;显式传入 Todo summary 的既有优先级保留。agent_open == 0且 Next Action 可执行时,产出requires_todo_expansion与可读修复建议。_with_synthetic_fork(架构测试 195 行):两个源模块、同名、不同值集,独立验证 baseline → baseline+1;单侧重命名后回落,保留现有 name-based inventory 的已知限制,没有声称解决 rename laundering。
对主干的风险
主要兼容性变化是手写宽泛标题(例如 Agent Backlog、用户待办)不再计数。已扫描交付文档和 fixtures,没有使用这些标题的模板;写入器也从未把它们当作对应区域。任意历史手写文档不在本次证明范围内。共享分类器仍使用有序 substring 匹配;将来增加短 marker 仍可能误判,应在同一个 owner 和回归表中处理,不能恢复另一套规则。
P3,非阻塞: vocabulary_v0.json 的 multi_value_forks_note 仍写 “4 counted forks”,而数值预算已是 2。应在现有语义文档对账工作中刷新这句说明;当前数值约束和 scope declaration 的执行一致,不影响本次正确性。
语义与 CI 对齐
这是复用已有语义,未新增词表、缩小扫描范围或通过改名隐藏 fork。三个预算实测为 2/2、1/1、6/6。最新 main 9118568dde6d07610dc3df4f5a8c5e9cac799fc4 另收紧了 conflicting budgets;已构造无冲突的合并树并验证 16/16、55/55 保留。
验证结果:
- 当前 head:Todo machine region、semantic vocabulary drift、state refresh projection 三套测试 89 passed。
- 与上述最新 main 的合并树:同组测试 90 passed,语义 smoke 通过,包含 main 新增的预算测试。
- 相同合成输入经过真实 count → gap → refresh 路径:归档、Codex、散文、completed、标题顺序、200 条无关 checkbox、显式 summary 与无 Next Action 分支均检查;旧实现违反 6 项独立预期,head 与合并树通过。隔离文件的 dry-run 读回确认状态未被写入。
- 仓库规定的 Ruff 范围通过;额外检查改动的 runtime/test 文件通过;配置内 Mypy 22 文件通过。
semantic-vocabulary-drift-smoke.py、state-projection-gap-smoke.py、docs-governance-smoke.py通过;按当前 CLI 生成 inventory 到临时文件再--check --output读回通过。canary premerge --from-git-diff:diff/compile、6 项 canary、公开边界扫描全部通过;5 个改动文件无私有内容,0 失败、0 manual hold。质量回执cqr_3e286bd951f3e4cdc2c8已精确校验有效,未应用额外 safe-fix。
初次隔离环境缺 Node parser 依赖导致测试失败,执行 npm ci --ignore-scripts 后完整重跑通过。额外将 examples 纳入 Ruff 的探针发现既有 unused render_inventory import;base/head 同样复现,examples 不在仓库规定的 lint 范围内,此处不声称扩展探针通过。没有跳过必需的本地验证。按当前评审配置,未查询或等待远端 CI;本次不是全平台发布验证。
可直接复现关键负例:
from loopx.state_projection import state_projection_gap_warning
state = "## Agent Todo Archive\n- [ ] archived\n\n## Next Action\n- Implement the reader and validate the result.\n"
gap = state_projection_gap_warning(state)
assert gap is not None and gap["agent_open_count"] == 0
assert gap["first_evidence"][0]["kind"] == "next_action_executable_without_agent_todo"examples/state-projection-gap-smoke.py 另覆盖真实 refresh 和下游 quota 路由。没有 SQL authority、frontend 或 Lark contract 变化,因此无需数据库迁移或新增 UI 配套;各入口消费同一既有诊断形状。
我的整体评价
APPROVE。 完整增量与当前目标一致:减少一个实际会隐藏工作缺口的规则分歧,同时保留持久的回归覆盖。已针对当前 head 重新执行证据,没有沿用旧批准代替验证。非阻塞风险是手写别名兼容性及上述陈旧说明;不存在新的写入或执行授权。
本记录用于用户已授权的维护者合并;实际合并仍以紧接合并前的 exact-head readiness 为准,合并动作绑定上述 head。
English verdict: APPROVE - d8eacce. Reusing the writer's canonical heading classifier fixes archived/prose counts and missing Codex Todos. Independent baseline/head refresh-path counterexamples pass at head; 89 exact-head and 90 integrated-tree tests pass, native lint/type checks and all selected premerge checks pass, and the exact-scope quality receipt is valid. Non-blocking risks: narrowed manual heading aliases and stale fork-count prose. CI was not consulted under the configured policy.
huangruiteng
left a comment
There was a problem hiding this comment.
Reviewed exact head: 88fabdfeb9ac84b2c00761c58b0800073002172a.
本次复审针对 GitHub 合入 main 后的新提交。分支已包含 main b7f66d3f59fe0fd6aea29b088613f6a3d55035d9;PR 自身仍为相同的 5 文件增量。比对确认受影响的 classifier、Todo caller、语义检查与测试文件和此前验证的合并树一致,并在新 head 重新执行 90 项测试、真实 refresh 反例、Ruff、Mypy、质量回执及 premerge。以下为当前 head 的完整结论。
动机
按 #4447 当前 Track A 的修正结论,这两份 Todo marker 分类的是同一份文档、同一种角色,应合并到已有 owner。早期“只补 scope declaration”的建议不足以修复这里已经存在的行为分歧。此 PR 完成一个独立有效的增量:计数不再把归档当在办、不再漏掉 Codex Todo,也不再把 Ownership 散文算作 User Todo;剩余 RAW_MATERIAL_KEY_HINTS 分类仍归原 tracker,不宣称整个语义收敛工作完成。
独立实测同一合成文档:base a96c9aa91c9d92454efcd7937483bac76e35fd4a 的 {'user': 1, 'agent': 4} → 当前 head 的 {'user': 0, 'agent': 3}。原 PR 正文的 before agent=5 应为 4。更重要的是:只有归档任务、但 Next Action 可执行时,真实 refresh_state_run 从无告警变为返回完整的缺失 Agent Todo 告警。
改动思路
计数复用 active_state_metadata.todo_role_for_heading,删除第二份分类规则。只同步两份 tuple 会保留未来漂移的来源;增加新分类框架也没有必要。调用方向仍是外层 projection → 既有 control-plane metadata owner,没有新增 capability、provider、状态表或权限。
本次保留已有私有 delegate:它没有独立决策,内联可以进一步省掉一跳,但当前形态只承载兼容性说明,无额外状态或维护分支,不构成阻塞。最有价值的简化已经由删除重复知识实现。
具体改动
完整范围为 5 文件、+160/-68:运行时改为共享分类器;registry 与 smoke anchor 同步收紧三个 fork 预算;Todo 回归测试保护三个实际缺陷;架构测试改用注入的双模块分歧样本,避免依赖正在消除的真实债务。
关键代码讲解
state_projection._role_for_heading(342 行):只调用 canonical classifier;归档优先级、user/agent marker 不再本地重述。summarize_state_todo_open_counts(439 行):遍历完整文档、按标题切换角色、排除已完成 checkbox,结果交给 gap warning。state_projection_gap_warning(755 行):缺省时采用上述计数;显式传入 Todo summary 的既有优先级保留。agent_open == 0且 Next Action 可执行时,产出requires_todo_expansion与可读修复建议。_with_synthetic_fork(架构测试 195 行):两个源模块、同名、不同值集,独立验证 baseline → baseline+1;单侧重命名后回落,保留现有 name-based inventory 的已知限制,没有声称解决 rename laundering。
对主干的风险
主要兼容性变化是手写宽泛标题(例如 Agent Backlog、用户待办)不再计数。已扫描交付文档和 fixtures,没有使用这些标题的模板;写入器也从未把它们当作对应区域。任意历史手写文档不在本次证明范围内。共享分类器仍使用有序 substring 匹配;将来增加短 marker 仍可能误判,应在同一个 owner 和回归表中处理,不能恢复另一套规则。
P3,非阻塞: vocabulary_v0.json 的 multi_value_forks_note 仍写 “4 counted forks”,而数值预算已是 2。应在现有语义文档对账工作中刷新这句说明;当前数值约束和 scope declaration 的执行一致,不影响本次正确性。
语义与 CI 对齐
这是复用已有语义,未新增词表、缩小扫描范围或通过改名隐藏 fork。三个预算实测为 2/2、1/1、6/6。最新 main 9118568dde6d07610dc3df4f5a8c5e9cac799fc4 另收紧了 conflicting budgets;已构造无冲突的合并树并验证 16/16、55/55 保留。
验证结果:
- 当前 head:Todo machine region、semantic vocabulary drift、state refresh projection 三套测试 90 passed。
- 与上述最新 main 的合并树:同组测试 90 passed,语义 smoke 通过,包含 main 新增的预算测试。
- 相同合成输入经过真实 count → gap → refresh 路径:归档、Codex、散文、completed、标题顺序、200 条无关 checkbox、显式 summary 与无 Next Action 分支均检查;旧实现违反 6 项独立预期,head 与合并树通过。隔离文件的 dry-run 读回确认状态未被写入。
- 仓库规定的 Ruff 范围通过;额外检查改动的 runtime/test 文件通过;配置内 Mypy 22 文件通过。
semantic-vocabulary-drift-smoke.py、state-projection-gap-smoke.py、docs-governance-smoke.py通过;按当前 CLI 生成 inventory 到临时文件再--check --output读回通过。canary premerge --from-git-diff:diff/compile、6 项 canary、公开边界扫描全部通过;5 个改动文件无私有内容,0 失败、0 manual hold。质量回执cqr_fe30402b17b911fd648c已精确校验有效,未应用额外 safe-fix。
初次隔离环境缺 Node parser 依赖导致测试失败,执行 npm ci --ignore-scripts 后完整重跑通过。额外将 examples 纳入 Ruff 的探针发现既有 unused render_inventory import;base/head 同样复现,examples 不在仓库规定的 lint 范围内,此处不声称扩展探针通过。没有跳过必需的本地验证。按当前评审配置,未查询或等待远端 CI;本次不是全平台发布验证。
可直接复现关键负例:
from loopx.state_projection import state_projection_gap_warning
state = "## Agent Todo Archive\n- [ ] archived\n\n## Next Action\n- Implement the reader and validate the result.\n"
gap = state_projection_gap_warning(state)
assert gap is not None and gap["agent_open_count"] == 0
assert gap["first_evidence"][0]["kind"] == "next_action_executable_without_agent_todo"examples/state-projection-gap-smoke.py 另覆盖真实 refresh 和下游 quota 路由。没有 SQL authority、frontend 或 Lark contract 变化,因此无需数据库迁移或新增 UI 配套;各入口消费同一既有诊断形状。
我的整体评价
APPROVE。 完整增量与当前目标一致:减少一个实际会隐藏工作缺口的规则分歧,同时保留持久的回归覆盖。已针对当前 head 重新执行证据,没有沿用旧批准代替验证。非阻塞风险是手写别名兼容性及上述陈旧说明;不存在新的写入或执行授权。
本记录用于用户已授权的维护者合并;实际合并仍以紧接合并前的 exact-head readiness 为准,合并动作绑定上述 head。
English verdict: APPROVE - 88fabdf. Reusing the writer's canonical heading classifier fixes archived/prose counts and missing Codex Todos. Independent baseline/head refresh-path counterexamples pass at head; 90 updated exact-head and 90 integrated-tree tests pass, native lint/type checks and all selected premerge checks pass, and the exact-scope quality receipt is valid. Non-blocking risks: narrowed manual heading aliases and stale fork-count prose. CI was not consulted under the configured policy.
|
Merged as Final reviewed head: Local evidence: 90 tests passed, real refresh-path counterexamples passed, native Ruff/Mypy passed, exact-scope quality receipt |
Refs #4447 — Track A, the first of the three undeclared forks the issue lists.
The two copies are one contract
loopx/state_projection.pycarried its own copy of the Todo header marker tuples;loopx/control_plane/goals/active_state_metadata.pycarried another. Both classify a heading from the same active-state document into the sameuser/agent/Noneroles —summarize_state_todo_open_countscounts Todos per role,find_todo_regionslocates the regions to write into. One contract, two implementations, and they had drifted.Measured over 23 realistic headings, the two disagree on 15.
Three of those are defects
state_projectionactive_state_metadataAgent Todo ArchiveagentNoneCodex TodoNoneagentOwnershipuserNoneownermarker matches prose sectionsThe counter had no archive guard although
TODO_ARCHIVE_HEADER_MARKERSexists for exactly that, and itsownermarker is a substring broad enough to catchOwnershipandDownstream Owner Notes.The consequence is silence, not a wrong number
The counts feed
state_projection_gap_warning, which reports that a Next Action is executable while no agent Todo stands behind it. An agent count inflated by archived work makesagent_open > 0, so that warning never fires.Measured on a state document with two live Todos, two archived, one
Codex Todoand oneOwnershipprose section:The fix
Delete both copies and call
todo_role_for_heading, so the counter and the region writer agree by construction rather than by two lists staying in sync.The markers this drops —
agent backlog,agent action,user action,用户,人工, bareowner— appear in no shipped template or fixture as a heading, and the writer never creates them. Counting them reported work the machinery could not see or act on.Ratchets
Two undeclared multi-value forks disappear with the copies, lowered in this diff on both the registry and
BUDGET_ANCHORsides as the equality check requires:RAW_MATERIAL_KEY_HINTS, the third fork the issue lists, is not touched here: its two definitions carry genuinely different value sets (body/chat/credential/dmagainstcredential/local_path/log/raw) and need their own classification.Validation
The regression test pins all three defects and was mutation-checked: reverting only
state_projection.pyfails it with{'user': 1} != {'user': 0}.🤖 Generated with Claude Code