Skip to content

fix(state): count Todo headings the way the region writer classifies them - #4643

Merged
huangruiteng merged 5 commits into
loopx-project:mainfrom
songoow:codex/single-source-todo-header-markers
Sep 17, 2026
Merged

huangruiteng merged 5 commits into
loopx-project:mainfrom
songoow:codex/single-source-todo-header-markers

Conversation

@songoow

@songoow songoow commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #4447 — Track A, the first of the three undeclared forks the issue lists.

The two copies are one contract

loopx/state_projection.py carried its own copy of the Todo header marker tuples; loopx/control_plane/goals/active_state_metadata.py carried another. Both classify a heading from the same active-state document into the same user / agent / None roles — summarize_state_todo_open_counts counts Todos per role, find_todo_regions locates the regions to write into. One contract, two implementations, and they had drifted.

Measured over 23 realistic headings, the two disagree on 15.

Three of those are defects

Heading state_projection active_state_metadata Why it matters
Agent Todo Archive agent None archived work counted as live
Codex Todo None agent a heading the writer creates, invisible to the count
Ownership user None bare owner marker matches prose sections

The counter had no archive guard although TODO_ARCHIVE_HEADER_MARKERS exists for exactly that, and its owner marker is a substring broad enough to catch Ownership and Downstream Owner Notes.

The consequence is silence, not a wrong number

The counts feed state_projection_gap_warning, which reports that a Next Action is executable while no agent Todo stands behind it. An agent count inflated by archived work makes agent_open > 0, so that warning never fires.

Measured on a state document with two live Todos, two archived, one Codex Todo and one Ownership prose section:

before:  {'user': 1, 'agent': 4}
after:   {'user': 0, 'agent': 3}      <- the truth

The fix

Delete both copies and call todo_role_for_heading, so the counter and the region writer agree by construction rather than by two lists staying in sync.

The markers this drops — agent backlog, agent action, user action, 用户, 人工, bare owner — appear in no shipped template or fixture as a heading, and the writer never creates them. Counting them reported work the machinery could not see or act on.

Ratchets

Two undeclared multi-value forks disappear with the copies, lowered in this diff on both the registry and BUDGET_ANCHOR sides as the equality check requires:

multi_value_forks            4 -> 2
multi_value_forks_semantic   3 -> 1
multi_value_fork_definitions 10 -> 6

RAW_MATERIAL_KEY_HINTS, the third fork the issue lists, is not touched here: its two definitions carry genuinely different value sets (body/chat/credential/dm against credential/local_path/log/raw) and need their own classification.

Validation

python3 examples/semantic-vocabulary-drift-smoke.py     # ok, multi_value_forks=2/2 semantic=1/1 definitions=6/6
python3 -m pytest tests/control_plane/test_todo_machine_region.py -q   # 20 passed

The regression test pins all three defects and was mutation-checked: reverting only state_projection.py fails it with {'user': 1} != {'user': 0}.

🤖 Generated with Claude Code

…them

`loopx/state_projection.py` carried its own copy of the Todo header marker
tuples, and `loopx/control_plane/goals/active_state_metadata.py` carried
another. Both classify a heading from the same active-state document into the
same `user` / `agent` / `None` roles, so this is one contract with two
implementations -- and they had drifted. Measured over 23 realistic headings,
the two disagree on 15.

Three of those disagreements are defects, not preferences:

- `Agent Todo Archive` counted as **live** agent Todos. The writer already
  guards archives through `TODO_ARCHIVE_HEADER_MARKERS`; the counter had no
  such guard.
- `Codex Todo` counted as nothing, although the region writer creates exactly
  that heading, so a whole class of agent Todos was invisible to the count.
- A bare `owner` marker made prose sections match: a section named `Ownership`
  or `Downstream Owner Notes` counted its bullets as user Todos.

The consequence is not a wrong number on a screen. The counts feed
`state_projection_gap_warning`, which reports that a Next Action is executable
while no agent Todo stands behind it. An agent count inflated by archived work
suppresses that warning, so the failure mode is silence. On a state document
with two live Todos, two archived ones, one `Codex Todo` and one `Ownership`
prose section, the counter returned `{'user': 1, 'agent': 5}` where the truth
is `{'user': 0, 'agent': 3}`.

Delete both copies and call `todo_role_for_heading`, so the counter and the
region writer agree by construction. The markers this drops -- `agent backlog`,
`agent action`, `user action`, `用户`, `人工`, bare `owner` -- appear in no
shipped template or fixture as a heading, and the writer never creates them, so
counting them reported work the machinery could not see or act on.

Two undeclared multi-value forks disappear with the copies:
`multi_value_forks` 4 -> 2, `multi_value_forks_semantic` 3 -> 1,
`multi_value_fork_definitions` 10 -> 6, lowered in this diff on both the
registry and `BUDGET_ANCHOR` sides as the equality check requires.

Refs loopx-project#4447

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

exact-head 复核(fcccae5ae)

这一刀修的是"计数和写入方各读一套标记"

summarize_state_todo_open_counts 与 find_todo_regions 解析同一份 active-state 文档、把标题分到同一套 user/agent/None 角色。一份契约两处实现,而它们已经漂移——23 个真实标题里 15 个分类不同。

其中三处是缺陷而非偏好:

标题 state_projection active_state_metadata 后果
Agent Todo Archive agent None 归档被当成在办
Codex Todo None agent 写入方创建的标题,计数看不见
Ownership user None 裸 owner 标记命中散文段落

后果是沉默,不是一个错数字

计数喂给 state_projection_gap_warning,后者报告「Next Action 可执行但背后没有 agent todo」。归档把 agent 计数抬高 → agent_open > 0 → 该告警永不触发。

实测(2 条在办 + 2 条归档 + 1 个 Codex Todo + 1 段 Ownership 散文):

修改前: {'user': 1, 'agent': 5}
修改后: {'user': 0, 'agent': 3}   <- 真值

修法

删掉两份副本,改调 todo_role_for_heading,让计数与写入方由构造保证一致,而不是靠两份列表保持同步。

被丢弃的标记(agent backlog、agent action、user action、用户、人工、裸 owner)在任何已交付模板或夹具中都不作为标题出现,写入方也从不创建它们。计入它们等于报告一个机制既看不见、也无法据以行动的工作量。

棘轮

两份副本消失后,两个未声明分叉随之消除,在同一 diff 内按相等检查的要求同时下调 registry 与 BUDGET_ANCHOR 两侧:

multi_value_forks            4 -> 2
multi_value_forks_semantic   3 -> 1
multi_value_fork_definitions 10 -> 6

RAW_MATERIAL_KEY_HINTS(issue 列的第三个分叉)本 PR 不碰:它的两份定义值集确实不同(body/chat/credential/dm vs credential/local_path/log/raw),需要单独分类。

验证(当前 exact head)

semantic-vocabulary-drift-smoke: ok   # multi_value_forks=2/2 semantic=1/1 definitions=6/6
pytest tests/control_plane/test_todo_machine_region.py  ->  20 passed

回归测试钉住三处缺陷,并做了变异验证:只回退 state_projection.py,测试失败并给出 {'user': 1} != {'user': 0}。

…odo-header-markers

loopx-project#4608 merged while this branch was open and both sides edit the same anchored
ratchet block. Resolved by taking the tighter side of each, as
budgets_only_decrease requires:

- same_runtime_forks_semantic 13 -> 11 from main, which loopx-project#4608 earned by
  importing the Todo task-class vocabulary from its owner.
- multi_value_forks 4 -> 2, multi_value_forks_semantic 3 -> 1 and
  multi_value_fork_definitions 10 -> 6 from this branch, which are the two
  undeclared forks the marker copies were.

Registry and BUDGET_ANCHOR carry the same values. Measured after the merge:
every counter is at budget, with multi_value_forks=2/2,
multi_value_forks_semantic=1/1 and multi_value_fork_definitions=6/6.

Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…ted one

PR loopx-project#4643 retired the AGENT/USER_TODO_HEADER_MARKERS forks, which four
drift tests from loopx-project#4614 had rented as samples for the rename-laundering
limit. Track A retires real forks one by one, so renting them makes
every retirement a test breakage -- this PR included.

The pins now inject their own synthetic fork (one name, two modules,
two disagreeing value sets) and measure against the live baseline
instead of hardcoded counts, so they pin the machinery -- the RFC
Section 9 known limit -- independently of the debt population:

- the laundering test measures base / base+1 / base instead of 3/2;
- the declaration test proves exclusion by removing SOURCE_SURFACES
  from the registry and watching the count rise by exactly one,
  instead of asserting today's undeclared count;
- the three-case rename boundary runs its undeclared cases on the
  synthetic fork; Case 1 (declared rename refused) still uses
  SOURCE_SURFACES, which is still declared and still 4-way divergent;
- the advisory test keeps one real assertion: RAW_MATERIAL_KEY_HINTS,
  the only undeclared fork surviving this PR, must stay listed, with a
  comment that the assertion retires with the fork.

Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

Self-review on exact head 0538faa3c — regression found before push, fixed in this head

What I found while answering "did the full suite pass?": this PR broke four tests in tests/architecture/test_semantic_vocabulary_drift.py — the pins merged in #4614 had rented AGENT_TODO_HEADER_MARKERS / USER_TODO_HEADER_MARKERS (the two forks this PR retires) as their samples for the rename-laundering limit. 4 failed / 57 passed locally. CI hadn't reached the Python shards when I caught it, so the pushed branch never advertised green on those tests.

Fix on this head (0538faa), not a sample swap: the pins now inject their own synthetic fork (one name, two modules, two disagreeing value sets) and measure against the live baseline instead of hardcoded counts (3 → base+1 / base):

  • test_renaming_one_side_of_a_fork_launders_the_semantic_budget — laundering measured as base/base+1/base; identical failure message and RFC §9 pin preserved.
  • test_bounded_context_scope_excludes_only_declared_multi_value_fork — now proves exclusion: remove SOURCE_SURFACES from the registry → undeclared count rises by exactly one. Stronger than the old == 3, which only recorded today's debt population.
  • test_rename_visibility_splits_into_three_cases — undeclared cases run on the synthetic fork; Case 1 (declared rename refused) still uses SOURCE_SURFACES, still declared and still 4-way divergent.
  • test_divergent_value_sets_lists_the_names_a_rename_would_hide — synthetic listed with 2 value sets; keeps one real assertion: RAW_MATERIAL_KEY_HINTS (the only undeclared fork surviving this PR) must stay listed, with a comment that the assertion retires with the fork.

Why synthetic rather than swapping the sample to RAW_MATERIAL_KEY_HINTS: Track A retires real forks one by one — renting the next one just schedules this same breakage for the next PR. The rename-laundering limit is a property of the inventory machinery, so the pin now owns its fixture.

Validation on this head:

  • tests/architecture/test_semantic_vocabulary_drift.py: 61 passed (was 4 failed / 57 passed on 5308c75)
  • tests/control_plane/test_todo_machine_region.py: 20 passed (unchanged, the PR's own regression tests from the original head, mutation-verified earlier)
  • examples/semantic-vocabulary-drift-smoke.py: exit 0, multi_value_forks=2/2 multi_value_forks_semantic=1/1 — all ratchets at floor
  • Injected-fork behavior verified directly: appears in multi_value_forks + divergent_value_sets with the real forks; partial rename removes it from both while real forks survive
  • Related suites (test_todo_next_action_settlement, test_todo_completion_validation, test_goal_objective_projection, test_todo_projection_recovery, test_todo_machine_section_projection, test_project_registry): 223 passed, 1 pre-existing main failure — test_todo_projection_recovery.py::test_production_scale_rebuild_retains_order_and_requires_private_declaration[sqlite] fails identically on clean origin/main (897e9ae), in the Node local_authority_provider.ts path; unrelated to this Python-only diff.

Surfaces changed: loopx/state_projection.py (delegating _role_for_heading to the writer's classifier), plus test/docs files. No behavior change beyond the counted-heading fix described in the PR body: open-role counts now match what the region writer classifies (Agent Todo Archive no longer counts as live; Codex Todo counts; prose Ownership no longer counts as user todos), which un-silences state_projection_gap_warning.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: 0538faa3c70fbe7985ad8d150eda3fd0c9bbfc9e (codex/todo-heading-role-single-source).

动机

#4447 里的第一个未声明 fork:loopx/state_projection.py 自带一份 Todo 标题标记元组,loopx/control_plane/goals/active_state_metadata.py 另有一份。两者都在把同一份 active-state 文档的标题分成 user / agent / None——前者喂 summarize_state_todo_open_counts,后者管区域写入——一份契约两套实现,而且已经漂移出三个真实缺陷:Agent Todo Archive 被当成在办的 agent 待办(没有归档守卫)、Codex Todo(区域写入器自己创建的标题)完全看不见、以及裸 owner 标记把 Ownership 这类散文小节算成 user 待办。

后果不是数字错,而是沉默:计数喂 state_projection_gap_warning,而它只在 agent_open == 0 时才会报出「Next Action 可执行、背后却没有 agent 待办」。

改动思路

入口是 summarize_state_todo_open_counts -> _role_for_heading,现在直接调用写入器的 todo_role_for_heading(先归档守卫、再 user、再 agent),两处副本元组删除。这是「删掉第二份实现」而不是「把两份手工同步」,正是 fork ratchet 想收敛的东西;machine_region.py:137 本来就用这个函数做角色回退,所以计数与区域写入器是按构造一致。registry 与 BUDGET_ANCHOR 同时下调三个 ratchet(multi_value_forks 4→2、multi_value_forks_semantic 3→1、multi_value_fork_definitions 10→6),等值 anchor 因此仍然成立。

具体改动

5 个文件、+160/-68:state_projection.py 删掉两份元组并改为委托(净减生产代码)、registry/anchor 的 ratchet、回归测试,以及 drift 测试相应更新;RAW_MATERIAL_KEY_HINTS(issue 里的第三个 fork)刻意不动,因为它的两份定义值集确实不同,需要单独分类。

我独立跑过前后对比(同一份文档:2 条在办 Agent Todo、2 条归档、1 条 Codex Todo、1 个 Ownership 散文小节):merge-base(origin/main 897e9ae){'user': 1, 'agent': 4} → head {'user': 0, 'agent': 3};pytest tests/control_plane/test_todo_machine_region.py -q → 20 passed;pytest tests/architecture/test_semantic_vocabulary_drift.py -q → 61 passed;smoke → ok,multi_value_forks=2/2 multi_value_forks_semantic=1/1 multi_value_fork_definitions=6/6。我也确认全仓已无对已删元组的引用(只剩 owner 模块自己)。

关键代码讲解

  • state_projection._role_for_heading(342):改为 return todo_role_for_heading(heading),docstring 把三个漂移与「计数被灌水会压掉警告」写清楚——把规则留在被强制执行的地方。
  • active_state_metadata.todo_role_for_heading(89):归档标记优先,其次 user(user todo / owner review reading queue / owner reading queue),其次 agent(agent todo / codex todo / project agent todo);这是唯一实现。
  • state_projection_gap_warning(755):if agent_open == 0 and executable: 才产出 next_action_executable_without_agent_todo 证据——我在 head 上读了这段,确认 PR 的「沉默」因果成立。
  • test_state_counts_classify_headings_exactly_as_the_region_writer:既断言 {'user': 0, 'agent': 3},又逐条断言五个标题的角色(含三个缺陷用例)。
  • inventory_ratchets:三个 fork 预算与 anchor 同步下调并锁死(smoke 输出 2/2、1/1、6/6)。

对主干的风险

最强回归是「旧状态文件里用了被删掉的拼法」:裸 owner、用户、人工、agent action、agent backlog、项目 agent、agent 待办 这些标题此前的角色判定会变成 None,计数下降。方向是 fail-closed(可能多报一次 gap 警告),且 PR 说明这些拼法不作为标题出现在任何已交付模板或 fixture 中、写入器也从不创建它们;我没有逐份历史 goal state 复算,因此把它记在 residual 里。第二个潜在风险是 substring 匹配本身较松,但这是写入器既有的规则,本次是把计数对齐到它、而不是新造规则。

一条 P3(F1,非阻塞):正文里的 before 数字在 merge-base 上复现不出来。正文写「同一个文档:before {'user': 1, 'agent': 5}」,但我用 merge-base 的真实函数量到的是 {'user': 1, 'agent': 4}——旧分类器的 agent 标记里没有 codex todo,所以那条 Codex Todo 贡献的是 0 而不是被算成 live;归档小节贡献 2。head 的 {'user': 0, 'agent': 3} 与三个缺陷(归档算在办、Codex Todo 不可见、Ownership 算 user)我都复核为真,只有 before 少算/多算一格。建议把正文数字改成实测值。

我的整体评价

结论 APPROVE。这是一次正确的去重:删掉第二份分类实现、让计数与区域写入器共用 owner,顺带把三处真实漂移修掉并下调对应 ratchet;生产代码净减少,回退成本是一个 commit。我独立复现了 merge-base 与 head 的计数差异、读了警告的抑制分支、跑了 20 + 61 个测试与 smoke,并确认删除的常量没有残留引用。

唯一 P3 是正文示例数字未复现,非阻塞;RAW_MATERIAL_KEY_HINTS 留作后续分类是恰当的边界。

English verdict: APPROVE - exact head 0538faa; state_projection now delegates to the writer's todo_role_for_heading instead of keeping a drifted copy, which removes two of the three undeclared forks (ratchets 4/3/10 -> 2/1/6, locked, smoke 2/2 1/1 6/6). I reproduced the behaviour on both trees for the described document (merge-base {'user': 1, 'agent': 4} -> head {'user': 0, 'agent': 3}), read the agent_open == 0 suppression branch in state_projection_gap_warning, confirmed no references to the deleted tuples remain, and ran 20 + 61 tests. One non-blocking P3: the PR body's illustrative before-count ({'user': 1, 'agent': 5}) does not reproduce - the merge-base measurement is agent 4, because the old classifier never matched 'codex todo'.

No conflict; the branch was only behind main.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow

songoow commented Sep 17, 2026

Copy link
Copy Markdown
Collaborator Author

Head moved: merged main, so the existing approval no longer covers this head (d8eacce19).

No conflict — the branch was only behind main.

Revalidated: semantic-vocabulary-drift-smoke ok, docs-governance-smoke ok, pytest tests/architecture/ 299 passed.

Per the exact-head rule in AGENTS.md, this needs a re-confirmation on d8eacce19.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: d8eacce197ee63f8ba5c123c97c2d71c2b770cfd (codex/single-source-todo-header-markers, re-review after the earlier head was superseded by a main merge).

动机

loopx/state_projection.py 和 loopx/control_plane/goals/active_state_metadata.py 各存了一份 Todo header marker 表,而它们分类的是同一份 active-state 文档、产出同样的 user/agent/None 角色:一个用于 summarize_state_todo_open_counts 计数,一个用于 find_todo_regions 定位写入区域。两份实现已经漂移:计数那份把裸 owner 当 user 标记(于是 Ownership 这种散文段落被计成 user todo)、没有 archive 守卫(Agent Todo Archive 被算成进行中)、还漏了 writer 会创建的 codex todo。后果是沉默而非数字错——agent 计数虚高使 state_projection_gap_warning(「Next Action 可执行但背后没有 agent Todo」)永远不触发。动机成立。

改动思路

不是把两份表对齐,而是删掉一份:计数改为调用 writer 的 todo_role_for_heading,让计数与写入区域按构造一致。删除副本带来的两个多值 fork 同步下调(registry 与 BUDGET_ANCHOR 两侧一起动,保持等值 anchor)。

具体改动

5 个文件、+160/-68:

  • loopx/state_projection.py:删除 USER_TODO_HEADER_MARKERS/AGENT_TODO_HEADER_MARKERS,_role_for_heading 改为委托 todo_role_for_heading(净减 34 行)。
  • loopx/semantics/vocabulary_v0.json + examples/semantic-vocabulary-drift-smoke.py:multi_value_forks 4→2、multi_value_forks_semantic 3→1、multi_value_fork_definitions 10→6。
  • tests/architecture/test_semantic_vocabulary_drift.py(+128)、tests/control_plane/test_todo_machine_region.py(+54):钉住「计数与 writer 同规则」与精确计数。

我复核的关键点(都在这个 head 上自己跑过):

  • before/after 复现:同一份文档(2 条进行中 agent todo + 2 条 archive + 1 条 Codex Todo + 1 段 Ownership)在 base a96c9aa91c 上得 {'user': 1, 'agent': 4},在 head 上得 {'user': 0, 'agent': 3}——archive 不再算进行中、Codex Todo 被计入、Ownership 不再算 user。
  • 分类器探针:Agent Todo Archive/Ownership/Downstream Owner Notes/用户/人工/Agent Backlog → None;Agent Todo/Codex Todo → agent;Owner Reading Queue → user。
  • pytest -q tests/control_plane/test_todo_machine_region.py → 20 passed;三个依赖 state_projection 的 control-plane 套件 → 51 passed;架构套件 → 67 passed;drift smoke → exit 0 且三个 ratchet 打印 2/2、1/1、6/6(等于实测)。
  • 真实数据回归扫描:对 5 份真实 ACTIVE_GOAL_STATE.md 的全部 ## 标题(67 个)用 base 与 head 两套分类器分别判定,差异 0——说明这次修的是潜在漂移,当前真实文档行为不变。

遗留问题(非阻塞,P3)

被删掉的标记里包含中文 用户/人工(archive 表里仍保留中文)。也就是说,手写 ## 用户待办 这类标题在 base 上会算作 user todo,在 head 上返回 None。我全仓与真实状态文档扫描确认:没有已发布模板/fixture 使用这些标题(唯一的 用户 出现在 README 正文),所以今天不改变任何行为。但如果非英文状态标题是受支持的用法,最小修法是把中文 user/agent 标记补进共享表;否则应在模块 docstring 与状态模板文档里写明「只有列出的英文标题会被识别」。

对主干的风险

最强回归是计数继续把归档/散文当工作进行中,或漏掉 writer 创建的标题,从而压掉 gap 警告。这次改动把两侧收敛到一个函数,方向正确;剩下的机制性风险是依旧靠有序子串匹配:这次通过把 owner 换成 owner reading queue 这类更长短语消除了 Ownership 误判,但将来若加入短标记,同类误判可以回来。未验证维度:中文/非英文标题(见 P3)。不涉及写路径、quota 或持久化状态;回退成本一个 commit。本次 head 只多了一次 main 合并,我对当前 main 重比了内容增量(5 文件、+160/-68)并重跑了全部证据。

我的整体评价

结论 APPROVE。这是「删掉重复知识而不是同步两份表」的正确做法:两侧按构造一致,三处实际缺陷(归档、Codex Todo、Ownership 散文)都被修正,且我用真实文档验证了当前行为不变。唯一 P3 是非英文标题的识别收窄,属文档/标记补充问题,不阻塞。

English verdict: APPROVE - exact head d8eacce (re-review after a main merge; content delta versus main unchanged at 5 files, +160/-68). The projection now calls the writer's todo_role_for_heading instead of keeping a second marker table: I reproduced the counts on base a96c9aa vs head ({'user': 1, 'agent': 4} -> {'user': 0, 'agent': 3}), the classifier rejects archived/prose headings and keeps Agent Todo, Codex Todo and Owner Reading Queue, 20 + 51 + 67 tests pass, the drift smoke exits 0 with multi_value_forks=2/2, multi_value_forks_semantic=1/1 and multi_value_fork_definitions=6/6, and all 67 headings across five live ACTIVE_GOAL_STATE.md files classify identically at base and head. One non-blocking P3: the removed Chinese user markers mean a hand-written heading such as 用户待办 no longer counts, which no shipped template or fixture exercises today.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: d8eacce197ee63f8ba5c123c97c2d71c2b770cfd.

动机

按 #4447 当前 Track A 的修正结论,这两份 Todo marker 分类的是同一份文档、同一种角色,应合并到已有 owner。早期“只补 scope declaration”的建议不足以修复这里已经存在的行为分歧。此 PR 完成一个独立有效的增量:计数不再把归档当在办、不再漏掉 Codex Todo,也不再把 Ownership 散文算作 User Todo;剩余 RAW_MATERIAL_KEY_HINTS 分类仍归原 tracker,不宣称整个语义收敛工作完成。

独立实测同一合成文档:base a96c9aa91c9d92454efcd7937483bac76e35fd4a 的 {'user': 1, 'agent': 4} → 当前 head 的 {'user': 0, 'agent': 3}。原 PR 正文的 before agent=5 应为 4。更重要的是:只有归档任务、但 Next Action 可执行时,真实 refresh_state_run 从无告警变为返回完整的缺失 Agent Todo 告警。

改动思路

计数复用 active_state_metadata.todo_role_for_heading,删除第二份分类规则。只同步两份 tuple 会保留未来漂移的来源;增加新分类框架也没有必要。调用方向仍是外层 projection → 既有 control-plane metadata owner,没有新增 capability、provider、状态表或权限。

本次保留已有私有 delegate:它没有独立决策,内联可以进一步省掉一跳,但当前形态只承载兼容性说明,无额外状态或维护分支,不构成阻塞。最有价值的简化已经由删除重复知识实现。

具体改动

完整范围为 5 文件、+160/-68:运行时改为共享分类器;registry 与 smoke anchor 同步收紧三个 fork 预算;Todo 回归测试保护三个实际缺陷;架构测试改用注入的双模块分歧样本,避免依赖正在消除的真实债务。

关键代码讲解

  • state_projection._role_for_heading(342 行):只调用 canonical classifier;归档优先级、user/agent marker 不再本地重述。
  • summarize_state_todo_open_counts(439 行):遍历完整文档、按标题切换角色、排除已完成 checkbox,结果交给 gap warning。
  • state_projection_gap_warning(755 行):缺省时采用上述计数;显式传入 Todo summary 的既有优先级保留。agent_open == 0 且 Next Action 可执行时,产出 requires_todo_expansion 与可读修复建议。
  • _with_synthetic_fork(架构测试 195 行):两个源模块、同名、不同值集,独立验证 baseline → baseline+1;单侧重命名后回落,保留现有 name-based inventory 的已知限制,没有声称解决 rename laundering。

对主干的风险

主要兼容性变化是手写宽泛标题(例如 Agent Backlog、用户待办)不再计数。已扫描交付文档和 fixtures,没有使用这些标题的模板;写入器也从未把它们当作对应区域。任意历史手写文档不在本次证明范围内。共享分类器仍使用有序 substring 匹配;将来增加短 marker 仍可能误判,应在同一个 owner 和回归表中处理,不能恢复另一套规则。

P3,非阻塞: vocabulary_v0.json 的 multi_value_forks_note 仍写 “4 counted forks”,而数值预算已是 2。应在现有语义文档对账工作中刷新这句说明;当前数值约束和 scope declaration 的执行一致,不影响本次正确性。

语义与 CI 对齐

这是复用已有语义,未新增词表、缩小扫描范围或通过改名隐藏 fork。三个预算实测为 2/2、1/1、6/6。最新 main 9118568dde6d07610dc3df4f5a8c5e9cac799fc4 另收紧了 conflicting budgets;已构造无冲突的合并树并验证 16/16、55/55 保留。

验证结果:

  • 当前 head:Todo machine region、semantic vocabulary drift、state refresh projection 三套测试 89 passed。
  • 与上述最新 main 的合并树:同组测试 90 passed,语义 smoke 通过,包含 main 新增的预算测试。
  • 相同合成输入经过真实 count → gap → refresh 路径:归档、Codex、散文、completed、标题顺序、200 条无关 checkbox、显式 summary 与无 Next Action 分支均检查;旧实现违反 6 项独立预期,head 与合并树通过。隔离文件的 dry-run 读回确认状态未被写入。
  • 仓库规定的 Ruff 范围通过;额外检查改动的 runtime/test 文件通过;配置内 Mypy 22 文件通过。
  • semantic-vocabulary-drift-smoke.py、state-projection-gap-smoke.py、docs-governance-smoke.py 通过;按当前 CLI 生成 inventory 到临时文件再 --check --output 读回通过。
  • canary premerge --from-git-diff:diff/compile、6 项 canary、公开边界扫描全部通过;5 个改动文件无私有内容,0 失败、0 manual hold。质量回执 cqr_3e286bd951f3e4cdc2c8 已精确校验有效,未应用额外 safe-fix。

初次隔离环境缺 Node parser 依赖导致测试失败,执行 npm ci --ignore-scripts 后完整重跑通过。额外将 examples 纳入 Ruff 的探针发现既有 unused render_inventory import;base/head 同样复现,examples 不在仓库规定的 lint 范围内,此处不声称扩展探针通过。没有跳过必需的本地验证。按当前评审配置,未查询或等待远端 CI;本次不是全平台发布验证。

可直接复现关键负例:

from loopx.state_projection import state_projection_gap_warning
state = "## Agent Todo Archive\n- [ ] archived\n\n## Next Action\n- Implement the reader and validate the result.\n"
gap = state_projection_gap_warning(state)
assert gap is not None and gap["agent_open_count"] == 0
assert gap["first_evidence"][0]["kind"] == "next_action_executable_without_agent_todo"

examples/state-projection-gap-smoke.py 另覆盖真实 refresh 和下游 quota 路由。没有 SQL authority、frontend 或 Lark contract 变化,因此无需数据库迁移或新增 UI 配套;各入口消费同一既有诊断形状。

我的整体评价

APPROVE。 完整增量与当前目标一致:减少一个实际会隐藏工作缺口的规则分歧,同时保留持久的回归覆盖。已针对当前 head 重新执行证据,没有沿用旧批准代替验证。非阻塞风险是手写别名兼容性及上述陈旧说明;不存在新的写入或执行授权。

本记录用于用户已授权的维护者合并;实际合并仍以紧接合并前的 exact-head readiness 为准,合并动作绑定上述 head。

English verdict: APPROVE - d8eacce. Reusing the writer's canonical heading classifier fixes archived/prose counts and missing Codex Todos. Independent baseline/head refresh-path counterexamples pass at head; 89 exact-head and 90 integrated-tree tests pass, native lint/type checks and all selected premerge checks pass, and the exact-scope quality receipt is valid. Non-blocking risks: narrowed manual heading aliases and stale fork-count prose. CI was not consulted under the configured policy.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head: 88fabdfeb9ac84b2c00761c58b0800073002172a.

本次复审针对 GitHub 合入 main 后的新提交。分支已包含 main b7f66d3f59fe0fd6aea29b088613f6a3d55035d9;PR 自身仍为相同的 5 文件增量。比对确认受影响的 classifier、Todo caller、语义检查与测试文件和此前验证的合并树一致,并在新 head 重新执行 90 项测试、真实 refresh 反例、Ruff、Mypy、质量回执及 premerge。以下为当前 head 的完整结论。

动机

按 #4447 当前 Track A 的修正结论,这两份 Todo marker 分类的是同一份文档、同一种角色,应合并到已有 owner。早期“只补 scope declaration”的建议不足以修复这里已经存在的行为分歧。此 PR 完成一个独立有效的增量:计数不再把归档当在办、不再漏掉 Codex Todo,也不再把 Ownership 散文算作 User Todo;剩余 RAW_MATERIAL_KEY_HINTS 分类仍归原 tracker,不宣称整个语义收敛工作完成。

独立实测同一合成文档:base a96c9aa91c9d92454efcd7937483bac76e35fd4a 的 {'user': 1, 'agent': 4} → 当前 head 的 {'user': 0, 'agent': 3}。原 PR 正文的 before agent=5 应为 4。更重要的是:只有归档任务、但 Next Action 可执行时,真实 refresh_state_run 从无告警变为返回完整的缺失 Agent Todo 告警。

改动思路

计数复用 active_state_metadata.todo_role_for_heading,删除第二份分类规则。只同步两份 tuple 会保留未来漂移的来源;增加新分类框架也没有必要。调用方向仍是外层 projection → 既有 control-plane metadata owner,没有新增 capability、provider、状态表或权限。

本次保留已有私有 delegate:它没有独立决策,内联可以进一步省掉一跳,但当前形态只承载兼容性说明,无额外状态或维护分支,不构成阻塞。最有价值的简化已经由删除重复知识实现。

具体改动

完整范围为 5 文件、+160/-68:运行时改为共享分类器;registry 与 smoke anchor 同步收紧三个 fork 预算;Todo 回归测试保护三个实际缺陷;架构测试改用注入的双模块分歧样本,避免依赖正在消除的真实债务。

关键代码讲解

  • state_projection._role_for_heading(342 行):只调用 canonical classifier;归档优先级、user/agent marker 不再本地重述。
  • summarize_state_todo_open_counts(439 行):遍历完整文档、按标题切换角色、排除已完成 checkbox,结果交给 gap warning。
  • state_projection_gap_warning(755 行):缺省时采用上述计数;显式传入 Todo summary 的既有优先级保留。agent_open == 0 且 Next Action 可执行时,产出 requires_todo_expansion 与可读修复建议。
  • _with_synthetic_fork(架构测试 195 行):两个源模块、同名、不同值集,独立验证 baseline → baseline+1;单侧重命名后回落,保留现有 name-based inventory 的已知限制,没有声称解决 rename laundering。

对主干的风险

主要兼容性变化是手写宽泛标题(例如 Agent Backlog、用户待办)不再计数。已扫描交付文档和 fixtures,没有使用这些标题的模板;写入器也从未把它们当作对应区域。任意历史手写文档不在本次证明范围内。共享分类器仍使用有序 substring 匹配;将来增加短 marker 仍可能误判,应在同一个 owner 和回归表中处理,不能恢复另一套规则。

P3,非阻塞: vocabulary_v0.json 的 multi_value_forks_note 仍写 “4 counted forks”,而数值预算已是 2。应在现有语义文档对账工作中刷新这句说明;当前数值约束和 scope declaration 的执行一致,不影响本次正确性。

语义与 CI 对齐

这是复用已有语义,未新增词表、缩小扫描范围或通过改名隐藏 fork。三个预算实测为 2/2、1/1、6/6。最新 main 9118568dde6d07610dc3df4f5a8c5e9cac799fc4 另收紧了 conflicting budgets;已构造无冲突的合并树并验证 16/16、55/55 保留。

验证结果:

  • 当前 head:Todo machine region、semantic vocabulary drift、state refresh projection 三套测试 90 passed。
  • 与上述最新 main 的合并树:同组测试 90 passed,语义 smoke 通过,包含 main 新增的预算测试。
  • 相同合成输入经过真实 count → gap → refresh 路径:归档、Codex、散文、completed、标题顺序、200 条无关 checkbox、显式 summary 与无 Next Action 分支均检查;旧实现违反 6 项独立预期,head 与合并树通过。隔离文件的 dry-run 读回确认状态未被写入。
  • 仓库规定的 Ruff 范围通过;额外检查改动的 runtime/test 文件通过;配置内 Mypy 22 文件通过。
  • semantic-vocabulary-drift-smoke.py、state-projection-gap-smoke.py、docs-governance-smoke.py 通过;按当前 CLI 生成 inventory 到临时文件再 --check --output 读回通过。
  • canary premerge --from-git-diff:diff/compile、6 项 canary、公开边界扫描全部通过;5 个改动文件无私有内容,0 失败、0 manual hold。质量回执 cqr_fe30402b17b911fd648c 已精确校验有效,未应用额外 safe-fix。

初次隔离环境缺 Node parser 依赖导致测试失败,执行 npm ci --ignore-scripts 后完整重跑通过。额外将 examples 纳入 Ruff 的探针发现既有 unused render_inventory import;base/head 同样复现,examples 不在仓库规定的 lint 范围内,此处不声称扩展探针通过。没有跳过必需的本地验证。按当前评审配置,未查询或等待远端 CI;本次不是全平台发布验证。

可直接复现关键负例:

from loopx.state_projection import state_projection_gap_warning
state = "## Agent Todo Archive\n- [ ] archived\n\n## Next Action\n- Implement the reader and validate the result.\n"
gap = state_projection_gap_warning(state)
assert gap is not None and gap["agent_open_count"] == 0
assert gap["first_evidence"][0]["kind"] == "next_action_executable_without_agent_todo"

examples/state-projection-gap-smoke.py 另覆盖真实 refresh 和下游 quota 路由。没有 SQL authority、frontend 或 Lark contract 变化,因此无需数据库迁移或新增 UI 配套;各入口消费同一既有诊断形状。

我的整体评价

APPROVE。 完整增量与当前目标一致:减少一个实际会隐藏工作缺口的规则分歧,同时保留持久的回归覆盖。已针对当前 head 重新执行证据,没有沿用旧批准代替验证。非阻塞风险是手写别名兼容性及上述陈旧说明;不存在新的写入或执行授权。

本记录用于用户已授权的维护者合并;实际合并仍以紧接合并前的 exact-head readiness 为准,合并动作绑定上述 head。

English verdict: APPROVE - 88fabdf. Reusing the writer's canonical heading classifier fixes archived/prose counts and missing Codex Todos. Independent baseline/head refresh-path counterexamples pass at head; 90 updated exact-head and 90 integrated-tree tests pass, native lint/type checks and all selected premerge checks pass, and the exact-scope quality receipt is valid. Non-blocking risks: narrowed manual heading aliases and stale fork-count prose. CI was not consulted under the configured policy.

@huangruiteng
huangruiteng merged commit ed1a826 into loopx-project:main Sep 17, 2026
19 of 20 checks passed
@huangruiteng

Copy link
Copy Markdown
Collaborator

Merged as ed1a8269ef2f12ad8863a04254d5d17021fa8100 after the explicitly requested maintainer review and merge.

Final reviewed head: 88fabdfeb9ac84b2c00761c58b0800073002172a; full exact-head approval. Immediately before merge, pr-review --check-merge-readiness returned ready=true, with a valid APPROVE and zero unresolved review threads. It also reported admin_bypass_required=true; the user-authorized self-merge used that maintainer bypass with an exact-head guard. The configured policy does not consult CI, so no remote CI result is represented as passed.

Local evidence: 90 tests passed, real refresh-path counterexamples passed, native Ruff/Mypy passed, exact-scope quality receipt cqr_fe30402b17b911fd648c valid, and all selected premerge checks passed with zero failures or manual holds. The original BEHIND result was resolved by updating the PR branch and repeating qualification on the new head before merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants