feat(manager): type declared evidence-source health in the turn window - #4810
Conversation
The manager evidence window declared every source with only source_id, status and scope, so a source that was declared but read nothing could only be narrated as a staleness caveat. Each declared source now carries a typed health row with freshness (current|stale|unknown), the typed reason, the coverage effect and the next action, reusing the read refusal vocabulary and the remote read path's next-action wording. A source that contributed nothing to the window therefore reads as a coverage fact - 'no evidence was read from this source; the answer must not present it as no progress' - instead of an unqualified progress reading, and the local source is typed the same way when the window read nothing. Successor slice of todo_386976ec9aae (steward RFC M2 anchor A11). Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
精确 head:6af5edb639478350a5c4b8c655f9bbaed631fd6f(base be7789fd9)。无阻断发现,建议合并;合并决定留给 maintainer。
动机
管家每轮的 evidence window 会声明每个证据源,但只用 source_id / status / scope 表达。所以「声明了却没读到」这件事在包体里没有类型化表示:declared_unread_sources: ["ssh:ark-devbox"] 只是把 id 列出来,没说什么;local 源无论这一窗口读没读,永远是 available。结果就是模型只有两条路:要么省略这个缺口,要么把它写成「数据可能过期」这类免责声明——而读者既看不出是哪个源没读到、覆盖了什么,也分不清「没读到」和「这个 Goal 没有进展」。
改动思路
给每个被声明的源加一条类型化的 health row,复用读失败(#4787)与远端读路径已有的四元组:源 id、类型化原因、覆盖率影响、下一步动作;并复用 ssh 路径的 source_freshness 词表,把新鲜度收敛成闭集 current | stale | unknown。渲染放在同一个 owner 内、由同一个 sources 列表派生,因此 source_health 与 declared_unread_sources 不可能互相矛盾。
具体改动
loopx/chat_manager_context.py:新增MANAGER_SOURCE_FRESHNESS_VALUES等 10 个常量(含那句关键的覆盖率影响文案)、纯函数_source_health_rows(sources, read_status=...),以及_evidence_window返回体里新增的source_health成员。- 分类规则:
available且本窗口读过 →current,reason/coverage_effect/next_action 全为 null;available但没读 →unknown+local_evidence_not_read_this_turn(窗口读了个空,不能被当成读到了);not_read→stale+declared_source_not_read_this_turn;not_configured→unknown+ 声明自带原因 + 注册别名并重试;任何未来状态落到unknown+source_status_<status>,不会静默变成 current。 tests/test_chat_manager_context.py:新增 3 个用例(读窗口与未读窗口的 row builder、远端源与缺别名源的类型化、turn context 里每个声明源都有一行且顺序一致)。
对主干的风险
- 纯增字段:
sources、declared_unread_sources、read_status、窗口边界与其它成员一律不变,唯一新增的是source_health;938 个 manager/steward/chat/coordination 测试保持全绿。 - 对现有读者的影响:只有对
evidence_window做整字典严格比较的读者才会观察到差异,仓库内的断言都按具名成员读取。 - 反证:把这个 head 的测试文件放到 base
be7789fd9上跑,两个新用例都以KeyError: 'source_health'失败,说明断言确实在测新增契约而不是现有输出。 - 诚实边界:这次改的是「包体里有类型化事实」,没有强制答案文本引用它们;父任务的「真实答案里原始 provider 文本的复测」仍然是验收项,未被这次改动关闭。
我的整体评价
正向且比例合适。它把一句本来只能靠模型自觉写的免责声明,换成包体里可被断言、可被引用的一条类型化事实,且没有引入新模块、能力、状态或 CLI 面;分类规则集中在一个纯函数里,词表与既有的读失败/远端源保持一致,非法状态(未来的新 status)也不会静默降级成 current。没有阻断问题。
English verdict: APPROVE - the manager evidence window now types every declared source's health (freshness, reason, coverage effect, next action) instead of leaving a declared-but-unread source untyped; the change is additive on an existing packet with one pure helper and three new cases, verified by a base/head counterfactual (KeyError: source_health on base), 24 passed in the changed suite, 52 in the adjacent suites and 938 in the manager/steward/chat/coordination selection with ruff clean. It does not yet force the answer text to cite the rows, so the parent row's live re-measurement stays open. Merging is a maintainer decision.
The manager turn context declares every evidence source, but only with
source_id,statusandscope. A source that was declared and read nothing therefore had no typed representation: the window could saydeclared_unread_sources: ["ssh:ark-devbox"]while nothing in the packet said what that means for the answer, so staleness could only be narrated as a caveat, and a stale or unread source was indistinguishable from a source that reported no progress.This adds one typed health row per declared source, reusing the vocabulary the read refusals already use (
source id, typed reason, coverage effect, next action) and the remote read path's next-action wording:{"source_id": "ssh:ark-devbox", "source_host": "ark-devbox", "status": "not_read", "freshness": "stale", "reason": "declared_source_not_read_this_turn", "coverage_effect": "no evidence was read from this source; the answer must not present it as no progress", "next_action": "Read this source with the manager evidence read tool, or let the next Turn's source rotation dial it."}freshnessis a closed set (current|stale|unknown) declared asMANAGER_SOURCE_FRESHNESS_VALUES.currentwith no reason, no coverage effect and no next action.localsource is typed the same way: reachable but unread in this window isunknownwithlocal_evidence_not_read_this_turn, so a window that read nothing cannot be answered as if it had.not_configuredkeeps the declaration's own reason (ssh_alias_not_configured) and gets the register-and-retry next action.The change is additive on the packet:
sources,declared_unread_sources,read_statusand every other existing field keep their exact shape, so no current reader changes behaviour. Owning boundary is the manager evidence window inloopx/chat_manager_context.py; no new capability, provider, state or CLI surface is added.Entry points: the manager/agent answer path consumes the packet directly, so no frontend or Lark companion change is needed; the packet is what the steward prompt already receives. Verified by reading the manager turn-context consumers in this repository.
Validation:
uv run --extra test python -m pytest tests/test_chat_manager_context.py -q-> 24 passed (3 new cases).uv run --extra test python -m pytest tests/test_manager_ssh_evidence.py tests/test_chat_project_coordination.py tests/test_chat_manager_inspection.py tests/test_manager_context_tracking.py -q-> 52 passed.uv run --extra test python -m pytest tests -q -k 'manager or steward or chat or coordination'-> 938 passed, 9879 deselected.uv run --extra test ruff check loopx/chat_manager_context.py tests/test_chat_manager_context.py-> all checks passed.Successor slice of the evidence-source-health row (steward RFC M2 anchor A11); the remaining live re-measurement of raw provider text in answers stays open on the parent row.
Control-plane change (
loopx/**): proposed for review and left for the maintainer to merge.