fix(chat): hold one managed executor per steward binding - #4481
Conversation
The steward channel's managed segment transport started a fresh DeepSeek Harness segment per Chat turn without checking whether the previous one was still running. An interrupted turn freed the session while its segment kept running in the host, so the follow-up turn started a second executor for the same binding -- the case the session-execution-modes RFC says must be refused with a typed error -- and the abandoned segment's answer was still appended to the channel's visible history. The adapter now owns one executor slot per binding: a start while the previous segment is still running waits a bounded hand-off window and then refuses with the typed managed_host_chat_segment_in_flight, and an interrupted segment's answer is discarded instead of entering visible history. The slot is released only when the segment's thread exits, including after a channel timeout, because the host cannot cancel a segment it has already started. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The manager evidence contract and both harness-selection mirrors now say that "exactly one segment per turn" is held (typed refusal on a second start, a discarded interrupted answer), and the M2-M3 row splits what has shipped from what has not. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed exact head: 78f7ebc2f26ce1b3cf7efaa43908aab5dcce670d
动机
管家 Chat 通道在托管宿主上按 turn 起一个 dsh segment。托管宿主没有 host-side
cancel:Chat runtime 取消一个 turn 只会丢弃它的结果,segment 仍在宿主里跑。
于是紧接着的下一轮 start_turn 会为同一个 binding 再起一个 segment——同一个
管家会话上出现两个执行器。docs/architecture/rfcs/harness-selection-dsh-pi-v0.md
的宿主模式 M2-M3 一行要求的正是「不引入第二执行器」并以 typed 错误拒绝,而当时的
实现既没有拒绝,也把被放弃 segment 的答案写进了通道可见的 history,让一个
已被取消的回答读起来像管家最新的话。
改动思路
把「一个 binding 一个托管执行器」从假设变成被持有的不变量,位置选在段边界——也就是
唯一知道 segment 生死的地方。三个决定值得说明:
- 只有占位交接串行。
start_turn在自己的线程里同步等 segment 结束,如果在
_claim_segment_slot()全程持锁,第二个调用方会卡在锁上直到第一个 segment 结束,
拿不到任何结论;现在它读到一次 typed refusal,这正是 RFC 要求的语义。 - 拒绝用 typed error(
MANAGED_HOST_CHAT_SEGMENT_IN_FLIGHT)而不是排队或静默并行,
并且 code 描述的是通道所处的状态,不是用户犯了错。 - 丢弃语义放在通道侧而不是 runtime 侧,因为通道是唯一知道自己这一轮被取消的边界。
段照常退出(它仍是该 binding 的唯一执行器),只是它的文字不再进入可见历史。
具体改动
loopx/chat_dsh.py- 新增
MANAGED_HOST_CHAT_SEGMENT_IN_FLIGHT与SEGMENT_HANDOFF_GRACE_SEC = 5.0:
后者是「已经在收尾的 segment」的交接窗口,不是 deadline;窗口过后仍活着的
segment 继续独占该 binding。 - 新增
_SegmentSlot(thread, interrupted)与 adapter 字段_slot/_slot_lock。 - 新增
_claim_segment_slot():上一个 segment 还活着时先join(窗口),仍活着则
抛 typed refusal;否则用新槽位占位。 start_turn()开头占位,创建后把线程登记进槽位;返回前若槽位被标记interrupted,
只解析并返回结果,不写 history、不发answer.delta/answer.final。interrupt_turn()在锁内把运行中的槽位标记为interrupted;注释说明托管宿主
没有可调用的取消,槽位到 segment 退出前一直占用。
- 新增
tests/test_chat_dsh_adapter.py:新增_blocking_adapter帮助函数(把第一个
segment 阻塞到测试放行,并把交接窗口改成 0.05s 以免测试变慢)与三个用例——
运行中拒绝第二次 start、被中断的答案不进入可见历史、上一个 segment 退出后下一轮
正常开始并保留完整 history。- 三处文档把这条契约写在它自己所在的边界上:
harness-selection-dsh-pi-v0.md与其
.zh-CN.md对应行由「未实现」改为「部分已实现」并列出仍未实现的部分;
manager-evidence-and-continuity-v0.md在托管段传输一节说明「exactly one」是被
持有的而不是被假设的。 loopx/semantics/inventory_v0.json:named_string_constants2050→2051,由新常量
引起,与产生它的提交同一个 commit,语义词汇漂移 smoke 因此仍绿。
对主干的风险
低。改动只落在管家 Chat 通道的托管段边界,且方向是收紧:原先静默起第二个执行器,
现在给出一个可读的 typed refusal。影响面是管家通道自己的 turn 生命周期,不动
Turn、CLI、profile 或权限状态,特征测试与既有 97 个 chat 用例覆盖了共享面。
两点如实记录,都不阻塞:
- 托管宿主目前没有取消能力,被中断的 segment 仍会跑到自己的超时为止并占用槽位。
这是宿主能力缺口,本 PR 只把它变成显式状态,代码注释与文档均已写明。 _claim_segment_slot()判定「上一个 segment」的依据是槽位里已登记线程,因此在
「占位成功」到「登记线程」之间的极窄窗口里,第二次 start 会看到空槽位并放行。
我没有把这段窗口改成「无线程即视为占用」,因为那样一旦 start 在两步之间抛错,
binding 会被永久卡死在 refusal 上;现在的取舍是窗口内偶发放行、异常可自愈,并且
runtime 的单活跃 turn 闸门把窗口压到两次相邻语句之间。若后续把取消能力补进宿主,
这里应一起收紧。
我的整体评价
APPROVE。这是把 RFC 已经写明的规则落到实现上的一小刀,改动集中在段边界,既没有
新模块、新 CLI、新 schema,也没有把管家通道的权威边界往外扩。校验在确切 head 上
跑过:10 个 adapter 用例、97 个 chat 面用例、托管管家真实传输 smoke(segment 沙箱
read-only 且写被拒)、ruff、docs governance smoke,以及 canary premerge 14 项全过、
manual_holds=0;change-quality receipt cqr_88a0e857815e5a1f65eb 对同一
scope_fingerprint 校验为 valid。反证也做了:旁路 typed refusal 后拒绝用例以
DID NOT RAISE 失败,旁路丢弃逻辑后可见历史用例失败,两处都已还原。
English verdict: APPROVE — head 78f7ebc2f26ce1b3cf7efaa43908aab5dcce670d holds one managed executor per steward binding, refuses a second start with the typed managed_host_chat_segment_in_flight, and keeps an interrupted segment's answer out of visible history; two non-blocking observations (no host-side cancel yet; a narrow claim-to-registration window) are recorded above. Validation at this head: 10 adapter tests, 97 chat-surface tests, the bundled steward managed-chat smoke (read-only segment sandbox enforced), ruff, docs governance, canary premerge 14/14 with 0 manual holds, and a valid change-quality receipt.
问题
管家 Chat 通道在 managed 宿主上执行时,每个 Chat turn 都会起一个 dsh segment。但 Chat runtime 在取消一个 turn 时会丢弃它的结果,而 managed 宿主没有 host-side cancel:被取消的 segment 仍在宿主里跑。于是下一轮
start_turn会再起一个 segment —— 同一个 binding 上出现两个执行器,正是docs/architecture/rfcs/harness-selection-dsh-pi-v0.md要求拒绝、并给出 typed error 的情形。同时被放弃的那个 segment 的答案仍会被写进通道可见的history。改动
loopx/chat_dsh.py:给 adapter 加一个 per-binding 的执行器槽位(_SegmentSlot+_slot_lock)。_claim_segment_slot()在起 segment 前占位:若上一个 segment 还活着,先在有限交接窗口内等待它退出;窗口过后仍活着,则抛出 typed errormanaged_host_chat_segment_in_flight,而不是悄悄起第二个执行器。start_turn里同步等自己的 segment,持锁会让第二个调用方卡在锁上,而不是读到那次 typed refusal。interrupt_turn()把正在跑的槽位标记为interrupted。segment 照常退出(它仍是该 binding 的唯一执行器),但它的答案不再折进可见 history —— 否则一个被放弃的回答会读起来像管家最新的话。start_turn在被标记后仍返回解析结果,只是不写 history、不发answer.delta/answer.final。tests/test_chat_dsh_adapter.py:三个新用例 —— 运行中拒绝第二次 start、被中断的答案不进入可见 history、上一个 segment 退出后下一轮可以正常开始。校验
python -m pytest tests/test_chat_dsh_adapter.py -q→ 10 passedpython -m pytest <chat surface suite> -q(chat_dsh / chat_session_active_turn / chat_turn_wait / chat_agent / chat_manager_context / chat_response_parsing)→ 97 passedpython examples/loopx-steward-managed-chat-smoke.py --timeout-seconds 180→ 通过;turn_status=completed,execution_profile=deepseek-v4-flash@high,segment 沙箱read-only且写被拒绝python -m ruff check loopx/chat_dsh.py tests/test_chat_dsh_adapter.py→ All checks passedpython examples/docs-governance-smoke.py→ okloopx canary premerge --from-git-diff --goal-id loopx-meta→ passed,selected=14,failures=0,manual_holds=0cqr_88a0e857815e5a1f65eb(scope_fingerprint=88a0e857815e5a1f65eb6301a525d1d87d10bf08513d54e452cb3e1917e85394,state=valid)test_a_second_segment_start_is_refused_while_one_is_running以DID NOT RAISE失败;把中断后丢弃答案旁路掉 →test_an_interrupted_segment_answer_never_becomes_visible_history失败。两处都已还原,测试重新通过。对主干的风险
低。改动只落在管家 Chat 通道的 managed 段边界上,且是收紧语义:原先会静默起第二个执行器,现在给出一个用户可读的 typed refusal。缺失的是 host 侧取消本身 —— managed 宿主目前没有取消能力,所以被中断的 segment 仍会跑到自己的超时为止;这一点在代码注释和文档里写明,等待 host 提供取消后收紧。