fix(steward): let the machine setting decide the manager channel's executor - #4510
Conversation
…ecutor A manager connection stored the executor endpoint as a decision when it was created, and the Lark route, the authorized-connection resolution and the answering Turn all read that record instead of the machine. A machine that later selected another steward executor therefore kept answering on the endpoint that was the default on the day of the connection: its capability readback reported the managed host while the channel still ran, and failed, on the interactive CLI endpoint. The connection write path had the same shape and no surface to change it, so the value could not be corrected at all. The machine configuration is now the one owner of that choice. The connection record keeps the resolved endpoint as an observation with its source; every read path re-resolves through the new `manager_connection_executor_endpoint` owner, and a connection write records the machine's current resolution while refusing a request that tries to override it. A Session left behind by a machine that changed its selection is refused with the typed `manager_channel_executor_rebind_required` receipt, and the reply names the one action that repairs it instead of the generic manager failure. Verified: the changed Lark, manager-channel, handoff and Lark-API suites pass (182 passed), including four new cases covering route precedence, the write record and refusal, authorized-session matching and the typed rebind reply; the steward channel-binding, steward managed-chat and managed-turn operator-flow smokes pass. The pre-existing failure of `test_every_production_steward_caller_passes_the_machine_defaults` reproduces unchanged on origin/main. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed exact head: bdca4e33cd9873b8301184d38c14c6941d109dfb (12 files, +278/-25).
动机
管家通道的执行器有两个所有者:machine configuration 的 steward_executor 命名空间,以及创建连接时写进 binding.routing.executor_endpoint_id 的那份值。读取路径只认后者,于是"本机选了 dsh"这件事对实际应答没有任何约束力。在本机上的真实表现正是如此:/api/chat/capabilities 报 executor_endpoint=dsh、source=machine_configuration、available=true,而管家群仍走 codex,并因为该端点属于已耗尽额度的个人订阅而返回 host_gate 失败。写入路径同形,且没有任何界面能改它——Dashboard 从不发送 executor_endpoint_id,编辑接口又会回读旧值,所以这个错误值无法被修正。
改动思路
删掉第二个决定,而不是加一层对账。machine configuration 成为唯一所有者;连接记录保留解析结果,但把它降级为观测值(连同新的 executor_endpoint_source)。所有读取路径统一走同一条解析,写入路径记录当前解析、并拒绝试图覆盖它的请求(点名所有者)。本机确实改了选择时,通道上已绑定的 Session 仍跑在旧端点——这一种情况现在给出类型化的 manager_channel_executor_rebind_required,回复直接写出唯一能修复它的动作,而不是笼统的"管家处理失败"。
具体改动
loopx/chat_manager.py:新增manager_connection_executor_endpoint(runtime_root),复用既有的load_effective_steward_executor_defaults与selected_manager_executor_endpoint,保持"本机文档 →LOOPX_MANAGER_ENDPOINT→ 出货默认"这一既有优先级。loopx/extensions/lark/manager_routing.py:decide_manager_event的路由端点改为该解析结果并附executor_endpoint_source;authorized_manager_goal_ids用同一解析与绑定 Session 比对。loopx/extensions/lark/goal_topic_edit.py:resolve_conversation_policy对 manager 返回机器解析与来源,并拒绝与其不一致的显式请求。loopx/extensions/lark/goal_topic_connections.py/goal_topic_batch.py/chat_lark_api.py:写入路径透传runtime_root,记录executor_endpoint_id+executor_endpoint_source;manager 连接的端点一律由机器解析决定。loopx/extensions/lark/goal_topic_runtime.py:answer_lark_goal_topic对 manager 不再读路由上的旧值;仅"同一受众、执行器不同"的失配给出类型化回执,其它失配(受众/Goal/已关闭)维持原有RuntimeError,避免误报。loopx/extensions/lark/manager_context.py:为新 code 增加用户可见文案。docs/architecture/rfcs/harness-selection-dsh-pi-v0.md/.zh-CN.md:写入该不变量(连接记录是观测值;改选择后需重新应用一次连接)。- 测试:新增 4 例——陈旧记录下路由仍走机器选择、写入记录并拒绝覆盖、授权会话必须与机器端点一致、陈旧会话得到类型化重绑回复。
对主干的风险
已披露的行为变更有两处,且都是本 PR 的目的本身:连接记录不再能压过 machine configuration;试图用连接级 executor_endpoint_id 覆盖机器选择的请求现在报错(invalid_lark_connection,点名所有者),而不是静默写入一个路由会忽略的值。未配置该命名空间的机器解析与行为完全不变(codex / product_default,由既有 shipped-default 测试与 steward channel-binding smoke 守住)。worker(goal)连接不受影响,仍用 agent_id。前端无需改动:可编辑该设置的面已经存在(machine-configuration-settings.tsx 渲染 steward_executor 命名空间字段,capability-localization.ts 中 executor_endpoint 即"管家执行器"),此前只是不生效;因此没有打包资产变动,也不需要 npm run build:chat 对齐重建。
我的整体评价
无阻断性问题,建议合并。这是一次删除重复决定的重构,而不是新增抽象:改动集中在唯一的渠道所有者与其读取路径,量级与问题相称。验证取自被改动的 Lark、manager-channel、handoff、Lark-API 套件(182 passed),以及三个 steward/managed smoke 与 canary premerge --tier standard(8/8 检查、公开边界扫描干净);唯一的失败 test_every_production_steward_caller_passes_the_machine_defaults 在干净 origin/main 上同样失败(已通过 stash 整个 diff 复现),属于既有的过期不变量测试,不在本 PR 范围内。线上证据:在受影响机器上把管家连接重新应用后,通道 Session 由 codex 迁到 dsh,真实 Turn 在 25s 内以 deepseek-v4-flash@high 完成。
残留风险与最强缺口:本 PR 之后,改了机器执行器的机器仍需手工重新应用一次连接才能让已绑定的 Session 跟上(现阶段靠类型化回执提示,而非自动重绑——因为读取路径不能写);另外 test_every_production_steward_caller_passes_the_machine_defaults 仍红在 main,应单独修掉,否则会持续掩盖真正的调用点回归。
English verdict: APPROVE - exact head bdca4e33c. The steward channel's executor now has one owner (the machine configuration); connection records store an observation, all read paths re-resolve through manager_connection_executor_endpoint, a write that tries to override the machine choice fails with the owner named, and a Session stranded on the previous endpoint is refused with the typed manager_channel_executor_rebind_required receipt instead of the generic manager failure. Validated by 182 passing tests in the changed suites plus four new cases, three steward/managed smokes, a clean canary premerge --tier standard run, and a live Turn on the repaired machine; the single suite failure is pre-existing on origin/main.
Motivation
A manager (steward) connection stored the executor endpoint as a decision when it was created. The Lark route, the authorized-connection resolution and the Turn that answers on the channel all read that stored value instead of the machine.
The consequence on a real machine: the operator's
steward_executormachine configuration selected the managed host (dsh),/api/chat/capabilitiesreportedexecutor_endpoint=dsh,source=machine_configuration,available=true— and the steward group still answered oncodex, failing with a host gate because that endpoint belonged to an exhausted personal subscription. The connection write path had the same shape (the stored value outranked the machine), and no surface could correct it: the Dashboard never sendsexecutor_endpoint_id, and the edit API re-read the stored value.Change
loopx/chat_manager.py::manager_connection_executor_endpoint(runtime_root)resolves the machine document, then the service environment, then the shipped default — the same precedence the capability readback already publishes.decide_manager_event, the Lark event decision andanswer_lark_goal_topicre-resolve through it, so a record written under an earlier default cannot decide a later Turn.authorized_manager_goal_idscompares the bound Session against the same resolution.executor_endpoint_idplus the newexecutor_endpoint_source) and refuses a request that asks for a different endpoint, naming the owner, instead of writing a value the route would have to ignore.manager_channel_executor_rebind_requiredreceipt, and the reply names the single repairing action (re-apply the connection) instead of the generic "manager failed" label.Scope and delivery
Backend only. The operator-facing control for this setting already exists and is unchanged: the Dashboard machine-configuration editor renders the
steward_executornamespace fields (machine-configuration-settings.tsx+capability-localization.ts, fieldexecutor_endpoint), with preview, apply and readback. Because the connection record no longer decides, that editor is now sufficient to move the steward channel; no packaged-frontend asset changes, and therefore nonpm run build:chatparity rebuild.Validation
pytest tests/extensions/test_lark_goal_topic_connections.py tests/extensions/test_lark_goal_topic_runtime.py tests/test_manager_channel_binding.py tests/test_manager_context_handoff.py tests/test_chat_lark_api_contract.py→ 182 passed, 1 failed.test_every_production_steward_caller_passes_the_machine_defaults, is pre-existing: it assertsloopx/chat_server.pystill holds a steward resolver call, and that call was extracted earlier. It reproduces identically on a cleanorigin/maincheckout (verified by stashing the whole diff).manager_failure_reply.examples/loopx-steward-channel-binding-smoke.py,examples/loopx-steward-managed-chat-smoke.py,examples/loopx-managed-turn-operator-flow-smoke.pyall pass.loopx canary premerge --from-git-diff --git-diff-base origin/main --tier standard→ 8 checks executed, 0 failures; public boundary scan clean.Live evidence
On the affected machine the same Telegram-free repair path (re-applying the manager connection via the edit API with the machine's endpoint restated) moved the channel Session from
codextodsh, and a real Turn then completed in 25s ondeepseek-v4-flash@high. This PR removes the possibility of that drift returning.Risk
Behaviour change, disclosed: a manager connection can no longer pin a different endpoint than the machine configuration, and such a request now fails with a typed error instead of silently writing a value the route ignores. Every affected connection would today be answering on an endpoint the readback does not report. Unconfigured machines keep resolving exactly as before (shipped default
codex).