feat(steward): make the steward executor a first-class machine setting - #4500
Conversation
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Reviewed exact head 0750380c31df7b56426cf28975d64f374d2aa116 (base f6a6d1139), 28 files, +1481/-86. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers.
动机
管家通道此前只有两层解析:Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / LOOPX_MANAGER_MODEL / LOOPX_MANAGER_REASONING_EFFORT)与出厂默认(codex + 厂商模型)。也就是说,一台机器上「管家跑在哪个执行器、哪个模型、哪个推理强度」这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config inspect 也读不到;一旦有人只改了 plist 或只改了别处,两边就会悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。
更小的修法也评估过并已采用:直接复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),本 PR 只新增一个命名空间,没有新增第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。
改动思路
新增 steward_executor 命名空间,schema steward_executor_machine_defaults_v0,三个字段:executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选了执行器时,模型与强度继续落到下层,不会顺带继承上一次读到的值。
权威状态是 <runtime_root>/machine/configuration.json 这一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,然后只归一化本命名空间,这样兄弟命名空间写坏不会连带改写一个合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default;命名空间非法 → configuration_invalid;store 不可读 → unavailable,各自带修复提示,人正在对话的通道不会因此断掉。回读侧新增两个附加字段 machine_defaults_status 与 machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予任何权限。
调用点统一收口:chat_server 的 capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、extensions/lark/goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取(表单提交不出命名空间会拒绝的值),并补齐 en / zh-CN 文案。
具体改动
生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行,归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余分散在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物与 6 处文档。
关键代码讲解
loopx/chat_manager.py:208_resolve_manager_endpoint—— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路返回,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。probe:机器dsh+ 环境codex→('dsh','machine_configuration');纯环境 →('codex','explicit_config');全无 →('codex','product_default')且原因为steward_channel_default。loopx/capabilities/steward_executor/machine_defaults.py:216effective_steward_executor_defaults—— None、无该命名空间、信封非法都归到absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。loopx/capabilities/steward_executor/machine_defaults.py:240load_effective_steward_executor_defaults—— 原样读文档、只归一化本命名空间;把OSError/TypeError/ValueError转成 typed 投影并附修复提示。loopx/capabilities/configuration_ui.py:146capability_configuration_editor/_steward_executor_editor_options—— 发布的编辑器supported_scopes/writable_scopes均为machine,Goal 不能覆盖;选项取自steward_executor_endpoints()/steward_reasoning_efforts(),未来新增执行器 id 会自动出现在表单里。
验证
本 head 实测:pytest tests/capabilities/test_steward_executor_machine_defaults.py tests/test_manager_channel_binding.py tests/test_chat_machine_configuration_api.py → 44 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿。tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改的基线检出上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。
接口层实测 tests/test_chat_machine_configuration_api.py::test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读值与提交值一致且不含本机路径。
对主干的风险
最强回归场景是「从未配置过的机器」和「没读机器文档的调用方」。前者:store 文件缺失时解析为 absent/capability_default,端点、模型、source 与默认原因与基线逐字段一致,未配置路径是字节级不变的。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,所以「通道总能说清自己在哪个执行器上、为什么」这一承诺在当前树成立,但它列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo 与 benchmark 面均未触碰。回滚即回到 env → 默认,运行时也可用既有 rollback 移除命名空间。默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它是「先于服务环境而非取代服务环境」,且默认仍为 codex。
default_off_isolation 为 isolated:注册表登记、编辑器存在、catalog 出现都只是可发现性,不会激活任何执行器;未使用时投影报 absent,不写任何新状态。authority_semantics 为 aligned:命名空间只记录一台机器上一个操作者的持久选择,不做 turn、不做委派、不授予权限。domain_neutrality 无新增义务文本。change_proportionality 为 proportionate:最小可行修法就是本案(既有 owner 内加一个命名空间 + 回读字段),没有新增命令、没有迁移、没有存储格式升版。
非阻断 finding(P3,已记录在复核结果中,建议后续一刀处理,不应阻塞本次合并):
- P3
loopx/chat_manager.py:348 / :499:未来调用方若漏传machine_defaults,回读会显示not_read与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加一条守护测试枚举生产调用点必须传参,或让管家路径上的该参数必填。 - P3
loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可以存下一个该执行器并不提供的模型,届时以运行时报错暴露。这与既有LOOPX_MANAGER_MODEL行为一致(两层保持同构)。最小修复:在回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。
我的整体评价
基线/head 对比:未配置机器与基线逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意的、已用真实 store 与真实解析函数反证过的变更,不是漂移。代码体量判定为 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,没有投机性框架,也没有为未来预留的未用字段。
结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;这一项已记入 residual risk,不构成阻断。
English verdict: Approved at exact head 0750380c — the steward executor/model/effort become one machine-scoped namespace that the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default, per field, with an unconfigured machine byte-identical to baseline and two recorded non-blocking P3 follow-ups.
The steward channel's executor, model and reasoning effort were selectable only through the Chat service environment, so a machine-local decision lived in a launch file no product surface could show or change. They are now one typed machine-configuration namespace, steward_executor, resolved as machine configuration, then service environment, then the shipped default. The namespace stores a decision and no credential: a configured credential still authenticates the selected executor instead of selecting one, the shipped default stays codex on every machine, and the choice grants no authority. Blank model and effort fields keep resolving from the lower layers, unknown fields and unlisted endpoints fail closed, and a malformed value or sibling namespace falls back with a typed reason instead of failing the surface a person talks to. Every steward entry point reads the same document: the Chat runtime controller owns the runtime root, the capabilities readback quotes the document's status and revision, and the Lark and managed-Turn paths resolve through the same owner. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The machine capability settings page lists the steward executor with the endpoints LoopX ships as channel executors, an optional model, and an optional reasoning effort, in English and Simplified Chinese. The packaged Chat bundle is rebuilt from the same source so the shipped asset matches. The browser acceptance fixture and its machine-settings scenario now carry the namespace end to end: the operator picks the executor in the form, the scenario asserts the exact previewed namespace configuration, and the apply keeps the reviewed plan revision. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The steward executor namespace is documented where its readers already look: the DSH/Pi selection RFC pair gains a dated increment and updates the steward row, the manager evidence protocol names the resolution order, and the interaction-pattern catalog lists the new built-in namespace so IP-030 stays complete. The channel binding smoke gains the machine-default section it now covers. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
0750380 to
69e3130
Compare
…mantic inventory The steward executor wiring adds call-site lines to three modules that sit at their recorded ceilings, so the ratchet's reviewed baseline records the current sizes the same way earlier feature commits did for the modules they touched. The semantic inventory is regenerated for the two new capability modules and their six named constants. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head 5038a6ec4854c5864030d30ec68090b501ae0d33, base origin/main @ 079643038. 29 files, +1493/-88. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers. This head supersedes the earlier card at 0750380c (rebased onto current main; the code diff is byte-identical) and adds the two bookkeeping repairs the premerge gate required.
动机
管家通道此前只能从 Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / _MODEL / _REASONING_EFFORT)或出厂默认(codex + 厂商模型)解析「跑在哪个执行器、哪个模型、哪个推理强度」。也就是说这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config describe / inspect 也读不到,两边还容易悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。
更小的修法评估过并已采用:复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),只新增一个命名空间,没有第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。
改动思路
新增 steward_executor 命名空间(schema steward_executor_machine_defaults_v0):executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选执行器时,模型与强度继续落到下层,不会继承上一次读到的值。
权威状态是 <runtime_root>/machine/configuration.json 一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,只归一化本命名空间,兄弟命名空间写坏不会连带改写合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default,命名空间非法 → configuration_invalid,store 不可读 → unavailable,各带修复提示,人正在对话的通道不会断。回读新增 machine_defaults_status / machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予权限。
调用点统一收口:chat_server capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取,并补齐 en / zh-CN 文案。
具体改动
生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行:归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物、6 处文档,以及 2 个仓库记账文件。
关键代码讲解
loopx/chat_manager.py:208_resolve_manager_endpoint—— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。实测:机器dsh+ 环境codex→('dsh','machine_configuration');纯环境 →('codex','explicit_config');全无 →('codex','product_default')且原因steward_channel_default。loopx/capabilities/steward_executor/machine_defaults.py:216effective_steward_executor_defaults—— None、无该命名空间、信封非法都归到absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。loopx/capabilities/steward_executor/machine_defaults.py:240load_effective_steward_executor_defaults—— 原样读文档、只归一化本命名空间,把OSError/TypeError/ValueError转成 typed 投影并附修复提示。loopx/capabilities/configuration_ui.py:146capability_configuration_editor/_steward_executor_editor_options—— 发布的编辑器supported_scopes/writable_scopes均为machine,Goal 不能覆盖;选项取自steward_executor_endpoints()/steward_reasoning_efforts()。
验证
本 head:pytest(steward 命名空间 + manager channel binding + chat machine-configuration API + capability configuration UI + periodic report machine store)→ 74 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿;tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改基线上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。
接口层实测 test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读与提交一致且不含本机路径。
仓库 premerge 门禁 loopx canary premerge --from-git-diff --tier standard:passed,merge_gate_passed=true,self_merge_allowed=true,19/19 选中断言 0 失败。它在本 head 之前报出两个真实失败,均已修复而非绕过:control_plane-maintainability-ratchet-smoke 指出三个被接线模块越过记录上限(chat_runtime 1502/1500、chat_server 1514/1513、goal_topic_runtime 1509/1500),按仓库既有做法把上限刷新到当前尺寸;semantic-vocabulary-drift-smoke 指出语义清单过期,重新生成(source_files 1181→1183,named_string_constants 2067→2075,与新增 2 个模块、8 个常量一致)。
对主干的风险
最强回归场景是「从未配置过的机器」与「没读机器文档的调用方」。前者:store 缺失时解析为 absent/capability_default,端点、模型、source 与默认原因逐字段与基线一致(基线 = 当前 main,本 PR 已在 rebase 后重跑)。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,但列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo、benchmark 面均未触碰。回滚即回到 env → 默认;运行时也可用既有 rollback 移除命名空间。
默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它「先于服务环境而非取代服务环境」,默认仍为 codex。default_off_isolation = isolated(注册表登记、编辑器存在、catalog 出现只是可发现性,不激活任何执行器;未使用时投影报 absent,不写新状态)。authority_semantics = aligned(只记录一台机器上一个操作者的持久选择,不做 turn、不委派、不授权)。domain_neutrality 无新增义务文本。change_proportionality = proportionate(最小可行修法就是本案,无新增命令、无迁移、无存储格式升版)。
非阻断 finding(P3,已记录在复核结果,建议后续单独一刀,不阻塞本次合并):
- P3
loopx/chat_manager.py:348 / :499:未来调用方若漏传machine_defaults,回读会显示not_read与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加守护测试枚举生产调用点必须传参,或让管家路径上该参数必填。 - P3
loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可存下一个该执行器并不提供的模型,届时以运行时报错暴露;与既有LOOPX_MANAGER_MODEL行为一致。最小修复:回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。
我的整体评价
基线/head 对比:未配置机器与当前 main 逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意变更,已用真实 store 与真实解析函数反证。代码体量判定 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,无投机性框架、无未来预留字段;两个记账文件是仓库自身要求的可见成本,不是隐藏。
结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;已记入 residual risk,不构成阻断。
English verdict: Approved at exact head 5038a6ec4 — the steward executor/model/effort become one machine-scoped namespace the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default, per field; an unconfigured machine is unchanged; the 19-check repository premerge gate passes at this head after the ratchet baseline and semantic inventory were repaired, and two non-blocking P3 follow-ups are recorded.
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
This card supersedes the two earlier cards at 0750380c and the prior 5038a6ec4 card: same exact-head evidence and same verdict, corrected so the English verdict line is machine-readable for the merge-readiness conclusion parser.
Exact head 5038a6ec4854c5864030d30ec68090b501ae0d33, base origin/main @ 079643038. 29 files, +1493/-88. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers. This head supersedes the earlier card at 0750380c (rebased onto current main; the code diff is byte-identical) and adds the two bookkeeping repairs the premerge gate required.
动机
管家通道此前只能从 Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / _MODEL / _REASONING_EFFORT)或出厂默认(codex + 厂商模型)解析「跑在哪个执行器、哪个模型、哪个推理强度」。也就是说这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config describe / inspect 也读不到,两边还容易悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。
更小的修法评估过并已采用:复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),只新增一个命名空间,没有第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。
改动思路
新增 steward_executor 命名空间(schema steward_executor_machine_defaults_v0):executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选执行器时,模型与强度继续落到下层,不会继承上一次读到的值。
权威状态是 <runtime_root>/machine/configuration.json 一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,只归一化本命名空间,兄弟命名空间写坏不会连带改写合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default,命名空间非法 → configuration_invalid,store 不可读 → unavailable,各带修复提示,人正在对话的通道不会断。回读新增 machine_defaults_status / machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予权限。
调用点统一收口:chat_server capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取,并补齐 en / zh-CN 文案。
具体改动
生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行:归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物、6 处文档,以及 2 个仓库记账文件。
关键代码讲解
loopx/chat_manager.py:208_resolve_manager_endpoint—— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。实测:机器dsh+ 环境codex→('dsh','machine_configuration');纯环境 →('codex','explicit_config');全无 →('codex','product_default')且原因steward_channel_default。loopx/capabilities/steward_executor/machine_defaults.py:216effective_steward_executor_defaults—— None、无该命名空间、信封非法都归到absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。loopx/capabilities/steward_executor/machine_defaults.py:240load_effective_steward_executor_defaults—— 原样读文档、只归一化本命名空间,把OSError/TypeError/ValueError转成 typed 投影并附修复提示。loopx/capabilities/configuration_ui.py:146capability_configuration_editor/_steward_executor_editor_options—— 发布的编辑器supported_scopes/writable_scopes均为machine,Goal 不能覆盖;选项取自steward_executor_endpoints()/steward_reasoning_efforts()。
验证
本 head:pytest(steward 命名空间 + manager channel binding + chat machine-configuration API + capability configuration UI + periodic report machine store)→ 74 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿;tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改基线上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。
接口层实测 test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读与提交一致且不含本机路径。
仓库 premerge 门禁 loopx canary premerge --from-git-diff --tier standard:passed,merge_gate_passed=true,self_merge_allowed=true,19/19 选中断言 0 失败。它在本 head 之前报出两个真实失败,均已修复而非绕过:control_plane-maintainability-ratchet-smoke 指出三个被接线模块越过记录上限(chat_runtime 1502/1500、chat_server 1514/1513、goal_topic_runtime 1509/1500),按仓库既有做法把上限刷新到当前尺寸;semantic-vocabulary-drift-smoke 指出语义清单过期,重新生成(source_files 1181→1183,named_string_constants 2067→2075,与新增 2 个模块、8 个常量一致)。
对主干的风险
最强回归场景是「从未配置过的机器」与「没读机器文档的调用方」。前者:store 缺失时解析为 absent/capability_default,端点、模型、source 与默认原因逐字段与基线一致(基线 = 当前 main,本 PR 已在 rebase 后重跑)。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,但列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo、benchmark 面均未触碰。回滚即回到 env → 默认;运行时也可用既有 rollback 移除命名空间。
默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它「先于服务环境而非取代服务环境」,默认仍为 codex。default_off_isolation = isolated(注册表登记、编辑器存在、catalog 出现只是可发现性,不激活任何执行器;未使用时投影报 absent,不写新状态)。authority_semantics = aligned(只记录一台机器上一个操作者的持久选择,不做 turn、不委派、不授权)。domain_neutrality 无新增义务文本。change_proportionality = proportionate(最小可行修法就是本案,无新增命令、无迁移、无存储格式升版)。
非阻断 finding(P3,已记录在复核结果,建议后续单独一刀,不阻塞本次合并):
- P3
loopx/chat_manager.py:348 / :499:未来调用方若漏传machine_defaults,回读会显示not_read与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加守护测试枚举生产调用点必须传参,或让管家路径上该参数必填。 - P3
loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可存下一个该执行器并不提供的模型,届时以运行时报错暴露;与既有LOOPX_MANAGER_MODEL行为一致。最小修复:回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。
我的整体评价
基线/head 对比:未配置机器与当前 main 逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意变更,已用真实 store 与真实解析函数反证。代码体量判定 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,无投机性框架、无未来预留字段;两个记账文件是仓库自身要求的可见成本,不是隐藏。
结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;已记入 residual risk,不构成阻断。
English verdict: APPROVE — one machine-scoped steward_executor namespace the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default per field; an unconfigured machine is unchanged; the 19-check premerge gate passes at this head; two non-blocking P3 follow-ups are recorded.
What changes
The steward channel's executor was selectable only through the Chat service
environment (
LOOPX_MANAGER_ENDPOINTand its model/effort siblings), so amachine-local decision lived in a launch file that no product surface could show
or change. This PR makes it a first-class machine setting.
A new typed machine-configuration namespace,
steward_executor, holds exactlythree fields:
{ "schema_version": "steward_executor_machine_defaults_v0", "executor_endpoint": "codex", "executor_model": null, "executor_reasoning_effort": null }loopx/chat_manager.py): machineconfiguration, then the service environment, then the shipped product default.
Each field is decided on its own, so a machine that selects only an executor
keeps resolving its model and effort from the lower layers.
capability_configuration_editorgains asteward_executordefinition (machine-only scope) whose option lists come fromthe owning namespace, plus English and Simplified Chinese copy. The packaged
Chat bundle is rebuilt from the same source.
loopx machine-config describepublishes thetemplate and documentation,
inspectreturns the stored revision, and thechannel binding adds
executor_endpoint_source: machine_configurationwith thedocument's
statusandconfiguration_revision.steward_executor_defaults()from the runtime root it already owns, so thecapabilities readback, the Codex App Chat adapter launch, the Lark manager
conversation, the Lark Topic reply, and the manager Turn context all resolve
from the same document instead of re-deriving the rule.
What it does not change
codexon every machine, and a configured credentialstill authenticates the selected executor without selecting one.
still needs its own credential and runtime, and
manager_runtimeremains aseparate machine decision.
accepts only the endpoints LoopX ships as channel executors), and an
unsupported reasoning effort fail closed before any effect. An operator who
needs an unlisted adapter still has
LOOPX_MANAGER_ENDPOINT.namespace resolves to the lower layers with a typed reason
(
configuration_invalid,unavailable) instead of failing the surface aperson talks to.
Validation
tests/capabilities/test_steward_executor_machine_defaults.py(new):registry/catalog shape, fail-closed normalization, blank-means-inherit, live
store round-trip, malformed value and malformed sibling fallback.
tests/test_manager_channel_binding.py: machine layer outranks theenvironment, per-field independence, the live controller read, and the
readback's status/revision.
tests/test_chat_machine_configuration_api.py: the namespace is editable andread back through the revision-locked API, and the machine capability catalog
lists it as machine-only.
tests/capabilities/test_capability_configuration_ui.py,tests/capabilities/test_periodic_report_machine_store.py: editor contract andthe built-in namespace inventory.
examples/loopx-steward-channel-binding-smoke.pygains the machine-defaultsection;
examples/interaction-pattern-catalog-smoke.pypasses with the newnamespace documented in IP-030.
examples/personal-workspace-browser-smoke.mjs(all fourscenarios, development mode) passes with new steps that select the steward
executor, assert the previewed namespace configuration, and apply the reviewed
plan revision.
tests -k "chat or manager or machine_configuration or steward": 468 passed.tests/capabilities: 1431 passed / 21 failures, all pre-existing in theenvironment on unmodified
main(benchmark/toolkit CLI tests that readinstalled package data), verified by running the same node ids on a clean tree.
ruff checkclean;scripts/generate_semantic_inventory.pyregenerated;packaged bundle rebuild is byte-reproducible (a fresh build of unmodified
origin/mainproduces no diff).Boundaries
No credential, local path, raw log, or private context appears in the diff,
tests, or docs. The namespace is machine-scoped: no Goal can override it.