Skip to content

feat(steward): make the steward executor a first-class machine setting - #4500

Merged
huangruiteng merged 4 commits into
mainfrom
codex/steward-executor-machine-config-20260916
Sep 16, 2026
Merged

huangruiteng merged 4 commits into
mainfrom
codex/steward-executor-machine-config-20260916

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

What changes

The steward channel's executor was selectable only through the Chat service
environment (LOOPX_MANAGER_ENDPOINT and its model/effort siblings), so a
machine-local decision lived in a launch file that no product surface could show
or change. This PR makes it a first-class machine setting.

A new typed machine-configuration namespace, steward_executor, holds exactly
three fields:

{
  "schema_version": "steward_executor_machine_defaults_v0",
  "executor_endpoint": "codex",
  "executor_model": null,
  "executor_reasoning_effort": null
}
  • Resolution order, stated once (loopx/chat_manager.py): machine
    configuration, then the service environment, then the shipped product default.
    Each field is decided on its own, so a machine that selects only an executor
    keeps resolving its model and effort from the lower layers.
  • Authored in the Dashboard: capability_configuration_editor gains a
    steward_executor definition (machine-only scope) whose option lists come from
    the owning namespace, plus English and Simplified Chinese copy. The packaged
    Chat bundle is rebuilt from the same source.
  • Read back everywhere: loopx machine-config describe publishes the
    template and documentation, inspect returns the stored revision, and the
    channel binding adds executor_endpoint_source: machine_configuration with the
    document's status and configuration_revision.
  • One owner for every entry point: the Chat runtime controller exposes
    steward_executor_defaults() from the runtime root it already owns, so the
    capabilities readback, the Codex App Chat adapter launch, the Lark manager
    conversation, the Lark Topic reply, and the manager Turn context all resolve
    from the same document instead of re-deriving the rule.

What it does not change

  • The shipped default stays codex on every machine, and a configured credential
    still authenticates the selected executor without selecting one.
  • The namespace stores no credential and grants no authority; the managed host
    still needs its own credential and runtime, and manager_runtime remains a
    separate machine decision.
  • Unknown fields, an unknown schema version, an unlisted endpoint (the namespace
    accepts only the endpoints LoopX ships as channel executors), and an
    unsupported reasoning effort fail closed before any effect. An operator who
    needs an unlisted adapter still has LOOPX_MANAGER_ENDPOINT.
  • A malformed namespace value, an unreadable store, or a malformed sibling
    namespace resolves to the lower layers with a typed reason
    (configuration_invalid, unavailable) instead of failing the surface a
    person talks to.

Validation

  • tests/capabilities/test_steward_executor_machine_defaults.py (new):
    registry/catalog shape, fail-closed normalization, blank-means-inherit, live
    store round-trip, malformed value and malformed sibling fallback.
  • tests/test_manager_channel_binding.py: machine layer outranks the
    environment, per-field independence, the live controller read, and the
    readback's status/revision.
  • tests/test_chat_machine_configuration_api.py: the namespace is editable and
    read back through the revision-locked API, and the machine capability catalog
    lists it as machine-only.
  • tests/capabilities/test_capability_configuration_ui.py,
    tests/capabilities/test_periodic_report_machine_store.py: editor contract and
    the built-in namespace inventory.
  • examples/loopx-steward-channel-binding-smoke.py gains the machine-default
    section; examples/interaction-pattern-catalog-smoke.py passes with the new
    namespace documented in IP-030.
  • Browser acceptance: examples/personal-workspace-browser-smoke.mjs (all four
    scenarios, development mode) passes with new steps that select the steward
    executor, assert the previewed namespace configuration, and apply the reviewed
    plan revision.
  • tests -k "chat or manager or machine_configuration or steward": 468 passed.
  • tests/capabilities: 1431 passed / 21 failures, all pre-existing in the
    environment on unmodified main (benchmark/toolkit CLI tests that read
    installed package data), verified by running the same node ids on a clean tree.
  • ruff check clean; scripts/generate_semantic_inventory.py regenerated;
    packaged bundle rebuild is byte-reproducible (a fresh build of unmodified
    origin/main produces no diff).

Boundaries

No credential, local path, raw log, or private context appears in the diff,
tests, or docs. The namespace is machine-scoped: no Goal can override it.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head 0750380c31df7b56426cf28975d64f374d2aa116 (base f6a6d1139), 28 files, +1481/-86. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers.

动机

管家通道此前只有两层解析:Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / LOOPX_MANAGER_MODEL / LOOPX_MANAGER_REASONING_EFFORT)与出厂默认(codex + 厂商模型)。也就是说,一台机器上「管家跑在哪个执行器、哪个模型、哪个推理强度」这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config inspect 也读不到;一旦有人只改了 plist 或只改了别处,两边就会悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。

更小的修法也评估过并已采用:直接复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),本 PR 只新增一个命名空间,没有新增第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。

改动思路

新增 steward_executor 命名空间,schema steward_executor_machine_defaults_v0,三个字段:executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选了执行器时,模型与强度继续落到下层,不会顺带继承上一次读到的值。

权威状态是 <runtime_root>/machine/configuration.json 这一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,然后只归一化本命名空间,这样兄弟命名空间写坏不会连带改写一个合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default;命名空间非法 → configuration_invalid;store 不可读 → unavailable,各自带修复提示,人正在对话的通道不会因此断掉。回读侧新增两个附加字段 machine_defaults_status 与 machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予任何权限。

调用点统一收口:chat_server 的 capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、extensions/lark/goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取(表单提交不出命名空间会拒绝的值),并补齐 en / zh-CN 文案。

具体改动

生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行,归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余分散在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物与 6 处文档。

关键代码讲解

  1. loopx/chat_manager.py:208 _resolve_manager_endpoint —— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路返回,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。probe:机器 dsh + 环境 codex → ('dsh','machine_configuration');纯环境 → ('codex','explicit_config');全无 → ('codex','product_default') 且原因为 steward_channel_default。
  2. loopx/capabilities/steward_executor/machine_defaults.py:216 effective_steward_executor_defaults —— None、无该命名空间、信封非法都归到 absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。
  3. loopx/capabilities/steward_executor/machine_defaults.py:240 load_effective_steward_executor_defaults —— 原样读文档、只归一化本命名空间;把 OSError/TypeError/ValueError 转成 typed 投影并附修复提示。
  4. loopx/capabilities/configuration_ui.py:146 capability_configuration_editor / _steward_executor_editor_options —— 发布的编辑器 supported_scopes/writable_scopes 均为 machine,Goal 不能覆盖;选项取自 steward_executor_endpoints() / steward_reasoning_efforts(),未来新增执行器 id 会自动出现在表单里。

验证

本 head 实测:pytest tests/capabilities/test_steward_executor_machine_defaults.py tests/test_manager_channel_binding.py tests/test_chat_machine_configuration_api.py → 44 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿。tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改的基线检出上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。

接口层实测 tests/test_chat_machine_configuration_api.py::test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读值与提交值一致且不含本机路径。

对主干的风险

最强回归场景是「从未配置过的机器」和「没读机器文档的调用方」。前者:store 文件缺失时解析为 absent/capability_default,端点、模型、source 与默认原因与基线逐字段一致,未配置路径是字节级不变的。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,所以「通道总能说清自己在哪个执行器上、为什么」这一承诺在当前树成立,但它列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo 与 benchmark 面均未触碰。回滚即回到 env → 默认,运行时也可用既有 rollback 移除命名空间。默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它是「先于服务环境而非取代服务环境」,且默认仍为 codex。

default_off_isolation 为 isolated:注册表登记、编辑器存在、catalog 出现都只是可发现性,不会激活任何执行器;未使用时投影报 absent,不写任何新状态。authority_semantics 为 aligned:命名空间只记录一台机器上一个操作者的持久选择,不做 turn、不做委派、不授予权限。domain_neutrality 无新增义务文本。change_proportionality 为 proportionate:最小可行修法就是本案(既有 owner 内加一个命名空间 + 回读字段),没有新增命令、没有迁移、没有存储格式升版。

非阻断 finding(P3,已记录在复核结果中,建议后续一刀处理,不应阻塞本次合并):

  • P3 loopx/chat_manager.py:348 / :499:未来调用方若漏传 machine_defaults,回读会显示 not_read 与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加一条守护测试枚举生产调用点必须传参,或让管家路径上的该参数必填。
  • P3 loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可以存下一个该执行器并不提供的模型,届时以运行时报错暴露。这与既有 LOOPX_MANAGER_MODEL 行为一致(两层保持同构)。最小修复:在回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。

我的整体评价

基线/head 对比:未配置机器与基线逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意的、已用真实 store 与真实解析函数反证过的变更,不是漂移。代码体量判定为 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,没有投机性框架,也没有为未来预留的未用字段。

结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;这一项已记入 residual risk,不构成阻断。

English verdict: Approved at exact head 0750380c — the steward executor/model/effort become one machine-scoped namespace that the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default, per field, with an unconfigured machine byte-identical to baseline and two recorded non-blocking P3 follow-ups.

The steward channel's executor, model and reasoning effort were selectable
only through the Chat service environment, so a machine-local decision lived
in a launch file no product surface could show or change. They are now one
typed machine-configuration namespace, steward_executor, resolved as machine
configuration, then service environment, then the shipped default.

The namespace stores a decision and no credential: a configured credential
still authenticates the selected executor instead of selecting one, the
shipped default stays codex on every machine, and the choice grants no
authority. Blank model and effort fields keep resolving from the lower
layers, unknown fields and unlisted endpoints fail closed, and a malformed
value or sibling namespace falls back with a typed reason instead of failing
the surface a person talks to.

Every steward entry point reads the same document: the Chat runtime
controller owns the runtime root, the capabilities readback quotes the
document's status and revision, and the Lark and managed-Turn paths resolve
through the same owner.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The machine capability settings page lists the steward executor with the
endpoints LoopX ships as channel executors, an optional model, and an
optional reasoning effort, in English and Simplified Chinese. The packaged
Chat bundle is rebuilt from the same source so the shipped asset matches.

The browser acceptance fixture and its machine-settings scenario now carry
the namespace end to end: the operator picks the executor in the form, the
scenario asserts the exact previewed namespace configuration, and the apply
keeps the reviewed plan revision.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The steward executor namespace is documented where its readers already look:
the DSH/Pi selection RFC pair gains a dated increment and updates the steward
row, the manager evidence protocol names the resolution order, and the
interaction-pattern catalog lists the new built-in namespace so IP-030 stays
complete. The channel binding smoke gains the machine-default section it now
covers.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng
huangruiteng force-pushed the codex/steward-executor-machine-config-20260916 branch from 0750380 to 69e3130 Compare September 16, 2026 05:01
…mantic inventory

The steward executor wiring adds call-site lines to three modules that sit at
their recorded ceilings, so the ratchet's reviewed baseline records the current
sizes the same way earlier feature commits did for the modules they touched.
The semantic inventory is regenerated for the two new capability modules and
their six named constants.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head 5038a6ec4854c5864030d30ec68090b501ae0d33, base origin/main @ 079643038. 29 files, +1493/-88. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers. This head supersedes the earlier card at 0750380c (rebased onto current main; the code diff is byte-identical) and adds the two bookkeeping repairs the premerge gate required.

动机

管家通道此前只能从 Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / _MODEL / _REASONING_EFFORT)或出厂默认(codex + 厂商模型)解析「跑在哪个执行器、哪个模型、哪个推理强度」。也就是说这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config describe / inspect 也读不到,两边还容易悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。

更小的修法评估过并已采用:复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),只新增一个命名空间,没有第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。

改动思路

新增 steward_executor 命名空间(schema steward_executor_machine_defaults_v0):executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选执行器时,模型与强度继续落到下层,不会继承上一次读到的值。

权威状态是 <runtime_root>/machine/configuration.json 一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,只归一化本命名空间,兄弟命名空间写坏不会连带改写合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default,命名空间非法 → configuration_invalid,store 不可读 → unavailable,各带修复提示,人正在对话的通道不会断。回读新增 machine_defaults_status / machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予权限。

调用点统一收口:chat_server capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取,并补齐 en / zh-CN 文案。

具体改动

生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行:归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物、6 处文档,以及 2 个仓库记账文件。

关键代码讲解

  1. loopx/chat_manager.py:208 _resolve_manager_endpoint —— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。实测:机器 dsh + 环境 codex → ('dsh','machine_configuration');纯环境 → ('codex','explicit_config');全无 → ('codex','product_default') 且原因 steward_channel_default。
  2. loopx/capabilities/steward_executor/machine_defaults.py:216 effective_steward_executor_defaults —— None、无该命名空间、信封非法都归到 absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。
  3. loopx/capabilities/steward_executor/machine_defaults.py:240 load_effective_steward_executor_defaults —— 原样读文档、只归一化本命名空间,把 OSError/TypeError/ValueError 转成 typed 投影并附修复提示。
  4. loopx/capabilities/configuration_ui.py:146 capability_configuration_editor / _steward_executor_editor_options —— 发布的编辑器 supported_scopes/writable_scopes 均为 machine,Goal 不能覆盖;选项取自 steward_executor_endpoints() / steward_reasoning_efforts()。

验证

本 head:pytest(steward 命名空间 + manager channel binding + chat machine-configuration API + capability configuration UI + periodic report machine store)→ 74 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿;tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改基线上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。

接口层实测 test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读与提交一致且不含本机路径。

仓库 premerge 门禁 loopx canary premerge --from-git-diff --tier standard:passed,merge_gate_passed=true,self_merge_allowed=true,19/19 选中断言 0 失败。它在本 head 之前报出两个真实失败,均已修复而非绕过:control_plane-maintainability-ratchet-smoke 指出三个被接线模块越过记录上限(chat_runtime 1502/1500、chat_server 1514/1513、goal_topic_runtime 1509/1500),按仓库既有做法把上限刷新到当前尺寸;semantic-vocabulary-drift-smoke 指出语义清单过期,重新生成(source_files 1181→1183,named_string_constants 2067→2075,与新增 2 个模块、8 个常量一致)。

对主干的风险

最强回归场景是「从未配置过的机器」与「没读机器文档的调用方」。前者:store 缺失时解析为 absent/capability_default,端点、模型、source 与默认原因逐字段与基线一致(基线 = 当前 main,本 PR 已在 rebase 后重跑)。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,但列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo、benchmark 面均未触碰。回滚即回到 env → 默认;运行时也可用既有 rollback 移除命名空间。

默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它「先于服务环境而非取代服务环境」,默认仍为 codex。default_off_isolation = isolated(注册表登记、编辑器存在、catalog 出现只是可发现性,不激活任何执行器;未使用时投影报 absent,不写新状态)。authority_semantics = aligned(只记录一台机器上一个操作者的持久选择,不做 turn、不委派、不授权)。domain_neutrality 无新增义务文本。change_proportionality = proportionate(最小可行修法就是本案,无新增命令、无迁移、无存储格式升版)。

非阻断 finding(P3,已记录在复核结果,建议后续单独一刀,不阻塞本次合并):

  • P3 loopx/chat_manager.py:348 / :499:未来调用方若漏传 machine_defaults,回读会显示 not_read 与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加守护测试枚举生产调用点必须传参,或让管家路径上该参数必填。
  • P3 loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可存下一个该执行器并不提供的模型,届时以运行时报错暴露;与既有 LOOPX_MANAGER_MODEL 行为一致。最小修复:回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。

我的整体评价

基线/head 对比:未配置机器与当前 main 逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意变更,已用真实 store 与真实解析函数反证。代码体量判定 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,无投机性框架、无未来预留字段;两个记账文件是仓库自身要求的可见成本,不是隐藏。

结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;已记入 residual risk,不构成阻断。

English verdict: Approved at exact head 5038a6ec4 — the steward executor/model/effort become one machine-scoped namespace the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default, per field; an unconfigured machine is unchanged; the 19-check repository premerge gate passes at this head after the ratchet baseline and semantic inventory were repaired, and two non-blocking P3 follow-ups are recorded.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

This card supersedes the two earlier cards at 0750380c and the prior 5038a6ec4 card: same exact-head evidence and same verdict, corrected so the English verdict line is machine-readable for the merge-readiness conclusion parser.
Exact head 5038a6ec4854c5864030d30ec68090b501ae0d33, base origin/main @ 079643038. 29 files, +1493/-88. Evidence plan executed from pull_request_review_execution_contract policy revision 5; pr-review --check-result returns ok=true, approval_consistent=true, no blockers. This head supersedes the earlier card at 0750380c (rebased onto current main; the code diff is byte-identical) and adds the two bookkeeping repairs the premerge gate required.

动机

管家通道此前只能从 Chat 服务环境变量(LOOPX_MANAGER_ENDPOINT / _MODEL / _REASONING_EFFORT)或出厂默认(codex + 厂商模型)解析「跑在哪个执行器、哪个模型、哪个推理强度」。也就是说这个决定只能写在 LaunchAgent 的 plist 里:产品前端看不到、改不了,machine-config describe / inspect 也读不到,两边还容易悄悄不一致。受影响的是运行本机 Chat 服务的人,以及所有解析该通道的入口(capabilities 回读、chat_runtime 的 codex/dsh 段构造、manager_turn_context 的 model_defaults、飞书管家路由)。

更小的修法评估过并已采用:复用既有 machine_configuration 能力(命名空间注册表、revision-locked 事务、公开 catalog、Dashboard 编辑器契约),只新增一个命名空间,没有第二套存储或第二套机制。没有更小的做法能同时满足「前端可改 + describe/inspect 可回读」。

改动思路

新增 steward_executor 命名空间(schema steward_executor_machine_defaults_v0):executor_endpoint(必填,取值来自 chat 层已发布的 MANAGER_ENDPOINT_KINDS)、executor_model(可空文本)、executor_reasoning_effort(可空,取值来自 MANAGER_REASONING_EFFORTS)。优先级在一处声明:machine_configuration > 服务环境 > 出厂默认,且按字段独立——机器只选执行器时,模型与强度继续落到下层,不会继承上一次读到的值。

权威状态是 <runtime_root>/machine/configuration.json 一个信封,决定权在 loopx/chat_manager.py 的 _resolve_manager_endpoint / manager_model_resolution / manager_model_config。读路径刻意分两层:store 原样读,只归一化本命名空间,兄弟命名空间写坏不会连带改写合法的管家选择。失败语义是 typed 而非抛错:store 缺失 → absent/capability_default,命名空间非法 → configuration_invalid,store 不可读 → unavailable,各带修复提示,人正在对话的通道不会断。回读新增 machine_defaults_status / machine_defaults_revision,让读通道的人能区分「机器决定」与「环境变量决定」。命名空间只存决定、不存发现:不含凭据、不读凭据、不授予权限。

调用点统一收口:chat_server capabilities、chat_runtime 的 codex 与 dsh 两条适配器构造、chat_lark_api、goal_topic_runtime、chat_manager_context 都经 steward_machine_defaults / steward_executor_defaults 读同一份文档。前端侧在 configuration_ui.py 增加 machine-only 编辑器定义,选项从命名空间自身读取,并补齐 en / zh-CN 文案。

具体改动

生产运行时代码约 500 行,最大热点是新增的 loopx/capabilities/steward_executor/machine_defaults.py(285 行:归一化 / 命名空间注册 / 投影 / 加载器同处一个内聚模块),其余在 chat_manager.py(+161/-14)、configuration_ui.py(+54)、6 个调用点、5 处测试、打包产物、6 处文档,以及 2 个仓库记账文件。

关键代码讲解

  1. loopx/chat_manager.py:208 _resolve_manager_endpoint —— 唯一优先级规则的所有者。机器值命中后在读环境变量之前短路,并把「出厂默认原因」置空,避免把机器决定报成被发现的默认值。实测:机器 dsh + 环境 codex → ('dsh','machine_configuration');纯环境 → ('codex','explicit_config');全无 → ('codex','product_default') 且原因 steward_channel_default。
  2. loopx/capabilities/steward_executor/machine_defaults.py:216 effective_steward_executor_defaults —— None、无该命名空间、信封非法都归到 absent/capability_default,不猜执行器;存在则必须满足闭集与强度词表。
  3. loopx/capabilities/steward_executor/machine_defaults.py:240 load_effective_steward_executor_defaults —— 原样读文档、只归一化本命名空间,把 OSError/TypeError/ValueError 转成 typed 投影并附修复提示。
  4. loopx/capabilities/configuration_ui.py:146 capability_configuration_editor / _steward_executor_editor_options —— 发布的编辑器 supported_scopes/writable_scopes 均为 machine,Goal 不能覆盖;选项取自 steward_executor_endpoints() / steward_reasoning_efforts()。

验证

本 head:pytest(steward 命名空间 + manager channel binding + chat machine-configuration API + capability configuration UI + periodic report machine store)→ 74 passed;pytest tests -k "chat or manager or machine_configuration or steward" → 468 passed;examples/loopx-steward-channel-binding-smoke.py 通过;打包前端在干净基线上重跑 build:chat 与提交产物 零差异;浏览器验收 4 个场景在 development 模式全绿;tests/capabilities 全量 1431 passed / 21 failed,这 21 个 node id 在未修改基线上以完全相同方式失败(benchmark / toolkit CLI 读已安装包数据),非本 head 回归。

接口层实测 test_the_steward_executor_namespace_is_editable_and_read_back:preview → apply → inspect 往返成立,编辑器选项为 ['codex','dsh'],回读与提交一致且不含本机路径。

仓库 premerge 门禁 loopx canary premerge --from-git-diff --tier standard:passed,merge_gate_passed=true,self_merge_allowed=true,19/19 选中断言 0 失败。它在本 head 之前报出两个真实失败,均已修复而非绕过:control_plane-maintainability-ratchet-smoke 指出三个被接线模块越过记录上限(chat_runtime 1502/1500、chat_server 1514/1513、goal_topic_runtime 1509/1500),按仓库既有做法把上限刷新到当前尺寸;semantic-vocabulary-drift-smoke 指出语义清单过期,重新生成(source_files 1181→1183,named_string_constants 2067→2075,与新增 2 个模块、8 个常量一致)。

对主干的风险

最强回归场景是「从未配置过的机器」与「没读机器文档的调用方」。前者:store 缺失时解析为 absent/capability_default,端点、模型、source 与默认原因逐字段与基线一致(基线 = 当前 main,本 PR 已在 rebase 后重跑)。后者:manager_channel_binding / manager_model_config 的 machine_defaults 默认 None,会得到 machine_defaults_status=not_read;我逐一搜索了测试之外的全部生产调用点,没有任何一个漏传,但列为非阻断 finding,因为未来新调用点可能重新引入该缺口。爆炸半径仅限管家通道:托管 Goal 执行、飞书投递、配额、todo、benchmark 面均未触碰。回滚即回到 env → 默认;运行时也可用既有 rollback 移除命名空间。

默认行为变化已披露:RFC 中英双版、manager evidence 协议、interaction-pattern catalog、命名空间描述与编辑器文案均注明它「先于服务环境而非取代服务环境」,默认仍为 codex。default_off_isolation = isolated(注册表登记、编辑器存在、catalog 出现只是可发现性,不激活任何执行器;未使用时投影报 absent,不写新状态)。authority_semantics = aligned(只记录一台机器上一个操作者的持久选择,不做 turn、不委派、不授权)。domain_neutrality 无新增义务文本。change_proportionality = proportionate(最小可行修法就是本案,无新增命令、无迁移、无存储格式升版)。

非阻断 finding(P3,已记录在复核结果,建议后续单独一刀,不阻塞本次合并):

  • P3 loopx/chat_manager.py:348 / :499:未来调用方若漏传 machine_defaults,回读会显示 not_read 与出厂默认,而机器实际选了别的;属回读低报而非误执行。最小修复:加守护测试枚举生产调用点必须传参,或让管家路径上该参数必填。
  • P3 loopx/capabilities/steward_executor/machine_defaults.py:74:只校验执行器与强度词表,不校验执行器/模型配对,机器可存下一个该执行器并不提供的模型,届时以运行时报错暴露;与既有 LOOPX_MANAGER_MODEL 行为一致。最小修复:回读中一并给出模型/强度以便看出不匹配,或在该配对已可用时校验。

我的整体评价

基线/head 对比:未配置机器与当前 main 逐字段等价(仅多一个 not_read/absent 状态位);已配置机器端点与模型 source 变为 machine_configuration,这是刻意变更,已用真实 store 与真实解析函数反证。代码体量判定 necessary,贴合度经验证:每个新增模块、字段与编辑器条目都有活跃调用点,无投机性框架、无未来预留字段;两个记账文件是仓库自身要求的可见成本,不是隐藏。

结论:APPROVE(作者自有 PR,GitHub 拒绝正式自批准,故以 COMMENTED review 记录同一结论)。合并前唯一需要的补充证据是一次针对本 head 的实时 Dashboard 点击验收——打包 UI 走的 HTTP 路径已由接口测试覆盖、打包产物已字节校验,但浏览器会话上次是在 development 模式跑通的;已记入 residual risk,不构成阻断。

English verdict: APPROVE — one machine-scoped steward_executor namespace the frontend edits and machine-config describe/inspect read back, resolved as machine_configuration > service environment > product default per field; an unconfigured machine is unchanged; the 19-check premerge gate passes at this head; two non-blocking P3 follow-ups are recorded.

@huangruiteng
huangruiteng merged commit f4ed58d into main Sep 16, 2026
22 of 24 checks passed
@huangruiteng
huangruiteng deleted the codex/steward-executor-machine-config-20260916 branch September 16, 2026 05:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant