Skip to content

feat(managed): default the steward and managed Turns to one execution profile - #4470

Merged
huangruiteng merged 4 commits into
mainfrom
codex/managed-execution-profile-defaults
Sep 15, 2026
Merged

huangruiteng merged 4 commits into
mainfrom
codex/managed-execution-profile-defaults

Conversation

@huangruiteng

@huangruiteng huangruiteng commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Managed execution profile, and a steward channel that can reach it

The steward channel and the governed Turn now resolve one managed execution
profile, and the steward channel's shipped executor becomes conditional on one
reported local fact.

execution profile   deepseek-official / deepseek-v4-flash / high
                    LOOPX_TURN_PROVIDER | LOOPX_TURN_MODEL | LOOPX_TURN_REASONING_EFFORT
                    then legacy DSH_PROVIDER | DSH_MODEL, then an explicit CLI argument
steward executor    operator credential configured -> dsh (managed)
                    otherwise                       -> codex (individual, personal CLI login)

Disclosure: this changes shipped defaults

  1. The steward channel's executor default is now conditional. The owner's
    earlier ruling was that the steward stays on codex. That is reversed here:
    with a configured operator credential the channel selects the managed host,
    and without one it keeps the interactive CLI endpoint so the steward stays
    reachable on a machine that has only a personal login. An explicit
    LOOPX_MANAGER_ENDPOINT still wins over both.
  2. A discovered credential never re-points an already-selected Turn host. The
    Turn default stays dsh regardless of the environment, as before. What the
    credential selects here is the channel's executor, which is a different
    surface with its own rule.
  3. The default model and reasoning effort for managed work change from the
    vendor CLI default to deepseek-v4-flash at reasoning effort high; both are
    overridable and the readback names which value applied.
  4. managed_host_chat_transport_unsupported is retired. It described a
    transport gap that this change closes, and keeping it would leave a working
    host unreachable. The reasons the channel can still report are the managed
    host's own launchability facts (dsh_runtime_unavailable,
    operator_credential_unconfigured, invalid_reasoning_effort).

The transport's limits are stated, not implied

loopx/chat_dsh.py runs one bounded DeepSeek Harness segment per Chat turn on
the resolved profile. It claims:

  • no partial streaming — the answer arrives as one final message;
  • no cross-turn host session — the visible history is Chat-side context, and each
    segment is fresh;
  • no tool authority — the channel pins DSH_PERMISSION_MODE=read-only, so dsh
    refuses a write itself instead of trusting a prompt.

The channel readback reports adapter_kind: deepseek_harness_segment,
trust_scope: read_only, sandbox_mode: read-only, and streaming: false.
Because the segment is not a session, the channel does not offer cross-turn host
continuity, and nothing here claims it does.

The agent-facing output budget stayed a contract

The profile first shipped as a nested object, and the base/head differential
refused it: a plan is a hot-path payload with a 64-character growth allowance per
row. It is now one line, deepseek-v4-flash@high, with the provider prepended
only when the resolved provider is not the shipped one — dropping it for a
deviating provider would make the line claim a profile the Turn would not use.
The field-by-field form with per-field sources belongs to the surfaces a person
reads while configuring LoopX.

Frontend

The manager header said the channel needed codex and sent a managed host to
loopx turn. That copy was true only while the gap existed. The header now names
the reason it actually received, and because the default is conditional it also
states which branch it took and why — a steward on the CLI endpoint looks
different when the operator chose it than when the machine simply has no
credential. The packaged chat bundle and its retention ring are rebuilt, and the
scenario covers the conditional default, a launchable managed host, and both
typed unavailability reasons.

Validation

Changed surfaces: managed execution profile resolution, the
managed_executor_binding_v0 readback, the steward channel binding and its chat
transport, the dsh host adapter, the dashboard header + packaged chat bundle,
and the protocol/RFC documents for all of the above.

  • loopx canary premerge --from-git-diff --timeout-seconds 300 — passed,
    19 checks, 0 failures, 0 advisory failures, 0 manual holds, 41 changed files.
  • examples/control_plane/cli-output-budget-regression-smoke.py — passed
    (base/head differential against origin/main, 102 rows).
  • python -m ruff check tests loopx/canary loopx/control_plane loopx/domain_packs loopx/presentation
    and python -m mypy — passed.
  • pytest over the affected suites (turn executor/driver, managed executor
    binding, default host binding, manager channel binding, manager context
    handoff, dsh goal mode, the new chat-dsh adapter tests, chat startup isolation,
    chat turn wait, kiro host surface) — 276 passed.
  • examples/loopx-steward-managed-chat-smoke.py — real bundled dsh against a
    local mock model endpoint: resolved binding dsh/managed/available, model and
    reasoning effort on the wire, answer persisted without the machine envelope,
    proposals parsed, and the read-only sandbox refused a write
    (permission_mode: read-only, write_denied: true).
  • examples/loopx-turn-dsh-real-e2e-smoke.py — passed, so the profile plumbing
    did not disturb the governed Turn.
  • examples/loopx-steward-channel-binding-smoke.py — passed.
  • npm run smoke:personal-workspace-packaged (packaged chat bundle, Chromium) —
    passed.
  • examples/docs-governance-smoke.py, examples/docs-asset-integrity-smoke.py —
    passed.

Known unrelated failure, present on the untouched base: two
tests/capabilities/test_benchmark_four_arm_contract.py cases fail identically
without this change.

One defect the first CI run caught

The first remote run of this PR failed in test-shard (4):
tests/test_manager_channel_binding.py::test_a_managed_host_that_cannot_launch_raises_a_typed_gate
asserted that the typed gate names DEEPSEEK_API_KEY, but which typed reason
appears depends on whether the optional deepseek-harness runtime is importable
on the machine running the suite. A developer environment that has the extra
resolves operator_credential_unconfigured; CI, which does not install it,
resolves dsh_runtime_unavailable. Both reasons are correct behavior — the
assertion was reading the environment instead of a pinned fact, and the same
assumption existed in examples/loopx-steward-channel-binding-smoke.py.

Both are now pinned in each direction: the credential gate is asserted with the
runtime present, the runtime-absent install guidance has its own case, and the
smoke pins the probe so its credential assertion carries the same fact
everywhere. The condition is reproducible locally by forcing
importlib.util.find_spec('deepseek_harness') to report the runtime absent, so
this no longer depends on luck. No production code changed in that fix.

Test plan for a reviewer

loopx turn plan --host dsh --format json | jq .managed_executor
loopx --format json chat capabilities | jq .manager.channel_binding

The first shows execution_profile: "deepseek-v4-flash@high"; the second shows
the resolved channel endpoint, its source, the reason the shipped default
applied, and the executor's availability.

… profile

The steward channel and the governed Turn now resolve one managed execution
profile: provider `deepseek-official`, model `deepseek-v4-flash` (DeepSeek V4.1
Flash) and reasoning effort `high`, overridable by `LOOPX_TURN_PROVIDER`,
`LOOPX_TURN_MODEL` and `LOOPX_TURN_REASONING_EFFORT`, then by the legacy
`DSH_PROVIDER`/`DSH_MODEL`, then by an explicit CLI argument. Every value reports
which of those decided it, and a credential only authenticates the profile; it
never chooses one.

The steward channel's shipped executor becomes conditional on exactly one
reported local fact: with an operator credential configured it selects the
managed host (`dsh`), and without one it stays on the interactive CLI endpoint
(`codex`). The readback names which branch applied, so a conditional default is
disclosed instead of being indistinguishable from an incidental environment
read. The channel remains reachable on a machine that has only a personal login,
and it can no longer land on a managed executor with an unrelated model.

The managed host answers through a segment transport
(`loopx/chat_dsh.py`): one bounded DeepSeek Harness segment per Chat turn, on the
resolved profile, with LoopX composing the bounded visible history and pinning
the segment sandbox read-only. The transport claims no partial streaming, no
cross-turn host session and no tool authority, and it reports those limits. The
`managed_host_chat_transport_unsupported` gate is retired with it, because it
described a gap that no longer exists; the reasons the channel can still report
are the managed host's own launchability facts.

Agent-facing payloads carry the resolved profile as one line
(`deepseek-v4-flash@high`, with the provider prepended only when it is not the
shipped one) because every plan carries it and the agent-facing output budget is
a contract.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…and why

The manager header said the channel needed `codex` and sent a managed host to
`loopx turn`. That copy was true only while the managed host had no chat
transport; the header now names the reason it actually received instead of one
hardcoded host.

Because the shipped executor is conditional, the chip also states which branch
it took and why: a steward on the interactive CLI endpoint looks different when
the operator chose it than when the machine has no credential, and the operator
can see the one fact that would move the channel to the managed host. An
unrecognized reason stays unclaimed.

The packaged chat bundle and its retention ring are rebuilt, and the scenario
covers the conditional default, the launchable managed host, and both typed
unavailability reasons, so the documentation assets are produced by the smoke
instead of being cropped by hand.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
… segment transport

The harness-selection RFC still described the steward default as `codex` and
gated promoting the managed host on a chat transport that now shipped. Both
language versions now state one conditional rule with its reason, the one
execution profile both managed surfaces resolve, and the segment transport's
typed limits, while keeping the retired-reason history legible.

The Turn protocol and the connector document the profile precedence, the
`invalid_reasoning_effort` refusal, and the one-line readback shape that keeps
the agent-facing output budget a contract. The manager protocol and the
documentation assets record what the header now renders.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head 02329d513c61b13b125bd9c30afc4a6e22ce1675 (base main), 41 files, +2350/-496.

动机

管家通道(manager channel)和受管 Turn(loopx turn / managed goal)此前各自决定自己跑哪个托管模型。基线 main 上,管家通道的执行器默认是常量 codex、模型默认是 MANAGER_MODEL_DEFAULT = "gpt-6-astra",与凭据无关;受管 Turn 的宿主虽然已经是 dsh,但 managed_executor_binding 根本不公布模型与 effort,dsh adapter 只从 DSH_PROVIDER/DSH_MODEL 读 provider/model,完全没有 reasoning effort 这个概念。

结果是两条独立的模型决策:交互的管家和它派发出去的受管工作可能落在两个不同的托管模型上,操作者无法判断某段工作到底是哪个 profile 产出的;而且管家默认根本跑不到"由操作者凭据计费"的执行器上,除非同时显式改 endpoint 再单独改模型。

改动后(本 head 实测):environ = {DEEPSEEK_API_KEY: 已配置} 时,manager_channel_binding 返回 executor_endpoint='dsh'、executor_kind='managed'、model='deepseek-v4-flash'、execution_profile='deepseek-v4-flash@high'、executor_endpoint_default_reason='operator_credential_configured';无凭据时返回 'codex' / 'gpt-6-astra' / 'operator_credential_absent',与基线逐字段一致。

更小的修法都存在但不够:把默认值直接改成常量 'dsh',会让只有个人 CLI 登录的机器上管家一开就撞 typed gate(交互通道反而变得不可用);只加一个管家专属模型覆盖,则两条独立决策原样保留。条件默认 + 单一 profile 是同时满足"没凭据仍可达"和"有凭据只有一个 profile"的最小形态。

改动思路

入口是两条:loopx chat 打开管家通道,以及 loopx turn plan/run-once。权威输入只有进程环境——操作者凭据的存在性(DEEPSEEK_API_KEY)、通道覆盖 LOOPX_MANAGER_ENDPOINT/LOOPX_MANAGER_MODEL/LOOPX_MANAGER_REASONING_EFFORT、以及 profile 覆盖 LOOPX_TURN_PROVIDER/LOOPX_TURN_MODEL/LOOPX_TURN_REASONING_EFFORT(旧名 DSH_PROVIDER/DSH_MODEL 降级保留)。决策边界被收紧成三个各自唯一的 owner:_resolve_manager_endpoint 决定通道选哪个执行器并给出 typed 原因;managed_execution_profile 决定 provider/model/effort 及每个字段的来源;managed_executor_binding 决定可用性裁决与单行读回。

正向路径:_resolve_manager_endpoint 先看显式覆盖,没有则用 operator_credential_configured 决定出厂默认 → manager_channel_binding 对 managed endpoint 直接引用 Turn 侧的 managed_executor_binding(而不是自己再推一遍可用性)→ chat_runtime._start_adapter 的 MANAGED_TURN_HOST 分支构造 DshChatAdapter → 每轮一个受 DSH_PERMISSION_MODE=read-only 约束的有界 segment,最终消息经 parse_agent_response 落回可见历史。

与既有实现的关系是"复用而不是新增权威":chat_manager、host_binding、chat_agent 的 endpoint 目录、chat_runtime._start_adapter、dsh_goal_mode/turn_host_adapter 都是被扩展;chat_endpoint_catalog.py 是从 chat_runtime.py 里机械抽出的 endpoint 行(为把该文件压回 1500 行棘轮之内,现为 1476);chat_dsh.py 是新模块,但它是既有 ChatRuntimeAdapter 协议在 dsh 上的第一个实现——在 loopx/chat_*.py 上有界负向搜索只找到 codex app-server、claude-code、kiro ACP、anthropic-api、openai-api 五种传输,没有托管宿主传输,因此这不是第二个决策 owner。

关于"新增/强化的状态":本 PR 没有引入任何权威状态,也没有引入需要手工同步的状态。managed_execution_profile 的每个字段、executor_endpoint_default_reason、execution_profile 单行串,全部是既有 canonical 事实(产品常量 + 环境 + 凭据存在性)的派生投影,且与选择结果在同一次解析里产生,读回不可能与 adapter 实际拿到的值漂移。

具体改动

运行时 15 个文件、测试/示例 8 个、公开文档 9 个、前端 4 个、构建配置 2 个、其他 3 个;其中 loopx/web/chat/ 下 5 个是 npm run build:chat 的生成产物,与 4 个源文件一一对应,重跑构建确认无差异。

关键代码讲解

  1. loopx/control_plane/turn_driver/execution_profile.py:88 managed_execution_profile —— 每字段三级优先级(显式参数 > 规范环境变量 > 旧名 > 出厂常量),并回报每个字段的来源与生效变量名。它从不抛异常:不支持的 effort 只体现为 reasoning_effort_supported=false,由 managed_profile_unavailable_reason 转成 invalid_reasoning_effort,避免"配置错误被当成运行时异常吞掉"。
  2. loopx/control_plane/turn_driver/host_binding.py:112 managed_executor_binding —— 新增 provider/model/reasoning_effort 入参与 execution_profile 输出。裁决优先级是 runtime 缺失 > 凭据未绑定 > profile 被拒,保证调用者拿到的是第一个真正拦住启动的事实;非 managed 执行器一律 execution_profile=None,不冒充。
  3. loopx/chat_manager.py:155 _resolve_manager_endpoint —— 返回 (endpoint, source, reason)。关键不变量:显式 LOOPX_MANAGER_ENDPOINT 永远压倒一切,凭据只影响出厂默认、且 reason 仅在 source 为 product_default 时非空,所以 UI 不会把"用户显式选的"渲染成"系统默认的"。
  4. loopx/chat_dsh.py:129 DshChatAdapter.start_turn —— 每轮恰好一个 segment,env 固定 DSH_PERMISSION_MODE=read-only;空回答或不可解析结果一律 fail closed(managed_host_chat_failed),超时抛 managed_host_chat_timeout,不会把"没答出来"当成答案。
  5. loopx/chat_runtime.py:395 _start_adapter 的 MANAGED_TURN_HOST 分支 —— 把选中的 endpoint 变成具体 adapter。两个易错点被显式处理:该宿主不把 history 塞进 objective(由 adapter 自己携带窗口,避免重复注入),以及管家通道保留自己的 LOOPX_MANAGER_MODEL/LOOPX_MANAGER_REASONING_EFFORT 覆盖并叠加在共享 profile 之上。

对主干的风险

最强的回归场景是:一台没有操作者凭据、但有个人 CLI 登录的机器,管家通道因为默认执行器从 codex 变成无法认证的托管宿主而直接不可用。该场景被结构性挡住——_resolve_manager_endpoint 先读 operator_credential_configured 才决定出厂默认;我在本 head 上实际执行了三种环境并逐字比对基线文件,无凭据路径的 endpoint/kind/model/source 与 main 完全一致。

其余阻断路径都是 typed 拒绝而非静默降级:显式选中 managed endpoint 但无凭据 → available=false、unavailable_reason='operator_credential_unconfigured',开 session 时抛 agent_endpoint_unavailable,next action 明确点名 DEEPSEEK_API_KEY;LOOPX_TURN_REASONING_EFFORT=ludicrous → reasoning_effort_supported=false,Turn 侧在花费一次运行之前抛 BuiltInHostError('dsh_execution_profile_rejected');dsh 运行时不可导入 → dsh_runtime_unavailable 并给出安装指引;segment 尝试写文件 → 被只读沙箱拒绝(实测 write_denied=true)。

影响面限于管家通道的 session 建立、受管 Turn 的读回,以及 dashboard 的 header 渲染;个人 Codex 路径、Claude/Kiro/直接模型 endpoint、goal pipeline、Turn 宿主选择默认值(仍为无条件 dsh)均未改动。回滚成本是一个环境变量或一个 merge commit。

可观测性上,executor_endpoint/executor_endpoint_source/executor_endpoint_default_reason/executor_kind/execution_profile/available/unavailable_reason 全部出现在 /api/chat capabilities 与 header;前端 schema 把新增字段标为 optional,旧 payload 仍可解析,因此默认关闭态保持了形状兼容。

验证:loopx canary premerge --from-git-diff --timeout-seconds 300 → passed,19/19 执行、0 失败、self_merge_allowed: true;ruff 全绿;mypy success(22 源文件);12 个受影响模块 276 passed;npm run build:chat 后 loopx/web/chat 无差异且打包版 personal-workspace smoke 通过;docs governance 与 asset integrity 通过。远端 CI 未等待(本 lane 的明确要求,改用本地 canary premerge 门禁);另有两例 tests/capabilities/test_benchmark_four_arm_contract.py 在未改动的基线 worktree 上同样失败,属既有基线失败而非本 head 回归。

我的整体评价

基线/head 对比共五行(无凭据、有凭据、显式覆盖、profile 覆盖、非法 effort),结论是"有意的默认变更 + 其余等价",而不是"决策码相同":唯一的意图变更就是"有凭据时管家通道的默认执行器与模型跟随共享 profile",且该变更在 header、PR 描述、两份协议文档和双语的 harness 选型 RFC 中同时披露。代码量判定为必要——新增的三块(profile 解析、托管宿主 chat 传输、endpoint 目录抽取)各自有活跃调用点,抽取部分还把 chat_runtime.py 压回了棘轮之内;proportionality 判定为相称,因为最省的那两种修法都会把问题留在原地或把交互通道弄坏。default-off 隔离判定为 isolated:无凭据路径与基线逐字段相同。authority_semantics 判定为 aligned:凭据既不选 provider 也不选模型,也不重指向已显式选定的执行器,adapter 明确声明 tool_calls: false / trust_scope: read_only。

残留风险如实记录:mock 模型 endpoint 无法证明厂商确实接受 deepseek-v4-flash@high,只证明了"该 profile 确实上到 wire 且答案被正确解析";chat_dsh.py 是首个真实流量的托管宿主聊天传输,只读沙箱、有界超时和空回答 fail-closed 把风险限制在"这一轮失败"而非越权动作。两条 P3 建议(adapter 级 resume 字段的语义、next-action 映射需与新原因同步增加)不构成阻断。结论:APPROVE。

English verdict: APPROVE at exact head 02329d513c61b13b125bd9c30afc4a6e22ce1675. The steward channel and the governed Turn now resolve one managed execution profile (deepseek-v4-flash@high), and the channel's shipped executor becomes dsh only when the operator credential is configured — otherwise it keeps the previous codex/gpt-6-astra default field-for-field. No blocking finding; two P3 notes only (the adapter-level resume capability name carries a different meaning than the Codex adapter's, and a typed reason without an entry in AGENT_ENDPOINT_NEXT_ACTIONS degrades to generic text). Validation: local canary premerge passed 19/19 with 0 failures, 276 targeted tests passed, the real-dsh managed-chat smoke returned execution_profile: deepseek-v4-flash@high with the read-only sandbox denying a write, and the regenerated chat bundle matches source; remote CI was deliberately not awaited for this lane, and two pre-existing benchmark-contract test failures reproduce unchanged on the untouched base.

The typed-gate assertions in the steward binding test and smoke read the
dsh runtime availability probe straight from the machine running them, so
the reason they asserted depended on whether that machine happened to have
the optional runtime installed. CI does not install it, and the credential
assertion failed there while passing locally.

Pin the probe in both directions instead: the credential gate is asserted
with the runtime present, the runtime-absent install step gets its own
pinned case, and the smoke pins the probe so its credential assertion
carries the same fact on every machine.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Reviewed exact head 9d69dbdedd7deb0e45b8e63b77357e0fff9d338a (base main), 41 files, +2401/-497.
上一轮评审的确切 head 是 02329d513c61b13b125bd9c30afc4a6e22ce1675,本轮变动只涉及两个证据文件,且是本轮 CI 抓出来的真实缺陷的修复(见"对主干的风险")。

动机

管家通道(manager channel)和受管 Turn(loopx turn)此前各自决定自己跑哪个托管模型。基线 main 上,管家通道的执行器默认是常量 codex、模型默认是 MANAGER_MODEL_DEFAULT = "gpt-6-astra",与凭据无关;受管 Turn 的宿主虽然已经是 dsh,但 managed_executor_binding 根本不公布模型与 effort,dsh adapter 只从 DSH_PROVIDER/DSH_MODEL 读 provider/model,完全没有 reasoning effort 这个概念。

结果是两条独立的模型决策:交互的管家和它派发出去的受管工作可能落在两个不同的托管模型上,操作者无法判断某段工作到底是哪个 profile 产出的;而且管家默认根本跑不到"由操作者凭据计费"的执行器上,除非同时显式改 endpoint 再单独改模型。

改动后(本 head 实测):配置了 DEEPSEEK_API_KEY 时,manager_channel_binding 返回 executor_endpoint='dsh'、executor_kind='managed'、model='deepseek-v4-flash'、execution_profile='deepseek-v4-flash@high'、executor_endpoint_default_reason='operator_credential_configured';无凭据时返回 'codex' / 'gpt-6-astra' / 'operator_credential_absent',与基线逐字段一致。

更小的修法都存在但不够:把默认值直接改成常量 'dsh',会让只有个人 CLI 登录的机器上管家一开就撞 typed gate;只加一个管家专属模型覆盖,则两条独立决策原样保留。条件默认 + 单一 profile 是同时满足"没凭据仍可达"与"有凭据只有一个 profile"的最小形态。

改动思路

入口是两条:loopx chat 打开管家通道,以及 loopx turn plan/run-once。权威输入只有进程环境——操作者凭据的存在性(DEEPSEEK_API_KEY)、通道覆盖 LOOPX_MANAGER_ENDPOINT/LOOPX_MANAGER_MODEL/LOOPX_MANAGER_REASONING_EFFORT、以及 profile 覆盖 LOOPX_TURN_PROVIDER/LOOPX_TURN_MODEL/LOOPX_TURN_REASONING_EFFORT(旧名 DSH_PROVIDER/DSH_MODEL 降级保留)。决策边界收紧成三个各自唯一的 owner:_resolve_manager_endpoint 决定通道选哪个执行器并给出 typed 原因;managed_execution_profile 决定 provider/model/effort 及每字段来源;managed_executor_binding 决定可用性裁决与单行读回。

正向路径:_resolve_manager_endpoint 先看显式覆盖,没有则用 operator_credential_configured 决定出厂默认 → manager_channel_binding 对 managed endpoint 直接引用 Turn 侧的 managed_executor_binding(而不是自己再推一遍可用性)→ chat_runtime._start_adapter 的 MANAGED_TURN_HOST 分支构造 DshChatAdapter → 每轮一个受 DSH_PERMISSION_MODE=read-only 约束的有界 segment,最终消息经 parse_agent_response 落回可见历史。

与既有实现的关系是"复用而不是新增权威":chat_manager、host_binding、chat_agent 的 endpoint 目录、chat_runtime._start_adapter、dsh_goal_mode/turn_host_adapter 都是被扩展;chat_endpoint_catalog.py 是从 chat_runtime.py 里机械抽出的 endpoint 行(为把该文件压回 1500 行棘轮之内,现为 1476);chat_dsh.py 是新模块,但它是既有 ChatRuntimeAdapter 协议在 dsh 上的第一个实现——在 loopx/chat_*.py 上有界负向搜索只找到 codex app-server、claude-code、kiro ACP、anthropic-api、openai-api 五种传输,没有托管宿主传输,因此不是第二个决策 owner。

关于"新增/强化的状态":本 PR 没有引入任何权威状态,也没有引入需要手工同步的状态。managed_execution_profile 的每个字段、executor_endpoint_default_reason、execution_profile 单行串,全部是既有 canonical 事实(产品常量 + 环境 + 凭据存在性)的派生投影,且与选择结果在同一次解析里产生,读回不可能与 adapter 实际拿到的值漂移。

具体改动

运行时 15 个文件、测试/示例 8 个、公开文档 9 个、前端 4 个、构建配置 2 个、其他 3 个;其中 loopx/web/chat/ 下 5 个是 npm run build:chat 的生成产物,与 4 个源文件一一对应,重跑构建确认无差异。

本轮(02329d513..9d69dbde)只改两个证据文件:tests/test_manager_channel_binding.py 把原来一条依赖环境的断言拆成两条各自固定可用性探针的用例,examples/loopx-steward-channel-binding-smoke.py 增加一个很小的包装函数把探针固定住。生产代码零改动。

关键代码讲解

  1. loopx/control_plane/turn_driver/execution_profile.py:88 managed_execution_profile —— 每字段三级优先级(显式参数 > 规范环境变量 > 旧名 > 出厂常量),回报每个字段来源与生效变量名。它从不抛异常:不支持的 effort 只体现为 reasoning_effort_supported=false,由 managed_profile_unavailable_reason 转成 invalid_reasoning_effort,避免"配置错误被当成运行时异常吞掉"。
  2. loopx/control_plane/turn_driver/host_binding.py:112 managed_executor_binding —— 新增 provider/model/reasoning_effort 入参与 execution_profile 输出。裁决优先级是 runtime 缺失 > 凭据未绑定 > profile 被拒,保证调用者拿到的是第一个真正拦住启动的事实;非 managed 执行器一律 execution_profile=None,不冒充。
  3. loopx/chat_manager.py:155 _resolve_manager_endpoint —— 返回 (endpoint, source, reason)。关键不变量:显式 LOOPX_MANAGER_ENDPOINT 永远压倒一切,凭据只影响出厂默认,且 reason 仅在 source 为 product_default 时非空,所以 UI 不会把"用户显式选的"渲染成"系统默认的"。
  4. loopx/chat_dsh.py:129 DshChatAdapter.start_turn —— 每轮恰好一个 segment,env 固定 DSH_PERMISSION_MODE=read-only;空回答或不可解析结果一律 fail closed(managed_host_chat_failed),超时抛 managed_host_chat_timeout。
  5. loopx/chat_runtime.py:395 _start_adapter 的 MANAGED_TURN_HOST 分支 —— 把选中的 endpoint 变成具体 adapter,两个易错点被显式处理:该宿主不把 history 塞进 objective(由 adapter 自己携带窗口),管家通道保留 LOOPX_MANAGER_MODEL/LOOPX_MANAGER_REASONING_EFFORT 覆盖并叠加在共享 profile 之上。

对主干的风险

上一轮 CI 抓到一个真实缺陷,这就是本轮改动的全部原因。 02329d513 的 test-shard (4) 失败在 tests/test_manager_channel_binding.py::test_a_managed_host_that_cannot_launch_raises_a_typed_gate:断言要求 gate 的 next_action 里出现 DEEPSEEK_API_KEY,但"哪个 typed reason 出现"取决于跑测试的机器是否装了可选的 dsh runtime——有 extra 的开发者机器上是 operator_credential_unconfigured(命中断言),CI 没装则变成 dsh_runtime_unavailable(断言失败)。同一处假设也存在于 examples/loopx-steward-channel-binding-smoke.py。这不是产品缺陷(两种 reason 都是正确的 typed 拒绝),而是证据缺陷:断言读的是评审者机器的环境,而不是一个被固定的前提。

修法是把探针在两个方向上都固定:凭据缺失的 gate 在"runtime 存在"下断言并点名 DEEPSEEK_API_KEY,"runtime 缺失"另立一条断言安装步骤;smoke 用 mock.patch.object(host_binding, "dsh_runtime_importable", ...) 固定为存在。验证方式是让 importlib.util.find_spec('deepseek_harness') 返回 None 复现 CI 条件——该 suite 与 smoke 在这个条件下都通过;正常条件下同样通过(成对反事实)。

除此之外,最强的回归场景仍被结构性挡住:一台没有操作者凭据、有个人 CLI 登录的机器上,管家通道默认仍是 codex;无凭据路径的 endpoint/kind/model/source 与 main 逐字段一致。其余阻断路径都是 typed 拒绝而非静默降级:无凭据的托管 endpoint → available=false / operator_credential_unconfigured;非法 effort → reasoning_effort_supported=false 且 Turn 侧在花费一次运行前抛 dsh_execution_profile_rejected;runtime 不可导入 → dsh_runtime_unavailable + 安装指引;segment 尝试写文件 → 被只读沙箱拒绝(实测 write_denied=true)。

影响面限于管家通道 session 建立、受管 Turn 读回与 dashboard header 渲染;个人 Codex 路径、Claude/Kiro/直接模型 endpoint、goal pipeline、Turn 宿主选择默认值(仍为无条件 dsh)均未改动。回滚成本是一个环境变量或一个 merge commit。可观测性上,executor_endpoint/executor_endpoint_source/executor_endpoint_default_reason/executor_kind/execution_profile/available/unavailable_reason 都出现在 /api/chat capabilities 与 header;前端 schema 把新增字段标为 optional。

验证:loopx canary premerge --from-git-diff --timeout-seconds 300 → passed,19/19 执行、0 失败、self_merge_allowed: true;ruff 在 canary 范围内全绿(examples 下另有 56 条既有 E402,均不在本 PR 触及的文件里);mypy success(22 源文件);四个直接相关 suite 102 passed;语义词表漂移 smoke ok;npm run build:chat 后 bundle 无差异且打包版 smoke 通过。远端检查在上一 head 上除上述一例外均为成功;新 head 的远端套件在本评审写作时尚未跑完,因此"新 head 远端全绿"是留待合并前确认的验证项。

我的整体评价

基线/head 对比共六行(无凭据、有凭据、显式覆盖、profile 覆盖、非法 effort、探针双向固定),结论是"有意的默认变更 + 其余等价":唯一意图变更就是"有凭据时管家通道的默认执行器与模型跟随共享 profile",且该变更在 header、PR 描述、两份协议文档和双语的 harness 选型 RFC 中同时披露。代码量判定为必要——新增的三块各有活跃调用点,抽取部分还把 chat_runtime.py 压回棘轮之内;proportionality 判定为相称;default-off 隔离判定为 isolated(无凭据路径与基线逐字段相同);authority_semantics 判定为 aligned(凭据既不选 provider 也不选模型,也不重指向已显式选定的执行器,adapter 明确 tool_calls: false / trust_scope: read_only)。

值得单独记一笔的是过程结论:本轮最初只有本地证据,而本地环境恰好装了 CI 没有的可选依赖,这类"环境巧合导致断言恒真"的问题只有远端运行能暴露。修复后我把该条件做成了可复现的本地路径(强制 find_spec 返回空),避免下次再靠运气。

残留风险如实记录:mock 模型 endpoint 无法证明厂商确实接受 deepseek-v4-flash@high,只证明了该 profile 确实上到 wire 且答案被正确解析;chat_dsh.py 是首个真实流量的托管宿主聊天传输,只读沙箱、有界超时与空回答 fail-closed 把风险限制在"这一轮失败"。两条 P3 建议(adapter 级 resume 字段语义、next-action 映射需随新 reason 同步增加)不构成阻断。结论:APPROVE。

English verdict: APPROVE at exact head 9d69dbdedd7deb0e45b8e63b77357e0fff9d338a. The steward channel and the governed Turn resolve one managed execution profile (deepseek-v4-flash@high), and the channel's shipped executor becomes dsh only when the operator credential is configured — otherwise it keeps the previous codex/gpt-6-astra default field-for-field. The delta since the previously reviewed head is only the repair of a real defect the first CI run caught: the typed-gate assertions read the optional dsh-runtime availability probe from the machine running them, so they held locally and failed in CI; both directions are now pinned, and the suite plus the smoke pass with the runtime reported absent (the CI condition) and present. No blocking finding remains; two P3 notes only (the adapter-level resume capability name carries a different meaning than the Codex adapter's, and a typed reason without an entry in AGENT_ENDPOINT_NEXT_ACTIONS degrades to generic text). Validation: local canary premerge passed 19/19 with 0 failures, 102 tests passed in the affected suites, the real-dsh managed-chat smoke returned execution_profile: deepseek-v4-flash@high with the read-only sandbox denying a write, and the regenerated chat bundle matches source; the new head's remote suites had not finished when this review was written, and two pre-existing benchmark-contract test failures reproduce unchanged on the untouched base.

@huangruiteng
huangruiteng merged commit 82b5ff5 into main Sep 15, 2026
31 checks passed
@huangruiteng
huangruiteng deleted the codex/managed-execution-profile-defaults branch September 15, 2026 19:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant