fix(connect): stop generating first-connect onboarding todos - #4465
Conversation
`loopx bootstrap` / `loopx connect` used to project a generated onboarding queue into a freshly connected goal: a user `onboarding_decision` gate, an agent `onboarding_todo_review` todo, optional candidate todos from a repository scan, and a `onboarding_connection_validation` todo for generic adapters. A default connect therefore landed in `state=operator_gate` with `should_run=false`, and autonomous callers had to pass `--accept-onboarding-agent-todos --begin-autonomous-advance --codex-app-heartbeat no` just to reach their own first delivery todo. An unattended benchmark run could spend its whole turn on that generated item. Connection is now a pure registration step: - `connect` / `bootstrap` write the registry entry and the active state only. The caller, or the connected domain adapter, owns the first delivery todo. - Removed the onboarding scan, candidate todos, generated user/agent todos, heartbeat opt-in gate, and `connection_validation` registry annotation, and deleted `loopx/onboarding.py`. - Removed the `--no-onboarding-scan`, `--accept-onboarding-agent-todos`, `--begin-autonomous-advance`, `--codex-app-heartbeat`, `--onboarding-connection-validation`, and `--onboarding-max-*` options, and the onboarding fields from the bootstrap payload and markdown. - `Next Action` now carries one neutral statement instead of an onboarding prompt, so a fresh goal has no hidden onboarding frontier. - Caller-facing prompts and the emitted bootstrap command pack no longer ask an agent to collect onboarding candidates; they ask for the first delivery todo instead. - Renamed the canary profile/surface from `new-user-onboarding-lifecycle` to `first-connect-contract` and regenerated the semantic inventory. Affected default-behavior lanes: fresh `connect`/`bootstrap` state projection, the quota entry gate for a newly connected goal, Codex App heartbeat opt-in, and the benchmark harnesses that worked around the gate. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Record the default-behavior change in the public docs: `connect`/`bootstrap` register the goal and write the active state only, and create no first-connect onboarding todo, owner-decision gate, or host-loop opt-in gate. The first delivery todo belongs to the caller or the connected domain adapter. - `docs/integration.md`: drop the removed `--no-onboarding-scan`/`--onboarding-connection-validation` provider example and state the registration-only contract. - Book chapter 05 (zh/en): replace the "next Agent Todo after connect" expectation and add the connection-contract note. - `docs/operations/new-project-codex-prompt.md`: replace the onboarding-scan and candidate-todo instructions with "collect the first delivery todo". - `docs/concepts/interaction-pattern-catalog.md`: point the state-projection pattern at the replacement first-connect smoke. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The SWE-marathon harnesses passed `--no-onboarding-scan` (v1/v2) or `--accept-onboarding-agent-todos --begin-autonomous-advance --codex-app-heartbeat no` (v3) to keep a native Codex goal off the generated first-connect onboarding todo. Upstream `connect`/`bootstrap` no longer writes any first-connect todo, user gate, candidate list, or opt-in gate, so those flags are gone and the harness no longer needs a bootstrap gate string. `_bootstrap_gates` now returns only the optional `--write-scope` argument, and the incident comments record that the workaround is obsolete rather than presenting it as the current recipe. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
…t smoke Delete the two smokes that encoded the removed first-connect onboarding queue (`examples/onboarding-connect-candidates-smoke.py`, `examples/project/onboarding-no-scan-projection-smoke.py`) and guard the new contract from the outside with `examples/project/first-connect-contract-smoke.py`: a real `connect` writes no onboarding field, marker, user todo, or agent todo, a fresh goal reports `should_run=true`/`effective_action=normal_run`, and the caller's own `[P0]` todo becomes the selected `recommended_action`. Update the remaining smokes and tests that passed the removed flags, and let the fixtures that needed a runnable frontier seed their own first agent todo instead of relying on connect to generate one. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Pre-merge validation evidence
Changed surfaces — Direct checks — Catalog canaries — Risk-profile smokes — 8 selected from the Public/private boundary — Focused validation beyond the gate
Failures / skips — none failing. Skipped locally: Manual holds — the gate reports |
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval).
Exact head reviewed: 145607e461de7b0ed2f39ed30cc5c60f010e923c (merged as 26eaefc21487294bb983ee084ad5879cadb54913)
动机
loopx bootstrap / loopx connect 过去会给每个刚接入的 goal 投影一整队生成的 onboarding 工作:一条 user_gate todo(task_class=user_gate、action_kind=onboarding_decision)、一条 agent todo(action_kind=onboarding_todo_review)、最多四条由仓库扫描派生的候选 todo,以及 generic adapter 专属的 onboarding_connection_validation agent todo;同时还会写 adapter.connection_validation 并把 onboarding_scan、accept_candidate_commands、heartbeat_opt_in_required 等字段回给调用方。
代价是可复现的,不是推测。我用同一份合成项目、同一组 connect 参数在 292b85e90 与本次 head 上各跑了一次真实 CLI:baseline 写入了 1 条 user_gate todo 加 1 条 onboarding_todo_review todo,## Next Action 变成 "Ask the user which proposed onboarding agent todos to accept…",随后 loopx quota should-run 返回 should_run=false、effective_action=operator_gate_notify、blocked_action_scope=gated_delivery、mode=user_gate。也就是说:无人值守的调用方要么先知道 --accept-onboarding-agent-todos --begin-autonomous-advance --codex-app-heartbeat no(或 --no-onboarding-scan)这套口令,要么就静默停在生成的门禁上。SWE-marathon harness 的注释记录了具体事故:一轮 trial 在生成出来的 onboarding 条目上耗掉约 900 秒,而任务文件一个字没改,gate 还是放行的、收据还是干净的。
更小的修法不够。只压掉那条 user_gate、保留扫描和候选清单,扫描派生的条目和连接期的 next action 仍然会和调用方自己的工作队列抢闸门——这正是被报告的问题本身。把那些开关留成 no-op 则保留了一个名不副实的 CLI 面。所以最小且能真正消掉干扰的做法就是删掉这个投影、把首发路由交还给 domain adapter,也就是这个 PR 做的事。
改动思路
入口是 loopx bootstrap|connect → loopx/cli_commands/bootstrap_connect.py → loopx.bootstrap.bootstrap_project。权威状态是 .loopx/registry.json 的 goal 条目加上 .codex/goals/<goal-id>/ACTIVE_GOAL_STATE.md,而闸门真正用来选工作的是状态文件里的 todo 队列,不是连接期的渲染文本。改动把决策边界明确成:connect 只负责登记,首发路由属于已接入的 domain adapter 与调用方自己的第一条 todo;bind_next_action_to_todo 的导入、shlex、add_todo_to_lines 与整个 loopx/onboarding.py 一起被删掉,## Next Action 退化成一个模块级常量。
我按 reviewer 的要求在 base 和 head 两棵树上做了资源与权限维度的复用搜索。最近的既有实现有三处:loopx/agent_onboarding.py(loopx_agent_onboarding_v0,由 loopx agent-onboard 暴露,负责 host skill 投递与 slash command 安装)、loopx/bootstrap_command_pack.py 的 host-loop activation 命令、以及 loopx/control_plane/work_items/backlog_hygiene.py 的隐藏积压检测。三者在资源(host skill 包而非 goal todo 队列)、数据范围(按 agent_type 与 host surface,不做仓库扫描)、权限(安装指引,不产生门禁或决策)、失败责任方(skill 投递复核)上都不同,所以保留它们是合理的边界,不是重复的权威。被删掉的那条规则在 head 上确实没有第二个实现:所有写入那些 todo kind、payload 字段和 registry 注解的代码都在被删掉的两个文件里;connection_validation、onboarding_decision、onboarding_todo_review 在 head 上只剩测试和守卫 smoke 里的"断言其不存在"。头部路径也随之收敛:连接后 quota should-run 返回 should_run=true、normal_delivery_allowed=true、effective_action=normal_run、mode=bounded_delivery;调用方写下自己的 [P0] 之后,recommended_action 与 agent_channel.primary_action 都是那一条。
具体改动
46 个文件、+401/-1558,净减 1157 行,是一份典型的删除型改动:production 13 个文件 +34/-906,其中 loopx/bootstrap.py -577/+3、被删除的 loopx/onboarding.py -214、loopx/cli_commands/bootstrap_connect.py -64 承载了几乎全部删除;tests 8 个文件 +20/-26;docs 5 个文件 +32/-40;smoke/example 14 个文件 +293/-554(删掉两个编码旧默认的 smoke 共 -519,新增 225 行契约 smoke);benchmark/demo 6 个文件 +22/-32 去掉已失效的 bootstrap 开关。
删掉的公开面必须点名:8 个 CLI 选项(--no-onboarding-scan、--onboarding-connection-validation、--accept-onboarding-agent-todos、--begin-autonomous-advance、--codex-app-heartbeat、--onboarding-max-{commits,status-paths,top-level-files})、10 个 bootstrap payload/markdown 字段(含 onboarding_scan、accept_candidate_commands、heartbeat_opt_in_required、host_loop_activation_required、onboarding_todos_written)、以及 registry 里的 adapter.connection_validation。canary profile new-user-onboarding-lifecycle 与 surfacing 条目一起改名为 first-connect-contract / first-connect。保留下来的 loopx agent-onboard/loopx_agent_onboarding_v0 是另一个契约(host skill 投递),名字与范围仍然吻合。
关键代码讲解
loopx/bootstrap.py:46DEFAULT_NEXT_ACTION—— 原先由onboarding_next_action()计算的多行任务式提示(或 connection-validation 动作)被替换成一个中性的静态常量:"Initial routing is owned by the connected domain adapter."。状态文件依然恒有且仅有一条 Next Action bullet,消费方是quota should-run的active_state_next_action与 status/diagnose 投影;这条常量同时被新 smoke 和改名后的tests/control_plane/test_todo_next_action_settlement.py钉住。loopx/bootstrap.py:118bootstrap_project—— 删除了build_onboarding_scan()调用、apply_onboarding_todos_to_state()写入、bind_next_action_to_todo()绑定、adapter["connection_validation"]注解,以及返回体里那一组onboarding_*字段;保留preserve_todos/force对既有状态的改写语义。调用方覆盖 CLI、chat_actions.py的 goal-create、control_plane/projects/registry.py、demo 与 benchmark harness。loopx/cli_commands/bootstrap_connect.py:118register_bootstrap_connect_command—— 八个 onboarding 参数整体从 parser 移除,因此旧调用会以 argparse 错误响亮失败,不会静默半写状态。examples/bootstrap-command-pack-smoke.py:225反向钉住了生成的connect_command_if_needed不再包含--no-onboarding-scan。loopx/project_prompt.py:72render_goal_start_bootstrap_command/render_codex_cli_bootstrap_message_text—— 面向调用方的文案从"展示候选 todo 并询问接受哪些 + autonomous/heartbeat"改成"只读核对后给出 1-3 个首个交付 todo 候选并等确认",同时保证渲染出的命令只使用 head 上仍然存在的选项。examples/project/first-connect-contract-smoke.py:1—— 取代被删的两个 smoke,用真实python -m loopx.cli驱动 generic 与 domain 两种 adapter:断言无 onboarding 标记、user/agent 队列为空、should_run=true/effective_action=normal_run、无user_gate、state_projection_gap为 None,并在调用方写入自己的[P0]后断言它成为被选中的recommended_action。
对主干的风险
最强的回归场景是外部集成:仍然传那 8 个已删选项、或解析已删 payload 字段的调用方会立刻失败。好消息是失败方式是响亮的——argparse 在任何状态写入之前就拒绝未知选项,不可能留下半写 goal;仓库内残余引用只有 benchmark 事故注释、deprecate/benchmark-legacy/,以及那条断言开关已消失的 smoke,没有任何过期的用户指引。回滚就是 revert 合并提交 26eaefc21,主题单一;两个方向都不需要状态迁移,因为 head 上没有任何读取被删字段的代码。
我独立跑了仓库原生验证。examples/project/first-connect-contract-smoke.py 通过;examples/control_plane/fresh-project-onboarding-regression-smoke.py 30/30 通过;agent-diagnose-packet、bootstrap-command-pack、canary/catalog-planner、fresh-clone-quickstart、bootstrap-force-preserve-todos、claude-goalmode-lifecycle、cli-bootstrap-rollout-helper-command-modularization、benchmark-native-goal-installed-profile、todo-claim-lease-roadmap、state-projection-gap 全部通过;被改动的 8 个测试文件跑出 254 passed / 4 skipped / 1 failed,唯一失败是 test_todo_projection_recovery.py::...[sqlite],原因明确写着本机 Node v25.5.0 + SQLite 3.51.2 不是被认证的运行时(要求 Node 22.22.3),与本次 diff 无关,且该文件在本机单独跑是 1 passed,评审头在 CI 上 29/29 全绿。git diff --check 干净。
几处需要说清楚的语义变化:连接期的 interaction_contract 从 mode=user_gate、must_attempt=false 变成 mode=bounded_delivery、agent_channel.must_attempt=true、quiet_noop_allowed=false,execution_obligation.must_attempt_work 为 true,义务内容落在 "run the first read-only adapter tick and save a compact run record",所以零 todo 的新连接是可完成的,不需要伪造队列项;这是机器强制的义务,PR 与文档没有把它说成"指引"。另一处未被断言覆盖的差异:新连接会在 quota/diagnose 里带上既有的 backlog_hygiene_warning(hidden_backlog_without_agent_todo、requires_agent_todo=true),因为它有 Next Action 文本而 agent todo 数为 0;它是建议性的、不门禁 should_run,并且作者已在 PR 正文里作为"后续独立决策"披露,所以我不把它算成本 PR 的新发现。另外 examples/shared-goal-authority-e2e/installed.py 与 SWE-marathon 的 Docker 路径我没有本地执行,属于我要明说的证据缺口。
我的整体评价
这是一次比例合适的删除。原始问题在每次连接、每个自动 trial 上都会发生,代价是一整轮预算的静默空转和干净的假收据;修法是删掉产生干扰的那个投影,净减 1157 行、删掉两个模块、没有引入任何新 CLI 面、新 schema 或新持久字段,也没有留下并行的第二权威。我在真实入口上做的基线/头对照把可见差异收敛到三处——todo 队列、Next Action 文本、闸门结果——与作者声明的意图一致,可复现、可回滚。风险主要是对外兼容性而非正确性:8 个 CLI 选项、一批 payload 字段,以及 canary profile 改名,这些都是破坏性的,必须在 1.0.5 的 release notes 里点名,而不是只躺在 PR 正文和文档里。结论:APPROVE。若后续要复审,需要的证据是 Docker 版 SWE-marathon bootstrap 路径和安装包路径的实测收据,以及那条 backlog_hygiene_warning 的处置决定。
English verdict: APPROVE. Exact head 145607e461de7b0ed2f39ed30cc5c60f010e923c (merged as 26eaefc21), 46 files +401/-1558. No blocking finding. I independently reproduced the reported defect at the pre-change baseline 292b85e90 (a fresh connect wrote a user_gate todo plus an onboarding_todo_review agent todo, and quota should-run returned should_run=false, effective_action=operator_gate_notify, mode=user_gate) and the fix at the reviewed head (zero todos, should_run=true, effective_action=normal_run, mode=bounded_delivery, and the caller's own [P0] todo selected as recommended_action and agent_channel.primary_action) by driving the real python -m loopx.cli entrypoint through both revisions with identical synthetic fixtures. Repository reuse, default-off isolation (not_applicable: this is a disclosed default change, not opt-in), authority semantics, typed-state, domain-neutrality, guidance-vs-obligation and behavior-change-disclosure evidence are all verified; the deletion leaves no parallel implementation and no remaining reader of the removed fields. Validation: examples/project/first-connect-contract-smoke.py, examples/control_plane/fresh-project-onboarding-regression-smoke.py (30/30), ten companion smokes, git diff --check, and the touched test files (254 passed / 4 skipped / 1 failed on an environment-only Node/SQLite qualification check that is green in CI; 29/29 checks passed on this head). Residual risk is compatibility, not correctness: the removal of eight public CLI options, the onboarding_* bootstrap payload fields and the new-user-onboarding-lifecycle canary profile rename are breaking for external callers and must be named in the 1.0.5 release notes, and the Docker-based harness path plus the installed-artifact path were not executed locally.
Why
loopx bootstrap/loopx connectused to project a generated onboarding queue into every freshly connected goal:user_gatetodo (task_class=user_gate,action_kind=onboarding_decision),action_kind=onboarding_todo_review,onboarding_connection_validationagent todo for generic adapters.So the default
connectlanded instate=operator_gatewithshould_run=false, and an autonomous caller had to pass--accept-onboarding-agent-todos --begin-autonomous-advance --codex-app-heartbeat no(or--no-onboarding-scan) just to reach its own first delivery todo. The SWE-marathon harness comments record the concrete incident: a run spent ~900s on the generated onboarding item instead of its task.Connection should be a registration step. The caller — or the connected domain adapter — owns the first delivery todo.
What changed
Runtime (
loopx/)connect/bootstrapnow write the registry entry and active state only. No onboarding scan, no candidate todos, no generated user/agent todos, no heartbeat opt-in gate, noconnection_validationadapter annotation. Deletedloopx/onboarding.py.--no-onboarding-scan,--accept-onboarding-agent-todos,--begin-autonomous-advance,--codex-app-heartbeat,--onboarding-connection-validation,--onboarding-max-{commits,status-paths,top-level-files}and the corresponding bootstrap payload/markdown fields.## Next Actionnow carries one neutral statement (Initial routing is owned by the connected domain adapter.) instead of an onboarding prompt.new-user-onboarding-lifecycle→first-connect-contract; semantic inventory regenerated.Benchmark harness — the SWE-marathon bootstrap gate string is gone;
_bootstrap_gatesnow returns only the optional--write-scope.Docs — integration guide, book chapter 05 (zh/en),
new-project-codex-prompt.md, and the interaction-pattern catalog record the registration-only contract.Validation — deleted the two smokes that encoded the removed queue and added
examples/project/first-connect-contract-smoke.py, which drives the real CLI and asserts the new contract from the outside.Affected default-behavior lanes (disclosure)
A default behavior change is intended here. Affected lanes:
connect/bootstrapstate projection (no generated user/agent todos),operator_gate→eligible,effective_action=normal_run),Validation
Focused, all on the rebased branch:
python3 examples/project/first-connect-contract-smoke.py— passpython3 examples/control_plane/fresh-project-onboarding-regression-smoke.py— pass (30/30)python3 examples/agent-diagnose-packet-smoke.py— passpython3 examples/bootstrap-command-pack-smoke.py,examples/project/project-prompt-smoke.py,examples/project/bootstrap-force-preserve-todos-smoke.py,examples/cli-bootstrap-rollout-helper-command-modularization-smoke.py,examples/control_plane/todo-claim-lease-roadmap-smoke.py,examples/claude-goalmode-lifecycle-smoke.py,examples/canary/catalog-planner-smoke.py— passpython3 examples/fresh-clone-quickstart-smoke.py— passpython3 examples/benchmark-native-goal-installed-profile-smoke.py --allow-dirty-source— passpython3 -m pytest tests/control_plane/test_goal_objective_projection.py tests/control_plane/test_todo_next_action_settlement.py tests/control_plane/test_goal_handoff_mode.py tests/control_plane/test_shadow_writer_boundaries.py tests/control_plane/test_start_goal_compact_projection.py— passpython3 -m pytest tests/control_plane/test_todo_projection_recovery.py— passpython3 -m pytest tests/control_plane/test_shadow_observable_e2e.py tests/control_plane/test_onboarding_model_behavior_qualification.py tests/control_plane/test_actual_default_model_behavior_portfolio.py tests/control_plane/test_doubao_onboarding_model_behavior_actor.py— 87 passedloopx canary premerge --from-git-diff --timeout-seconds 400— 0 failures across diff hygiene, changed-filepy_compile, catalog canaries, risk-profile smokes, and boundary checks; gate reportsmanual_review_required(benchmark-sensitive surface).Skips / holds:
examples/shared-goal-authority-e2e/installed.pywas updated for the removed flags but not executed locally (it requires a built wheel/sdist artifact and an isolated venv; covered bypy_compileand the installed-profile smoke).benchmark-native-goal-installed-profile-smokeneeds--allow-dirty-sourcebefore the branch is committed; it should be rerun clean by CI.examples/install-local-smoke.pyexceeds the 120s default canary timeout in this environment (~2m07s on a pristineorigin/maincheckout too), which is why the premerge run above used--timeout-seconds 400.Follow-up (not in this PR)
The connection-time
Next Actionplaceholder is a plain bullet, sodiagnosestill surfaces the genericbacklog_hygiene_warning(hidden_backlog_without_agent_todo) until the caller writes its first agent todo. This is pre-existing behavior for scan-disabled connections and is advisory only — it does not gateshould_run— so it is left as a separate decision.