Skip to content

fix(chat): settle queued requests when runtime preparation fails - #5145

Merged
huangruiteng merged 1 commit into
mainfrom
codex/steward-message-continuity
Sep 27, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/steward-message-continuity

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

An accepted queued Chat request could remain runnable after workspace or adapter preparation raised an exception. The queue thread died before claiming the request, so clients waited until their answer timeout and reported a generic failure without settling the queued work.

This change settles preparation failures through the shared fenced Turn lifecycle, preserves cancellation, releases the active claim and prevents same-ingress redelivery from rerunning failed work. A fresh request can proceed after runtime repair. Lark reports runtime initialization/restoration failures distinctly; unexpected exception details stay in local diagnostics.

The existing Chat runtime owns this behavior for all queue callers. No manager-only scheduler, new capability or parallel TS authority is introduced. The S1/S10 roadmap and conversation RFC now include pre-dispatch failure and recovery acceptance. The existing App event consumer already renders turn.failed messages; its schema and packaged assets are unchanged. Successful owner selection and receiver adoption remain separate acceptance work.

Validation:

  • 115 tests pass across queue preparation, active-turn concurrency, startup isolation and Lark delivery suites.
  • The removed-release regression fails on the base with a dead worker and timeout, and passes on this head without starting a provider.
  • Focused Ruff, configured mypy and six-path public-boundary scan pass.
  • All 18 selected premerge checks and direct checks pass. The first run lacked npm dev dependencies; those were installed and the checks rerun. The aggregate then reported a stale quality receipt because origin/main advanced during validation; the two intervening commits have no changed-file overlap, and exact-scope quality was re-reviewed, recorded and verified against the new base. No checks were waived.

No live model call or external group message was used for validation. Maintainer merge and rollout remain pending.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

English verdict: APPROVE - The exact head settles accepted preparation failures through the existing fenced Chat lifecycle and preserves cancellation, ingress replay and fresh recovery.

Reviewed head: 0aec8a7b38b1feba2fd18144c466c23814bfe43c; immutable pre-change baseline: cd32b28567216c9621bb86061ab1ae1b97724491. No blocking finding.

动机

这是一个有用、完整的准备阶段修复:请求已被接受,但工作区或执行器准备抛错时,旧代码的队列线程会退出,请求仍然可运行,用户只能等到回答超时。安装恢复后,旧请求还可能被下一次唤醒意外执行。本次改动让每个已接受请求获得真实的终态,并明确“修复环境后发起新请求”的恢复路径。它完成的是共享 Chat 准备边界,不代表管家的 owner 选择、receiver 接收或完整 A24 旅程已经验收。

改动思路

入口仍是 enqueue_turn 与 resume_session_queue,事实来源仍是持久化的 Session、Turn 和 ingress identity。准备失败后,队列先使用原有 FIFO claim,再委托现有失败 owner 写入终态、事件和释放 claim;恢复已有 active Turn 也走同一 owner。只捕获并打印异常会留下 runnable 请求,另建管家状态机则增加重复 authority;这里扩展现有 owner 是足够小的修复。健康路径继续使用原有执行逻辑,普通执行失败仍只允许 starting/running,准备失败才显式包含 queued。

具体改动

生产代码涉及共享 Chat 队列和 Lark 错误呈现,测试覆盖准备失败、并发终态和用户可读恢复;RFC 与 roadmap 记录这次可独立回滚的交付边界。全量 diff 为 6 个文件 +233/-13,其中生产 +48/-12、测试 +161、文档 +24/-1,没有新 CLI、存储格式或能力开关。

关键代码讲解

  • ChatRuntimeController._drain_session_queue(loopx/chat_runtime.py:1096):在恢复 active Turn 与准备新队列请求两处捕获异常;新请求仍先由 canonical FIFO claim,随后逐个失败并继续处理队列,避免接受后无结果。
  • _fail_queue_preparation(同文件 1184):保留 typed provider error code 与 gate;意外本地错误使用可操作的 runtime_unavailable,完整 traceback 留在本地。两个真实调用分支共用这个小的错误边界。
  • _fail_turn(同文件 1417):复用 store 的状态 fence。默认允许状态保持原样,准备调用显式加入 queued;若并发取消或完成已经获胜,更新返回空,后续 failure event、error message 和 release 都不会执行。
  • ChatSessionStore.release_active_turn(loopx/chat_store.py:467,未改动):只释放匹配当前 Turn 的 claim。这是复用的状态 owner,不是新增管家 scheduler。
  • manager_failure_reply(loopx/extensions/lark/manager_context.py:37):增加本地环境准备失败和原会话恢复失败两条中文修复提示,未知异常的原有呈现边界保留。

Dashboard 的现有 data/chat.ts 消费 turn.failed 并抛出含完整 payload 的 ChatApiError;本 PR 保持事件格式,因此无需新 frontend schema 或设置入口。这里确认了实际 consumer 源码和共享回归路径,没有将源码检查算作打包 UI 实测。

对主干的风险

最强反例是“测试注入的异常通过,但真实 release 文件丢失仍使线程退出”,以及“失败有回执,但修复后同一 ingress 又执行”。我用隔离的真实 Chat store/controller 删除实际 manager runtime asset:head 在 provider 启动前写出 failed/runtime_unavailable,queue 与 active claim 清空,provider 调用为零;修复后旧 ingress 保持 failed,仅新请求执行一次。旧代码在同一输入下仍为 queued,恢复后会执行旧请求和新请求。四个 FIFO 请求、已有 active claim、另一健康 Session、private error、取消/完成竞态也覆盖了。

115 项原生 Chat/Lark 回归通过,8 个独立场景全部通过;同一脚本在 immutable base 上 5 个变化场景失败、3 个健康/终态控制通过。Ruff 覆盖全部 4 个改动 Python 文件,仓库配置的 mypy 20 个源文件通过。风险检查执行 18 项选择检查和 5 项直接检查,全部通过,无 skip 或 manual hold;全 6 个公开路径扫描无泄露命中。未查询或等待远端 CI。

语义与 CI 对齐

此变更复用原有 Turn 状态、typed error 和 claim owner;RFC 明确披露准备失败由“线程退出、请求滞留”变为“持久化终态”。没有新增可选能力、默认关闭声明或更广 actor/authority 协议,因此无须构造虚假的 feature-off 验收。并发 terminal fence 与健康 FIFO 的 base/head 比较支持兼容判断。风险检查中的语义词汇检查通过;更广的 TS turn-driver 迁移仍须保留这套矩阵。

我的整体评价

APPROVE,未发现阻断问题。长期推进改善在于失败请求不再积累并在之后突然重放;用户体验改善在于及时、真实的失败和明确的新请求恢复。失败即终态意味着修复后需要显式新请求,这符合已接受的身份重放契约,独立场景也实际验证了恢复。当前 diff 没有额外持久状态或平行 decoder;未来改动考虑已落实为共享 fenced failure 的复用和两个准备分支的局部整合,进一步迁移不应塞入这个修复。

质量回执 cqr_909fd99f22f6084173f2 与全 6 文件的 base/head 匹配,scope fingerprint 为 909fd99f22f6084173f287d7971617bc2923c3559191a36b5a42a4f15dfe8478,没有未解决风险。验证边界是隔离真实状态和准备路径,成功 provider transport 为合成输入;实际服务重装、真实模型、Windows 和打包浏览器恢复尚未资格化,仍属 原 RFC 的 operational/full journey acceptance。本评审只批准该准备边界,不关闭父 milestone;运行时变更交由 maintainer 合并。新 head、共享调用方或交付主张变化需要重新评审。

@huangruiteng
huangruiteng merged commit 13c614a into main Sep 27, 2026
27 of 33 checks passed
@huangruiteng
huangruiteng deleted the codex/steward-message-continuity branch September 27, 2026 03:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant