Skip to content

test(dashboard): let the steward picker assertion settle and name its resolution - #4607

Merged
huangruiteng merged 1 commit into
mainfrom
codex/dashboard-acceptance-picker-settle
Sep 16, 2026
Merged

huangruiteng merged 1 commit into
mainfrom
codex/dashboard-acceptance-picker-settle

Conversation

@huangruiteng

Copy link
Copy Markdown
Collaborator

What this changes

One file: examples/personal-workspace-browser/execution-chip.mjs.

The steward-picker assertion read the Goal-composer runtime label once, immediately after the page opened:

const pickerLabel = (await stewardPicker.page
  .locator("div.personal-agent-select button.personal-select-trigger")
  .innerText()).replace(/\s+/g, " ").trim();
if (!pickerLabel.includes("DeepSeek Harness (managed)")) {
  throw new Error(`Chat runtime picker ignored the declared steward executor: ${pickerLabel}`);
}

That label is resolved from the same fetchChatCapabilities() response that carries the channel binding, and discoveredAgents falls back to a single Codex entry while runtimeAgents is still empty. The execution-chip assertions never observe the fallback because the chip is only rendered once the binding lands (!selectedGoal && managerChannelBinding), so waitFor(".personal-execution-chip") implicitly waits for the response; the picker renders immediately and has no such wait.

The assertion now waits for the declared value and, if it never arrives, fails with the three facts that decide the label plus a first-viewport screenshot:

Error: Chat runtime picker never settled: label=Chat Codex declared=dsh \
  adapters=[codex:available, claude-code:available, offline-agent:unavailable, dsh:available] capabilities=200

Why

The required dashboard-acceptance check fails on main at exactly that line (Chat runtime picker ignored the declared steward executor: Chat Codex), and that failure fails checks, pytest and merge-gate for every open PR. The current message carries neither the adapters the page resolved nor a screenshot, so the CI-only failure could not be attributed from the log: the full personal-workspace catalog passes on this machine at e66615d33 with and without coverage, and the same failure does not reproduce locally at all.

This does not weaken the assertion — it still requires the same resolved value, and it still rejects a picker that advertises a discovered CLI. It stops asserting on a state the page never promised was settled, and it makes the next CI failure self-describing instead of requiring a local reproduction that may not exist.

Validation

$ LOOPX_DASHBOARD_COVERAGE=1 node examples/personal-workspace-browser-smoke.mjs
scenarios=navigation-sorting,chat-recovery,typed-actions,team-plan,execution-chip,progressive-loading
personal-workspace-browser-smoke (development): ok

$ LOOPX_PERSONAL_WORKSPACE_SCENARIO=execution-chip ...        # expectation forced to never match
Error: Chat runtime picker never settled: label=Chat DeepSeek Harness (managed) declared=dsh
  adapters=[codex:available, claude-code:available, offline-agent:unavailable, dsh:available] capabilities=200

The second run is a deliberate negative check that the new diagnostics path reports the real resolution instead of a bare label.

Boundary

Smoke-only change to a public example: no runtime, scoring, permission or scoring-semantics change, no private state, no launch of any benchmark job.

…esolution

The steward-picker assertion in the personal-workspace browser smoke read the
Goal-composer runtime label once, right after the page opened. That label is
resolved from the same capabilities response that carries the channel binding,
so a single read can still observe the pre-fetch `Chat Codex` fallback. The
execution-chip assertions never hit this, because the chip only renders once the
binding lands while the picker renders immediately.

The assertion now waits for the declared value, and when it never arrives it
fails with the three facts that decide the label -- the rendered label, the
declared endpoint and the adapters the page actually received -- plus a
first-viewport screenshot. The required `dashboard-acceptance` check fails on
this line in CI (`Chat runtime picker ignored the declared steward executor:
Chat Codex`) and that message carried neither the resolved adapters nor an
artifact, so it could not be attributed from the log alone.

Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
@huangruiteng

Copy link
Copy Markdown
Collaborator Author

Review record — exact head 0c7d235d3

Change under review. One file, examples/personal-workspace-browser/execution-chip.mjs (+65/−12). It replaces two one-shot picker-label reads with waitForPickerLabel, and adds pickerResolution so a non-settling picker fails with the facts that decide the label instead of a bare string. The assertions themselves are unchanged in strength: the declared steward executor is still required, a discovered CLI is still rejected, and an undeclared machine must still resolve to the shipped Chat Codex default.

Root cause this establishes. The picker label is resolved from the same fetchChatCapabilities() response that carries manager.channel_binding, and discoveredAgents falls back to a single Codex entry while runtimeAgents is empty. The execution-chip assertions never observed that fallback because the chip only renders once the binding lands (!selectedGoal && managerChannelBinding), so waitFor(".personal-execution-chip") implicitly waits for the response. The picker renders immediately and had no such wait, which is exactly why the last two main runs failed on this line while the same tree passes locally.

Evidence.

Where Result
main e66615d33 CI (runs 35131695171, 35129353025) dashboard-acceptance fails at execution-chip.mjs:221, Chat Codex
main e66615d33 local, full catalog, with and without coverage pass
this head, full catalog with coverage pass
this head, CI dashboard-acceptance pass (was failing on main for the same tree)
this head, CI checks / pytest / merge-gate pass
forced negative check (expectation made unsatisfiable) Chat runtime picker never settled: label=Chat DeepSeek Harness (managed) declared=dsh adapters=[codex:available, claude-code:available, offline-agent:unavailable, dsh:available] capabilities=200
loopx canary premerge --from-git-diff ok — 4 catalog canaries, public boundary, diff checks

Failures and holds. None. presentation is skipped by the workflow, as on main. The smoke still asserts the same resolved value, so it cannot mask a picker that reports the wrong runtime; the only behaviour change on failure is that the message is now attributable without a machine that reproduces it.

Why this coverage is enough. The change is confined to one public browser-smoke example; there is no runtime, permission, scoring, benchmark-scoring or submission behaviour, no private state and no benchmark launch. The decisive validation is the required check itself going green on an unchanged main tree, which is also what the repository's merge-gate consumes.

Approval conclusion: ready to merge at 0c7d235d3. Verdict: single-file smoke repair with the required check repaired on the same tree; unblocks every open PR whose merge-gate was inheriting this failure.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self-review record on exact head 0c7d235d3235f3a0068592185eececeb3e990a57.

Approval conclusion: ready to merge at this head. Single-file, test-only repair to examples/personal-workspace-browser/execution-chip.mjs (+65/-12): the steward-picker assertions now wait for the declared value instead of reading once, and a picker that never settles fails with the rendered label, the declared endpoint, the adapters the page actually received and a first-viewport screenshot. Assertion strength is unchanged.

Decisive evidence: the required dashboard-acceptance job failed on main e66615d33 at this exact line in two consecutive runs (35131695171, 35129353025) while the same tree passes locally, and it passes on this head. checks, pytest and merge-gate — which were failing only by inheritance from that job — are green on this head, and loopx canary premerge --from-git-diff returned ok (4 catalog canaries, public boundary, diff checks). No failures, no skips beyond the workflow's own presentation skip, no manual holds.

Boundary: public browser-smoke example only — no runtime, permission, scoring, submission or benchmark-launch change, and no private state. This is what unblocks every open PR whose merge-gate was inheriting the failure.

Verdict: approved for merge at 0c7d235d3235f3a0068592185eececeb3e990a57; the required check is repaired on the same tree it was failing on.

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approval conclusion (author-owned PR; GitHub blocks formal self-approval)

Exact head: 0c7d235

动机

main 的必需检查 dashboard-acceptance 在连续两次运行里红在 examples/personal-workspace-browser/execution-chip.mjs:221(Chat runtime picker ignored the declared steward executor: Chat Codex)。因为 checks 继承该结果,pytest 与 merge-gate 随之对每个 PR 全红,而同一棵树在本机跑完整 personal-workspace 目录是绿的。原失败信息只带一个 label,既没有页面实际解析到的 adapters,也没有首屏截图,所以在拿不到可复现环境时无法归因。

改动思路

断言不应读取一个页面从未承诺已经稳定的状态。picker 的 label 由携带 manager.channel_binding 的同一次 fetchChatCapabilities() 响应解析;当 runtimeAgents 仍为空时,discoveredAgents 会退化成单个 Codex 条目。execution chip 从未观察到该回退,是因为 chip 只在 binding 到位后才渲染(!selectedGoal && managerChannelBinding),于是 waitFor(".personal-execution-chip") 等价于等待该响应;picker 则立即渲染,缺少这层等待。因此把一次性读取改为等待声明的取值,并让"始终不收敛"以可归因的方式失败。

具体改动

仅 examples/personal-workspace-browser/execution-chip.mjs(+65/−12):

  • 新增 pickerSelector / readPickerLabel,把选择器收敛到一处;
  • 新增 waitForPickerLabel(page, settled, timeoutMs = 15_000):在 15s 内轮询直到收敛,超时抛出 Chat runtime picker never settled: ...;
  • 新增 pickerResolution(page):在页面内重新读取 /api/chat/capabilities,把已渲染 label、声明的 endpoint、页面实际收到的 adapters(含 available/unavailable)与响应状态拼进失败信息,并落一张首屏截图 execution-chip-picker-unresolved.png;
  • 两个 picker 断言改用它:声明 dsh 的机器必须收敛到 DeepSeek Harness (managed) 且不得出现 Codex;未声明 steward executor 的机器必须收敛到出货默认 Chat Codex。断言强度不变。

对主干的风险

风险面是 examples/ 下的一个公开浏览器 smoke,不触碰 runtime、权限、评分、提交或 benchmark 启动路径,也不含私有状态。唯一行为变化是失败信息更具体、以及等待页面收敛;它无法掩盖"picker 报告错误 runtime",因为要求的取值没有放宽。真正需要单独评审的运行时行为为零。

我的整体评价

Approve。这修在同一棵树、同一行原本失败的位置:dashboard-acceptance 在 main e66615d33 上红、在该 head 上绿,checks / pytest / merge-gate 随之全绿,loopx canary premerge --from-git-diff ok(4 项 catalog canary + public boundary + diff 检查)。本地完整目录(含与不含 coverage)通过,另做过一次负向验证,确认新的失败信息会打印真实解析结果而非裸 label。它同时解开所有因继承该失败而卡住的 PR。

English verdict: APPROVE

@huangruiteng
huangruiteng merged commit ecd1df5 into main Sep 16, 2026
23 checks passed
@huangruiteng
huangruiteng deleted the codex/dashboard-acceptance-picker-settle branch September 16, 2026 21:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant