test(steward): walk the frontend steward journey and record its gaps - #4588
Conversation
The steward lane had a CLI-level qualification record but no frontend-first journey: nothing proved what an owner actually sees from the first screen through asking the steward for work, reading the admitted team plan card, and confirming it. Add a `steward-journey` scenario to the existing personal-workspace browser smoke catalog that walks those beats in order on the workspace first screen and asserts the frontend-visible facts: the Goal board lanes, the Goal conversation accepting the owner's ask, the plan card naming each lane's Agent, first bounded Todo, acceptance and staffing gap, and confirming sending exactly one apply with one durable write. Beats the surfaces do not yet prove are recorded as typed gaps with the probe that looked for them instead of being asserted away: the steward prompt set is defined in the client model but not reachable from the conversation, the confirmation outcome does not distinguish committed/partial/all-gap/stale per lane, and no per-lane readiness ladder, lane-level correction, blocker ownership, or completion-by-return is visible. The scenario emits `steward-journey-report.json` for the follow-up product case. Everything is synthetic: the fixture substitutes the agent turn, and no live Goal, agent, workspace or local path is read or captured. Validation: - `LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey node examples/personal-workspace-browser-smoke.mjs` - packaged: same with `LOOPX_PERSONAL_WORKSPACE_PACKAGED=1` - full catalog: `node examples/personal-workspace-browser-smoke.mjs` Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
The lane's acceptance is that the steward journey is reproducible from the
*packaged* frontend on synthetic data, but the report could not tell a packaged
run from a development-server run: both produced the same
`{beats, gaps, scenario}` file and the same console line. A packaged
acceptance therefore rested on who happened to run it rather than on an
artifact.
The report now carries `mode`, `served_root` and the served `url`, the console
line prints `mode=packaged|development`, and a packaged run asserts it served
the built bundle (`/chat/`) so a packaged run that silently served source
fails instead of passing quietly.
Recorded on this branch: `mode=packaged` and `mode=development` both report
`beats=3` with gaps `2-steward-prompts,3-confirm,4-readiness,5-correction,
6-recovery,7-return`, and the full development suite still passes all seven
scenarios.
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
8b2d3bd to
e0e15a1
Compare
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
Exact head: a9bcd65
动机
owner 要求有一条从前端出发的 steward 旅程验收:把已验证的环节证明出来,把未证明的环节如实记成 typed gap,而不是靠聊天记忆。这个 PR 是那条验收的可执行一半;产品侧案例文档由 #4589 提供(已合并)。
改动思路
在既有的个人工作区浏览器 smoke 里新增一个确定性场景,跑在合成数据上、由 fixture 代替 agent turn,因此它断言的是工作区真实渲染出来的事实。场景同时记录 mode / served_root / url,使得 packaged 模式能证明自己服务的是构建产物而不是源码;未证明的 beat 以探针找过什么、实际发现了什么的方式记录,而不是写成未实现。
具体改动
只改 examples 下两个文件:新增 scenarios/personal-workspace-browser/steward-journey.mjs,并在 personal-workspace-browser-smoke.mjs 的目录里登记它。
对主干的风险
测试专用:不触碰 runtime、权限、评分、提交或 benchmark 启动路径,不含私有状态、凭据或内部上下文。它只读工作区渲染结果,因此不可能改变产品行为;真正的运行时行为需要单独评审的部分为零。
我的整体评价
Approve。本 head 只多加了一条把已更新 main 合入的、带 sign-off 的合并提交(a9bcd6501),场景内容不变。证据:本机在该 head 上实测 development 与 packaged 两种模式均通过,三条已证明 beat 与五条 typed gap 一致;CI 该 head 23 项检查全绿。
English verdict: APPROVE
Goal And Delivered Outcome
Source: local steward lane (codex-local-steward) under the overall roadmap, G0/S1/S12; owner direction on 2026-09-17 asks for a frontend-first steward journey plus product-side usage patterns.
Before: the steward lane had a CLI-level qualification record only. Nothing proved what an owner sees from the workspace first screen through asking the steward for work, reading the admitted team plan card, and confirming it.
After: the personal-workspace browser smoke catalog gains a
steward-journeyscenario that walks those beats in order on the real workspace surfaces and asserts the frontend-visible facts.Scope And Continuation
Changed surfaces:
examples/personal-workspace-browser/steward-journey.mjs(new scenario) and the scenario catalog line inexamples/personal-workspace-browser-smoke.mjs. No product/runtime code, no protocol, no scoring, no permission change.Asserted beats:
Beats the surfaces do not yet prove are recorded as typed gaps with the probe that looked for them, not asserted away:
stewardPrompts) but not reachable from the conversation;The scenario writes
steward-journey-report.json(beats + gaps + probe evidence) into the smoke output dir for the follow-up product case (English canonical + zh-CN mirror, next slice).Public And Private Boundary
Synthetic only: the fixture substitutes the agent turn, screenshots and the report are written under the gitignored
output/playwright/path, and no live Goal, agent, workspace, credential or local path is read or captured. A marker scan over both changed files is clean.Validation
LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey node examples/personal-workspace-browser-smoke.mjs→ok, beats=3, gaps reportedLOOPX_PERSONAL_WORKSPACE_PACKAGED=1 LOOPX_PERSONAL_WORKSPACE_SCENARIO=steward-journey node examples/personal-workspace-browser-smoke.mjs→oknode examples/personal-workspace-browser-smoke.mjs→ok(navigation-sorting, chat-recovery, typed-actions, team-plan, steward-journey, execution-chip, progressive-loading)Untested: the real agent turn behind the intake (the fixture substitutes it), Lark surfaces, and any live cohort.