fix(reward-memory): preserve verified context delivery across assessment - #5158
Conversation
Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
评审结论:APPROVE,无阻塞发现。精确 head:1d686a9bb1bbbbc0aa6a31e8556a465a4059f1f4;基线:420782f03725bf9b7603be481f2b0525beff5807。这是修复已验证交付事实丢失的 SDK 增量,不表示记忆增益、全部宿主接入或 Goal 完成。
动机
现有 query-ready SDK 已把 context delivery 和语义判断分开,但后续 assessment 从默认 false 的 packet 重建投影,覆盖先前已验证的交付。应用、忽略、反驳都触发这一错误。使用相同 SDK 测试和公开 fixture,在隔离的基线运行时得到 4 个该断言失败、19 个既有分支通过;本 head 的全部 27 个 SDK 测试通过。修复后调用方能诚实保留交付事实并继续判断,不需重新召回,也不把 applied 当成 utility。
改动思路
最小修复应留在现有 TS projection owner,而不是让 Python 复制决策,或新增存储、迁移、配置和重查 provider。这里复用已有 receipt 绑定约束,私有结果保留原 context receipt;Python 仅传现有白名单字段,TS 重新核对应用、当前产物、surface 与经验摘要。交付是先前 callback 的事实,disposition 是后续判断,两者独立。失败判断保持原输出并允许使用同一条目重试。旧 hook、公开 packet 格式和配置 owner 不变。#5154 的输入 checkpoint/诊断修复不属于这条语义链,也不是本 PR 的依赖。
具体改动
全量差异为五个文件:生产 Python 17 增/7 删,生产 TS 25 增/10 删;SDK 测试增加 40 行,TS 测试增加 47 行,双语文档增加 15 行。没有生成代码或迁移机制。
关键代码讲解
boundReceiptDigests(TS 第 29 行)集中原有 schema、application/artifact/surface、outcome、readback/current-artifact 和最多 8 个摘要的校验。同一规则用于当前语义 receipt 和历史 context receipt,没有第二套放宽规则。projectRewardMemoryDecision(TS 第 73 行)在语义分支独立计算交付事实。当前归因必须属于原交付条目,错误应用、产物、surface 或经验不能借用交付 proof;缺失 proof 的直接 semantic callback 仍可正常完成其判断,但 context 标记为 false。无效判断只保留合法的历史交付,不宣称消费完成。RewardMemoryDecisionResult与_project(Python 第 24、78 行)增加 caller-private 原回执槽位,复用紧凑 allowlist;正文、查询、推理和基线不进入 TS/public packet。字段默认 None,不能用老公开 packet 自动补造交付历史。run_reward_memory_decision/assess_reward_memory_decision(Python 第 95、169 行)在 TS 确认初次交付后保留原回执,随后评估仍使用原 qualified items 和累计计数。完整结果 replay 不再调用模型;不完整评估后的有效重试也不再查询 provider。文档明确此 proof 仅覆盖 SDK callback,不覆盖前端、飞书或收益。
对主干的风险
最危险的反例是给合法语义判断附上其他应用或其他条目的交付回执,造成 false positive。TS 负例检查九类绑定偏差及畸形对象;缺失/不匹配 proof 不会产生交付标记,原合法直接判断保持可用。SDK 测试同时执行失败→重试→精确 replay,保留累计双 corpus 2 次调用、1 条过滤,单 corpus 重试仍为 1 次。关闭配置仍返回 None、没有 TS/provider 调用;preview、缺策略、空/过滤/不可用、错误 scope、模型失败与 TS 故障保持原 fail-open 输出和实际计数。
原应用 owner 仍核对回调引用来自实际召回、ignored/refuted 不改基线;此 PR 不签发任何新行动权限。TS 故障仍为 incomplete,不在 Python 推断成功。原回执只在 caller-private 结果中保留,不新增跨进程持久化。无需新前端控件:现有配置未变,这次实际入口是已导出的可选 SDK;Dashboard/Lark 原状态投影与页面包未改,本评审也不冒充验证了通用 UI/Lark 消费。
语义与 CI 对齐
复用现有 delivery、semantic disposition、utility 词汇及 TS 单一权威,修正 projection 丢失事实这一显式语义差异,不把它归一化成 baseline parity。按当前 Goal 评审契约不读取或等待远端 CI;本地 SDK 27 项、TS 8 项、两组相邻回归 76/80 项、TS typecheck、Ruff 和五路径公开扫描通过,预合并 17 项无失败/人工 hold。无 authority/store 改动,不要求无关的 provider 三存储晋级验证。
我的整体评价
long_horizon 改善:后续判断和失败重试保留真实交付及计数,避免误导重查;user_experience 改善:SDK 投影不再自相矛盾,未增加配置步骤。生产净增加 25 行,保持初次 plan/project 两次与 assessment 一次传输,没有额外 provider 调用;相比新 cache 或 Python 决策分支,成本与缺陷相称。缺失原私有回执时仍保守不补历史,是有证据的兼容边界,不是隐含新版 decoder。
该独立 SDK 修复可以合入,但运行时改动仍需维护者审核,不自合并。剩余缺口是 caller-owned 跨进程恢复、通用宿主回传及真实配对效果验收;没有 live OpenViking、live model、金融 utility 或升级资格的结论。合入后再升级,并由原调用方沿实际接入路径验收,不重写旧回执。
English verdict: APPROVE - head 1d686a9; verified SDK context-delivery lineage survives assessment and retry without extra recall, while foreign proof and utility promotion remain rejected. SDK 27, TS 8, adjacent 76/80 and premerge 17 checks passed; live provider/model and universal UI/Lark delivery are not claimed.
Goal And Delivered Outcome
context_delivery_verifiedwithfalse.420782f03725bf9b7603be481f2b0525beff5807; the fixed head preserves the separately verified delivery while retaining the actual disposition, original counters andutility_verified=false.main. Related fix(reward-memory): share surface checkpoints and typed input diagnostics #5154 addresses input checkpoints/diagnostics and is not required by this fix.Scope And Continuation
Validation
1d686a9bb1bbbbc0aa6a31e8556a465a4059f1f4reward-memory-scoped-feedback-ingest.public.json, SHA256ed0fe10fc638d2f2f7a88f25c63cfc15fa7892ff31c54f49dfda4e2d5a215061.uv run --extra test python -m pytest -q tests/capabilities/test_reward_memory_decision.py— 27 passed; real production TS transport and SDK, synthetic metered provider. Covers two-corpus counters, exact replay, disabled isolation, fail-open and failed-assessment recovery without recall.node --experimental-strip-types --test tests/control_plane_ts/reward_memory_decision.test.ts— 8 passed; identity/artifact/surface/lesson mismatches, malformed proof, no-proof direct assessment and incomplete assessment.uv run --extra test loopx canary premerge --from-git-diff— 17 selected checks passed: 8 catalog canaries, 8 risk-profile smokes and 1 public-boundary check; zero failures or manual holds. Passing validation does not grant merge authority.Frontend / Visual Evidence
Type of Change
LoopX Area / Technical Direction
reward_memory.decision.projectin TypeScript. Python adapter +10 net LOC; TS +15 net LOC; no parallel Python policy. Initial SDK path remains two transport calls, assessment remains one; no extra provider call or migration scaffolding.Shared-authority RFC fixture impact
N/A: no coordination schema, lease/quota/authority provider, runtime routing or promotion change. Existing shared-authority fixture and store arms are unaffected.
Boundary Checklist