Skip to content

Commit 70b3cca

Browse files
authored
Merge pull request #5154 from loopx-project/codex/reward-memory-surface-checkpoint-20260927
fix(reward-memory): share surface checkpoints and typed input diagnostics
2 parents af3e7f1 + a875085 commit 70b3cca

14 files changed

Lines changed: 449 additions & 51 deletions

File tree

‎docs/reference/reward-memory-decision-consumption.md‎

Lines changed: 60 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -26,7 +26,11 @@ Read-authority checkpoints must match the exact consumer surface and corpus;
2626
a turn-admission checkpoint cannot authorize a different review surface.
2727
`freshness_context.age_seconds`, when supplied, is a nonnegative integer.
2828
Rejected requests expose only the original hook's allowlisted
29-
`boundary_reason_code`, never exception text or private input values.
29+
`boundary_reason_code` and, for typed input errors, `boundary_detail_code`,
30+
never exception text or private input values. The details distinguish
31+
`freshness_age_invalid`, `freshness_context_invalid`,
32+
`read_authority_checkpoint_missing` and `read_authority_checkpoint_invalid`.
33+
Existing reason codes, ValueError compatibility, validation order and gates remain unchanged.
3034

3135
使用上述导出入口和 `resolve_reward_memory_experiment` 的原配置读回;不可用时
3236
不能拿未经验证的配置替代。原 hook 的范围、revision、问题、时点、时效/冲突、
@@ -36,7 +40,44 @@ Rejected requests expose only the original hook's allowlisted
3640

3741
读授权 checkpoint 必须匹配本次 surface/corpus,不能拿 Turn 准入的 checkpoint
3842
授权另一评审入口;age_seconds 如提供,须为非负整数。拒绝回执仅投影原 hook
39-
白名单内的 boundary_reason_code,不暴露异常正文或私有参数。
43+
白名单内的 boundary_reason_code,以及输入错误的 boundary_detail_code;细分年龄非法、
44+
时效上下文非法、读授权缺失和读授权格式非法,不暴露异常正文或私有参数。
45+
保留原错误码、ValueError 兼容、校验顺序与门禁。
46+
47+
Use `build_reward_memory_surface_read_authority_checkpoints(config, surface_id,
48+
verified=original_proof_verified, source_ref=original_read_authority_source)`
49+
from the same package. It selects only that surface's configured corpora through
50+
the existing configuration owner, then TS assembles the exact workspace/project,
51+
optional user/peer/session, read-authority and surface references. The caller must
52+
actually verify its original read authority: an enabled config or ingest policy
53+
alone is not read proof. `verified=False` stays false and blocks recall. It does
54+
not infer a proof source, enable the capability or contact a provider. The Turn
55+
wrapper uses this same projection and retains its verified registry source.
56+
57+
通用 helper 按实际 surface 和原配置选择 corpus,由 TS 组装精确范围;调用方仍须
58+
真实核验原读权限并显式传入 verified/source_ref,不能把开关或写入 policy 当作读授权。
59+
False 不会升级为 True;不推断授权来源、不启用能力、不调用 provider。原 Turn wrapper
60+
复用该投影并保留 registry 来源。不要以生成了 checkpoint 为由宣称授权核验已完成。
61+
62+
Checkpoint transport failure remains optional-enrichment failure: managed Turn
63+
admission returns its existing fail-open `runtime_unavailable` packet. The explicit
64+
`agent-turn-recall --execute` CLI returns a safe `runtime_unavailable` packet and
65+
exit code 2. Neither path calls the provider or writes a successful same-Turn
66+
receipt when checkpoint construction fails; a later healthy retry uses the same
67+
Turn identity. These zero-call guarantees apply before provider invocation only.
68+
69+
checkpoint 传输失败不成为普通 Turn 的新门禁:managed 准入沿用原 fail-open
70+
`runtime_unavailable`;显式 CLI 返回安全的同类 packet 和退出码 2。构建失败时
71+
均不调用 provider、不写成功的同 Turn 回执;恢复后沿用原 Turn 身份重试。
72+
零调用保证仅适用于 provider 调用前的构建失败,不覆盖调用后的异常。
73+
74+
If computing age from timestamps, first reject an observation in the future;
75+
then round elapsed seconds upward to an integer. Never clamp a negative age,
76+
refresh the original observation time, or change policy to make recall pass.
77+
This helper intentionally does not calculate or correct age for the caller.
78+
79+
由时间戳计算年龄时,先拒绝未来观察,再将经过秒数向上取整;不能截断负值、
80+
刷新原观察时间或改 policy 来过门。helper 不替调用方计算或纠正年龄。
4081

4182
TypeScript owns admission and completion (`reward_memory.decision.plan/project`);
4283
Python adapts the existing provider/applier and retains transient private values.
@@ -48,6 +89,14 @@ TS 负责准入和完成语义;Python 只适配现有 provider/applier 并保
4889
TS 只收到引用、状态、计数与摘要,不收到问题、经验正文、原产物或模型判断内容;
4990
不新增存储、SDK、密钥、开关或行动授权。
5091

92+
This slice does not migrate the existing Python SDK's scope/freshness validation;
93+
it adds no second TS admission rule for those checks. Python remains the original
94+
configuration/provider adapter and input-error source; TS owns the shared
95+
checkpoint projection and allowlisted decision diagnostics.
96+
97+
此切片不迁移原 Python SDK 的范围/时效校验,也不在 TS 复制准入规则。Python 保留
98+
原配置/provider 适配与输入错误来源,TS 持有共享 checkpoint 投影及白名单诊断。
99+
51100
| Mode / 模式 | Provider / 调用 | Meaning / 意义 |
52101
| --- | --- | --- |
53102
| Disabled/unconfigured / 未配置或关闭 | Zero; returns `None` / 零调用,无新 packet | Original path unchanged / 原路径不变 |
@@ -111,6 +160,12 @@ returns `replay_request_mismatch`. Reassessment uses retained qualified items an
111160
the original **cumulative** multi-corpus counters, not a second query. This is
112161
caller-retained replay, not automatic cross-process persistence or a new cache.
113162

163+
Retain the complete private result, not just context/public_packet/application
164+
receipt. `assess_reward_memory_decision` needs the exact recall session and
165+
attribution. A lost session after EOF/restart is incomplete, even if context was
166+
delivered; do not re-query or fabricate semantic completion. There is currently
167+
no supported cross-process restore API. Caller-owned persistence and a future
168+
validated restore contract remain separate from this in-process replay API.
114169
The private result retains the **original context-delivery receipt** separately
115170
from the later semantic receipt. TypeScript revalidates its application, artifact,
116171
surface and lesson attribution, so assessment (including an incomplete assessment)
@@ -124,6 +179,9 @@ cannot recreate that private lineage or upgrade historical receipts.
124179
`previous_result` 仅复用配置和输入均匹配的请求,变化则拒绝复用。后续判断使用
125180
原条目和累计多 corpus 遥测,不重复查询。这不是自动跨进程存储或新的缓存。
126181

182+
需保留完整私有 result,不能只存 context/public_packet/application receipt。
183+
EOF/重启丢失 recall_session 时,交付过上下文也不能完成 assessment;不重查、不补造
184+
语义完成。当前没有受支持的跨进程恢复 API,持久化与后续验证恢复合同是独立缺口。
127185
私有结果分别保留原上下文交付回执和后续语义回执,TS 对应用、产物、surface 与
128186
经验归因重新核验;评估成功或不完整均不抹掉此前已验证的交付。直接语义 callback
129187
没有该回执时仍为 `context_delivery_verified=false`,语义判断与效果另行记录。

‎loopx/capabilities/agent_turn_recall/cli.py‎

Lines changed: 21 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -244,13 +244,31 @@ def handle_agent_turn_recall_command(
244244
"experiment": experiment_status,
245245
}
246246
else:
247+
try:
248+
read_checkpoints = _read_authority_checkpoints(config, args.goal_id)
249+
except RuntimeError:
250+
# This failure precedes the provider; do not fabricate a
251+
# zero-call receipt for errors after recall has begun.
252+
payload = {
253+
"ok": False,
254+
"schema_version": AGENT_TURN_RECALL_SCHEMA_VERSION,
255+
"status": "runtime_unavailable",
256+
"reason_code": "automatic_recall_runtime_failed",
257+
"goal_id": args.goal_id,
258+
"agent_id": args.agent_id,
259+
"provider_call_count": 0,
260+
"grants_new_action_authority": False,
261+
"quota_spend_performed": False,
262+
"external_writes_performed": False,
263+
"suppress_external_sinks": True,
264+
}
265+
print_payload(payload, output_format(args), _render)
266+
return 2
247267
payload = run_agent_turn_recall(
248268
config,
249269
situation,
250270
observed_at=datetime.now(timezone.utc).isoformat(),
251-
read_authority_checkpoints=_read_authority_checkpoints(
252-
config, args.goal_id
253-
),
271+
read_authority_checkpoints=read_checkpoints,
254272
) | {
255273
"goal_id": args.goal_id,
256274
"agent_id": args.agent_id,

‎loopx/capabilities/agent_turn_recall/runtime.py‎

Lines changed: 4 additions & 21 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,7 @@
1717
resolve_reward_memory_experiment,
1818
resolve_reward_memory_surface_config,
1919
)
20+
from ..reward_memory.read_authority import build_reward_memory_surface_read_authority_checkpoints
2021
from .core import (
2122
AGENT_TURN_RECALL_SCHEMA_VERSION,
2223
AGENT_TURN_RECALL_SURFACE_ID,
@@ -149,28 +150,10 @@ def resolve_reward_memory_turn_session_ref(
149150
def reward_memory_turn_read_authority_checkpoints(
150151
config: Mapping[str, Any], goal_id: str
151152
) -> dict[str, dict[str, Any]]:
152-
route = resolve_reward_memory_surface_config(
153-
config,
154-
AGENT_TURN_RECALL_SURFACE_ID,
153+
return build_reward_memory_surface_read_authority_checkpoints(
154+
config, AGENT_TURN_RECALL_SURFACE_ID,
155+
verified=True, source_ref=f"registry:{goal_id}:reward-memory",
155156
)
156-
checkpoints: dict[str, dict[str, Any]] = {}
157-
for item in route["recall_corpora"]:
158-
corpus = item["corpus"]
159-
scope = corpus["scope"]
160-
checkpoint = {
161-
"verified": True,
162-
"corpus_id": corpus["corpus_id"],
163-
"workspace_ref": scope["workspace_ref"],
164-
"project_ref": scope["project_ref"],
165-
"surface_id": AGENT_TURN_RECALL_SURFACE_ID,
166-
"read_authority": corpus["read_authority"],
167-
"source_ref": f"registry:{goal_id}:reward-memory",
168-
}
169-
for field in ("user_ref", "peer_ref", "session_ref"):
170-
if scope.get(field):
171-
checkpoint[field] = scope[field]
172-
checkpoints[corpus["corpus_id"]] = checkpoint
173-
return checkpoints
174157

175158

176159
def deduplicated_agent_turn_recall_payload(

‎loopx/capabilities/reward_memory/__init__.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -61,6 +61,7 @@
6161
run_reward_memory_automatic_ingest_hook,
6262
run_reward_memory_automatic_recall_hook,
6363
)
64+
from .read_authority import build_reward_memory_surface_read_authority_checkpoints
6465
from .outcome_lifecycle import (
6566
reconcile_pending_turn_outcome_ingests,
6667
reconcile_pending_turn_outcome_ingests_fail_open,
@@ -73,6 +74,7 @@
7374
"RewardMemoryDecisionResult",
7475
"assess_reward_memory_decision",
7576
"run_reward_memory_decision",
77+
"build_reward_memory_surface_read_authority_checkpoints",
7678
"RewardMemoryFilteredRecallItem",
7779
"RewardMemoryRecallItem",
7880
"RewardMemoryRecallSession",

‎loopx/capabilities/reward_memory/application.py‎

Lines changed: 60 additions & 23 deletions
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77
from collections.abc import Callable, Mapping, Sequence
88
from dataclasses import dataclass
99
from datetime import datetime
10-
from typing import Any
10+
from typing import Any, Literal, get_args
1111

1212
from ...control_plane.runtime.public_safety import public_safe_compact_text
1313
from ..context_providers import build_context_provider
@@ -98,6 +98,22 @@ class _ActiveItemDecision:
9898
]
9999

100100

101+
RecallInputErrorCode = Literal[
102+
"freshness_age_invalid", "freshness_context_invalid",
103+
"read_authority_checkpoint_missing", "read_authority_checkpoint_invalid",
104+
]
105+
106+
107+
class RewardMemoryRecallInputError(ValueError):
108+
"""An existing SDK input rejection with an allowlisted, non-content code."""
109+
110+
def __init__(self, reason_code: RecallInputErrorCode, message: str) -> None:
111+
if reason_code not in get_args(RecallInputErrorCode):
112+
raise ValueError("unsupported recall input error code")
113+
super().__init__(message)
114+
self.reason_code = reason_code
115+
116+
101117
def _token(value: object, label: str) -> str:
102118
result = str(value or "").strip()
103119
if not TOKEN_RE.fullmatch(result):
@@ -243,22 +259,17 @@ def _authority_checkpoint(
243259
raw: object, *, corpus: Mapping[str, Any], request: Mapping[str, Any]
244260
) -> tuple[dict[str, Any], list[str]]:
245261
if not isinstance(raw, Mapping):
246-
raise ValueError("read_authority_checkpoint must be an object")
247-
checkpoint = {
248-
"verified": _boolean(raw, "verified"),
249-
"corpus_id": _token(raw.get("corpus_id"), "checkpoint.corpus_id"),
250-
"workspace_ref": _token(raw.get("workspace_ref"), "checkpoint.workspace_ref"),
251-
"project_ref": _token(raw.get("project_ref"), "checkpoint.project_ref"),
252-
"surface_id": _token(raw.get("surface_id"), "checkpoint.surface_id"),
253-
"read_authority": _token(
254-
raw.get("read_authority"), "checkpoint.read_authority"
255-
),
256-
"source_ref": _optional_token(raw.get("source_ref"), "checkpoint.source_ref"),
257-
}
258-
for field in IDENTITY_SCOPE_FIELDS:
259-
expected_scope = corpus["scope"].get(field)
260-
if expected_scope:
261-
checkpoint[field] = _optional_token(raw.get(field), f"checkpoint.{field}")
262+
raise RewardMemoryRecallInputError(
263+
"read_authority_checkpoint_missing" if raw is None else "read_authority_checkpoint_invalid",
264+
"read_authority_checkpoint must be an object",
265+
)
266+
try:
267+
checkpoint = _normalize_authority_checkpoint(raw, corpus=corpus)
268+
except ValueError as exc:
269+
raise RewardMemoryRecallInputError(
270+
"read_authority_checkpoint_missing" if not raw else "read_authority_checkpoint_invalid",
271+
str(exc),
272+
) from exc
262273
reasons: list[str] = []
263274
expected = {
264275
"corpus_id": corpus["corpus_id"],
@@ -283,22 +294,48 @@ def _authority_checkpoint(
283294
return checkpoint, reasons
284295

285296

297+
def _normalize_authority_checkpoint(
298+
raw: Mapping[str, Any], *, corpus: Mapping[str, Any],
299+
) -> dict[str, Any]:
300+
checkpoint = {
301+
"verified": _boolean(raw, "verified"),
302+
"corpus_id": _token(raw.get("corpus_id"), "checkpoint.corpus_id"),
303+
"workspace_ref": _token(raw.get("workspace_ref"), "checkpoint.workspace_ref"),
304+
"project_ref": _token(raw.get("project_ref"), "checkpoint.project_ref"),
305+
"surface_id": _token(raw.get("surface_id"), "checkpoint.surface_id"),
306+
"read_authority": _token(
307+
raw.get("read_authority"), "checkpoint.read_authority"
308+
),
309+
"source_ref": _optional_token(raw.get("source_ref"), "checkpoint.source_ref"),
310+
}
311+
for field in IDENTITY_SCOPE_FIELDS:
312+
expected_scope = corpus["scope"].get(field)
313+
if expected_scope:
314+
checkpoint[field] = _optional_token(raw.get(field), f"checkpoint.{field}")
315+
return checkpoint
316+
317+
286318
def _freshness_reasons(
287319
corpus: Mapping[str, Any], freshness: Mapping[str, Any]
288320
) -> list[str]:
289321
reasons: list[str] = []
290322
mode = corpus["freshness"]["mode"]
291-
source_truth_current = _boolean(freshness, "source_truth_current")
292-
source_revision = _optional_token(
293-
freshness.get("source_revision"), "freshness_context.source_revision"
294-
)
323+
try:
324+
source_truth_current = _boolean(freshness, "source_truth_current")
325+
source_revision = _optional_token(
326+
freshness.get("source_revision"), "freshness_context.source_revision"
327+
)
328+
except ValueError as exc:
329+
raise RewardMemoryRecallInputError("freshness_context_invalid", str(exc)) from exc
295330
age_seconds = freshness.get("age_seconds")
296331
if age_seconds is not None and (
297332
isinstance(age_seconds, bool)
298333
or not isinstance(age_seconds, int)
299334
or age_seconds < 0
300335
):
301-
raise ValueError("freshness_context.age_seconds must be a non-negative integer")
336+
raise RewardMemoryRecallInputError(
337+
"freshness_age_invalid", "freshness_context.age_seconds must be a non-negative integer",
338+
)
302339
if mode in {"source_truth_bound", "execution_bound"} and not source_truth_current:
303340
reasons.append("source_truth_not_current")
304341
if mode in {"revision_bound", "session_archive_bound"} and (
@@ -391,7 +428,7 @@ def build_reward_memory_recall_request(
391428
):
392429
raise ValueError(f"limit must be between 1 and {MAX_RESULTS}")
393430
if not isinstance(request.get("freshness_context"), Mapping):
394-
raise ValueError("freshness_context must be an object")
431+
raise RewardMemoryRecallInputError("freshness_context_invalid", "freshness_context must be an object")
395432
if _boolean(request, "raw_content_captured"):
396433
raise ValueError("recall requests must not capture raw content")
397434

‎loopx/capabilities/reward_memory/decision.py‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -62,6 +62,7 @@ def _recall_telemetry(hook: Mapping[str, Any]) -> dict[str, Any]:
6262
"filtered_count": sum(item.get("filtered_item_count", 0) for item in attempts),
6363
"recall_status": attempts[-1].get("status") if attempts else None,
6464
"boundary_reason_code": hook.get("reason_code"),
65+
"boundary_detail_code": hook.get("boundary_detail_code"),
6566
}
6667

6768

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,28 @@
1+
"""Original configuration IO adapter for the shared TS checkpoint projection."""
2+
from __future__ import annotations
3+
4+
from collections.abc import Mapping
5+
from typing import Any, cast
6+
7+
from ...control_plane.effect_runtime import effect_runtime_result
8+
from .experiment import resolve_reward_memory_surface_config
9+
10+
11+
def build_reward_memory_surface_read_authority_checkpoints(
12+
config: Mapping[str, Any], surface_id: str, *, verified: bool, source_ref: str,
13+
) -> dict[str, dict[str, Any]]:
14+
"""Project only the configured surface; the caller supplies existing read proof.
15+
16+
Enabled configuration is not proof. False remains false. This does not call
17+
a provider, select a policy source, or verify/expand the caller's authority.
18+
"""
19+
route = resolve_reward_memory_surface_config(config, surface_id)
20+
result = effect_runtime_result("reward_memory.read_authority.surface_checkpoints", {
21+
"surface_id": surface_id, "verified": verified, "source_ref": source_ref,
22+
"corpora": [{"corpus_id": item["corpus"]["corpus_id"],
23+
"read_authority": item["corpus"]["read_authority"],
24+
"scope": {key: item["corpus"]["scope"].get(key) for key in (
25+
"workspace_ref", "project_ref", "user_ref", "peer_ref", "session_ref",
26+
)}} for item in route["recall_corpora"]],
27+
})
28+
return cast(dict[str, dict[str, Any]], result["checkpoints"])

‎loopx/capabilities/reward_memory/runtime_hooks.py‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@
77
from .application import (
88
RewardMemoryApplier,
99
RewardMemoryRecallSession,
10+
RewardMemoryRecallInputError,
1011
apply_reward_memory_recall,
1112
build_reward_memory_recall_request,
1213
execute_reward_memory_recall,
@@ -168,6 +169,14 @@ def run_reward_memory_automatic_recall_hook(
168169
provider_binding=corpus_route["provider_binding"],
169170
provider=provider,
170171
)
172+
except RewardMemoryRecallInputError as exc:
173+
return base | {
174+
"status": "guard_rejected",
175+
"reason_code": "exact_corpus_request_invalid",
176+
"boundary_detail_code": exc.reason_code,
177+
"recall_attempts": attempts,
178+
"telemetry": telemetry,
179+
}
171180
except (KeyError, OSError, RuntimeError, TypeError, ValueError):
172181
return base | {
173182
"status": "guard_rejected",

0 commit comments

Comments
 (0)