Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
8f3340e
fix(replan): share semantic discharge and vision input projection in …
huangruiteng Sep 19, 2026
91e0046
fix(qualification): exercise required vision through bound Turn closeout
huangruiteng Sep 19, 2026
c7038eb
docs(replan): clarify outcome projection and qualification boundaries
huangruiteng Sep 19, 2026
93a081a
fix(qualification): support confined JSON authoring and own vision sc…
huangruiteng Sep 19, 2026
4a0b67b
fix(qualification): align bounded read host and recoverable tool errors
huangruiteng Sep 19, 2026
cb9d800
docs(qualification): distinguish host execution from semantic admission
huangruiteng Sep 19, 2026
0e999c7
fix(qualification): return unexecuted host rejections to the model
huangruiteng Sep 19, 2026
39d001c
docs(qualification): preserve bounded recovery without extending grammar
huangruiteng Sep 19, 2026
e626beb
fix(qualification): budget complete vision closeout independently
huangruiteng Sep 19, 2026
cdde518
docs(qualification): distinguish scenario budgets from semantic accep…
huangruiteng Sep 19, 2026
6a6ef5d
fix(qualification): return pre-execution authoring rejection as tool …
huangruiteng Sep 19, 2026
2918526
docs(qualification): clarify rejected authoring and post-effect failures
huangruiteng Sep 19, 2026
10e3722
fix(qualification): ground evidence references and admit literal CLI …
huangruiteng Sep 19, 2026
7624100
docs(qualification): define source reference and literal argv boundaries
huangruiteng Sep 19, 2026
41fe524
fix(testing): let vision actors correct rejected drafts
huangruiteng Sep 19, 2026
a420717
docs(testing): define recoverable vision draft errors
huangruiteng Sep 19, 2026
597b3ff
fix(testing): return pre-admission workspace rejection to actor
huangruiteng Sep 19, 2026
67a9e6f
fix(testing): budget complete closeout with validation recovery
huangruiteng Sep 19, 2026
edf5566
refactor(testing): qualify vision through an isolated native shell
huangruiteng Sep 19, 2026
4e6f003
docs(testing): replace command grammar with native execution boundaries
huangruiteng Sep 19, 2026
0404a1b
fix(testing): prepare native shell isolation on Linux
huangruiteng Sep 19, 2026
2084d0a
test(replan): align blocked wait with vision writeback
huangruiteng Sep 19, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions .github/workflows/python-tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -284,6 +284,23 @@ jobs:
run: |
python -m pip install --disable-pip-version-check -e ".[test]"
npm ci --ignore-scripts
- name: Prepare native shell isolation
# Ubuntu 24.04 may need an executable-scoped userns profile for bwrap.
# Do not disable AppArmor or the system-wide unprivileged-userns guard.
run: |
sudo apt-get update
sudo apt-get install -y bubblewrap
if ! bwrap --unshare-all --ro-bind / / /bin/true; then
sudo tee /etc/apparmor.d/loopx-qualification-bwrap >/dev/null <<'PROFILE'
abi <abi/4.0>,
include <tunables/global>
profile loopx-qualification-bwrap /usr/bin/bwrap flags=(unconfined) {
userns,
}
PROFILE
sudo apparmor_parser -r /etc/apparmor.d/loopx-qualification-bwrap
fi
bwrap --unshare-all --ro-bind / / /bin/true
- name: Run test shard
# Split the whole collection, not a hand-maintained list of directories.
# Without timing history least_duration alternates equal-weight tests.
Expand Down
11 changes: 11 additions & 0 deletions docs/architecture/rfcs/typescript-control-plane-migration-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -320,6 +320,17 @@ not wait for a PostgreSQL service and never expires receipts at day ten.

### Delivery semantics: correctness before migration

The replan obligation outcome policy now lives in
`work_items/replan_semantics.ts`: required-outcome selection, vision-path and
terminal consistency matching, and the matching refresh input projection share
one owner. Python retains progress normalization/novelty and persistence adapters,
but no longer duplicates obligation-to-outcome matching. This is a bounded rule
convergence, not a settlement-writer or store migration. Existing outcome
characterization precedes the move; the intentional correction is executable
authoring for all vision triggers, with real bound CLI closeout/readback and
negative qualification-scope cases. Checkpoint recovery and in-flight rules
remain in their existing owners; no new capability, provider or setting is added.

The delivery-history boundary now treats `classification`, `health_check`, and
`recommended_action` as narrative. They cannot create or discharge a
follow-through obligation, prove an outcome, or classify delivery scale.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -257,6 +257,14 @@ receipt 过期。

### 交付语义:先修正规则,再迁移

Replan 的义务结果规则现收敛到 `work_items/replan_semantics.ts`:接受结果选择、
vision path/terminal 一致性校验与对应 refresh 输入投影共用同一 owner。
Python 保留 progress 归一化/新颖性与持久化适配,不再重复义务匹配规则。
这是有界规则收敛,不是 settlement writer 或存储迁移。先刻画既有接受语义,
再修正所有 vision trigger 的可执行写入投影,并验证真实绑定 CLI 闭环、回读及
资格范围错配反例。Checkpoint 恢复与 in-flight 规则仍由既有边界负责,
不新增 capability、provider 或设置。

交付历史边界将 `classification`、`health_check` 与 `recommended_action` 视为
叙述文本。它们不能生成或解除 follow-through obligation,不能证明 outcome,也
不能判定交付规模。例如,`unblocked after dependency update` 不构成 blocker
Expand Down
55 changes: 49 additions & 6 deletions docs/development/testing-and-quality.md
Original file line number Diff line number Diff line change
Expand Up @@ -745,8 +745,8 @@ action without calling the tool, or issuing an unallowlisted command fails.

replan semantic action 另有一条 function-tool 行为资格门,因为 no-tool JSON 决策不能
证明模型会使用覆盖账本选择新方向并完成真实写回。资格门创建一个包含两个等价 typed
progress observation 的隔离、public-safe 临时 Goal;真实模型只看到正式 thin Codex App
heartbeat task body 和普通 `exec_command` tool。真实 quota 必须投影 host coverage context
progress observation 的隔离、public-safe 临时 Goal;模型接收正式 thin Codex App
heartbeat task body、受限执行环境说明和 `exec_command` tool。真实 quota 必须投影 host coverage context
与最小 action packet,模型随后提交的 typed semantic delta 还要通过独立语义判定和真实
写时闸门。若模型选择新 successor,资格门要求它以当前 `obligation_id` 调用真实
`todo add`,验证 Todo 原子 receipt 与 `host_action=end_current_heartbeat`,且不得在同一
Expand Down Expand Up @@ -787,13 +787,56 @@ python3 scripts/qualify-doubao-capability-monitor-repair-tool-live.py \
```

The regular live suite is
`actual_default_model_behavior_portfolio_v0`: nineteen one-arm scenarios and two
`actual_default_model_behavior_portfolio_v0`: twenty-one one-arm scenarios and two
attempts each. Its selected-Todo case starts from a production thin heartbeat,
executes real quota, and requires the model to perform the selected Todo's
read-only target action. Its required-vision replan case independently builds a
hermetic missing-vision state, executes real quota, and requires the model to
use host-projected frontier/work-source context and submit a typed semantic
action through the real write path. The other turn cases remain
hermetic required-profile/missing-vision state with a future monitor and
peer-owned work. The actor must author an evidence-linked vision, execute the
projected bound refresh and spend, and pass durable checkpoint, one-spend and
next-Turn readback. Readback must clear the original missing-baseline obligation;
a legitimate new successor requirement is allowed. Source alignment includes trigger kinds, accepted outcomes
and qualification scope; a successful ordinary `typed_progress_repeat` refresh
cannot qualify this journey. The narrow semantic-action gate remains useful
but does not prove full closeout. Run this focused journey with
`uv run --extra test python scripts/qualify-doubao-replan-semantic-action-live.py --required-vision --qualification-id <public-safe-run-id>`.
The complete required-vision journey has a 32-call bound; the narrow
single-semantic-action qualifier retains seven. The increased budget covers
evidence discovery, JSON authoring, refresh, settlement and bounded recovery,
including multiple field-validation corrections before a final spend;
it does not authorize more effects or weaken the checkpoint/readback oracle.
Each receipt records `tool_call_limit` beside actual usage. Earlier seven- and sixteen-call
failures remain failures; the revised budget defines a new qualification, not
a retrospective pass. Limit exhaustion after refresh but before spend still
fails.

Required-vision qualification runs a normal shell in an OS-isolated fixture,
using `sandbox-exec` on macOS or `bubblewrap` on Linux. These are execution
prerequisites: missing isolation fails explicitly rather than falling back to an
unrestricted host. Shell variables, pipelines, compound commands, Python/JSON
validation, and draft rewrites are ordinary operations, not an allowlisted
command language. Read-only inspection may precede quota; writeback still needs
the current quota-derived binding. The source must be observed in returned
tool data, but no particular read command or metadata-read sequence is required.

The shell may write project drafts and its private `$TMPDIR`. Original inputs,
authority stores, host-private data and external network access are protected
by the execution boundary. A fixture-local `loopx` command forwards expanded
argv to the existing real CLI executor; only the task's quota, help, vision
refresh and settlement operations have authority. The shell cannot forge run
history or receipts by writing the store directly. Turn identity is supplied
as normal host context, never inferred from a model's incorrect binding.

Shell and real CLI failures return output and exit status for correction within
the same call budget. Returning an error is not semantic acceptance or rollback.
The model authors its own decision; the host does not fill semantic fields or
change evidence. An observed source's evidence id and exact source reference
identify the same evidence. Checkpoint, one-spend and following-Turn verification
remain independent of shell success. Receipts disclose native-shell execution
without persisting raw conversations. Other actors retain their existing host
and seven-call contracts. Earlier bounded-grammar failures remain failed
qualifications and are not evidence of core semantic rejection.
The other turn cases remain
bounded packet-interpretation checks. Nine core-contract scenarios cover
onboarding, agent identity and goal selection, selected todo, peer identity
routing, same-agent continuation, final human gate, healthy continuation, and
Expand Down
15 changes: 15 additions & 0 deletions docs/reference/protocols/goal-vision-replan-contract-v0.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,21 @@ pathless packet fails instead of recording a partial closure. A matching typed
semantic ACK settles the vision-derived duty even while the original acceptance
gap remains visible in the source projection.

Quota's replan writeback projection and write-time outcome matching share
`work_items/replan_semantics.ts`. Every vision-derived trigger, including a
missing required baseline, projects evidence-linked JSON authoring through
`replan_action_packet.writeback_contract`; limits come from the existing vision
validator. This is guidance for a valid refresh path, not a new obligation or
the removal of typed successor/blocker/terminal alternatives. A new surface id
alone cannot satisfy a vision obligation. Execute the current settlement binding
exactly once; accepted semantic writeback, satisfied checkpoint, settled Turn
and Goal completion remain separate facts. Existing missing-checkpoint recovery
stays on the original Turn, and in-flight continuation remains unchanged.

投影与写入校验共用 TS 语义规则;required-vision 不再投影只有普通进度标识的模板。
JSON 写作契约复用 vision 校验器,不新增 ACK 仪式,也不改变既有 successor、blocker、
terminal 出口。语义接受、checkpoint 满足、Turn 结算与 Goal 完成仍须分别验证。

Inline vision writes require `--agent-id`. JSON packets must also resolve to
the same `agent_id` as the refresh run. This keeps `research-executor`,
`evaluator-promoter`, and other roles from overwriting or satisfying each
Expand Down
2 changes: 2 additions & 0 deletions loopx/control_plane/effect_runtime_handlers.ts
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@ import {
} from "./turn_driver/delivery_continuity.ts";
import { reduceTurnSettlementTransaction } from "./turn_driver/settlement.ts";
import { evaluateHostTodoCompletion } from "./turn_driver/host_todo_completion.ts";
import { projectReplanSemantics } from "./work_items/replan_semantics.ts";
import {
projectReplanSettlementContract,
projectTodoLifecycleSettlementReentry,
Expand Down Expand Up @@ -736,6 +737,7 @@ export function createEffectRuntimeHandlers(
["turn.settlement.reduce", reduceTurnSettlementTransaction],
["turn.host_todo_completion.evaluate", evaluateHostTodoCompletion],
["work_item.replan_settlement.project", projectReplanSettlementContract],
["work_item.replan_semantics.project", projectReplanSemantics],
[
"work_item.replan_settlement.reentry",
projectTodoLifecycleSettlementReentry,
Expand Down
2 changes: 2 additions & 0 deletions loopx/control_plane/quota/turn_envelope.ts
Original file line number Diff line number Diff line change
Expand Up @@ -257,6 +257,8 @@ function replanActionPacket(payload: JsonObject): JsonObject | null {
"schema_version", "decision", "obligation_id", "uncovered_frontier",
"required_outcome", "allowed_terminal", "bounded_frontier",
]);
const writeback = object(source.writeback_contract);
if (writeback.vision_authoring) compact.writeback_contract = writeback;
return Object.keys(compact).length > 0 ? compact : null;
}

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@
SELECTED_TODO_TOOL_FIXTURE_ACTION_TEXT,
SELECTED_TODO_TOOL_FIXTURE_TODO_ID,
)
from .replan_vision_closeout_behavior import required_vision_scenario_contract

ACTUAL_DEFAULT_MODEL_BEHAVIOR_PORTFOLIO_SCHEMA_VERSION = (
"actual_default_model_behavior_portfolio_v0"
Expand Down Expand Up @@ -843,47 +844,6 @@
raise ValueError("compaction regression must preserve the selected todo")


def _validate_required_vision_replan_scenario(
source_packet: Mapping[str, Any],
contract: Mapping[str, Any],
) -> None:
semantics = model_behavior_semantic_contract_from_packet(
source_packet,
arm="full_packet",
)
vision = semantics["vision_continuation"]
trigger_kinds = set(vision.get("trigger_kinds", []))
required = {
"selected_todo_id": None,
"user_action_required": False,
"must_attempt_work": True,
"quiet_noop_allowed": False,
}
if any(contract.get(field) != value for field, value in required.items()):
raise ValueError("required-vision scenario must execute before quiet wait")
if vision.get("required") is not True or (
"required_agent_vision_missing" not in trigger_kinds
):
raise ValueError("required-vision scenario must preserve the profile gap")
if semantics["required_reads"]:
raise ValueError("required-vision replan must not require a model read ritual")
action_packet = source_packet.get("replan_action_packet")
obligation = source_packet.get("autonomous_replan_obligation")
if not (
isinstance(action_packet, Mapping)
and isinstance(obligation, Mapping)
and action_packet.get("decision") == "replan_required"
and action_packet.get("obligation_id") == obligation.get("obligation_id")
and dict(obligation.get("replan_context") or {}).get("delivery")
== "host_projected"
):
raise ValueError(
"required-vision scenario must preserve host-delivered replan context"
)
if semantics["scheduler_action"].get("action") != "run_now":
raise ValueError("required-vision scenario must remain immediately runnable")


def _validate_planning_horizon_model_scenario(
source_packet: Mapping[str, Any],
) -> None:
Expand Down Expand Up @@ -991,10 +951,8 @@
def _validate_control_plane_composition_scenario(
spec: _ScenarioSpec,
source_packet: Mapping[str, Any],
contract: Mapping[str, Any],

Check warning on line 954 in loopx/control_plane/testing/actual_default_model_behavior_portfolio.py

View check run for this annotation

SonarQubeCloud / SonarCloud Code Analysis

Remove the unused function parameter "contract".

See more on https://sonarcloud.io/project/issues?id=huangruiteng_loopx&issues=AaC5jFylQT3MxfJY-OFS&open=AaC5jFylQT3MxfJY-OFS&pullRequest=4730
) -> None:
if spec.scenario_id == "turn_required_vision_replan":
_validate_required_vision_replan_scenario(source_packet, contract)
if spec.scenario_id == "turn_scoped_gate_successor_replan":
signature = quota_action_signature_document(source_packet)
action = dict(signature.get("action") or {})
Expand Down Expand Up @@ -1111,6 +1069,8 @@
_validate_identity_scenario_contract(spec, source_packet, contract)
_validate_planning_context_scenario(spec, source_packet, contract)
_validate_control_plane_composition_scenario(spec, source_packet, contract)
if spec.scenario_id == "turn_required_vision_replan":
contract.update(required_vision_scenario_contract(source_packet, contract))
_validate_compaction_scenario(spec, source_packet, actor_packet, contract)
if spec.scenario_family == "diagnostic_authority_boundary":
diagnostic = dict(actor_packet.get("agent_todo_summary") or {})
Expand Down Expand Up @@ -1170,7 +1130,7 @@
)


def _receipt_alignment(

Check failure on line 1133 in loopx/control_plane/testing/actual_default_model_behavior_portfolio.py

View check run for this annotation

SonarQubeCloud / SonarCloud Code Analysis

Refactor this function to reduce its Cognitive Complexity from 17 to the 15 allowed.

See more on https://sonarcloud.io/project/issues?id=huangruiteng_loopx&issues=AaC5jFylQT3MxfJY-OFT&open=AaC5jFylQT3MxfJY-OFT&pullRequest=4730
spec: _ScenarioSpec,
receipt: Mapping[str, Any],
expected: Mapping[str, Any],
Expand All @@ -1183,6 +1143,11 @@
if receipt.get(field) != expected[field]
]
mismatches.extend(str(item) for item in receipt.get("safety_violations") or [])
if spec.scenario_id == "turn_required_vision_replan" and not (
receipt.get("semantic_action_accepted") is True
and set(receipt.get("selected_semantic_outcomes") or []).intersection(expected["required_semantic_outcomes"])
):
mismatches.append("required_vision_outcome_not_accepted")
if (
spec.semantic_contract_fields
and receipt.get("semantic_contract_complete") is not True
Expand Down
26 changes: 19 additions & 7 deletions loopx/control_plane/testing/model_tool_behavior.py
Original file line number Diff line number Diff line change
Expand Up @@ -225,17 +225,24 @@ def actor_ref(self) -> str:
def next_tool_call(
self,
messages: list[dict[str, Any]],
*,
tool_description: str | None = None,
) -> ExecToolCall | None:
return self.next_step(messages).tool_call
return self.next_step(messages, tool_description=tool_description).tool_call

def next_step(
self,
messages: list[dict[str, Any]],
*,
tool_description: str | None = None,
) -> ExecToolStep:
body = {
"model": self._model,
"messages": messages,
"tools": [EXEC_COMMAND_TOOL],
"tools": [EXEC_COMMAND_TOOL if tool_description is None else {
**EXEC_COMMAND_TOOL,
"function": {**EXEC_COMMAND_TOOL["function"], "description": tool_description},
}],
"tool_choice": "auto",
"thinking": {"type": "disabled"},
"temperature": 0,
Expand Down Expand Up @@ -527,7 +534,7 @@ def execute_loopx_cli(
else os.pathsep.join((str(source_root), existing_pythonpath))
)
completed = subprocess.run(
[sys.executable, "-m", "loopx.cli", *argv],
[sys.executable, "-P", "-m", "loopx.cli", *argv],
cwd=project_root,
env=env,
check=False,
Expand All @@ -540,8 +547,13 @@ def execute_loopx_cli(
bounded_detail = (
detail if len(detail) <= 1_000 else detail[:500] + "\n...\n" + detail[-500:]
)
raise RuntimeError(
"LoopX CLI command failed with "
f"exit={completed.returncode}: {bounded_detail}"
)
raise LoopxCliExecutionError(completed.returncode, bounded_detail)
return completed.stdout


class LoopxCliExecutionError(RuntimeError):
"""An executed CLI returned nonzero; this does not imply state rollback."""

def __init__(self, returncode: int, detail: str) -> None:
super().__init__(f"LoopX CLI command failed with exit={returncode}: {detail}")
self.returncode = returncode
Loading
Loading