Archive original DeepSWE GPT xhigh v1 code and results - #4498
gwh6669999 wants to merge 3 commits into
Conversation
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
There was a problem hiding this comment.
🟡 Changes recommended
Unresolved critical delivery and launcher issues, plus moderate verification and admission-check concerns, must be addressed before approval.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
Archives the original DeepSWE GPT xhigh v1 harness and historical 113-task results, with provenance documentation and read-only verification.
Changes:
- Adds the immutable 14-file benchmark archive.
- Documents provenance, dependencies, limitations, and result caveats.
- Adds archive identity and syntax checks.
File summaries
| File | Review summary |
|---|---|
benchmark/deepswe-gptxhigh-versions.md |
Documents archive provenance and execution boundaries. |
benchmark/deepswe-gptxhigh-v1/workspace_delivery.py |
Workspace delivery logic; critical mutation risk identified. |
benchmark/deepswe-gptxhigh-v1/run_loopx_rerun_54_20260908.sh |
Archived 54-task launcher. |
benchmark/deepswe-gptxhigh-v1/run_five_arms_remaining59_20260910.sh |
Five-arm launcher; critical endpoint-port isolation issue identified. |
benchmark/deepswe-gptxhigh-v1/README.md |
Preserves benchmark methodology and historical results. |
benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py |
Admission checks; revision-pinning and source-heuristic concerns identified. |
benchmark/deepswe-gptxhigh-v1/pier_cn.py |
Archived environment and arm setup. |
benchmark/deepswe-gptxhigh-v1/loopx_wen_native_runner.py |
Archived native Goal runner. |
benchmark/deepswe-gptxhigh-v1/loopx_turn_runner.py |
Archived LoopX turn runner. |
benchmark/deepswe-gptxhigh-v1/loopx_native_codex.py |
Archived native Codex integration. |
benchmark/deepswe-gptxhigh-v1/loopx_heartbeat_supervisor.py |
Archived heartbeat supervisor. |
benchmark/deepswe-gptxhigh-v1/loopx_codex_cli_runner.py |
Archived Codex CLI runner. |
benchmark/deepswe-gptxhigh-v1/goal_codex.py |
Archived Codex harness. |
benchmark/deepswe-gptxhigh-v1/goal_claude.py |
Archived Claude harness. |
benchmark/deepswe-gptxhigh-v1/codex_nosandbox_wrapper.py |
Archived Codex wrapper. |
benchmark/check_deepswe_v1.py |
Archive verifier; embedded-program syntax coverage is incomplete. |
Review details
Suppressed comments (3)
benchmark/check_deepswe_v1.py:40
- The verifier only compiles string assignments named
_RUNNERand_BOOTSTRAP. The archive also generates executable Python invalidator_command()and in the two shell heredocs, but those programs are never parsed, so a syntax error there can still produce anok: truearchive check. Compile every embedded program or narrow the documented verification claim to the two constants.
for node in ast.parse(raw).body:
if not isinstance(node, ast.Assign) or not isinstance(node.value, ast.Constant):
continue
names = {target.id for target in node.targets if isinstance(target, ast.Name)}
if names & {"_RUNNER", "_BOOTSTRAP"} and isinstance(node.value.value, str):
source = node.value.value
if "_RUNNER" in names:
source = source.format(remote_dir="/tmp/loopx-goal")
compile(source, f"{path.name}:embedded", "exec")
embedded_count += 1
benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py:34
- This admission check does not actually pin the evaluated LoopX revision:
MR_EXPECTED_LOOPX_REVISIONcan be set to whateverHEADis, so a clean checkout at a different revision can still be admitted. That undermines the documented fixed2cef51d...reproduction boundary and can make results look comparable when they are not; keep the expected SHA immutable for reproduction runs, or make any override explicitly non-comparable and record that mode.
EXPECTED_LOOPX_REVISION = os.environ.get(
"MR_EXPECTED_LOOPX_REVISION", "2cef51d08b2a0103f4ba026bf47fd70dc8acee30"
)
benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py:185
- Admission is based on raw substring membership across the entire runner source. A comment or docstring can satisfy a required marker or trigger the forbidden
"app-server"check, while equivalent code built dynamically can evade the check, so source-only edits can admit or reject an execution surface without changing its behavior. Validate a typed/runtime contract or run an import/dispatch probe instead of using these text heuristics.
if any(marker not in source for marker in contract["required"]):
errors.append("runner_contract_marker_missing")
if any(marker in source for marker in contract["forbidden"]):
errors.append("runner_uses_wrong_surface")
- Files reviewed: 16/16 changed files
- Comments generated: 2
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| PIER_CUSTOM_NETWORKS=1 MR_MODELONLY_NET=1 MR_MODELONLY_HOST=127.0.0.1 MR_MODELONLY_PORT=4250 \ | ||
| MR_API_BASE=http://127.0.0.1:4250/v1 MR_AGENT=codex MR_MODEL="$MODEL" MR_EFFORT="$EFFORT" MR_REASONING_EFFORT="$EFFORT" \ |
| def _candidate(path: Path, base_sha: str) -> Candidate | None: | ||
| _commit_pending(path) |
huangruiteng
left a comment
There was a problem hiding this comment.
Exact head reviewed: 50bf824d20041a10a1fcec57485e5eb573193e42
动机
历史 DeepSWE GPT xhigh v1 的 harness 和 113 任务结果在公开仓库里没有落脚点:v1 的出处、运行前置条件、已知缺陷都没有记录,读者看到历史结果表也无法对着产生它的代码去读。PR 的动机是把这份不可变快照连同出处、限制说明和一个只读校验命令一起归档,并且明确拒绝"为了让它能跑而改写 v1"。这个诉求是真的:没有冻结副本,历史结论既不可引用也无法复核;而 v1 与正在运行的 benchmark/swe-marathon/ 属于不同世代,不能被静默改写。影响面是维护者和外部读者,没有任何运行时调用者。更小的做法(只发文档 + 树哈希,不发字节)会失去独立复核字节的能力,所以作者选择连字节一起归档是可以理解的——但这并不要求把归档放进活跃工作区,这正是下面 P1 要说的。
改动思路
入口只有一个:python3 benchmark/check_deepswe_v1.py(只读)。权威输入是 benchmark/deepswe-gptxhigh-v1/ 下 14 个文件的字节与可执行位;决策所有者是校验命令自己——它按 sorted entries、mode 与 blob sha 重新计算 git tree 哈希,与 EXPECTED_TREE 常量比对。没有副作用:不写盘、不调模型、不跑任务代码。失败与重试所有者是校验命令(任何内容、mode 或语法偏差都抛错退出 1)与归档里的启动脚本(后者明确不承诺可运行)。复用判断上,归档本身保持与活跃 owner 分离是合理的:不可变导出不应被并入、更不应被静默更新;真正不成立的是落点——benchmark/README.md 规则 3 写"只有通用化代码与公开安全结论进入本目录",并在同页写明"历史 runner 与带日期的 packet 归档在 deprecate/benchmark-legacy/",deprecate/benchmark-legacy/README.md 也把自己定义为 pre-reset benchmark 实现的考古目录。PR 没有改写这条规则,而是绕过它,于是文本与布局开始互相矛盾。状态模型上没有新增持久状态,唯一新契约是被校验命令强制的树哈希 pin。
具体改动
新增 16 个文件、4688 行、0 删除,全部是新增面:14 个冻结归档文件(loopx_turn_runner.py 623 行、goal_codex.py 589 行、loopx_native_codex.py 470 行等)、benchmark/check_deepswe_v1.py(73 行)与 benchmark/deepswe-gptxhigh-versions.md。仓库中没有任何 loopx/、CLI、scheduler、examples/ 或 .github/ 引用归档或校验命令,归档是惰性的;examples/run-smokes.py 与工作流也都未被触碰(checks.total = 0,此 head 无 CI rollup)。边界扫描(16 个新增文件全部)干净:没有绝对本地路径、没有内部域名、没有凭据、没有内部链接,只出现 127.0.0.1 占位符。
关键代码讲解
benchmark/check_deepswe_v1.py:18 check_archive:按mode + name + blob sha重建 git tree 哈希并与EXPECTED_TREE比对,同时compile()全部 Python(含内嵌 runner/bootstrap 程序)并对 shell 脚本跑bash -n;任何非普通文件、符号链接或哈希偏差都直接抛错。这是本 PR 唯一有持久价值的机制:它让"字节未被改动"从声明变成可执行检查。benchmark/deepswe-gptxhigh-v1/workspace_delivery.py:181 normalize_delivery:归档里真正的交付门——从 linked worktree 把 agent 产出恢复进任务工作区,先证明 patch 可应用再安装,空 patch 视为失败。它解释了归档 README 所说的"非空提交 patch"。benchmark/deepswe-gptxhigh-v1/loopx_turn_runner.py:232 run_turn:LoopX 臂的一次 Turn 驱动,关键不变量是"没有移动 HEAD 就没有交付"(head_sha在:224),空 patch 不能算进展。benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py:67 delivery_self_test:准入前先证明投递路径可用(digest pin 的 LoopX revision、端口、任务清单),不通过就不进任务。
对主干的风险
最强回归场景不是运行时报错,而是"归档被当成当前实践":贡献者把 benchmark/deepswe-gptxhigh-v1/ 视作活跃 benchmark 工作区的一部分,接进 CI 或把归档里的启动器模式抄进新工作,仓库对"什么算活跃"的唯一权威随之消失。触发状态正是把一份带日期(2026-09-08/09-10)的历史包放进那个"只进通用化代码"的目录。可观测点是仓库布局和 benchmark/README.md;回退/恢复成本极低——一次 git mv。关键在于不可变性不受影响:pin 的是内容寻址的 tree 哈希,与路径无关,且校验命令本身支持 --archive,因此搬迁后同一条命令、同一个哈希仍然通过。最小修复有两条:(1) 把 14 个文件搬到 deprecate/benchmark-legacy/ 下并同步文档路径;(2) 若作者坚持放在活跃工作区,则必须在同一 PR 内改写 benchmark/README.md 的落点规则,让规则和布局只有一个 owner。可观测性/对照上,这个 PR 的所有测试无论归档放在哪个目录都会全绿——校验命令通过恰恰不能证明落点正确,这正是必须去读 owning README 的原因。另外还有两条非阻断发现:归档 README 的 v1 五臂结果表(含 ssh-goal、codex-cli 两臂)与当前 benchmark/swe-marathon/README.md 撤回待复验的口径并存,撤回说明目前只写在同伴文档的 Known Limits 里(P2);出处声明点名的 98c262a 在本仓库不可达,读者只能验证树哈希(P3)。
我的整体评价
结论:REQUEST_CHANGES。 归档内容本身是干净、公开安全、惰性且可校验的:我用真实命令复核了 python3 benchmark/check_deepswe_v1.py = ok true、archive_tree 1bc5d2b3b74761a97d34ba3f3612e977fd610340(与 git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1 一致),并做了两条负向验证——追加一个字节、去掉一个可执行位,两次都以 archive contents or executable modes differ from the original v1 snapshot 退出 1。阻断项只有一个,且修复很轻:落点与 benchmark/README.md 的规则冲突(P1)。repository_reuse = separation_justified(不可变导出与实际 owner 分离是对的),observable_semantics = intentional_change_validated(与当前 main 的差异是有意保留的 v1 缺陷),change_proportionality = disproportionate(机制对归档目标合适,落点让 4688 行惰性历史代码穿过活跃工作区),authority_semantics = not_applicable,default_off_isolation = not_applicable(无开关)。残留风险:只验证了树哈希,源码 commit 不可达,归档代码未执行、结果未复验。复审条件:把归档搬到 deprecate/benchmark-legacy/(或同 PR 改写落点规则),并把撤回口径放到同伴文档开头;P3 的出处表述可一并修正。
English verdict: REQUEST_CHANGES. Exact head 50bf824d20041a10a1fcec57485e5eb573193e42. The archive itself is clean, inert and verifiable: python3 benchmark/check_deepswe_v1.py returns ok with the pinned tree 1bc5d2b3b74761a97d34ba3f3612e977fd610340, and both a mutated byte and a removed execute bit fail closed. The blocking finding is placement: benchmark/README.md:67 states that only generalized code and public-safe conclusions enter the active workspace and that legacy runners and dated packets are archived under deprecate/benchmark-legacy/, yet this PR adds a dated historical package to benchmark/. Minimum repair: git mv the archive to the legacy home (the tree hash is content-addressed, so the immutability proof and the checker survive unchanged), or change the documented rule in the same PR. Two non-blocking findings: the archived five-arm table republishes the ssh-goal and codex-cli rates whose data the current SWE-Marathon publication withdraws pending revalidation, with the correction only in the companion doc's Known Limits (P2), and the declared source commit 98c262a is not reachable in this repository, so only the tree hash is independently verifiable (P3).
|
【在同一 PR 内改写 benchmark/README.md 的落点规则,让规则和布局只有一个 owner】这样改吧 |
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
|
已按最后确认的方案在同一 PR 修改落点规则,未搬迁或改写 v1 快照。
验证: Implemented the maintainer-selected placement-policy option in this PR. The canonical README now distinguishes inert immutable experiment archives from retired implementations; AGENTS defers to it. Withdrawal caveats are prominent, provenance limits and checker coverage are explicit, and the original v1 tree remains byte-for-byte unchanged. No benchmark execution or score revalidation is claimed. |
huangruiteng
left a comment
There was a problem hiding this comment.
动机
历史 DeepSWE GPT xhigh v1 的 harness 和 113 任务结果一直没有公开落脚点:读者看到结果表却读不到产生它的代码,也无法判断这份历史文件后来有没有被改过。这个 PR 把 v1 原始 14 个文件作为不可变快照归档,并配一份写明出处、运行前置条件、已知缺陷和撤回口径的同伴文档,再加一个只读校验命令;同时把"这类快照可以放在哪"写成规则。上一轮我要求改动时就说过:连字节一起归档是合理的,只是当时它落在活跃工作区里而规则不允许——这一轮作者补上了规则。
诉求本身成立。影响面是维护者、评审者和外部读者,没有任何运行时调用者;不做这件事的代价是历史结论既不可引用也无法复核。反过来,把 4700 行历史实验代码放进仓库而不声明归属和边界,代价是贡献者可能把冻结代码当成当前实践、甚至接进 CI,所以"归档 + 边界 + 可校验"这三件一起做是对的。
为什么更小的修法不够:上一轮给出的替代方案(只发树哈希和文档、不发字节)会丢掉"把结果对着产生它的代码读"的能力,也就是本 PR 的目的;而本 PR 现在剩下的最小修法不是缩小归档,而是让它的两条验收命令能共存——这是几行代码的问题,不是重新设计。
改动思路
入口只有一个:python3 benchmark/check_deepswe_v1.py(只读,文档写在同伴文档的 Executability Checks 一节)。权威输入是 benchmark/deepswe-gptxhigh-v1/ 下 14 个提交对象的字节与可执行位,pin 在 EXPECTED_TREE = "1bc5d2b3b74761a97d34ba3f3612e977fd610340"。决策所有者是校验命令自己:它按 mode + name + blob sha1 重建 git tree 哈希并比对,任何多余条目、符号链接、模式差异、语法错误都直接抛错退出 1;成功时明确报告 benchmark_executed: false、standalone_runnable: false、runtime_validation: "not_performed",把自身覆盖面讲清楚,而不是暗示"实验跑通了"。
正路径是:读 benchmark/README.md#archive-placement 的四条条件 → 读同伴文档(撤回提示就是文档开头的第一个引用块)→ 跑校验命令 → 得到 ok 和树哈希。没有任何副作用,不写盘、不联网、不调模型、不执行任务代码。
与既有实现的关系:仓库里既有的两个 owner 分别是活跃研究区 benchmark/ 和考古区 deprecate/benchmark-legacy/,loopx/canary/maintainability_ratchet.py::_is_benchmark_module_path 早已把两者同等对待(都不计入模块债务),所以把快照放在 benchmark/ 下并不新增一类债务;上一轮提供的另一条修法是把 14 个文件搬到考古区,作者选了"把落点规则写清楚"这条同样被允许的修法。不可变导出本就不该并入活跃实现、也不该被静默更新,所以"与活跃 owner 分离"是正确的复用判断;真正需要修的是规则的一致性——现在规则有一个声明 owner,却还有两处旧表述(见下)。
具体改动
18 个文件、+4752/−7。按角色拆开:14 个冻结归档文件(11 个 .py、2 个 .sh、1 个 README.md)共 4688 行、0 删除;benchmark/check_deepswe_v1.py +73 行,是整份改动里唯一的可执行行为;benchmark/deepswe-gptxhigh-versions.md +115 行;benchmark/README.md +44/−6;AGENTS.md +6/−1。第二个提交 1e7556555 只动这三份文档,把"不可变快照"这类归档的落点规则写清楚,并把上一轮的撤回口径挪到了同伴文档开头。
关键代码讲解
check_archive(benchmark/check_deepswe_v1.py:18):按 mode + name + blob sha1 重建 git tree 哈希,与 EXPECTED_TREE 比对,同时 compile 11 个 Python 文件、两个内嵌 _RUNNER/_BOOTSTRAP 常量,并对两个启动脚本跑 bash -n。这是本 PR 唯一有持久价值的机制:它让"字节未被改动"从声明变成可执行检查。我用 git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1 独立复核,得到同一个树哈希。
条目枚举(benchmark/check_deepswe_v1.py:21):对 archive.iterdir() 排序后要求每个条目都是普通文件、非符号链接,否则报 unexpected archive entry。这条守卫能抓住增删文件,但它不区分"人为新增的文件"和"工具顺手写下的字节码目录"——这正是第一条阻断项的落点。
normalize_delivery(benchmark/deepswe-gptxhigh-v1/workspace_delivery.py:181):归档里真正的交付门,从 linked worktree 把 agent 产出恢复进任务工作区,空 patch 视为失败。它解释了冻结 README 里"必须有非空提交 patch"的有效性口径,也说明了为什么结果表不能只靠 Goal/Todo 生命周期自证。
delivery_self_test(benchmark/deepswe-gptxhigh-v1/preflight_loopx_rerun.py:67):准入自检,不通过就不进任务;两个归档启动脚本都会先调它。归档里的这些逻辑在本仓库全部惰性:全仓检索 deepswe-gptxhigh 与 check_deepswe_v1,除了同伴文档、benchmark/README.md 和校验命令本身没有其他引用。
AGENTS.md:458-463 的归档归属规则:从"历史版本一律归 deprecate/benchmark-legacy/"改成"按 benchmark/README.md 的落点规则:退役实现与带日期 packet 归考古区,明确标识的不可变快照可在满足该文档条件的前提下留在 benchmark/,两类都不属于活跃 CI benchmark 执行"。这是一条 agent 指令面的策略放宽,PR 描述、benchmark/README.md 与同伴文档都做了披露;它唯一的自动后果是快照落在 benchmark/ 下,而 loopx/canary/premerge.py:72 早已把该前缀视为 benchmark-sensitive 并转为人工评审 hold。
对主干的风险
P2(阻断)— 仓库自己的 premerge 门会让本 PR 的验收命令变红。 loopx canary premerge --from-git-diff 的 changed_python_py_compile 步骤执行 python3 -m py_compile,对本 PR 就是编译归档里的 11 个文件,CPython 会把 __pycache__/ 写进那个冻结目录;随后文档指定的验收命令 python3 benchmark/check_deepswe_v1.py 返回 {"ok": false, "error": "unexpected archive entry: __pycache__"},退出码 1。我在真实提交内容上端到端复现:从 git archive 取干净字节 → 校验 ok → 跑门禁那条编译命令 → 校验失败;评审用的 worktree 在跑过门禁之后也确实处于这个 RED 状态;python -B 挡不住(py_compile 自己会写)。后果:按仓库文档顺序执行的维护者会在快照毫无改动的情况下看到红,而顺手可得的"修法"(给归档加 ignore、或删掉固定条目守卫)都会削弱这条身份检查。最小修法:让 check_archive 跳过被 gitignore 的字节码(__pycache__/、*.pyc),或改成遍历固定期望文件名——这正是仓库既有约定(loopx/semantics/inventory.py:24 的 SKIP_PARTS、loopx/contract.py:162 等);因为 .gitignore:7 本来就忽略 __pycache__/,跳过它不可能掩盖任何被跟踪字节的改动,树哈希仍是最终权威。回归建议:补一条"先跑门禁的编译命令、再断言校验仍为 ok"的用例。
P2(阻断)— pin 的是字节,但没有任何东西固定检出的字节。 校验命令读的是工作区文件(path.read_bytes()),而 .gitattributes 里只有一条 Vite 空行规则,没有覆盖这份归档。用 core.autocrlf=true(Git for Windows 的默认)检出时,归档文件被转成 CRLF,校验命令报的是 Command '['bash', '-n', .../run_five_arms_remaining59_20260910.sh']' returned non-zero exit status 2——一个针对未改动文件的误导性语法错误,甚至还没走到树哈希比对。我在本轮用 git -c core.autocrlf=true worktree add 复现。仓库有 windows-latest 作业(.github/workflows/python-tests.yml:459),所以这不是假想平台;同伴文档"读者可以对着本目录验证该身份"的那句在 Windows 检出上不成立,而新规则的第 3 条条件恰好要求这条检查能验证"被保留的字节"。最小修法:给 .gitattributes 加 benchmark/deepswe-gptxhigh-v1/** text eol=lf(或 -text)。我逐 blob 检查过 14 个文件都含 0 个 CR 字节,所以 eol=lf 不会改动任何冻结内容。回归建议:断言 git check-attr text -- <归档文件> 的解析结果,或像本轮一样从 autocrlf 检出跑一次。
P3(非阻断)— 落点规则还有两处旧回声。 新章节自称拥有归档落点,AGENTS.md 也指向它,但 loopx/capabilities/benchmark_toolkit/README.md:1362-1365 仍写"历史 runner 与带日期 research packet 保留在 deprecate/benchmark-legacy/",docs/development/documentation-layout.md:41 仍写"superseded runner 归档到 deprecate/benchmark-legacy/"。读这两份文档的人会得出"本 PR 的快照该进考古区"的相反结论,也就是上一轮要求消除的"规则与布局不一致"。最小修法:两处各加一句指向 benchmark/README.md#archive-placement,或把旧句收窄为"退役实现与被取代的 runner"。
残余风险与证据边界。 归档本身干净:14 个文件没有绝对本地路径、内部域名、凭据、轨迹或私有上下文,只有 127.0.0.1 占位符;loopx canary premerge --from-git-diff 19 项全通过(0 失败,1 个 benchmark-sensitive 人工 hold)、docs governance 通过、合入当前主干 origin/main(079643038)是干净合并。但历史结论仍是作者自述:出处提交 98c262a 在本仓库不可达(文档已披露),归档代码我一次也没有执行,文档列的已知缺陷只做了抽查(硬编码解释器路径 .venv-user-395647、启动脚本与任务清单口径),作者提到的 datacurve-pier 0.3.1 + Python 3.12.13 的 import/dispatch smoke 在本环境无法复现。本轮按解析出的策略不取用远程 CI(wait_for_ci=false,来源也没有给出该 head 的 rollup);分支处于 BEHIND,任何 rebase 之后这次的合并结论和上述复现都需要重跑。
我的整体评价
我用真实命令复核而不是只读 diff:干净字节下校验 ok、archive_tree 与 git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1 完全一致;对副本改一个字节、去掉一个可执行位,两次都以 archive contents or executable modes differ from the original v1 snapshot fail closed;仓库门禁 19 项全过、docs governance 通过、合入主干干净。上一轮的两条非阻断项也确认修好:撤回提示已经是同伴文档开头的第一个引用块,出处不可达这一点已写明为 provenance metadata。
所以我不怀疑这份快照的正确性和边界,也不认为归档体积本身有问题——它由"发布字节"这个目的决定,且完全不落在产品运行时。阻断我的是另外一件事:这条身份检查在两个普通环境里会对未改动的快照报红,一个是仓库自己要求的 premerge 门之后,一个是 Windows 默认检出的行尾转换。两处修法都是几行(跳过被 gitignore 的字节码;.gitattributes 里 pin 住行尾),修完这次改动就只剩 P3 的文档回声。下一次评审请从新的 exact head 开始。
English verdict: REQUEST_CHANGES — exact head 1e7556555527b66a867c40666232424c866d83ab of #4498. The archive is clean, public-safe, inert and independently verifiable (no absolute host paths, internal hosts, credentials or trajectories in the 14 files; git rev-parse HEAD:benchmark/deepswe-gptxhigh-v1 equals the pinned tree 1bc5d2b3b74761a97d34ba3f3612e977fd610340; a mutated byte and a dropped execute bit both fail closed; 19/19 pre-merge checks pass with one benchmark-sensitive manual hold; docs governance ok; clean merge into current main). Two blocking P2s, both about the verification keeping its promise rather than about the frozen bytes: (1) loopx canary premerge --from-git-diff compiles the changed archive files with python3 -m py_compile, CPython writes __pycache__/ into the frozen directory, and the documented acceptance command then returns {"ok": false, "error": "unexpected archive entry: __pycache__"} (reproduced end to end from a clean git archive extract; -B does not prevent it) — minimum repair: skip gitignored bytecode in check_archive per the repository convention in loopx/semantics/inventory.py:24; (2) nothing pins the checked-out bytes, so a core.autocrlf=true checkout converts the archive to CRLF and the check fails with a misleading bash -n syntax error on an unchanged file (reproduced with git -c core.autocrlf=true worktree add; .gitattributes has no rule for the archive and .github/workflows/python-tests.yml:459 runs a Windows job) — minimum repair: benchmark/deepswe-gptxhigh-v1/** text eol=lf, safe because all 14 blobs contain zero CR bytes. One non-blocking P3: loopx/capabilities/benchmark_toolkit/README.md:1362-1365 and docs/development/documentation-layout.md:41 still state the pre-change placement rule, so the newly declared owner of archive placement has two contradicting echoes. Residual risk: the provenance commit 98c262a is unreachable upstream (disclosed), the archived code was never executed here, the author-reported datacurve-pier import/dispatch smoke is not reproducible in this environment, remote CI was intentionally not consulted (wait_for_ci=false), and the branch is BEHIND main, so the merge check and all reproductions must be repeated after any rebase.
Signed-off-by: gwh6669999 <gwh2860667743@gmail.com>
Archive the original DeepSWE GPT xhigh v1 code and historical 113-task summary under
benchmark/deepswe-gptxhigh-v1/. This is a clean replacement for #4495, based on currentmain, without the later revised package or its development commits.本 PR 仅归档旧 v1 的原始代码与结果。原始 14 个文件(含 README)的内容和权限完全不变,不包含正在运行的新 benchmark,也不包含
v1-revised执行逻辑修复。Scope:
1bc5d2b3b74761a97d34ba3f3612e977fd610340, matching98c262a:benchmark/deepswe-five-armexactly apart from the containing directory name.python3 benchmark/check_deepswe_v1.pychecks archive contents/modes and Python/Bash syntax without model calls or task execution.Validation passed: original tree identity, syntax for 11 Python files, 2 embedded Python programs and 2 Bash scripts, and whitespace checks. The commit includes a DCO Signed-off-by trailer.
This is an immutable historical export, not a standalone runnable release. External runner/task/network/provider dependencies and known execution defects remain documented. No end-to-end benchmark run or revalidation of scores is claimed. The existing SSH Goal / Codex CLI withdrawal caveat remains applicable.