You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: apps/presentation/site/public/blog/application-scenarios/index.html
+24-4Lines changed: 24 additions & 4 deletions
Original file line number
Diff line number
Diff line change
@@ -83,8 +83,27 @@ <h2>01 / Finish a complex task</h2>
83
83
</section>
84
84
85
85
<sectionid="benchmark">
86
-
<h2>Benchmarks: three studies, three signals worth investigating</h2>
87
-
<p>The studies below observe continuation, delivery, and validation behavior under different tasks, model settings, and measurement rules. Each needs to be read on its own terms. The current evidence does not establish a universal LoopX gain.</p>
86
+
<h2>Benchmarks: start with LHTB, then three supporting signals</h2>
87
+
<p>Start with recovery and regression across long-horizon domains, then use three software-engineering studies to examine continuation, delivery, and validation. The four studies differ in tasks, model settings, and measurement rules. Each needs to be read on its own terms; the current evidence does not establish a universal LoopX gain.</p>
88
+
89
+
<divclass="study-signal" id="benchmark-lhtb">
90
+
<h3>LHTB: recover progress and preserve working results</h3>
91
+
<pclass="study-setting">GPT-5.6 Sol / max · 46 matched tasks, five execution mechanisms · one effective run per task-arm cell</p>
92
+
<p><strong>What LHTB tests:</strong> Long-Horizon Terminal-Bench places an agent in a stateful container and asks it to sustain work across hundreds of dependent terminal actions. Its official introduction uses nine categories: software and reverse engineering; scientific computing and simulation; earth, climate, and energy; multimodal and imaging analysis; research reproduction and ML; systems, performance, and security; interactive games; APEX professional workflows; and logic and constraint puzzles.</p>
93
+
<p>Tasks include migrating an old framework, recovering data from scientific figures, reproducing paper experiments, handling investment-banking or legal matters, playing 2048 turn by turn, and searching for puzzle solutions. <ahref="../../benchmarks/lhtb/?lang=en#categories">Nine categories and representative tasks →</a></p>
94
+
<p><strong>How it scores:</strong> hidden verifiers check final artifacts or replayable outcomes and award continuous reward from 0 to 1. This study counts reward ≥ 0.95 as solved, keeping partial progress distinct from full acceptance.</p>
95
+
<divclass="table-scroll"><tableclass="benchmark-table"><caption>LHTB five-arm results: mean reward and strict solves</caption><thead><tr><thscope="col">Execution mechanism</th><thscope="col">Mean reward</th><thscope="col">Solved / 46</th></tr></thead><tbody>
<p><strong>Observation:</strong> New Heartbeat has the highest mean reward, 0.0154 above Legacy Heartbeat. Its strict solve count matches Plain and Legacy at 7/46. Higher average progress has not translated into more fully solved tasks.</p>
103
+
<p><strong>Insight:</strong> each new Heartbeat wake starts a fresh executor, recovering progress from the workspace, registry, and Todos. The brief reports a consulting task resumed at S27 after interruption, alongside a DuckDB regression where later optimization broke correctness. The next questions are whether externalized state reliably enables recovery and whether checkpoints and rollback preserve working results.</p>
104
+
<pclass="signal-boundary"><strong>Boundary:</strong> transport, session lifetime, LoopX version, and replanning change together. Effective runs include replacement trials, some with longer time limits. There are no repeated seeds, and cost telemetry differs, so this does not isolate a mechanism or establish an equal-budget efficiency advantage.</p>
105
+
<pclass="study-link"><ahref="../../benchmarks/lhtb/?lang=en">Full study: five mechanisms, gains and losses, and the 46-task matrix →</a></p>
106
+
</div>
88
107
89
108
<divclass="study-signal" id="benchmark-marathon">
90
109
<h3>SWE-Marathon: continuation must close real gaps</h3>
@@ -113,8 +132,8 @@ <h3>DeepSWE × V4 Flash max: let counterexamples change the implementation</h3>
113
132
<pclass="study-link"><ahref="../../benchmarks/deepswe/behavior-discovery/">Full study: requirement coverage, behavior cases, and the long-duration slice →</a></p>
114
133
</div>
115
134
116
-
<pclass="takeaway">The current signal is that useful continuation, reliable delivery, and counterexample-driven repair deserve further study.</p>
117
-
<p>The studies offer local positive observations while exposing cost, failure, and attribution problems. Matched budgets and repeated experiments are needed to determine which gains reproduce reliably.</p>
135
+
<pclass="takeaway">Recover progress, preserve working results, and check whether continuation closes acceptance gaps.</p>
136
+
<p>The four studies offer local positive observations while exposing regressions, cost, and attribution problems. Matched budgets and repeated experiments should examine recovery, rollback, delivery, and counterexample-driven repair separately.</p>
118
137
<pclass="small-note">Withdrawn SSH Goal and Codex CLI scores and conclusions from SWE-Marathon and DeepSWE × Sol are not used in these comparisons.</p>
119
138
</section>
120
139
@@ -182,6 +201,7 @@ <h2>Evolution: add one verifiable capability at a time</h2>
182
201
<h2>Public sources and further reading</h2>
183
202
<pclass="small-note">The technical content is derived only from public repository material. Source and data links are pinned to the revisions read for this article; historical experiments retain their own versions. This page introduces no new experimental results.</p>
184
203
<olclass="source-list">
204
+
<li><ahref="../../benchmarks/lhtb/?lang=en">LHTB: five long-horizon execution mechanisms</a>; <ahref="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/data.json">public aggregates for 46 tasks</a>; <ahref="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/benchmark/LHTB/studies/five-arm-gpt56sol-max/README.md">setup and evidence boundary</a>; <ahref="https://github.com/huangruiteng/loopx/blob/d60813451609ad28b2855fe7acda9a4701740c40/apps/presentation/site/src/lhtb-copy.json">mechanisms and cases reported in the brief</a>; <ahref="https://zli12321.github.io/LHTB/index.html">official benchmark</a>. See the study for contributor attribution.</li>
185
205
<li><ahref="../../benchmarks/swe-marathon/">SWE-Marathon: continuous self-verification</a>; <ahref="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/README.md">setup, positive and negative cases, and limitations</a>; <ahref="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/swe-marathon/data.json">public aggregate data</a>. See the study for contributors and case provenance.</li>
186
206
<li><ahref="../../benchmarks/deepswe-sol/">DeepSWE × Sol: from continued execution to valid delivery</a> (Chinese). The standalone brief covers historical results over 113 tasks, mechanism diagrams, and pinned primary sources; research archive contribution: <ahref="https://github.com/huangruiteng/loopx/pull/4502">@gwh6669999, #4502</a>.</li>
187
207
<li><ahref="../../benchmarks/deepswe/behavior-discovery/">DeepSWE × V4 Flash max: from hints to behavior</a>; <ahref="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/README.md">disclosure boundary</a>; <ahref="https://github.com/huangruiteng/loopx/blob/56fa3f11b5bbe108e1bf4de254fcbdf74a84a505/benchmark/deepswe/behavior-discovery/index.html">charts, cases, and metrics at the pinned revision</a>.</li>
0 commit comments