[EMNLP 2026] WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
-
Updated
Aug 22, 2026 - Python
[EMNLP 2026] WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Reliability-first agent harness with cross-vendor agent-as-judge review. Author with one CLI, review with another. No SDK lock-in.
To associate your repository with the agent-as-judge topic, visit your repo's landing page and select "manage topics."