A curated list of benchmarks for evaluating multimodal agents — systems that plan, use tools, and act over inputs/outputs beyond plain text.
This list focuses strictly on benchmarks and evaluation suites. For agents and models, see the sibling lists:
- awesome-video-agents — agentic video editing / production systems
- awesome-any2any-models — unified any-to-any multimodal models
A benchmark qualifies if it evaluates agents that operate over multimodal inputs and/or outputs. It must satisfy both:
- Agentic — tasks require multi-step planning, tool use, environment interaction, or long-horizon reasoning. Single-turn QA does not qualify.
- Multimodal — inputs and/or outputs include modalities beyond text: image, video, audio, GUI screenshots, 3D scenes, robot sensors, etc.
Excluded (with rationale):
- Pure-text agent benchmarks (AgentBench, SWE-bench, ToolBench) — agentic but not multimodal. Listed under Related only.
- Pure multimodal QA benchmarks (MMBench, MME, MMMU, MMVet) — multimodal but not agentic.
- Pure video/image perception benchmarks (Video-MME, EgoSchema, LVBench when used as static QA) — unless cast as an agentic task.
Included by judgment call: GUI / web / mobile benchmarks (screenshots are a primary multimodal input), embodied & robotic suites, and visual tool-use benchmarks.
- Comparison Table
- GUI & Computer-Use Agents
- Web Agents
- Mobile / Android Agents
- Embodied & Robotic Agents
- Video Understanding Agents
- Document & Multimodal Reasoning Agents
- General-Purpose Multimodal Agent Suites
- Tool-Use & Function-Calling (Multimodal)
- Domain-Specific (Medical, Enterprise, Science)
- Related Surveys & Aggregators
- Related (Text-Only or Non-Agentic) for Context
- Contributing
- License
Entries within each section are sorted newest first by initial arXiv / publication year.
The most-cited / most-used benchmarks at a glance. Sizes are approximate.
| Name | Modalities | Tasks | Size | Year |
|---|---|---|---|---|
| OSWorld | Screenshot + A11y tree | Real OS computer-use | 369 tasks | 2024 |
| WindowsAgentArena | Screenshot + A11y tree | Windows OS tasks | 150+ tasks | 2024 |
| WebArena | HTML + Screenshot | Realistic web nav | 812 tasks | 2023 |
| VisualWebArena | Screenshot + HTML | Visually grounded web | 910 tasks | 2024 |
| Mind2Web | HTML + Screenshot | Generalist web nav | ~2k tasks, 137 sites | 2023 |
| Online-Mind2Web | Live web + Screenshot | Live web nav | 300 tasks, 136 sites | 2025 |
| Mind2Web 2 | Live web + Screenshot | Agentic search | 130 tasks | 2025 |
| AndroidWorld | Android screenshots | Mobile app control | 116 tasks, 20 apps | 2024 |
| SPA-Bench | Android screenshots | Smartphone agents | 340 tasks | 2024 |
| GAIA | Text + Image + Files | General assistant | 466 questions | 2023 |
| AssistantBench | Web + Screenshot | Realistic web tasks | 214 tasks | 2024 |
| VisualAgentBench | Mixed (GUI/Embodied/Design) | Visual foundation agent | 5 envs | 2024 |
| EmbodiedBench | RGB + State | Embodied MLLM eval | 1,128 tasks, 4 envs | 2025 |
| VLABench | RGB + Lang | Long-horizon manipulation | 100 task categories | 2024 |
| GTA | Image + Text | Implicit tool use | 229 tasks | 2024 |
- OSWorld-Human — Gold human trajectories for every OSWorld task; measures agent step-efficiency.
papercode - OSWorld-G / Jedi — Fine-grained GUI grounding suite (564 samples) paired with the Jedi training set.
papercode - ScreenSpot-Pro — High-res professional GUI grounding across 23 pro apps, 5 domains, 3 OSes.
papercode - AgentStudio — Toolkit + benchmarks (GroundUI, IDMBench) for building general virtual agents.
papercode - WindowsAgentArena — Windows-OS extension of OSWorld with 150+ tasks across Office / browser / system.
papercode - OSWorld — Live Ubuntu / Windows / macOS environment with 369 execution-graded tasks.
papercodeleaderboard - Spider2-V — 494 enterprise data-engineering tasks across 20 pro apps; GUI + SQL + Python.
papercode - CRAB — Cross-environment (Ubuntu + Android) benchmark with graph-based eval, 120 tasks.
papercode - OmniACT — Desktop + web action grounding dataset for generalist multimodal agents.
papercode - ScreenAgent / ScreenSpot — Original GUI element grounding benchmark (text + icon targets).
papercode
- Mind2Web 2 — 130 long-horizon agentic-search tasks evaluated with an Agent-as-a-Judge rubric.
papercode - MM-BrowseComp — Multimodal browsing benchmark; reverse image search and image-grounded retrieval.
papercode - Online-Mind2Web — 300 live-web tasks across 136 sites; reveals overstated progress on static benchmarks.
papercode - ST-WebAgentBench — Safety- and trustworthiness-focused web agent benchmark across 6 risk dims.
papercode - WorkArena++ — 682 compositional ServiceNow tasks built atop WorkArena atoms.
papercode - AssistantBench — 214 realistic, time-consuming, multi-site web tasks.
papercode - MMInA — Multi-hop multimodal internet agent benchmark over live websites.
papercode - VisualWebArena — 910 visually grounded tasks across self-hosted Classifieds / Shopping / Reddit.
papercodeleaderboard - WorkArena — ServiceNow enterprise web tasks; introduced with the BrowserGym environment.
papercode - WebArena — Reproducible self-hosted web environment (shopping / forum / GitLab / CMS), 812 tasks.
papercodeleaderboard - Mind2Web — First generalist web-agent benchmark; ~2k tasks across 137 real websites.
papercode
- MobileWorld — Cross-app, MCP-augmented mobile agent benchmark; ~2x AndroidWorld step count.
papercode - A3 (Android Agent Arena) — Open Android agent arena spanning 21 third-party apps.
papercode - SPA-Bench — Multilingual cross-app smartphone benchmark with interactive scoring + efficiency metrics.
papercode - AndroidLab — XML + SoM Android agent benchmark across 9 apps, 138 tasks.
papercode - AndroidWorld — Dynamic Android benchmark with programmatic reward functions, 116 tasks, 20 real apps.
papercodeleaderboard - B-MoCA — Benchmarks mobile control agents under randomized device configurations.
papercode
- EmbodiedBench — 1,128 tasks across EB-ALFRED / EB-Habitat / EB-Navigation / EB-Manipulation for MLLM agents.
papercode - VLABench — 100 categories of long-horizon language-conditioned manipulation; tests VLAs + VLM/LLM workflows.
papercode - BEHAVIOR-1K — 1,000 everyday human activities in OmniGibson with realistic physics.
papercode - ALFWorld — Aligned text + ALFRED embodied environments; canonical sim-to-language transfer.
papercode - Habitat 2.0 / 3.0 — Photorealistic interactive simulator for navigation + social rearrangement.
papercode
- LVBench — Extreme long-video understanding requiring agentic retrieval + reasoning over hours.
papercode - VideoAgent benchmarks (NExT-QA / EgoSchema agentic split) — Long-form video understanding with LLM-as-agent over frames.
papercode - Video-MME (agentic subset) — Used as a probe for tool-using video agents; full set is QA, agentic systems are the standard SOTA.
paperleaderboard
Note: pure video-QA (EgoSchema, MVBench, Video-MME used statically) is excluded per the criteria above; entries here are commonly run with agentic frame-selection / retrieval pipelines.
- ChartAgent / ChartBench — Visually grounded tool-augmented agent benchmark over complex charts.
papercode - AEC-Bench — Multimodal benchmark for agentic systems in Architecture / Engineering / Construction.
paper
- TheAgentCompany — 175 enterprise tasks in a simulated software company; mixes web, GUI, code, comms.
papercodeleaderboard - VisualAgentBench (VAB) — Unified benchmark across Embodied, GUI, and Visual Design tasks for LMM agents.
papercode - BrowserGym Ecosystem — Unifies MiniWoB++, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX under one API.
papercode - GAIA — 466 multi-step general-assistant questions with image / audio / file attachments.
paperleaderboard
- M3-Bench — Multi-modal, multi-hop, multi-threaded tool-using MLLM agent benchmark.
paper - GTA — 229 real user queries with image attachments; tool + step choices are implicit.
papercodeleaderboard - m&m's — Tool-use benchmark with 4k+ multi-step plans across 33 tools (vision, audio, text).
papercode
- AgentClinic — Simulated clinical environments with multimodal images (radiology, pathology) + dialogue.
papercode
- GUI-Agents-Paper-List — Continuously updated catalog of GUI-agent papers and benchmarks.
- Awesome-Agent-Papers — Broad agent survey covering methodology, applications, and benchmarks.
- HAL Leaderboards — Princeton's holistic agent leaderboards (GAIA, SWE-bench, etc.).
- An Illusion of Progress? Assessing the Current State of Web Agents — Critical review; introduces Online-Mind2Web.
These are widely referenced but fall outside this list's criteria. Linked for orientation only.
- AgentBench — Text-only agent benchmark across 8 environments. Multimodal extensions exist but the core is text.
- SWE-bench / SWE-bench Multimodal — Code agents; SWE-bench-M adds screenshot-grounded bugs and is borderline-eligible.
- ToolBench / API-Bank — Text tool-use benchmarks.
- MMBench / MME / MMMU / MMVet — Multimodal QA, no agent loop.
- EgoSchema / Video-MME — Long-form video QA; commonly used by video agents but the benchmark itself is non-agentic.
See CONTRIBUTING.md. PRs welcome — please verify every link and respect the inclusion criteria.
To the extent possible under law, the maintainers have waived all copyright and related rights to this work.
