Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Awesome Multimodal Agent Benchmarks Awesome

A curated list of benchmarks for evaluating multimodal agents — systems that plan, use tools, and act over inputs/outputs beyond plain text.

This list focuses strictly on benchmarks and evaluation suites. For agents and models, see the sibling lists:

Scope & Inclusion Criteria

A benchmark qualifies if it evaluates agents that operate over multimodal inputs and/or outputs. It must satisfy both:

  1. Agentic — tasks require multi-step planning, tool use, environment interaction, or long-horizon reasoning. Single-turn QA does not qualify.
  2. Multimodal — inputs and/or outputs include modalities beyond text: image, video, audio, GUI screenshots, 3D scenes, robot sensors, etc.

Excluded (with rationale):

  • Pure-text agent benchmarks (AgentBench, SWE-bench, ToolBench) — agentic but not multimodal. Listed under Related only.
  • Pure multimodal QA benchmarks (MMBench, MME, MMMU, MMVet) — multimodal but not agentic.
  • Pure video/image perception benchmarks (Video-MME, EgoSchema, LVBench when used as static QA) — unless cast as an agentic task.

Included by judgment call: GUI / web / mobile benchmarks (screenshots are a primary multimodal input), embodied & robotic suites, and visual tool-use benchmarks.

Contents

Entries within each section are sorted newest first by initial arXiv / publication year.

Comparison Table

The most-cited / most-used benchmarks at a glance. Sizes are approximate.

Name Modalities Tasks Size Year
OSWorld Screenshot + A11y tree Real OS computer-use 369 tasks 2024
WindowsAgentArena Screenshot + A11y tree Windows OS tasks 150+ tasks 2024
WebArena HTML + Screenshot Realistic web nav 812 tasks 2023
VisualWebArena Screenshot + HTML Visually grounded web 910 tasks 2024
Mind2Web HTML + Screenshot Generalist web nav ~2k tasks, 137 sites 2023
Online-Mind2Web Live web + Screenshot Live web nav 300 tasks, 136 sites 2025
Mind2Web 2 Live web + Screenshot Agentic search 130 tasks 2025
AndroidWorld Android screenshots Mobile app control 116 tasks, 20 apps 2024
SPA-Bench Android screenshots Smartphone agents 340 tasks 2024
GAIA Text + Image + Files General assistant 466 questions 2023
AssistantBench Web + Screenshot Realistic web tasks 214 tasks 2024
VisualAgentBench Mixed (GUI/Embodied/Design) Visual foundation agent 5 envs 2024
EmbodiedBench RGB + State Embodied MLLM eval 1,128 tasks, 4 envs 2025
VLABench RGB + Lang Long-horizon manipulation 100 task categories 2024
GTA Image + Text Implicit tool use 229 tasks 2024

GUI & Computer-Use Agents

  • OSWorld-Human — Gold human trajectories for every OSWorld task; measures agent step-efficiency. paper code
  • OSWorld-G / Jedi — Fine-grained GUI grounding suite (564 samples) paired with the Jedi training set. paper code
  • ScreenSpot-Pro — High-res professional GUI grounding across 23 pro apps, 5 domains, 3 OSes. paper code
  • AgentStudio — Toolkit + benchmarks (GroundUI, IDMBench) for building general virtual agents. paper code
  • WindowsAgentArena — Windows-OS extension of OSWorld with 150+ tasks across Office / browser / system. paper code
  • OSWorld — Live Ubuntu / Windows / macOS environment with 369 execution-graded tasks. paper code leaderboard
  • Spider2-V — 494 enterprise data-engineering tasks across 20 pro apps; GUI + SQL + Python. paper code
  • CRAB — Cross-environment (Ubuntu + Android) benchmark with graph-based eval, 120 tasks. paper code
  • OmniACT — Desktop + web action grounding dataset for generalist multimodal agents. paper code
  • ScreenAgent / ScreenSpot — Original GUI element grounding benchmark (text + icon targets). paper code

Web Agents

  • Mind2Web 2 — 130 long-horizon agentic-search tasks evaluated with an Agent-as-a-Judge rubric. paper code
  • MM-BrowseComp — Multimodal browsing benchmark; reverse image search and image-grounded retrieval. paper code
  • Online-Mind2Web — 300 live-web tasks across 136 sites; reveals overstated progress on static benchmarks. paper code
  • ST-WebAgentBench — Safety- and trustworthiness-focused web agent benchmark across 6 risk dims. paper code
  • WorkArena++ — 682 compositional ServiceNow tasks built atop WorkArena atoms. paper code
  • AssistantBench — 214 realistic, time-consuming, multi-site web tasks. paper code
  • MMInA — Multi-hop multimodal internet agent benchmark over live websites. paper code
  • VisualWebArena — 910 visually grounded tasks across self-hosted Classifieds / Shopping / Reddit. paper code leaderboard
  • WorkArena — ServiceNow enterprise web tasks; introduced with the BrowserGym environment. paper code
  • WebArena — Reproducible self-hosted web environment (shopping / forum / GitLab / CMS), 812 tasks. paper code leaderboard
  • Mind2Web — First generalist web-agent benchmark; ~2k tasks across 137 real websites. paper code

Mobile / Android Agents

  • MobileWorld — Cross-app, MCP-augmented mobile agent benchmark; ~2x AndroidWorld step count. paper code
  • A3 (Android Agent Arena) — Open Android agent arena spanning 21 third-party apps. paper code
  • SPA-Bench — Multilingual cross-app smartphone benchmark with interactive scoring + efficiency metrics. paper code
  • AndroidLab — XML + SoM Android agent benchmark across 9 apps, 138 tasks. paper code
  • AndroidWorld — Dynamic Android benchmark with programmatic reward functions, 116 tasks, 20 real apps. paper code leaderboard
  • B-MoCA — Benchmarks mobile control agents under randomized device configurations. paper code

Embodied & Robotic Agents

  • EmbodiedBench — 1,128 tasks across EB-ALFRED / EB-Habitat / EB-Navigation / EB-Manipulation for MLLM agents. paper code
  • VLABench — 100 categories of long-horizon language-conditioned manipulation; tests VLAs + VLM/LLM workflows. paper code
  • BEHAVIOR-1K — 1,000 everyday human activities in OmniGibson with realistic physics. paper code
  • ALFWorld — Aligned text + ALFRED embodied environments; canonical sim-to-language transfer. paper code
  • Habitat 2.0 / 3.0 — Photorealistic interactive simulator for navigation + social rearrangement. paper code

Video Understanding Agents

Note: pure video-QA (EgoSchema, MVBench, Video-MME used statically) is excluded per the criteria above; entries here are commonly run with agentic frame-selection / retrieval pipelines.

Document & Multimodal Reasoning Agents

  • ChartAgent / ChartBench — Visually grounded tool-augmented agent benchmark over complex charts. paper code
  • AEC-Bench — Multimodal benchmark for agentic systems in Architecture / Engineering / Construction. paper

General-Purpose Multimodal Agent Suites

  • TheAgentCompany — 175 enterprise tasks in a simulated software company; mixes web, GUI, code, comms. paper code leaderboard
  • VisualAgentBench (VAB) — Unified benchmark across Embodied, GUI, and Visual Design tasks for LMM agents. paper code
  • BrowserGym Ecosystem — Unifies MiniWoB++, WebArena, VisualWebArena, WorkArena, AssistantBench, WebLINX under one API. paper code
  • GAIA — 466 multi-step general-assistant questions with image / audio / file attachments. paper leaderboard

Tool-Use & Function-Calling (Multimodal)

  • M3-Bench — Multi-modal, multi-hop, multi-threaded tool-using MLLM agent benchmark. paper
  • GTA — 229 real user queries with image attachments; tool + step choices are implicit. paper code leaderboard
  • m&m's — Tool-use benchmark with 4k+ multi-step plans across 33 tools (vision, audio, text). paper code

Domain-Specific (Medical, Enterprise, Science)

  • AgentClinic — Simulated clinical environments with multimodal images (radiology, pathology) + dialogue. paper code

Related Surveys & Aggregators

Related (Text-Only or Non-Agentic) for Context

These are widely referenced but fall outside this list's criteria. Linked for orientation only.

Contributing

See CONTRIBUTING.md. PRs welcome — please verify every link and respect the inclusion criteria.

License

CC0

To the extent possible under law, the maintainers have waived all copyright and related rights to this work.

About

A curated list of benchmarks for evaluating multimodal agents.

Resources

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors