Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella et al.
A benchmark for comprehensively evaluating multimodal LLM performance in vision-driven embodied agents, offering 1,128 tasks across 4 environments and 6 capability evaluation sets.
Compared to language-centric embodied agents, MLLM-based embodied agents lack comprehensive evaluation frameworks, making it difficult to assess their actual performance and limitations.
We designed 1,128 testing tasks covering high-level semantic tasks to low-level atomic actions across diverse environments (e.g., household, navigation, manipulation). We also curated six evaluation subsets for essential capabilities: commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning.
Evaluation of 24 state-of-the-art MLLMs shows that GPT-4o achieves the best average score of 28.9%, but all models struggle significantly with low-level manipulation. This benchmark provides a standardized evaluation platform for MLLM-based embodied agents and offers insights for future development.