EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Rui Yang, Hanyang Chen, Junyu Zhang, Mark Zhao, Cheng Qian, Kangrui Wang, Qineng Wang, Teja Venkat Koripella, M. Movahedi, Manling Li, Heng Ji, Huan Zhang, Tong Zhang

arXiv:2502.09560 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation embodied-agents navigation commonsense-reasoning long-term-planning manipulation multi-modal-llm spatial-awareness

Abstract

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents remain underexplored due to the lack of comprehensive evaluation frameworks. To bridge this gap, we introduce EmbodiedBench, an extensive benchmark designed to evaluate vision-driven embodied agents. EmbodiedBench features: (1) a diverse set of 1,128 testing tasks across four environments, ranging from high-level semantic tasks (e.g., household) to low-level tasks involving atomic actions (e.g., navigation and manipulation); and (2) six meticulously curated subsets evaluating essential agent capabilities like commonsense reasoning, complex instruction understanding, spatial awareness, visual perception, and long-term planning. Through extensive experiments, we evaluated 24 leading proprietary and open-source MLLMs within EmbodiedBench. Our findings reveal that: MLLMs excel at high-level tasks but struggle with low-level manipulation, with the best model, GPT-4o, scoring only 28.9\% on average. EmbodiedBench provides a multifaceted standardized evaluation platform that not only highlights existing challenges but also offers valuable insights to advance MLLM-based embodied agents. Our code and dataset are available at https://embodiedbench.github.io.

한국어 요약

한 줄 요약

EmbodiedBench는 시각 기반 에미바디드 에이전트의 종합적 평가를 위한 벤치마크로, 1,128개의 다양한 태스크와 6개의 능력 중심 서브셋을 포함한다.

핵심 기여도

핵심 아이디어

기존 연구는 주로 언어 중심 에이전트에 집중했으나, MLLM 기반 에이전트는 평가 프레임워크 부족으로 연구가 제한적이었다. EmbodiedBench는 이 격차를 메우기 위해 설계된 벤치마크로, 고수준(예: 집안일)과 저수준(예: 네비게이션, 조작) 태스크를 모두 포함한다. 특히, 에이전트가 수행해야 할 핵심 능력을 6개의 서브셋으로 구분하여 평가함으로써, 단순 정확도보다 더 깊은 분석이 가능하다. 연구팀은 또한 에이전트 프레임워크에 시각 인식, 인-컨텍스트 학습, 환경 피드백 등을 통합하여 MLLM의 잠재력을 최대화했다.

기술적 접근법

주요 결과

의의 및 한계

EmbodiedBench는 MLLM 기반 에이전트의 평가를 체계화함으로써, 연구자들이 모델의 강점과 약점을 명확히 파악할 수 있도록 지원한다. 특히, 시각 정보의 중요성과 저수준 태스크 처리 능력의 부족을 드러내어, 향후 연구 방향을 제시한다. 그러나 EmbodiedBench는 특정 환경(예: EB-ALFRED, EB-Habitat)에 집중하며, 실생활 환경과의 직접적인 연계성은 명시되지 않음. 또한, 평가 대상 모델은 주로 영어 기반 MLLM에 한정되어 다국어 모델 평가는 제한적이다.

실용적 활용

EmbodiedBench는 로봇, 서비스 에이전트, 가상 환경 내 자율 시스템 등에 적용 가능한 평가 기준을 제공한다. 특히, MLLM 기반 에이전트의 시각 인식 및 장기 계획 능력을 개선하는 데 활용할 수 있으며, 산업 현장에서의 실용성 향상에 기여할 수 있다.