Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, K. Koishida, A. Bucker, Lawrence Jang, Zack Hui
arXiv:2409.08264 · 2026-07-27 공개 · arXiv · PDF
multi-modal llm-agent task-planning tool-usage navi-agent azure windows-agent os-agent
Abstract
Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena. Webpage: https://microsoft.github.io/WindowsAgentArena Code: https://github.com/microsoft/WindowsAgentArena
한국어 요약
한 줄 요약
Windows Agent Arena는 Windows OS 내에서 다중 모달 에이전트를 대규모로 평가하는 벤치마크 환경으로, Navi 에이전트의 성능을 19.5%의 성공률로 평가한다.
핵심 기여도
- Windows Agent Arena라는 새로운 벤치마크 환경을 소개, 154개의 다중 모달 Windows OS 기반 태스크를 포함.
- Azure 클라우드를 활용한 평가 병렬화로 전체 평가 시간을 20분 이내로 단축.
- 새로운 다중 모달 에이전트 Navi를 제시, Windows 도메인에서 19.5% 성공률 달성.
- Mind2Web 벤치마크에서도 Navi의 경쟁력 있는 성능을 검증.
핵심 아이디어
기존 에이전트 평가 시스템은 특정 도메인(예: 텍스트, 웹 탐색)에 제한적이며, 복잡한 다단계 작업의 평가가 느리고 비효율적이었다. Windows Agent Arena는 실제 Windows OS 환경에서 작동하는 에이전트를 평가하기 위해 설계되었으며, OSWorld 프레임워크를 기반으로 154개의 다양한 태스크를 구축했다. 이는 계획 수립, 화면 이해, 도구 사용 등 인간과 유사한 능력을 평가할 수 있도록 설계되었다. 또한, Azure 클라우드 기반의 병렬 처리를 통해 평가 시간을 기존 일 단위에서 20분 이내로 단축하는 것이 핵심 기술적 혁신이다. Navi 에이전트는 Set-of-Marks 프롬프팅과 시스템 접근성 트리, 픽셀 기반 요소 탐지기를 결합하여 성능을 향상시켰다.
기술적 접근법
- **Windows Agent Arena**: Windows OS 기반의 실제 환경에서 작동하는 에이전트 평가 플랫폼.
- **OSWorld 프레임워크 활용**: 154개의 다중 모달 태스크를 생성.
- **Azure 클라우드 기반 병렬 평가**: Docker 컨테이너를 사용한 Azure VM에서 실행, 평가 시간 20분 이내.
- **Navi 에이전트**: Set-of-Marks 프롬프팅, 시스템 접근성 트리, 픽셀 기반 요소 탐지기 사용.
- **성능 평가**: WindowsAgentArena 내에서 19.5% 성공률, Mind2Web에서 경쟁력 있는 성능.
주요 결과
- **WindowsAgentArena**: 154개의 Windows OS 기반 다중 모달 태스크 포함.
- **Navi 성능**: Windows 도메인에서 19.5% 성공률 (베이스라인인 무보조 인간 대비 74.5% 대비 +19.5% 개선).
- **Mind2Web 성능**: Navi는 웹 기반 벤치마크에서도 경쟁력 있는 성능을 보임.
- **평가 시간**: Azure 클라우드 기반 병렬 처리로 20분 이내에 전체 평가 가능.
의의 및 한계
Windows Agent Arena는 실제 Windows 환경에서 작동하는 에이전트의 평가를 가능하게 하며, 대규모 병렬 평가를 통해 연구 효율성을 높이는 데 기여한다. 또한, Navi 에이전트는 Set-of-Marks 프롬프팅과 접근성 트리 기반의 화면 이해 기법을 통해 Windows 도메인에서 일정 수준의 성능을 보인다. 그러나 19.5%의 성공률은 인간 수준(74.5%)에 비해 여전히 낮으며, 복잡한 다단계 작업 처리 능력이 한계가 있다. 또한, 평가 시스템의 확장성과 에이전트의 자율성, 안전성 문제는 추가 연구가 필요하다.
실용적 활용
Windows Agent Arena는 연구자들이 Windows 환경에서 작동하는 에이전트를 개발하고 평가하는 데 유용하며, 특히 다중 모달 작업 처리 능력을 테스트하는 데 적합하다. Navi 에이전트는 프롬프팅 기법과 접근성 트리 기반의 화면 이해 기법을 활용한 에이전트 개발에 참고가 될 수 있으며, 클라우드 기반 평가 시스템은 대규모 실험을 효율적으로 수행할 수 있는 기반을 제공한다.