Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, K. Koishida, A. Bucker, Lawrence Jang, Zack Hui

arXiv:2409.08264 · 2026-07-27 공개 · arXiv · PDF

multi-modal llm-agent task-planning tool-usage navi-agent azure windows-agent os-agent

Abstract

Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in realistic environments remains a challenge since: (i) most benchmarks are limited to specific modalities or domains (e.g. text-only, web navigation, Q&A, coding) and (ii) full benchmark evaluations are slow (on order of magnitude of days) given the multi-step sequential nature of tasks. To address these challenges, we introduce the Windows Agent Arena: a reproducible, general environment focusing exclusively on the Windows operating system (OS) where agents can operate freely within a real Windows OS and use the same wide range of applications, tools, and web browsers available to human users when solving tasks. We adapt the OSWorld framework (Xie et al., 2024) to create 150+ diverse Windows tasks across representative domains that require agent abilities in planning, screen understanding, and tool usage. Our benchmark is scalable and can be seamlessly parallelized in Azure for a full benchmark evaluation in as little as 20 minutes. To demonstrate Windows Agent Arena's capabilities, we also introduce a new multi-modal agent, Navi. Our agent achieves a success rate of 19.5% in the Windows domain, compared to 74.5% performance of an unassisted human. Navi also demonstrates strong performance on another popular web-based benchmark, Mind2Web. We offer extensive quantitative and qualitative analysis of Navi's performance, and provide insights into the opportunities for future research in agent development and data generation using Windows Agent Arena. Webpage: https://microsoft.github.io/WindowsAgentArena Code: https://github.com/microsoft/WindowsAgentArena

한국어 요약

한 줄 요약

Windows Agent Arena는 Windows OS 내에서 다중 모달 에이전트를 대규모로 평가하는 벤치마크 환경으로, Navi 에이전트의 성능을 19.5%의 성공률로 평가한다.

핵심 기여도

핵심 아이디어

기존 에이전트 평가 시스템은 특정 도메인(예: 텍스트, 웹 탐색)에 제한적이며, 복잡한 다단계 작업의 평가가 느리고 비효율적이었다. Windows Agent Arena는 실제 Windows OS 환경에서 작동하는 에이전트를 평가하기 위해 설계되었으며, OSWorld 프레임워크를 기반으로 154개의 다양한 태스크를 구축했다. 이는 계획 수립, 화면 이해, 도구 사용 등 인간과 유사한 능력을 평가할 수 있도록 설계되었다. 또한, Azure 클라우드 기반의 병렬 처리를 통해 평가 시간을 기존 일 단위에서 20분 이내로 단축하는 것이 핵심 기술적 혁신이다. Navi 에이전트는 Set-of-Marks 프롬프팅과 시스템 접근성 트리, 픽셀 기반 요소 탐지기를 결합하여 성능을 향상시켰다.

기술적 접근법

주요 결과

의의 및 한계

Windows Agent Arena는 실제 Windows 환경에서 작동하는 에이전트의 평가를 가능하게 하며, 대규모 병렬 평가를 통해 연구 효율성을 높이는 데 기여한다. 또한, Navi 에이전트는 Set-of-Marks 프롬프팅과 접근성 트리 기반의 화면 이해 기법을 통해 Windows 도메인에서 일정 수준의 성능을 보인다. 그러나 19.5%의 성공률은 인간 수준(74.5%)에 비해 여전히 낮으며, 복잡한 다단계 작업 처리 능력이 한계가 있다. 또한, 평가 시스템의 확장성과 에이전트의 자율성, 안전성 문제는 추가 연구가 필요하다.

실용적 활용

Windows Agent Arena는 연구자들이 Windows 환경에서 작동하는 에이전트를 개발하고 평가하는 데 유용하며, 특히 다중 모달 작업 처리 능력을 테스트하는 데 적합하다. Navi 에이전트는 프롬프팅 기법과 접근성 트리 기반의 화면 이해 기법을 활용한 에이전트 개발에 참고가 될 수 있으며, 클라우드 기반 평가 시스템은 대규모 실험을 효율적으로 수행할 수 있는 기반을 제공한다.