PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

arXiv:2608.13552 · 2026-08-14 공개 · arXiv · PDF

long-horizon world-models video-quality geometry-consistency interaction-fidelity agent-players playworld out-of-sight-evolution

Abstract

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

한국어 요약

한 줄 요약

PlayWorld는 장기적 목표 기반의 대체적 에이전트 플레이어를 통해 비디오 월드 모델을 평가하는 벤치마크를 제시한다.

핵심 기여도

핵심 아이디어

기존 월드 모델 평가 방식은 사용자 행동에 따라 정의된 고정된 경로를 기반으로 하여, 모델 간 비교에 한계가 있었다. 예를 들어, 360도 회전을 위한 동일한 3개의 명령어가 다른 모델에서는 부분 회전만 수행할 수 있어, 기하학적 일관성 평가가 불가능하다. PlayWorld는 이 문제를 해결하기 위해 Agent Player라는 에이전트를 도입하여, 사용자와 유사한 방식으로 장기적 목표를 추구하도록 설계되었다. Agent Player는 각 단계에서 생성된 프레임과 행동 기록을 관찰하고, Keep, Stop, Extend, Correct, End와 같은 결정을 통해 행동 수와 지속 시간을 모델에 맞게 조정함으로써, 동일한 목표 하에서 공정한 비교가 가능하도록 한다.

기술적 접근법

주요 결과

의의 및 한계

PlayWorld는 장기적 상호작용 목표 기반의 월드 모델 평가를 가능하게 하여, 기존 고정 경로 기반 평가의 한계를 극복한다. 특히, Agent Player를 통해 모델별 행동 반응에 맞춘 적응적 평가가 가능하며, 4가지 핵심 차원에서 체계적인 평가를 제공한다. 그러나, 모든 월드 모델이 동일한 인터페이스를 제공하지 않거나, 일부는 웹 기반으로 접근이 제한되어 있어, 평가 범위 확장에 한계가 있을 수 있다. 또한, Agent Player의 결정 로직이 인간 사용자와 완전히 일치하지 않을 가능성도 존재한다.

실용적 활용

PlayWorld는 월드 모델의 장기적 상호작용 능력을 평가하는 데 활용될 수 있으며, 게임 개발, VR/AR 환경 구축, 인공지능 기반 시뮬레이션 등에서 모델 신뢰도를 검증하는 데 유용하다. 특히, 사용자와 유사한 방식으로 모델을 평가함으로써, 실제 적용 시 예상되는 문제를 사전에 파악할 수 있다.