Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, Hongyang Li

arXiv:2405.17398 · 2026-07-27 공개 · arXiv · PDF

video-generation autonomous-driving generalization high-fidelity reward-learning fidelity-measurement driving-world-model latent-replacement

Abstract

World models can foresee the outcomes of different actions, which is of paramount importance for autonomous driving. Nevertheless, existing driving world models still have limitations in generalization to unseen environments, prediction fidelity of critical details, and action controllability for flexible application. In this paper, we present Vista, a generalizable driving world model with high fidelity and versatile controllability. Based on a systematic diagnosis of existing methods, we introduce several key ingredients to address these limitations. To accurately predict real-world dynamics at high resolution, we propose two novel losses to promote the learning of moving instances and structural information. We also devise an effective latent replacement approach to inject historical frames as priors for coherent long-horizon rollouts. For action controllability, we incorporate a versatile set of controls from high-level intentions (command, goal point) to low-level maneuvers (trajectory, angle, and speed) through an efficient learning strategy. After large-scale training, the capabilities of Vista can seamlessly generalize to different scenarios. Extensive experiments on multiple datasets show that Vista outperforms the most advanced general-purpose video generator in over 70% of comparisons and surpasses the best-performing driving world model by 55% in FID and 27% in FVD. Moreover, for the first time, we utilize the capacity of Vista itself to establish a generalizable reward for real-world action evaluation without accessing the ground truth actions.

한국어 요약

한 줄 요약

Vista는 고해상도 예측과 다중 조작 가능성을 갖춘 일반화 가능한 자율 주행 월드 모델로, 기존 모델 대비 55% 개선된 FID 성능을 보인다.

핵심 기여도

핵심 아이디어

기존 월드 모델은 일반화 능력, 예측 정확도, 조작 가능성 측면에서 한계가 있었다. Vista는 이 문제를 해결하기 위해 시스템적 진단을 바탕으로 세 가지 핵심 요소를 도입했다. 첫째, 고해상도 예측을 위해 움직임 인스턴스와 구조 정보를 학습하는 두 개의 새로운 손실 함수(‘dynamics enhancement loss’, ‘structure preservation loss’)를 제안한다. 둘째, 장기 예측 일관성을 위해 과거 프레임을 prior로 삼는 ‘latent replacement’ 방법을 사용한다. 셋째, 다양한 조작 조건(command, goal point, trajectory, angle, speed)을 통합한 조작 인터페이스를 설계하여, 고수준 의도부터 저수준 조작까지 유연하게 반영할 수 있도록 했다.

기술적 접근법

주요 결과

의의 및 한계

Vista는 고해상도, 장기 예측, 다중 조작 가능성을 갖춘 월드 모델로서, 자율 주행 분야에서의 일반화 능력과 안정성을 높이는 데 기여한다. 특히, 기존 모델이 단일 조작 조건만 지원하는 반면, Vista는 다양한 조작 조건을 통합한 인터페이스를 제공하여 유연한 적용이 가능하다. 또한, Vista 자체를 reward function으로 활용할 수 있어, 실제 행동 평가에 새로운 가능성을 열었다. 그러나 계산 효율성, 훈련 규모, 품질 유지 측면에서 여전히 개선이 필요하며, 이는 향후 연구 주제로 제시된다.

실용적 활용

Vista는 자율 주행 시스템의 장기 예측 및 행동 평가에 활용될 수 있으며, 특히 다양한 환경에서 일반화 가능한 월드 모델이 필요한 연구 및 산업 분야에서 유용할 것으로 기대된다. 또한, planning 알고리즘과의 호환성을 고려한 조작 인터페이스는 실제 차량 제어 시스템에 직접 적용 가능하다.