Epona: Autoregressive Diffusion World Model for Autonomous Driving

Kaiwen Zhang, Zhenyu Tang, Xiaotao Hu, Xingang Pan, Xiaoyang Guo, Yuan Liu, Jing Huang, Li Yuan, Qian Zhang, Xiaoxiao Long, Xun Cao, Wei Yin

arXiv:2506.24113 · 2026-07-27 공개 · arXiv · PDF

autonomous-driving world-model trajectory-planning motion-planning autoregressive-diffusion long-horizon-prediction high-resolution-generation navsim-benchmark

Abstract

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is because conventional video diffusion models rely on global joint distribution modeling of fixed-length frame sequences rather than sequentially constructing localized distributions at each timestep. In this work, we propose Epona, an autoregressive diffusion world model that enables localized spatiotemporal distribution modeling through two key innovations: 1) Decoupled spatiotemporal factorization that separates temporal dynamics modeling from fine-grained future world generation, and 2) Modular trajectory and video prediction that seamlessly integrate motion planning with visual modeling in an end-toend framework. Our architecture enables high-resolution, long-duration generation while introducing a novel chain-of-forward training strategy to address error accumulation in autoregressive loops. Experimental results demonstrate state-of-the-art performance with 7.4 % FVD improvement and minutes longer prediction duration compared to prior works. The learned world model further serves as a realtime motion planner, outperforming strong end-to-end planners on NAVSIM benchmarks.

한국어 요약

한 줄 요약

Epona는 자율주행을 위한 고해상도, 장기 예측을 지원하는 오토회귀 확산 월드 모델로, 7.4% FVD 개선과 2분 이상 예측을 달성한다.

핵심 기여도

핵심 아이디어

기존 비디오 확산 모델은 고정 길이의 프레임 시퀀스에 대한 전역 결합 분포를 모델링하므로, 유연한 길이와 장기 예측에 한계가 있었다. Epona는 이 문제를 해결하기 위해 **spacetime-disentangled processing**을 도입하여, 시간 동역학은 GPT-style 트랜스포머가, 공간 렌더링과 운동 생성은 별도의 DiT(Diffusion Transformer)가 처리하도록 분리한다. 이는 **causal attention**을 활용한 시간적 인과성을 강화하면서, 각 모듈이 최적화된 방식으로 작동하게 한다.

또한, **asynchronous multi-modal generation**을 통해 3초 운동 경로와 다음 프레임을 병렬적으로 생성하며, **flow-matching objectives**를 통해 모달 간 일관성을 유지한다. 이는 기존의 토큰화 기반 모델이 시각 품질을 저하시키는 문제를 회피하면서, 운동 계획과 시각 예측을 동시에 수행할 수 있는 구조를 구축한다.

기술적 접근법

주요 결과

의의 및 한계

Epona는 기존 확산 모델의 장기 예측 한계와 GPT-style 모델의 시각 품질 저하 문제를 동시에 해결하며, **장기 운동 계획과 시각 예측을 통합한 end-to-end 월드 모델**로서의 가능성을 제시한다. 특히, **자율주행 시스템에서 실시간 운동 계획**(최대 20Hz)이 가능하다는 점에서 실용적 가치가 높다.

그러나, **모델 훈련 시 복잡한 계산 요구**와 **10프레임 이하의 조건 입력 제한**은 확장성에 한계가 있다. 또한, **다중 뷰 입력 없이 단일 카메라만 사용**한다는 점에서 일부 상황에서 성능 저하가 발생할 수 있다.

실용적 활용

Epona는 **자율주행 시스템**에서 미래 운동 경로와 주변 환경을 고해상도로 예측하는 데 활용 가능하다. 특히, **실시간 운동 계획**(20Hz)이 가능하므로, **차량 제어 및 안전 시스템**에 즉각적으로 적용할 수 있다. 또한, **자연스러운 장기 운동 시뮬레이션**을 통해 **자율주행 알고리즘 테스트 및 교육**에도 유용하게 사용될 수 있다.