EchoWM: Open and Enterable Omnimodal World Models

Songchun Zhang, Yaowei Li, Junhao Zhuang, Weiyang Jin, Haoyu Wang, Xin Lu, Yilang Sun, Shiyi Zhang, Haoran Li, Xiaoxiao Ma, Yuming Li, Yijun Liu, Yaofeng Su, Yanwen Ma, Haoyu Wu, Zihan Su, Yue Ma, Lvmin Zhang, Haoyang Huang, Zeyue Xue, Anyi Rao, Nan Duan

arXiv:2608.23189 · 2026-08-25 공개 · arXiv · PDF

video-generation audio-generation trajectory-control omnimodal-world-models first-person autoregressive-training world-model-benchmarks enterable-media

Abstract

We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.

한국어 요약

한 줄 요약

EchoWM은 720p 영상, 환경음, 음악, 음성을 동시 생성하며, 1인칭 및 3인칭 시점의 연속 이동에 반응하는 옴니모달 월드 모델이다.

핵심 기여도

핵심 아이디어

EchoWM은 사용자의 **카메라 인텐트**(Camera Intent)를 중심으로 상호작용을 구성한다. 이는 1인칭 시점에서는 관찰자의 이동을, 3인칭 시점에서는 카메라-캐릭터 동역학을 학습하여, 뷰에 종속되지 않은 통일된 인터페이스를 제공한다. 이 인텐트는 **공유 상대 6-DoF 궤적**(Shared Metric-Scale Relative 6-DoF Trajectory)으로 표현되며, **Dataset-level Calibration**을 통해 이질적인 데이터 출처 간 이동의 크기를 일관되게 유지한다.

EchoWM의 핵심 통찰은 **관찰자-주체 관계**(Observer-Subject Relationship)를 데이터에서 학습하여, 1인칭/3인칭 시점을 구분 없이 동일한 인터페이스로 처리할 수 있다는 점이다. 이는 기존 시스템에서 요구되는 뷰-특화된 컨트롤러나 카메라 릿(Camera Rig) 없이도 가능하다.

기술적 접근법

주요 결과

의의 및 한계

EchoWM은 **생성 미디어**(Generative Media)와 **인터랙티브 월드 모델**(Interactive World Models)을 결합한 첫 번째 시도로, **1인칭/3인칭 시점의 통일된 인터페이스**, **멀티모달 생성**, **장기 생성 일관성**을 동시에 달성한 점에서 학술적·실용적 가치가 있다.

하지만, **사용자 행동**(예: 캐릭터의 임의 행동)은 카메라 이동으로 축소 표현되며, **별도의 음성-음향 컨트롤러**(Explicit Action-to-Sound Controller)는 제공되지 않아, **사운드 제어의 정확도**는 모델 내부의 **교차 모달 결합**(Cross-modal Coupling)에 의존한다는 한계가 있다.

실용적 활용

EchoWM은 **게임 개발**, **VR/AR 콘텐츠 생성**, **멀티모달 인터랙티브 미디어 플랫폼** 등에서 활용 가능하다. 특히, **사용자의 연속 이동에 반응하는 생성 환경**이 필요한 시나리오에서, **동시 생성되는 영상, 음향, 음성**을 통해 몰입감 있는 경험을 제공할 수 있다.