Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang
arXiv:2608.23383 · 2026-08-29 공개 · arXiv · PDF
long-horizon world-model audio-visual-generation self-gradient-forcing wbench sana-wm-bench cross-shot-memory geometry-aware-conditioning
Abstract
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.
한국어 요약
한 줄 요약
JoyAI-Echo-1.5는 장기적 영상 생성과 인터랙티브 월드 모델링을 위한 통합 오디오-비주얼 생성 시스템으로, 메모리, 기하학적 제어, 롤아웃 기반 학습을 결합하여 일관성과 제어력을 향상시킨다.
핵심 기여도
- 장거리 영상 생성을 위한 **composable cross-shot memory** 도입: 다중 샷에서 시각적 증거와 음성 필터링을 통해 캐릭터 외형과 음성 일관성을 유지.
- 인터랙티브 월드 모델링을 위한 **6-DoF 카메라 트레젝토리** 생성: 이질적인 네비게이션 입력을 정확한 기하학적 경로로 변환.
- **Self-Gradient Forcing** 기법을 사용한 롤아웃 기반 학습: 짧고 긴 시간 범위에서 생성 안정성 향상.
- WBench에서 평균 81.7점으로 1위 성적 달성.
핵심 아이디어
JoyAI-Echo-1.5는 단일 클립 생성에서 벗어나, 장기적 스토리와 인터랙티브 월드를 생성하기 위해 **메모리 기반 학습**과 **기하학적 제어**를 결합한 새로운 접근법을 제안한다. 장거리 영상 생성 모듈은 **cross-shot memory**를 통해 이전 샷의 시각적 정보와 음성 필터링을 활용하여 캐릭터 일관성을 유지하며, 월드 모델링 모듈은 **6-DoF 카메라 트레젝토리**를 생성하여 다양한 환경에서 정밀한 제어를 가능하게 한다. 또한, **Self-Gradient Forcing**을 통해 모델이 스스로 생성한 롤아웃에서 학습함으로써 장기적 안정성을 확보한다.
기술적 접근법
- **Long-video variant**:
- **Composable cross-shot memory**를 도입하여 여러 샷의 시각적 정보를 집계.
- **Speech-filtered full-shot audio**를 사용한 음성 피쳐 추출.
- 텍스트, 이미지, 메모리 조건을 유연하게 결합.
- **World-model variant**:
- 이질적인 네비게이션 입력을 **6-DoF 카메라 트레젝토리**로 변환.
- **Geometry-aware conditioning pathway**를 통해 정밀한 카메라 제어.
- **Training**:
- **Bidirectional audio-visual backbone**을 **causal few-step generator**로 변환.
- **Progressive teacher forcing**과 **short- 및 long-horizon Self-Gradient Forcing**을 사용한 롤아웃 기반 학습.
주요 결과
- **Long-video generation**:
- 기존 베이스라인 대비 **cross-shot consistency**, **visual quality**, **text alignment**, **speech fidelity** 개선.
- **World modeling**:
- **WBench**에서 평균 81.7점으로 1위.
- **SANA-WM-Bench**에서 **visual quality**와 **long-horizon persistence**에서 우수한 성능.
- **Trajectory control**:
- **Rotation error (R)**, **relative translation error (T)**, **CMC**, **VBench Overall** 등 지표에서 높은 정확도.
의의 및 한계
JoyAI-Echo-1.5는 장거리 영상 생성과 인터랙티브 월드 모델링을 위한 실용적인 기반을 제공하며, **메모리**, **기하학적 제어**, **롤아웃 기반 학습**이 상호보완적으로 작용함을 입증한다. 특히, **WBench**와 **SANA-WM-Bench**에서의 높은 성능은 이 시스템이 다양한 응용 분야에서 유용함을 보여준다. 그러나 **복잡한 장거리 트레젝토리에서의 회전 드리프트**는 주요한 한계로, 향후 연구에서는 **기하학적 일관성**과 **지속적 월드 상태 표현**이 필요하다.
실용적 활용
JoyAI-Echo-1.5는 영화 제작, 게임 개발, 가상 현실(VR) 및 증강현실(AR) 환경 구축 등에서 활용 가능하다. 특히, **유연한 텍스트-이미지-메모리 조건**을 지원하는 장거리 영상 생성 기능은 스토리텔링 콘텐츠 제작에 적합하며, **6-DoF 카메라 제어**는 인터랙티브 환경에서의 정밀한 사용자 경험을 제공한다.