Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models

Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh

arXiv:2608.19556 · 2026-08-30 공개 · arXiv · PDF

video-generation autoregressive-diffusion motion-prior perceptual-anchor geometry-drift ar-video reconstruction-reward stream4d

Abstract

Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction. However, a single rigid 3d reconstruction cannot model a dynamic scene, so this critic penalizes genuine object motion as reconstruction error and is maximized by freezing the video. This shortcut is especially detrimental in the AR setting, where each chunk can propagate an already-static configuration. In this work, we propose Stream4D, which replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards. To further guide motion magnitude and quality, we add a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts. Our final recipe combines these two terms with a lightweight perceptual anchor. Across various autoregressive video backbones and various generation horizons, Stream4D improves 4D reconstruction quality, preserves motion more effectively, and achieves higher human-aligned preference. Project page: https://banyuanhao.github.io/Stream4D/

한국어 요약

한 줄 요약

Stream4D는 4D 일관성을 강화한 강화학습 기반 스트리밍 자동회귀 확산 영상 모델 훈련 프레임워크로, 4D-PSNR 6.76 dB 개선을 달성한다.

핵심 기여도

핵심 아이디어

기존 스트리밍 자동회귀 확산 모델은 3D Gaussian-Splatting 기반의 정적 비판자(critic)를 사용해 운동을 제약하거나 완전히 제거하는 경향이 있었다. 이는 특히 자동회귀 설정에서 이전 프레임의 정적 구성을 다음 청크로 전파하는 단점을 초래한다. Stream4D는 이 문제를 해결하기 위해 **4D Gaussian-Splatting (4D-GS)** 모델을 기반으로 한 **feed-forward 4D reconstruction reward**를 도입한다. 이는 시간과 공간에서 일관된 운동을 보상하며, **coherent motion**을 유도한다. 또한, **Gaussian motion reward**를 통해 자연스러운 운동 크기와 **jitter**, **non-rigid artifact**를 제어하며, **perceptual anchor**를 통해 시각적 품질을 보존한다.

기술적 접근법

주요 결과

의의 및 한계

Stream4D는 자동회귀 영상 생성 모델에서 **4D 일관성**을 강화하는 첫 번째 강화학습 기반 프레임워크로, **움직임 보존**과 **시각적 일관성**을 동시에 달성한다. 기존 3D 기반 비판자의 한계를 극복하고, 다양한 백본으로의 전이 가능성을 보여준다. 그러나 **4D-GS 모델의 재구성 능력**에 의존하므로, 복잡한 다물체 운동이나 장기적 일관성 유지에는 한계가 있을 수 있다. 또한, **4D-GS 모델의 학습 데이터와 품질**이 최종 성능에 큰 영향을 미친다.

실용적 활용

Stream4D는 실시간 스트리밍 영상 생성, 시뮬레이션 환경, 에이전트 기반 인터랙티브 애플리케이션에 적용 가능하다. 특히, **장기 운동 일관성**과 **시각적 품질**이 중요한 VR, AR, 게임 개발 분야에서 유용하게 활용될 수 있다.