Gen3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas Muller, Alexander Keller, Sanja Fidler, Jun Gao

arXiv:2503.03751 · 2026-07-27 공개 · arXiv · PDF

video-generation novel-view-synthesis camera-control generative-model monocular-video point-clouds sparse-view depth-prediction

Abstract

We present Gen3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if implemented at all, is imprecise, because camera parameters are mere inputs to the neural network which must then infer how the video depends on the camera. In contrast, Gen3C is guided by a 3D cache: point clouds obtained by predicting the pixel-wise depth of seed images or previously generated frames. When generating the next frames, Gen3C is conditioned on the 2D renderings of the 3D cache with the new camera trajectory provided by the user. Crucially, this means that Gen3C neither has to remember what it previously generated nor does it have to infer the image structure from the camera pose. The model, instead, can focus all its generative power on previously unobserved regions, as well as advancing the scene state to the next frame. Our results demonstrate more precise camera control than prior work, as well as state-of-the-art results in sparse-view novel view synthesis, even in challenging settings such as driving scenes and monocular dynamic video. Results are best viewed in videos. Check out our webpage! 1

한국어 요약

한 줄 요약

Gen3C는 3D 캐시를 기반으로 정밀 카메라 제어와 시간적 3D 일관성을 달성하는 비디오 생성 모델이다.

핵심 기여도

핵심 아이디어

Gen3C는 기존 비디오 생성 모델이 3D 정보를 거의 활용하지 않아 발생하는 일관성 문제를 해결하기 위해 3D 캐시를 도입한다. 이 캐시는 입력 이미지 또는 이전 생성 프레임의 픽셀별 깊이를 추정하여 생성된 포인트 클라우드 형태로 저장된다. 이후 사용자가 제공한 카메라 트래젝토리에 따라 3D 캐시를 2D 렌더링하여 비디오 생성 모델의 조건으로 활용한다. 이 방식은 모델이 이전 생성 내용을 기억하거나 카메라 포즈에서 이미지 구조를 추론할 필요가 없게 하여, 생성 능력을 미관측 영역과 장면 상태 전진에 집중시킨다.

핵심적인 통찰은 3D 캐시가 명시적인 기하학적 정보를 제공함으로써, 비디오 생성 모델이 카메라 제어와 장면 일관성을 향상시킬 수 있다는 점이다. 이는 기존의 텍스트나 이미지 조건만을 사용하는 방식과 구별된다.

기술적 접근법

주요 결과

의의 및 한계

Gen3C는 비디오 생성 모델에 명시적인 3D 기하학 정보를 도입함으로써, 장면 일관성과 카메라 제어를 획기적으로 향상시킨다. 특히, 드라이빙 시뮬레이션과 같은 복잡한 환경에서도 뛰어난 성능을 보이며, 디지털 콘텐츠 제작 및 시뮬레이션 환경에서의 실용성을 입증한다. 또한, 3D 캐시를 기반으로 객체 제거 및 장면 편집과 같은 고급 작업이 가능하다는 점에서 기존 모델과 차별화된다.

하지만, 동적 콘텐츠 생성 시에는 사전 생성된 비디오에 의존하여 객체의 움직임을 제공해야 하므로, 이 자체가 별도의 도전 과제가 된다. 향후 텍스트 조건을 통한 움직임 제어를 통합하는 것이 유망한 확장 방향으로 제시된다.

실용적 활용

Gen3C는 영화 제작, VR/AR, 로봇 시뮬레이션, 게임 개발 등에서 사용자 지정 카메라 트래젝토리와 장면 일관성을 요구하는 비디오 생성 작업에 적용 가능하다. 특히, 단일 이미지나 제한된 뷰 입력으로도 고해상도 3D 일관성 비디오를 생성할 수 있어, 콘텐츠 제작 효율성을 높이는 데 기여할 수 있다.