Recammaster: Camera-Controlled Generative Rendering From a Single Video

Jianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang, Lianrui Mu, Jinwen Cao, Zuozhu Liu, Haoji Hu, Xiang Bai, Pengfei Wan, Di Zhang

arXiv:2503.11647 · 2026-07-27 공개 · arXiv · PDF

text-to-video camera-control generative-model super-resolution unreal-engine multi-camera video-rendering video-stabilization

Abstract

Camera control has been actively studied in text or image conditioned video generation tasks. However, altering camera trajectories of a given video remains under-explored, despite its importance in the field of video creation. It is non-trivial due to the extra constraints of maintaining multiple-frame appearance and dynamic synchronization. To address this, we present ReCamMaster, a cameracontrolled generative video re-rendering framework that reproduces the dynamic scene of an input video at novel camera trajectories. The core innovation lies in harnessing the generative capabilities of pre-trained text-to-video models through a simple yet powerful video conditioning mecha-nism-its capability is often overlooked in current research. To overcome the scarcity of qualified training data, we construct a comprehensive multi-camera synchronized video dataset using Unreal Engine 5, which is carefully curated to follow real-world filming characteristics, covering diverse scenes and camera movements. It helps the model generalize to in-the-wild videos. Lastly, we further improve the robustness to diverse inputs through a meticulously designed training strategy. Extensive experiments show that our method substantially outperforms existing state-of-theart approaches. Our method also finds promising applications in video stabilization, super-resolution, and outpainting. Our code and dataset are publicly available at: https://github.com/KwaiVGI/ReCamMaster.

한국어 요약

한 줄 요약

ReCamMaster는 단일 영상에서 카메라 제어를 통해 새로운 시점의 동영상을 생성하는 생성적 렌더링 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 연구는 텍스트 또는 이미지 조건부 영상 생성에 집중했으나, 주어진 영상의 카메라 궤적을 변경하는 작업은 여전히 미비한 상태였다. ReCamMaster는 이 문제를 해결하기 위해 **사전 학습된 텍스트-영상 모델**(text-to-video model)의 생성 능력을 활용하는 **단순하지만 강력한 비디오 조건부 메커니즘**을 도입했다. 이 메커니즘은 **소스 영상과 타겟 카메라 궤적**을 결합하여 새로운 시점의 영상을 생성하며, 기존 연구에서 간과되었던 **동영상 조건부 생성**(conditional video generation)의 잠재력을 발휘한다. 특히, **프레임 차원의 토큰 연결**(frame concatenation)은 기존의 채널 연결(channel concatenation) 방식보다 **동영상 동기화 및 시각적 일관성**을 향상시킨다.

기술적 접근법

주요 결과

의의 및 한계

ReCamMaster는 카메라 궤적을 자유롭게 조정하면서도 **다중 프레임의 외관과 동적 일관성**을 유지하는 데 성공했다. 이는 영상 제작, 편집, 안정화 등 다양한 분야에서 실용적 활용 가능성을 열어준다. 특히, **Unreal Engine 5 기반의 대규모 데이터셋**은 연구자들이 카메라 제어 기반 영상 생성을 보다 효과적으로 연구할 수 있도록 기반을 제공한다. 그러나 **소스와 타겟 토큰 연결**은 성능 향상과 함께 **계산 비용 증가**를 동반하며, **사전 학습된 T2V 모델의 한계**(예: 손 생성의 부족)를 그대로 상속한다는 점이 한계로 지적된다.

실용적 활용

ReCamMaster는 **영상 안정화**(video stabilization), **초해상도**(super-resolution), **영상 확장**(outpainting) 등 다양한 영상 편집 작업에 적용 가능하다. 특히, **하드웨어 제한이 있는 아마추어 촬영**에서 **전문적인 카메라 움직임을 사후 조정**할 수 있어 영상 제작의 품질을 대폭 향상시킬 수 있다.