DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

Wenbo Hu, Xiangjun Gao, Xiaoyu Li, Sijie Zhao, Xiaodong Cun, Yong Zhang, Long Quan, Ying Shan

arXiv:2409.02095 · 2026-07-27 공개 · arXiv · PDF

diffusion-models zero-shot-learning image-to-video synthetic-datasets depth-based-visual-effects conditional-video-generation segment-wise-inference video-depth-estimation

Abstract

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. The generalization ability to open-world videos is achieved by training the video-to-depth model from a pretrained image-to-video diffusion model, through our meticulously designed three-stage training strategy. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that can process extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation.

한국어 요약

한 줄 요약

DepthCrafter는 110프레임 길이의 개방세계 동영상에 대해 시간적으로 일관된 고해상도 깊이 시퀀스를 생성하는 새로운 방법이다.

핵심 기여도

핵심 아이디어

기존의 영상 깊이 추정 방법은 카메라 포즈나 광학 흐름을 필요로 하며, 특히 개방세계 영상의 다양한 움직임과 길이를 처리하기 어려웠다. DepthCrafter는 이러한 문제를 해결하기 위해 이미지-영상 확산 모델(예: SVD)을 기반으로 영상-깊이 모델을 훈련시키는 새로운 접근법을 제안한다. 이 모델은 합성 데이터셋(정밀한 깊이 정보)과 실제 데이터셋(다양한 콘텐츠)을 결합하여, 세 단계 훈련 전략을 통해 110프레임 길이의 깊이 시퀀스를 생성할 수 있도록 설계되었다. 또한, 초장 동영상 처리를 위해 분할 추론과 이음 전략을 도입하여 시간적 일관성을 유지한다.

기술적 접근법

주요 결과

의의 및 한계

DepthCrafter는 개방세계 영상에서 카메라 포즈나 광학 흐름 없이도 시간적으로 일관된 깊이 시퀀스를 생성할 수 있어, 영상 기반 3D 재구성, 혼합현실, 자율주행 등 다양한 분야에 활용 가능하다. 특히, 110프레임 길이의 시퀀스 생성과 초장 영상 처리 능력은 기존 방법을 크게 앞선다. 그러나 학습에 사용된 데이터셋의 범위가 제한적이기 때문에, 특정 유형의 영상에서는 일반화 능력이 떨어질 수 있다. 또한, 실제 세계에서의 대규모 배포 가능성은 아직 검증되지 않았다.

실용적 활용

DepthCrafter는 혼합현실(MR) 및 가상현실(VR)에서의 3D 환경 구축, 자율주행 차량의 주변 환경 인식, AI 기반 콘텐츠 생성 등에 활용될 수 있다. 특히, 조건부 영상 생성과 깊이 기반 시각 효과(VFX) 분야에서 높은 잠재력을 보인다.