Video Depth Anything: Consistent Depth Estimation for Super-Long Videos

Sili Chen, Hengkai Guo, Shengnan Zhu, Feihu Zhang, Zilong Huang, Jiashi Feng, Bingyi Kang

arXiv:2501.12375 · 2026-07-27 공개 · arXiv · PDF

depth-estimation zero-shot-learning long-video temporal-consistency real-time-processing video-benchmarks depth-anything video-depth

Abstract

Depth Anything has achieved remarkable success in monocular depth estimation with strong generalization ability. However, it suffers from temporal inconsistency in videos, hindering its practical applications. Various methods have been proposed to alleviate this issue by leveraging video generation models or introducing priors from optical flow and camera poses. Nonetheless, these methods are only applicable to short videos (< 10 seconds) and require a trade-off between quality and computational efficiency. We propose Video Depth Anything for high-quality, consistent depth estimation in super-long videos (over several minutes) without sacrificing efficiency. We base our model on Depth Anything V2 and replace its head with an efficient spatial-temporal head. We design a straightforward yet effective temporal consistency loss by constraining the temporal depth gradient, eliminating the need for additional geometric priors. The model is trained on a joint dataset of video depth and unlabeled images, similar to Depth Anything V2. Moreover, a novel key-frame-based strategy is developed for long video inference. Experiments show that our model can be applied to arbitrarily long videos without compromising quality, consistency, or generalization ability. Comprehensive evaluations on multiple video benchmarks demonstrate that our approach sets a new state-of-the-art in zero-shot video depth estimation. We offer models of different scales to support a range of scenarios, with our smallest model capable of real-time performance at 30 FPS.

한국어 요약

한 줄 요약

Video Depth Anything는 초장시간 동영상에서도 깊이 추정의 일관성을 유지하면서 계산 효율성을 보장하는 모델이다.

핵심 기여도

핵심 아이디어

기존의 Depth Anything V2는 단일 이미지에서 뛰어난 성능을 보이지만, 동영상에서는 시간에 따른 깊이 일관성이 부족하다. 이를 해결하기 위해, Video Depth Anything는 기존 모델의 헤드를 **Spatial-Temporal Head (STH)**로 교체하여 시간 정보를 모델링한다. STH는 4개의 **Temporal Attention Layer**를 포함하며, 각 공간 위치에서 시간 축을 따라 어텐션을 적용하여 시간 정보를 효과적으로 통합한다. 이는 기존 이미지 인코더가 시간 정보를 무시하는 한계를 보완한다.

또한, 기하학적 가정이나 영상 생성 모델을 사용하지 않고, **Temporal Gradient Matching Loss**를 통해 시간 축 상의 깊이 변화를 제약함으로써 일관성을 확보한다. 이 손실 함수는 기존의 Scale-Shift-Invariant Loss와 Spatial Gradient Matching Loss와 함께 최적화되어, 시간 일관성을 향상시키는 동시에 공간 정확도를 유지한다.

기술적 접근법

주요 결과

의의 및 한계

Video Depth Anything는 기존 모델의 일반화 능력과 계산 효율성을 유지하면서, 초장시간 동영상에서도 깊이 일관성을 보장하는 **새로운 기준**을 제시한다. 특히, 기하학적 가정이나 영상 생성 모델 없이도 시간 일관성을 달성한 점은 학술적으로 중요한 기여이다. 또한, 다양한 크기의 모델을 제공하여 실시간 처리부터 고정밀 추정까지 다양한 시나리오에 적용 가능하다는 실용적 가치를 가진다.

한편, **500프레임 이하의 데이터셋만 평가**했기 때문에, 더 긴 영상에서의 일관성 유지 능력은 추가 실험 필요. 또한, **중첩 프레임 기반 추론 전략**은 메모리 사용량 증가를 유발할 수 있다.

실용적 활용