Depth Anything V2

Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao, Xiaogang Xu, Jiashi Feng, Hengshuang Zhao

arXiv:2406.09414 · 2026-07-27 공개 · arXiv · PDF

model-scaling stable-diffusion monocular-depth-estimation synthetic-images teacher-student-training generalization-capability evaluation-benchmark depth-anything-v2

Abstract

This work presents Depth Anything V2. Without pursuing fancy techniques, we aim to reveal crucial findings to pave the way towards building a powerful monocular depth estimation model. Notably, compared with V1, this version produces much finer and more robust depth predictions through three key practices: 1) replacing all labeled real images with synthetic images, 2) scaling up the capacity of our teacher model, and 3) teaching student models via the bridge of large-scale pseudo-labeled real images. Compared with the latest models built on Stable Diffusion, our models are significantly more efficient (more than 10x faster) and more accurate. We offer models of different scales (ranging from 25M to 1.3B params) to support extensive scenarios. Benefiting from their strong generalization capability, we fine-tune them with metric depth labels to obtain our metric depth models. In addition to our models, considering the limited diversity and frequent noise in current test sets, we construct a versatile evaluation benchmark with precise annotations and diverse scenes to facilitate future research.

한국어 요약

한 줄 요약

Depth Anything V2는 합성 이미지와 대규모 의사 라벨을 활용해 단일 카메라 깊이 추정 모델의 성능과 효율성을 획기적으로 향상시킨다.

핵심 기여도

핵심 아이디어

Depth Anything V2는 복잡한 기술 대신, 학습 데이터의 질과 양을 개선함으로써 모델 성능을 극대화하는 접근법을 제안한다. 기존 모델(V1)이 실제 라벨 데이터에 의존했다면, V2는 합성 이미지를 사용해 학습 데이터의 다양성과 정확도를 높였다. 특히, 합성 이미지의 단점을 보완하기 위해 대규모 의사 라벨(real images)을 활용한 두 단계 학습 방식(교사-학생 모델)을 도입했다. 이는 합성 이미지의 장점을 유지하면서 실제 이미지의 특성을 반영하는 데 기여한다. 또한, 기존 테스트셋의 노이즈와 다양성 부족을 인식하고, DA-2K라는 새로운 평가 벤치마크를 제안함으로써 연구의 신뢰성을 높였다.

기술적 접근법

주요 결과

의의 및 한계

Depth Anything V2는 단일 카메라 깊이 추정 분야에서 높은 정밀도와 효율성을 동시에 달성한 기반 모델로, 다양한 규모의 모델 제공과 미세 조정 가능성으로 실용성과 연구 가치를 높였다. 특히, 합성 이미지와 의사 라벨을 활용한 데이터 전략은 향후 MDE 연구에 중요한 인사이트를 제공한다. 그러나 합성 데이터의 한계와 의사 라벨 생성 과정에서의 오류 가능성은 여전히 개선이 필요한 부분이다. 또한, 투명 물체나 반사 표면에 대한 예측 정확도는 여전히 한계가 있을 수 있다.

실용적 활용

Depth Anything V2는 3D 재구성, 자율 주행, AI 생성 콘텐츠(AIGC) 등 다양한 산업 분야에서 활용 가능하다. 특히, 다양한 모델 규모(25M~1.3B 파라미터)를 제공함으로써, 실시간 처리가 필요한 애플리케이션부터 고성능 서버 기반의 복잡한 작업까지 폭넓게 사용할 수 있다. 또한, 미세 조정(fine-tuning)을 통해 특정 업무에 최적화된 모델로 전환할 수 있어 연구 및 산업 현장 모두에서 유용하다.