DepthSplat: Connecting Gaussian Splatting and Depth

Haofei Xu, Songyou Peng, Fangjinhua Wang, Hermann Blum, Dániel Baráth, Andreas Geiger, M. Pollefeys

arXiv:2410.13862 · 2026-07-27 공개 · arXiv · PDF

novel-view-synthesis depth-estimation multi-view gaussian-splatting scan-net dl3dv real-estate-10k feed-forward-reconstruction

Abstract

Gaussian splatting and single-view depth estimation are typically studied in isolation. In this paper, we present Depth-Splat to connect Gaussian splatting and depth estimation and study their interactions. More specifically, we first contribute a robust multi-view depth model by leveraging pretrained monocular depth features, leading to high-quality feed-forward 3D Gaussian splatting reconstructions. We also show that Gaussian splatting can serve as an unsupervised pre-training objective for learning powerful depth models from large-scale multi-view posed datasets. We validate the synergy between Gaussian splatting and depth estimation through extensive ablation and cross-task transfer experiments. Our DepthSplat achieves state-of-the-art performance on ScanNet, RealEstate10K and DL3DV datasets in terms of both depth estimation and novel view synthesis, demonstrating the mutual benefits of connecting both tasks. In addition, DepthSplat enables feed-forward reconstruction from 12 input views (512 × 960 resolutions) in 0.6 seconds.

한국어 요약

한 줄 요약

DepthSplat은 Gaussian splatting과 depth estimation을 결합하여 ScanNet, RealEstate10K, DL3DV에서 최고 성능을 달성한 통합 모델이다.

핵심 기여도

핵심 아이디어

기존 연구는 Gaussian splatting과 single-view depth estimation을 독립적으로 다루었으나, 본 연구는 두 기법의 상호작용을 통해 **성능 향상과 상호 보완**을 추구한다. 핵심 아이디어는 **monocular depth feature**와 **multi-view feature matching**을 결합하여 robust한 depth 추정을 실현하는 동시에, 이를 Gaussian splatting에 활용해 높은 품질의 3D 재구성을 가능하게 하는 것이다.

DepthSplat은 **Depth Anything V2** 사전 학습 모델을 기반으로 monocular depth feature를 multi-view branch에 결합하여, 기존 multi-view depth 방법이 어려운 **texture-less 영역, 반사 표면, 가림 현상**에서도 안정적인 depth 추정을 달성한다. 또한, Gaussian splatting 모듈은 **fully differentiable**하게 설계되어, **photometric supervision만으로도 depth 모델을 unsupervised pre-training**할 수 있다. 이는 ground truth geometry 없이도 대규모 multi-view 데이터셋에서 depth 모델을 학습할 수 있는 새로운 접근법이다.

기술적 접근법

주요 결과

의의 및 한계

DepthSplat은 Gaussian splatting과 depth estimation의 **상호보완적 결합**을 통해, 기존 한계를 극복하고 높은 정확도와 빠른 속도를 동시에 달성한 점에서 학술적·실용적 가치가 크다. 특히, **사전 학습된 monocular depth feature를 활용한 multi-view depth 모델**은 기존 복잡한 아키텍처를 대체하며, **unsupervised pre-training**을 통해 대규모 데이터셋 활용 가능성을 열었다.

하지만, **입력 뷰가 매우 적은 경우** camera pose 추정이 어려워 성능 저하가 발생할 수 있다. 또한, **입력 뷰 수가 많아질수록 Gaussian 수가 급증**하여 계산 복잡도가 증가하는 문제가 있으며, 이는 향후 연구 주제로 제시된다.

실용적 활용

DepthSplat은 **실내/실외 3D 재구성**, **자율주행**, **증강현실(AR)** 등에서 빠른 feed-forward 성능과 높은 정확도를 요구하는 분야에 적용 가능하다. 특히, **사전 학습된 모델 기반의 unsupervised 학습**은 라벨 데이터가 부족한 상황에서 유용하며, **대규모 multi-view 데이터셋을 활용한 depth 모델 학습**에도 활용할 수 있다.