MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Ruicheng Wang, Sicheng Xu, Cassie Dai, Jianfeng Xiang, Yu Deng, Xin Tong, Jiaolong Yang

arXiv:2410.19115 · 2026-07-27 공개 · arXiv · PDF

monocular-reconstruction monocular-geometry affine-invariant point-cloud-alignment depth-map-estimation camera-field-of-view geometry-supervision open-domain-images

Abstract

We present MoGe, a powerful model for recovering 3D geometry from monocular open-domain images. Given a single image, our model directly predicts a 3D point map of the captured scene with an affine-invariant representation, which is agnostic to true global scale and shift. This new representation precludes ambiguous supervision in training and facilitates effective geometry learning. Furthermore, we propose a set of novel global and local geometry supervision techniques that empower the model to learn high-quality geometry. These include a robust, optimal, and efficient point cloud alignment solver for accurate global shape learning, and a multi-scale local geometry loss promoting precise local geometry supervision. We train our model on a large, mixed dataset and demonstrate its strong generalizability and high accuracy. In our comprehensive evaluation on diverse unseen datasets, our model significantly outperforms state-of-the-art methods across all tasks, including monocular estimation of 3D point map, depth map, and camera field of view.

한국어 요약

한 줄 요약

MoGe는 단일 이미지에서 정확한 3D 기하학을 추정하는 모델로, affine-invariant 표현과 ROE 정렬 알고리즘, multi-scale 로스를 통해 기존 방법 대비 35% 이상의 오류 감소를 달성한다.

핵심 기여도

핵심 아이디어

기존 단일 이미지 기하학 추정 방법은 정확한 카메라 파라미터 추정이 어려워 기하학 왜곡이 발생한다. MoGe는 카메라 파라미터 추정 없이 3D point map을 직접 예측하는 direct approach를 채택한다. 이 모델은 affine-invariant 표현을 사용하여 전역 스케일과 이동에 무관하게 3D 점을 예측함으로써 훈련 시 모호한 지도를 제거한다. 이는 특히 focal-distance ambiguity 문제를 해결하는 데 기여한다. 또한, 기존의 global alignment 방법이 outlier에 민감하거나 근사치 기반으로 정확도가 낮았던 문제를 해결하기 위해 ROE 정렬 알고리즘을 도입한다. 지역적 기하학 정밀도를 높이기 위해 multi-scale local geometry loss를 사용하여 지역별 3D 점 구름의 차이를 독립적인 affine 정렬 하에 페널티를 부여한다.

기술적 접근법

주요 결과

의의 및 한계

MoGe는 단일 이미지에서 3D 기하학을 추정하는 기존 방법들의 한계를 극복하고, 훈련 지도 설계의 중요성을 강조한다. affine-invariant 표현과 ROE 정렬 알고리즘은 모호한 지도 문제를 해결하고, multi-scale 로스는 지역적 정밀도를 향상시켜 실용적이고 정확한 기하학 추정을 가능하게 한다. 그러나 본 연구는 open-domain 이미지에 초점을 맞추고 있어, 특정 도메인에 최적화된 모델과의 비교는 제한적이다. 또한, 훈련 데이터셋의 다양성과 질이 모델 성능에 큰 영향을 미친다는 점에서, 더 다양한 데이터셋에 대한 실험 필요성이 제기된다.

실용적 활용

MoGe는 3D-aware 이미지 편집, depth-to-image synthesis, novel view synthesis, 3D scene understanding 등 다양한 3D 비주얼 분야에 활용 가능하다. 또한, 비디오나 다중 뷰 기반 3D 재구성에서 초기 기하학 prior를 제공하는 데 유용할 수 있다.