Global Structure-from-Motion Revisited

Linfei Pan, Dániel Baráth, M. Pollefeys, Johannes L. Schönberger

arXiv:2407.20219 · 2026-07-27 공개 · arXiv · PDF

open-source computer-vision efficiency structure-from-motion camera-motion colmap global-sfm incremental-sfm

Abstract

Recovering 3D structure and camera motion from images has been a long-standing focus of computer vision research and is known as Structure-from-Motion (SfM). Solutions to this problem are categorized into incremental and global approaches. Until now, the most popular systems follow the incremental paradigm due to its superior accuracy and robustness, while global approaches are drastically more scalable and efficient. With this work, we revisit the problem of global SfM and propose GLOMAP as a new general-purpose system that outperforms the state of the art in global SfM. In terms of accuracy and robustness, we achieve results on-par or superior to COLMAP, the most widely used incremental SfM, while being orders of magnitude faster. We share our system as an open-source implementation at {https://github.com/colmap/glomap}.

한국어 요약

한 줄 요약

GLOMAP은 전역 Structure-from-Motion 문제를 해결하는 새로운 시스템으로, COLMAP과 유사한 정확도를 유지하면서 수십 배 빠른 처리 속도를 제공한다.

핵심 기여도

핵심 아이디어

기존 전역 SfM 시스템은 **rotation averaging**과 **translation averaging**을 별도 단계로 수행하며, 특히 translation averaging 단계에서 정확도와 안정성 문제가 발생했다. 이는 **스케일 모호성**, **정확한 카메라 내부 파라미터 부재**, **거의 일직선상의 카메라 운동**(near-collinear motion) 등으로 인해 발생한다. GLOMAP은 이러한 문제를 해결하기 위해 **카메라 위치와 3D 점 위치를 동시에 추정하는 Global Positioning** 단계를 도입한다. 이는 기존의 translation averaging과 별도의 triangulation 단계를 통합하여, **전체적인 최적화 과정을 단일 단계로 줄임**.

기술적 접근법

주요 결과

의의 및 한계

GLOMAP은 전역 SfM의 효율성과 정확도 간의 갈등을 해결한 **새로운 접근법**을 제시한다. 기존 전역 SfM 시스템은 정확도가 낮고, 증분식 시스템은 느리다는 한계를 극복함. 특히, **Global Positioning** 모듈은 기존 translation averaging 단계의 불안정성을 제거함. 그러나, **정확한 2-view 기하 추정**(feature matching)이 여전히 필수적이라는 점은 한계로 작용할 수 있다. 또한, **COLMAP과 동일한 feature matches를 사용**한 실험만 수행되어, 전체적인 성능 비교의 완전성은 제한적일 수 있다.

실용적 활용

GLOMAP은 **실시간 3D 재구성**, **클라우드 기반 매핑 및 로컬라이제이션**, **드론 또는 자율주행차의 시퀀셜 이미지 처리** 등에 적용 가능하다. 특히, **인터넷 사진**(unknown intrinsics)을 포함한 데이터셋 처리에서도 유용하며, **대규모 이미지 데이터셋**에서의 빠른 처리 속도가 실용적 가치를 높인다.