Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction

Chin-Yang Lin, Yang-Che Sun, Cheng Sun, Fu-En Yang, Min-Hung Chen, Yen-Yu Lin, Wei-Chen Chiu, Yu-Lun Liu

arXiv:2609.04201 · 2026-09-04 공개 · arXiv · PDF

scan-net loop-closure kitti online-reconstruction relative-pose keyframe-querying pose-graph-optimization asymmetric-attention

Abstract

Online 3D reconstruction models perform poorly on long videos. This happens because regressing poses relative to a fixed first-frame anchor forces extrapolation far beyond the training distribution. Small drifts accumulate and amplify into significant geometric collapse. However, we observe that per-frame depth remains stable throughout this failure. The backbone's local geometry remains intact; only the global pose head breaks down. Motivated by this decoupling, we introduce Scal3R. This approach reformulates online reconstruction as multi-reference relative pose querying. We use lightweight learnable tokens, which make up about ~1% of the parameters, and inject them into a completely frozen backbone via asymmetric attention. This setup queries poses relative to multiple past keyframes. An online pose-graph optimization system with loop closure suppresses long-range drift. Scal3R reaches convergence in 8 hours on a single GPU. It reduces the average ATE by over 60% on KITTI compared to the online baseline. It also achieves state-of-the-art performance across Virtual KITTI, Sintel, TUM-Dynamic, ScanNet, and 7-Scenes. Project page: https://linjohnss.github.io/scal3r/

한국어 요약

한 줄 요약

Scal3R은 KITTI 데이터셋에서 평균 ATE를 60% 이상 개선하며, 1% 미만의 파라미터로 온라인 3D 재구성을 안정적으로 수행하는 다중 참조 상대 자세 쿼리 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 온라인 3D 재구성 모델은 첫 번째 프레임을 기준으로 전역 자세를 회귀하는 방식을 사용하지만, 이는 긴 영상에서 예측 분포를 벗어나게 되어 누적 드리프트와 기하학적 붕괴를 유발한다. Scal3R은 이 문제를 해결하기 위해 전역 자세 회귀를 다중 참조 상대 자세 쿼리로 재구성한다. 각 프레임의 로컬 기하학은 안정적이기 때문에, 이 정보를 활용해 과거 키프레임에 대한 상대 자세를 추정하는 방식을 채택한다.

이를 위해 Scal3R은 동결된 백본에 약 1%의 파라미터를 차지하는 가벼운 학습 가능한 토큰을 주입하고, 비대칭 어텐션을 통해 이미지 특징을 쿼리한다. 이는 백본의 기하학적 표현력을 유지하면서도, 자세 추정 과정에서 로컬 영역에 집중하도록 유도한다. 또한, 온라인 PGO와 루프 클로저를 결합하여 전역 일관성을 유지한다.

기술적 접근법

주요 결과

의의 및 한계

Scal3R은 기존 온라인 3D 재구성 모델의 전역 자세 회귀 방식의 한계를 극복하며, 긴 영상에서도 안정적인 재구성을 가능하게 한다. 특히, 동결된 백본과 가벼운 토큰을 활용한 비대칭 어텐션은 모델의 기하학적 표현력을 유지하면서도 학습 비용을 줄이는 데 기여한다.

그러나, Scal3R은 훈련 데이터가 제한된 환경에서의 일반화 능력을 명시하지 않았으며, 루프 클로저의 성능이 외부 영상 스트림의 복잡성에 따라 변동할 수 있다는 한계가 있다.

실용적 활용

Scal3R은 자율주행, 드론 네비게이션, AR/VR 등에서 실시간 3D 재구성을 필요로 하는 산업에 적용 가능하다. 특히, 긴 영상에서도 드리프트를 억제하는 특성으로, 대규모 환경의 지속적인 모니터링 및 매핑에 유용하다.