Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction

Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian

arXiv:2608.27529 · 2026-08-31 공개 · arXiv · PDF

long-horizon pose-estimation visual-context streaming-reconstruction local-context sequence-benchmark ate-rpe abot-recon

Abstract

Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation. Early streaming models achieve causal, bounded-cost inference using finite context buffers or compact recurrent states, yet their estimates often deteriorate as sequences grow. Recent methods improve long-horizon stability by coupling short-range context with persistent or multi-level long-range memory. We pursue a different route: we keep the learned temporal state strictly local and formulate predictions whose targets remain independent of sequence length. We present ABot-Recon, a simple streaming model that caches KV features from only the preceding 11 frames. It predicts a point map in the current camera coordinate system together with an adjacent-frame relative pose. These predictions remain equivariant under changes of reference frame, and global poses and geometry are recovered through sequential composition. To reduce accumulated drift, a lightweight temporal refiner improves relative rotations using recent visual and motion context, while a composition-aware pose loss supervises multi-step pose composition. Extensive evaluations on challenging long-sequence benchmarks demonstrate the superior long-horizon performance of our local-context approach. On Oxford Spires, ABot-Recon achieves an ATE of 4.35 m and an RPE-R of $0.12^\circ$, reducing both errors by approximately 40\% relative to the best prior results.

한국어 요약

한 줄 요약

ABot-Recon은 11프레임의 로컬 컨텍스트만 사용해 장거리 스트리밍 3D 재구성을 안정적으로 수행하는 모델로, Oxford Spires에서 ATE 4.35m, RPE-R 0.12°를 달성한다.

핵심 기여도

핵심 아이디어

ABot-Recon은 기존 스트리밍 3D 재구성 모델들이 장거리 정보를 저장하거나 재귀적 상태를 유지하는 방식과 달리, **로컬 컨텍스트만을 사용**하는 새로운 접근법을 제안한다. 이는 장거리 스트림에서도 예측 범위가 고정되어 있기 때문에, **예측 문제의 복잡도가 증가하지 않도록 유지**할 수 있다.

구체적으로, ABot-Recon은 **현재 카메라 좌표계에서 포인트 맵 $P_i$와 인접 프레임 간 상대 자세 $T_{i-1\leftarrow i}$를 예측**하며, 이 예측은 **등변형(equivariant)**을 유지한다. 즉, 참조 프레임이 바뀌어도 예측이 일관되게 변환되도록 설계되어, 전역 자세는 순차적으로 조합하여 복원된다.

이러한 구조는 **12프레임 내에서만 작동**하며, **11개의 이전 프레임의 KV 피처만 캐시**하고, **지속적인 장거리 상태를 유지하지 않는다**. 이는 메모리와 계산 비용을 고정시키며, 장거리 스트림에서도 안정적인 재구성을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

ABot-Recon은 장거리 스트리밍 3D 재구성에서 **지속적인 장거리 메모리가 필요하지 않다는 점에서 혁신적**이며, **로컬 예측을 기반으로 전역 재구성을 안정적으로 수행**할 수 있음을 보여준다. 특히, **등변형 예측과 조합 인식 손실**을 통해 장거리 누적 오류를 효과적으로 제어하며, **SLAM 모듈과의 호환성**도 높아 실용적이다.

그러나, **공간이 제한된 실내 환경**(예: 7Scenes)에서는 지속적인 장거리 정보가 유용할 수 있으며, 이에 비해 ABot-Recon의 성능 개선은 상대적으로 적다. 또한, **루프 클로저가 없을 경우 장거리 드리프트가 발생할 수 있으나**, 이는 별도의 모듈로 해결 가능하며 학습 과정에는 영향을 주지 않는다.

실용적 활용

ABot-Recon은 **로봇, 자율 주행, 임베디드 시스템** 등에서 실시간 3D 재구성과 카메라 추적을 필요로 하는 상황에 적합하다. 특히, **메모리와 계산 자원이 제한된 장치**에서 유용하며, **SLAM 모듈과의 결합**을 통해 추가적인 정확도 향상을 기대할 수 있다. 또한, **장거리 스트리밍 환경**(예: 도시 규모의 자율 주행)에서 안정적인 성능을 보이므로, **대규모 3D 맵 생성**에도 활용 가능하다.