Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

Jianing Yang, Alexander Sax, Kevin J. Liang, Mikael Henaff, Hao Tang, Ang Cao, Joyce Chai, Franziska Meier, Matt Feiszli

arXiv:2501.13928 · 2026-07-27 공개 · arXiv · PDF

transformer multi-view camera-pose inference-speed dust3r fast3r scalable-reconstruction

Abstract

Multi-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R’s Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy.

한국어 요약

한 줄 요약

Fast3R은 1,000개 이상의 이미지를 단일 순방향 패스로 처리하는 Transformer 기반 3D 재구성 모델로, DUSt3R 대비 14배 오류 감소와 250 FPS 이상의 추론 속도를 달성한다.

핵심 기여도

핵심 아이디어

Fast3R은 기존 3D 재구성 방법에서 일반적인 쌍별 처리와 글로벌 정렬 과정을 제거하고, Transformer 기반의 병렬 처리를 도입함으로써 효율적이고 확장 가능한 3D 재구성을 가능하게 한다. 기존 DUSt3R은 쌍별 이미지 처리를 기반으로 하여 N개의 이미지에 대해 O(N²)의 연산이 필요하며, 글로벌 정렬이 필수적이었다. Fast3R은 이를 Transformer의 all-to-all self-attention을 통해 N개 이미지를 동시에 처리하고, 순차적 처리 없이 포인트맵을 생성한다. 이는 오류 누적을 줄이고, 추론 속도를 크게 향상시키는 핵심 아이디어이다.

Fast3R은 이미지 인코딩, 퓨전 트랜스포머, 포인트맵 디코딩의 세 단계로 구성되며, 각 이미지 패치는 이미지 인덱스에 따른 위치 임베딩을 추가하여 전역 좌표계를 유지한다. 특히, Position Interpolation 기법을 통해 훈련 시 N=20개 이미지만 사용해도 추론 시 N=1,000개 이미지를 처리할 수 있도록 확장성을 확보하였다.

기술적 접근법

주요 결과

의의 및 한계

Fast3R은 기존 SfM 및 MVS 파이프라인의 복잡성과 느린 처리 속도를 극복하고, 단일 순방향 패스로 N개 이미지를 처리함으로써 3D 재구성의 효율성과 확장성을 획기적으로 향상시킨다. 특히, Transformer 기반의 all-to-all self-attention을 통해 오류 누적을 줄이고, Position Interpolation 기법을 통해 훈련-추론 간의 이미지 수 차이를 해결하였다. 이는 실시간 3D 재구성, 대규모 장면 스캔 등 다양한 응용 분야에서 유용하다.

그러나, Fast3R의 성능은 훈련 데이터의 정확성과 양에 크게 의존하며, 이는 확장성의 한계로 작용할 수 있다. 또한, 현재는 주로 정적 장면에 초점을 맞추고 있어, 동적 장면(4D)에 대한 처리는 추가 연구가 필요하다.

실용적 활용

Fast3R은 자율주행, 증강현실(AR), 로봇 비전 등에서 대규모 이미지 집합을 빠르고 정확하게 3D 재구성해야 하는 상황에 적합하다. 특히