GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting

Kai Zhang, Sai Bi, Hao Tan, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, Zexiang Xu

arXiv:2404.19702 · 2026-07-27 공개 · arXiv · PDF

transformer vision-transformer scene-reconstruction differentiable-rendering realestate10k reconstruction-model objaverse large-reconstruction-model

Abstract

We propose GS-LRM, a scalable large reconstruction model that can predict high-quality 3D Gaussian primitives from 2-4 posed sparse images in 0.23 seconds on single A100 GPU. Our model features a very simple transformer-based architecture; we patchify input posed images, pass the concatenated multi-view image tokens through a sequence of transformer blocks, and decode final per-pixel Gaussian parameters directly from these tokens for differentiable rendering. In contrast to previous LRMs that can only reconstruct objects, by predicting per-pixel Gaussians, GS-LRM naturally handles scenes with large variations in scale and complexity. We show that our model can work on both object and scene captures by training it on Objaverse and RealEstate10K respectively. In both scenarios, the models outperform state-of-the-art baselines by a wide margin. We also demonstrate applications of our model in downstream 3D generation tasks. Our project webpage is available at: https://sai-bi.github.io/project/gs-lrm/ .

한국어 요약

한 줄 요약

GS-LRM은 2~4장의 투영 이미지에서 0.23초 내 3D 가우시안을 생성하는 빠르고 확장 가능한 트랜스포머 기반 모델이다.

핵심 기여도

핵심 아이디어

GS-LRM은 기존의 트리플레인 NeRF 기반 LRM과 달리, 3D 가우시안을 직접 예측함으로써 복잡한 볼륨 렌더링을 생략하고 빠른 렌더링을 가능하게 한다. 기존 LRM은 입력 이미지와 출력 3D 표현 간의 불일치로 인해 디자인 복잡도가 높았지만, GS-LRM은 입력 이미지와 출력 가우시안을 같은 픽셀 공간에 정렬하여 학습 효율성을 높인다. 이는 각 픽셀에 대응하는 3D 가우시안을 예측함으로써, 고주파 세부 정보를 보존할 수 있게 한다. 또한, 트랜스포머 블록 내의 self-attention 메커니즘을 통해 다중 뷰 간의 픽셀 정보를 효과적으로 집약하여, 정확한 3D 복원을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

GS-LRM은 3D 복원 분야에서 빠른 추론 속도와 높은 정확도를 동시에 달성한 첫 트랜스포머 기반 모델로, 객체 및 장면 복원 모두에서 SOTA 성능을 보인다. 특히, 기존 NeRF 기반 모델의 볼륨 렌더링 복잡도와 고해상도 처리 한계를 극복한 점이 주목할 만하다. 그러나, 2~4장의 투영 이미지만을 기반으로 복원하기 때문에, 입력 이미지의 품질과 각도에 따라 결과가 크게 달라질 수 있는 한계가 있다. 또한, 현재는 Objaverse와 RealEstate10K 데이터셋에만 제한적으로 훈련되었으며, 더 다양한 도메인으로 확장 가능성은 추가 연구가 필요하다.

실용적 활용

GS-LRM은 실시간 3D 콘텐츠 생성, AR/VR, 자율주행 시스템의 환경 인식 등에 활용 가능하다. 특히, 고해상도 3D 복원이 필요한 산업 현장에서 빠른 추론 속도와 높은 정확도를 통해 실용적 가치가 높다.