ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion

Cho In, Jeonghwan Cho, Mijin Yoo, Gim Hee Lee, Seon Joo Kim

arXiv:2607.20417 · 2026-07-25 공개 · arXiv · PDF

novel-view-synthesis feed-forward dl3dv compact-representation realrestate10k adaptive-token-expansion sparse-to-adaptive gaussian-offsets

Abstract

3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather than scene complexity, resulting in dense and often redundant Gaussian sets. We present ATSplat, a feed-forward 3DGS framework that restores the adaptive allocation capability of 3DGS optimization through Adaptive 3D Tokens. ATSplat first lifts coarse patch-level depth and camera cues into sparse 3D anchor tokens, forming a compact scaffold of the scene. Each token is then regressed into local Gaussians with learnable 3D offsets, decoupling primitive placement from input image grids. An Adaptive Token Expansion module predicts a token-level uncertainty score, supervised by rendering error maps, and selectively expands high-uncertainty tokens through learnable expansion layers. This sparse-to-adaptive formulation enables ATSplat to concentrate primitives in challenging regions while maintaining a compact representation. Experiments on two representative datasets, RealEstate10K and DL3DV, show that ATSplat achieves state-of-the-art rendering quality while reducing the number of Gaussians by more than 5.7times compared with dense feed-forward 3DGS methods. From 12 input images at 512 times 960 resolution, ATSplat completes reconstruction in less than a second using a single commercial GPU, and renders high-quality novel views at 1136 FPS (512 times 960) with only 311K Gaussians.

한국어 요약

한 줄 요약

ATSplat은 적응형 토큰 확장을 통해 3D Gaussian Splatting의 장점을 복원한 실시간 렌더링 성능을 갖는 컴팩트한 피드포워드 프레임워크이다.

핵심 기여도

핵심 아이디어

기존 피드포워드 3DGS는 입력 이미지 그리드에 Gaussian을 할당하여, 렌더링 효율성과 표현력 사이의 균형을 잃는다. ATSplat은 이 문제를 해결하기 위해 **Adaptive 3D Anchor Tokens**를 도입하여, 입력 이미지 구조가 아닌 **장면 복잡도에 따라 Gaussian 배치를 결정**한다.

ATSplat은 먼저 **패치 단위의 깊이와 카메라 정보를 3D anchor tokens으로 변환**하여 장면의 컴팩트한 구조를 형성한다. 이후, 각 토큰은 **3D 오프셋을 학습하여 로컬 Gaussian으로 디코딩**되며, 입력 그리드와 독립적으로 배치된다.

핵심적인 **Adaptive Token Expansion 모듈**은 렌더링 오류 맵을 기반으로 토큰별 불확실도 점수를 학습하고, 어려운 영역의 토큰을 선택적으로 확장하여 표현력 집중. 이는 3DGS 최적화의 핵심 원리인 **적응형 용량 할당**을 피드포워드 환경에서 복원한다.

기술적 접근법

주요 결과

의의 및 한계

ATSplat은 피드포워드 3DGS에서 **3DGS 최적화의 핵심 원리인 적응형 용량 할당**을 복원함으로써, 표현력과 효율성을 동시에 달성한 사례로, 학술적으로 중요한 기여를 한다. 또한, **실시간 렌더링과 낮은 메모리 사용량**은 산업적 활용 가능성을 높인다.

하지만, ATSplat은 **토큰 확장은 하지만 불필요한 토큰은 제거하지 않음**으로써, 효율성 개선의 한계가 존재. 또한, **더 큰 규모의 장면, 더 많은 입력 뷰, 더 높은 해상도**에 대한 확장성도 아직 검증되지 않았다.

실용적 활용

ATSplat은 **실시간 렌더링이 필요한 AR/VR, 콘텐츠 생성, 자율 주행** 등에서 유용하게 활용될 수 있다. 특히, **입력 이미지 수가 제한된 환경**에서 높은 품질의 3D 재구성을 제공하며, **GPU 자원이 제한된 모바일 장치**에서도 적용 가능성이 있다.