LLaVA-3D: A Simple Yet Effective Pathway to Empowering LMMs with 3D Capabilities

Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, Xihui Liu

arXiv:2409.18125 · 2026-07-27 공개 · arXiv · PDF

vision-language instruction-tuning multimodal-models large-scale-datasets scene-understanding llava-3d position-embeddings

Abstract

Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the development of LMMs with 3D scene understanding capabilities has been hindered by the lack of large-scale 3D vision-language datasets and powerful 3D encoders. In this paper, we introduce a simple yet effective framework called LLaVA-3D. Leveraging the strong 2D visual understanding priors from LLaVA, our LLaVA-3D efficiently adapts LLaVA for 3D scene understanding without compromising 2D understanding capabilities. To achieve this, we utilize the 3D position embeddings to enhance the 2D CLIP Patches with 3D spatial context information and construct 3D patches. By integrating the 3D position embeddings into 2D LMMs and employing joint 2D and 3D vision-language instruction tuning, we establish a unified architecture for both 2D visual understanding and 3D scene understanding. In contrast to previous 3D LMMs, LLaVA-3D supports decoding accurate 3D spatial perception outputs, e.g., 3D bounding boxes, directly from these 3D patches, without relying on the time-consuming off-the-shelf 3D segmentors. Experimental results show that LLaVA-3D converges $3.5 \times$ faster than existing 3D LMMs when trained on 3D vision-language datasets. Moreover, LLaVA-3D not only achieves state-of-the-art performance across various 3D tasks but also maintains comparable 2D visual understanding and vision-language conversation capabilities with LLaVA.

한국어 요약

한 줄 요약

LLaVA-3D는 3D 포지션 임베딩과 3D 패치를 활용해 2D LMM인 LLaVA를 3D 시나리오로 확장한 모델로, 3D 객체 인식 및 추론 성능을 향상시키며 3.5배 빠른 수렴 속도를 보인다.

핵심 기여도

핵심 아이디어

LLaVA-3D는 기존 2D LMM인 LLaVA의 강력한 2D 시각 이해 능력을 기반으로, 3D 공간 인지 능력을 추가하는 새로운 접근법을 제시한다. 기존 3D LMM은 3D 포인트 클라우드와 별도의 3D 세그멘터에 의존하지만, LLaVA-3D는 3D 포지션 임베딩을 2D CLIP 패치에 결합하여 3D 패치를 생성함으로써 3D 정보를 자연스럽게 통합한다. 이는 3D 공간 정보를 2D 이미지 기반으로 처리할 수 있도록 하며, 별도의 3D 인코더나 세그멘터 없이도 3D 객체 인식 및 추론이 가능하다는 점에서 혁신적이다. 또한, 2D와 3D 데이터셋을 함께 사용한 조인트 인스트럭션 튜닝을 통해 2D 시각 이해 능력을 유지하면서 3D 능력을 추가하는 통합 아키텍처를 구축한다.

기술적 접근법

주요 결과

의의 및 한계

LLaVA-3D는 3D 시각-언어 데이터셋 부족과 3D 인코더 부재라는 주요 장벽을 극복한 첫 번째 통합 접근법으로, 2D LMM을 기반으로 3D 능력을 추가하는 간단하면서도 효과적인 확장 방식을 제시한다. 특히, 오프라인 3D 세그멘터 없이도 3D 바운딩 박스를 디코딩할 수 있다는 점에서 실용적 가치가 높다. 그러나 본 연구는 기존 3D 데이터셋(예: ScanNet, Matterport3D)을 기반으로 훈련되었으며, 대규모 3D 데이터셋이 부족한 상황에서는 성능 한계가 있을 수 있다. 또한, 3D 포지션 임베딩의 정확도가 3D 인식 성능에 직접적인 영향을 미치므로, 이에 대한 최적화가 필요할 수 있다.

실용적 활용

LLaVA-3D는 로봇 제어, 3D 객체 인식, AR/VR 환경에서의 시각-언어 상호작용 등 다양한 3D 시나리오에 적용 가능하다. 특히, 2D LMM을 기반으로 3D 능력을 추가할 수 있는 확장성 덕분에, 다양한 2D LMM(예: LLaVA, LLaVA-Video)을 기반으로 3D 능력을 추가하는 데 활용할 수 있다.