3D-VLA: A 3D Vision-Language-Action Generative World Model

Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, Chuang Gan

arXiv:2403.09631 · 2026-07-27 공개 · arXiv · PDF

diffusion-models vision-language-action multimodal-generation generative-world-model embodied-foundation-model instruction-dataset robotics-datasets reasoning-planning

Abstract

Recent vision-language-action (VLA) models rely on 2D inputs, lacking integration with the broader realm of the 3D physical world. Furthermore, they perform action prediction by learning a direct mapping from perception to action, neglecting the vast dynamics of the world and the relations between actions and dynamics. In contrast, human beings are endowed with world models that depict imagination about future scenarios to plan actions accordingly. To this end, we propose 3D-VLA by introducing a new family of embodied foundation models that seamlessly link 3D perception, reasoning, and action through a generative world model. Specifically, 3D-VLA is built on top of a 3D-based large language model (LLM), and a set of interaction tokens is introduced to engage with the embodied environment. Furthermore, to inject generation abilities into the model, we train a series of embodied diffusion models and align them into the LLM for predicting the goal images and point clouds. To train our 3D-VLA, we curate a large-scale 3D embodied instruction dataset by extracting vast 3D-related information from existing robotics datasets. Our experiments on held-in datasets demonstrate that 3D-VLA significantly improves the reasoning, multimodal generation, and planning capabilities in embodied environments, showcasing its potential in real-world applications.

한국어 요약

한 줄 요약

3D-VLA는 3D 비전-언어-행동 생성 세계 모델로, 2D 기반 VLA 모델의 한계를 극복하고 3D 환경에서의 추론과 행동 계획을 강화한다.

핵심 기여도

핵심 아이디어

기존 VLA 모델은 2D 입력에 의존하며, 3D 물리적 세계와의 통합이 부족하다. 또한, 행동 예측 시 환경의 동역학을 고려하지 않고 단순히 인식에서 행동으로의 매핑을 학습한다. 인간은 3D 내부 표현을 기반으로 미래 시나리오를 상상하며 행동을 계획하는 반면, 기존 모델은 이러한 능력을 구현하지 못한다. 이를 해결하기 위해 3D-VLA는 3D-LLM에 기반한 생성 세계 모델을 제안하며, 환경과 상호작용할 수 있는 interaction tokens을 도입한다. 또한, 디퓨전 모델을 통해 3D 목표 이미지 및 포인트 클라우드를 생성하고, 이를 LLM과 연결하여 다중 모달 생성 능력을 강화한다.

기술적 접근법

주요 결과

의의 및 한계

3D-VLA는 3D 환경에서의 추론, 생성, 행동 계획을 통합한 첫 번째 모델로, 실제 로봇 조작 및 인공지능 시스템 개발에 기여할 수 있다. 또한, 2D 기반 모델의 한계를 극복하며, 3D 정보를 활용한 새로운 임베디드 작업이 가능해진다. 그러나, 3D 데이터 수집 및 처리는 여전히 어려움이 있으며, 대규모 3D 데이터셋 구축은 추가 연구가 필요하다. 또한, 디퓨전 모델과 LLM 간의 정확한 정렬 및 생성 품질 향상도 개선 포인트로 남아 있다.

실용적 활용

3D-VLA는 로봇 조작, 자율 주행, 가상 현실(VR) 및 증강현실(AR) 시스템 등 3D 환경에서의 인지 및 행동이 필요한 다양한 산업 분야에 적용 가능하다. 특히, 로봇이 3D 환경에서의 목표를 생성하고, 이를 기반으로 행동을 계획하는 데 유용하며, 실제 세계에서의 인공지능 시스템 개발에 기여할 수 있다.