Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy

arXiv:2609.04196 · 2026-09-04 공개 · arXiv · PDF

multimodal-model closed-loop-control physics-simulation geometry-reconstruction trajectory-data omni-camera vision-language-camera puffin-16m

Abstract

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

한국어 요약

한 줄 요약

Puffin-World는 물리, 기하, 외관을 통합한 3D 월드 생성 및 재구성을 위한 단일 프레임워크로, Omni-Camera와 Physics Propagation을 통해 물리적으로 일관되고 시각적으로 안정적인 월드 모델링을 실현한다.

핵심 기여도

핵심 아이디어

Puffin-World는 기존 2D 중심의 모델과는 달리, 3D 월드를 물리적, 기하적, 시각적 상태를 통합하여 모델링하는 새로운 패러다임을 제시한다. 이는 단일 프레임워크 내에서 카메라 물리 이해, 자유 시점 공간 시뮬레이션, 3D 월드 생성 및 재구성을 가능하게 한다. 핵심은 **Omni-Camera 표현**과 **Physics Propagation**이다. Omni-Camera는 중력 기반 절대 방향과 레이 기반 상대 기하학을 결합하여, 카메라의 절대 물리 상태를 유지하면서 자유로운 시점 제어를 가능하게 한다. Physics Propagation은 기준 뷰에서의 절대 물리 상태를 미래 프레임으로 전파하여, 복잡한 카메라 움직임에서도 시각적 안정성과 물리적 일관성을 유지한다.

기술적 접근법

주요 결과

의의 및 한계

Puffin-World는 물리적, 기하적, 시각적 상태를 통합한 단일 프레임워크로, 복잡한 월드 모델링과 시뮬레이션을 가능하게 한다. 특히, **Physics Propagation**을 통해 중력 기반의 일관된 표현을 유지함으로써, 기존 카메라 표현 방식(예: Plücker embedding)의 한계를 극복한다. 또한, **Puffin-16M** 데이터셋은 다양한 움직임과 라벨을 통해 모델 확장을 가능하게 한다. 그러나 현재는 **정적 장면에만 적용**되며, **동적 장면**이나 **더 긴 시간 범위**에 대한 확장은 아직 연구 중이다. 또한, **하이퍼파라미터 세부 사항**은 명시되지 않았다.

실용적 활용

Puffin-World는 AR/VR, 로봇 시뮬레이션, 자율 주행, 3D 콘텐츠 생성 등에서 활용 가능하다. 특히, **자유 시점 제어**와 **물리적으로 일관된 월드 생성**이 필요한 산업 현장에서 유용하며, **Puffin-16M** 데이터셋은 3D 월드 모델링 연구를 촉진하는 기반이 될 수 있다.