Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

arXiv:2608.02711 · 2026-08-05 공개 · arXiv · PDF

diffusion-model multimodal-model text-to-3d nano3d-v2 hunyuan3d-vlm hunyuan3d-dit part-generation

Abstract

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/

한국어 요약

한 줄 요약

Hunyuan3D-Buffalo 1.0은 87M 규모의 3D 다중모달 데이터셋을 기반으로 3D 이해, 생성, 편집을 통합한 첫 단일 아키텍처다.

핵심 기여도

핵심 아이디어

기존 3D 생성 및 편집 모델은 데이터 부족, 특히 대규모 기하학적 일관성 있는 편집 데이터 부족으로 인해 분리된 시스템으로 개발되어 왔다. 본 연구는 이 문제를 해결하기 위해 **Nano3D-v2** 알고리즘을 도입하여 대규모 3D 편집 데이터를 생성하고, **Hunyuan3D-VLM**과 **3D-DiT**를 결합한 통합 아키텍처를 제안한다. Hunyuan3D-VLM은 3D 객체의 세분화된 의미, 구조, 공간적 이해를 담당하며, 3D-DiT는 고해상도 3D 생성을 담당한다. 이 두 모듈은 **MLP-Connector**를 통해 조건부 공간에 정렬되어, 생성 및 편집 과정에서 상위 수준의 추론이 확산 과정을 효과적으로 가이드할 수 있도록 설계되었다.

기술적 접근법

주요 결과

의의 및 한계

Hunyuan3D-Buffalo 1.0은 3D 생성, 이해, 편집을 단일 아키텍처로 통합함으로써, 3D 모델링 분야에서의 다중모달 학습의 가능성을 입증한다. 특히, 생성과 이해가 편집 성능 향상에 기여한다는 점에서, 통합 학습의 효과를 실증적으로 보여준다. 그러나 **단일 단계 고질량 기하 표현**, **텍스트-3D 데이터셋의 캡션 품질**, **텍스처 편집**, **편집 데이터 생성 파이프라인의 안정성**, **새로운 아키텍처 탐색**, **데이터 확장** 등의 문제는 여전히 해결되지 않은 한계점이다.

실용적 활용

Hunyuan3D-Buffalo 1.0은 3D 콘텐츠 생성, 게임 개발, 제품 디자인, AR/VR 등에서 사용 가능한 통합 솔루션으로 활용될 수 있다. 특히, 텍스트 기반의 3D 편집 기능은 디자이너가 복잡한 3D 모델을 쉽게 수정할 수 있도록 지원하며, 부품 수준의 생성 기능은 제품 설계 및 맞춤화 분야에서 활용 가능하다.