ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes

Mingda Lin, Weijie Wang, Zeyu Zhang, Bowen Cui, Yefei He, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang

arXiv:2609.01740 · 2026-09-03 공개 · arXiv · PDF

transformer high-fidelity latent-representation shape-net trellis nested-dropout cod-vae token-prefix

Abstract

Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding $32\times$ and $8\times$ shorter token sequences, respectively.

한국어 요약

한 줄 요약

ZipTok3D는 극단적으로 짧은 토큰 시퀀스로도 고정밀 3D 재구성을 가능하게 하는 3D 토크나이저로, ShapeNet과 TRELLIS 데이터셋에서 각각 32배와 8배 짧은 토큰 수를 달성했다.

핵심 기여도

핵심 아이디어

ZipTok3D는 기존 3D 토크나이저가 극단적으로 짧은 토큰 수에서 재구성 품질이 급격히 저하되는 문제를 해결하기 위해, **진행적 정보를 포함한 글로벌 토큰 프리픽스**를 구성하고, 이를 **반복적 디코딩**을 통해 풀어내는 방식을 채택했다.

기존 토크나이저는 고정된 토큰 수로 학습하며, 정보 분포를 암묵적으로 결정하지만, ZipTok3D는 **네스티드 드롭아웃**(nested dropout)을 통해 훈련 중 임의의 길이의 프리픽스만 사용하게 하여, **초기 토큰이 전체 기하 구조를 포함하도록 강제**한다. 이는 토큰 수가 줄어들 때도 핵심 정보가 유지되도록 한다.

디코더는 **파라미터 공유 Transformer 블록**(parameter-shared Transformer block)을 **5단계 반복**하여, 토큰 내의 정보를 점진적으로 풀어내며 세부 기하 구조를 복원한다. 이 방식은 별도의 샘플링 단계 없이도 고정밀 재구성을 가능하게 한다.

기술적 접근법

주요 결과

의의 및 한계

ZipTok3D는 극단적으로 짧은 토큰 수에서도 고정밀 3D 재구성을 가능하게 하며, 토큰 수를 줄이는 방식으로 **다운스트림 생성 모델링의 비용을 절감**할 수 있다. 특히, **네스티드 프리픽스 훈련**과 **반복적 디코딩**의 결합은 기존 토크나이저의 한계를 극복하는 새로운 접근법을 제시한다.

그러나, **입력 별로 적응적인 토큰 수 및 디코딩 깊이 조절**은 아직 연구되지 않았으며, **생성 모델과의 통합**도 향후 연구 주제로 남아 있다. 또한, **복잡한 객체나 대규모 데이터셋**에서의 성능은 추가 실험 필요.

실용적 활용

ZipTok3D는 3D 생성 모델에서 **토큰 수를 줄여 처리 비용을 절감**할 수 있어, **실시간 3D 생성**, **모바일/엣지 기기**에서의 활용, **대규모 3D 데이터 저장 및 전송 효율화**에 유용하게 사용될 수 있다. 특히, **클래스 조건 생성**(class-conditioned generation)에서도 성능이 뛰어나, **산업 디자인**, **게임 콘텐츠 생성**, **의료 영상 분석** 등에 적용 가능하다.