ZipTok3D: 간결한 토큰 접두사를 활용한 고품질 3D 토큰화
ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes
September 1, 2026
저자: Mingda Lin, Weijie Wang, Zeyu Zhang, Bowen Cui, Yefei He, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang
cs.AI
초록
간결한 토큰 시퀀스는 효율적인 3D 생성에 필수적이다. 그러나 기존의 3D 토크나이저들은 일반적으로 잠재 표현을 공간 영역별로 정리하거나 고정 크기의 전역 토큰 집합으로 구성하며, 두 접근 방식 모두 토큰 예산이 매우 낮은 수준으로 압축되면 재구성 품질이 급격히 저하된다. 본 논문은 극도로 짧은 토큰 시퀀스에서도 고충실도의 재구성을 수행하도록 설계된 3D 토크나이저 ZipTok3D를 제시한다. 그 핵심 아이디어는 객체의 기하학적 형태를 점진적으로 정보량이 풍부한 전역 토큰 프리픽스들로 구성하고, 반복적 디코딩을 통해 이렇게 압축된 표현을 펼쳐내는 것이다. 구체적으로, 중첩 드롭아웃이 훈련 중 인코딩 후 잠재 시퀀스를 무작위로 잘라내며, 이때 유지된 각 프리픽스가 전체 객체를 재구성하도록 강제함으로써 필수적인 기하 정보가 앞쪽 토큰들에 우선적으로 자리 잡게 된다. 이후 디코더는 파라미터가 공유된 트랜스포머 블록을 반복 적용하며, 별도의 생성 샘플링 단계 없이 각 프리픽스로부터 정밀한 기하학적 세부 정보를 복원한다. 동일한 토큰 차원을 사용할 때 ZipTok3D는 ShapeNet에서 단 1개의 토큰, TRELLIS에서 4개의 토큰만을 사용하여 32토큰 기반의 COD-VAE 기준선과 필적하는 재구성 품질을 달성하고, 이는 각각 32배 및 8배 짧은 토큰 시퀀스를 제공한다.
English
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding 32times and 8times shorter token sequences, respectively.