ChatPaper.aiChatPaper

Block3D: 블록 단위 확산 기반 효율적 텍스트-3D 생성

Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion

August 20, 2026
저자: Bowen Cui, Weijie Wang, Zeyu Zhang, Yefei He, Mingda Lin, Haoyu Zhao, Yuanyu He, Donny Y. Chen, Feng Chen, Bohan Zhuang
cs.AI

초록

텍스트-3D 생성은 빠르게 발전해 왔지만, 낮은 추론 비용으로 높은 기하학적 충실도를 달성하는 것은 여전히 어려운 과제이다. 기존 텍스트-3D 방법들은 이산 형상 토큰을 자기회귀적으로 디코딩하거나 확산 또는 플로우 모델을 사용해 전역 3D 표현을 반복적으로 정제한다. 그러나 자기회귀 디코딩은 순차적이어서 오류를 수정할 수 없으며, 확산 및 플로우 매칭 모델은 전체 표현을 반복적으로 처리하므로 고품질 생성의 비용이 점점 증가한다. 본 논문에서는 이산 형상 토큰 시퀀스를 연속 블록으로 분할하고, 블록을 자기회귀적으로 생성하면서 현재 블록 내 모든 토큰을 공동으로 노이즈 제거하는 블록 단위 확산 프레임워크인 Block3D를 제안한다. 오류 누적을 완화하기 위해 각 블록이 확정되기 전에 낮은 신뢰도의 토큰을 수정하는 신뢰도 기반 블록 내 보정을 도입한다. TRELLIS-500K의 홀드아웃 세트에서 Block3D는 평균 종단 간 생성 시간을 25.71초에서 4.99초로 단축하여, 기하학적 충실도를 희생하지 않으면서 미세 조정된 자기회귀 기준 모델 대비 5.15배 속도 향상을 달성한다.
English
While text-to-3D generation has advanced rapidly, achieving high geometric fidelity at low inference cost remains challenging. Existing text-to-3D methods either decode discrete shape tokens autoregressively or iteratively refine global 3D representations with diffusion or flow models. However, autoregressive decoding is sequential and cannot revise errors, whereas diffusion and flow-matching models repeatedly process the full representation, making high-quality generation increasingly expensive. In this paper, we propose Block3D, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block. To alleviate error accumulation, we introduce confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized. On a held-out set from TRELLIS-500K, Block3D reduces mean end-to-end generation time from 25.71 seconds to 4.99 seconds, achieving a 5.15times speedup over the fine-tuned autoregressive baseline without sacrificing geometric fidelity.