ChatPaper.aiChatPaper

Hunyuan3D-Buffalo 1.0:スケーラブルな3D生成、理解、編集のための統合マルチモーダルモデル

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

August 3, 2026
著者: Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
cs.AI

要旨

近年の画像生成における進歩は、理解・生成・編集を統合したマルチモーダルモデルの可能性を実証している。しかし、統合された3Dモデリングは、マルチモーダルデータの不足、特に大規模かつ幾何学的に一貫性のある編集データの欠如によって依然として制約を受けている。この限界に対処するため、我々はHunyuan3D-Buffalo 1.0を提案する。これは、3D理解、テキストからの3D生成、指示に基づく3D編集、テキストに基づくパーツ生成を単一のアーキテクチャ内でサポートする統合フレームワークである。スケーラブルな学習を可能にするため、我々は8700万規模の3Dマルチモーダルコーパスを構築した。これは、2500万の理解サンプル、5000万のテキストから3Dへのペア、およびNano3D-v2を用いて生成された1200万の編集ペアで構成される。アーキテクチャ上、本フレームワークは意味的・構造的・空間的理解のためのHunyuan3D-VLMと、高忠実度の3D合成のためのHunyuan3D DiTを組み合わせている。VLMは生成のためのマルチモーダルな意味的条件を提供し、編集とパーツ生成はさらに、ソースオブジェクトの表現を拡散プロセスに条件付けして、その全体的な構造と未編集領域を維持する。大規模な実験により、Hunyuan3D-Buffalo 1.0はテキストからの3D生成および3D編集ベンチマークにおいて最先端またはトップクラスの性能を達成し、強力な理解能力とパーツ生成能力を示すことが実証された。さらに、我々の分析は、生成と理解の両方が編集を改善することを示しており、統合された3Dマルチモーダル学習の有効性を実証している。プロジェクトページ: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
English
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/