ChatPaper.aiChatPaper

Hunyuan3D-Buffalo 1.0:用于可扩展3D生成、理解与编辑的统一多模态模型

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

August 3, 2026
作者: Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
cs.AI

摘要

近期图像生成领域的进展展示了集成理解、生成与编辑的统一多模态模型的巨大潜力。然而,统一的3D建模仍受限于多模态数据的稀缺,尤其是缺乏大规模且几何一致的编辑数据。为解决这一局限,我们提出了Hunyuan3D-Buffalo 1.0——一个在单一架构内支持3D理解、文本到3D生成、指令引导的3D编辑以及文本驱动的部件生成的统一框架。为实现可扩展的训练,我们构建了8700万规模的3D多模态语料库,包含2500万个理解样本、5000万对文本到3D数据对,以及使用Nano3D-v2生成的1200万对编辑数据对。在架构上,该框架将Hunyuan3D-VLM(用于语义、结构和空间理解)与Hunyuan3D DiT(用于高保真3D合成)相结合。VLM为生成提供多模态语义条件,而编辑与部件生成任务额外以源对象表示作为扩散过程的条件,以保留对象的整体结构和未编辑区域。大量实验表明,Hunyuan3D-Buffalo 1.0在文本到3D生成和3D编辑基准上取得了最先进或领先的性能,同时展现出强大的理解和部件生成能力。进一步分析表明,生成和理解均能促进编辑性能提升,验证了统一3D多模态训练的有效性。项目主页:https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
English
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/