Hunyuan3D-Buffalo 1.0: 확장 가능한 3D 생성, 이해 및 편집을 위한 통합 멀티모달 모델
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
August 3, 2026
저자: Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li, Yansong Qu, Chunshi Wang, Jingwei Xu, Yunhan Yang, Zibo Zhao, Jiachen Xu, Jiaao Yu, Lifu Wang, Zhihao Liang, Zhuo Chen, Chunchao Guo
cs.AI
초록
최근 이미지 생성 분야의 발전은 이해, 생성, 편집을 아우르는 통합 멀티모달 모델의 가능성을 입증했습니다. 그러나 통합 3D 모델링은 여전히 희소한 멀티모달 데이터, 특히 대규모이면서 기하학적으로 일관된 편집 데이터의 부족으로 인해 제약을 받고 있습니다. 이러한 한계를 해결하기 위해, 우리는 단일 아키텍처 내에서 3D 이해, 텍스트-3D 생성, 지시 기반 3D 편집, 텍스트 기반 부분 생성을 지원하는 통합 프레임워크인 Hunyuan3D-Buffalo 1.0을 제안합니다. 확장 가능한 학습을 위해, 우리는 Nano3D-v2를 사용해 생성한 2,500만 개의 이해 샘플, 5,000만 개의 텍스트-3D 쌍, 1,200만 개의 편집 쌍으로 구성된 8,700만 규모의 3D 멀티모달 코퍼스를 구축했습니다. 구조적으로, 이 프레임워크는 의미론적, 구조적, 공간적 이해를 위한 Hunyuan3D-VLM과 고충실도 3D 합성을 위한 Hunyuan3D DiT를 결합합니다. VLM은 생성을 위한 멀티모달 의미 조건을 제공하며, 편집 및 부분 생성은 소스 객체 표현에 따라 확산 과정을 추가로 조건화하여 전체 구조와 편집되지 않은 영역을 보존합니다. 광범위한 실험을 통해 Hunyuan3D-Buffalo 1.0이 텍스트-3D 생성 및 3D 편집 벤치마크에서 최첨단 또는 선도적인 성능을 달성하고, 강력한 이해 및 부분 생성 능력을 갖추었음을 입증했습니다. 추가 분석은 생성과 이해가 모두 편집 성능을 향상시킨다는 것을 보여주며, 통합 3D 멀티모달 학습의 효과성을 확인합니다. 프로젝트 페이지: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/
English
Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimodal data, particularly the lack of large-scale and geometrically consistent editing data. To address this limitation, we propose Hunyuan3D-Buffalo 1.0, a unified framework supporting 3D understanding, text-to-3D generation, instruction-guided 3D editing, and text-grounded part generation within a single architecture. To enable scalable training, we construct an 87M-scale 3D multimodal corpus, comprising 25M understanding samples, 50M text-to-3D pairs, and 12M editing pairs generated using Nano3D-v2. Architecturally, the framework combines Hunyuan3D-VLM for semantic, structural, and spatial understanding with Hunyuan3D DiT for high-fidelity 3D synthesis. The VLM provides multimodal semantic conditions for generation, while editing and part generation additionally condition the diffusion process on the source object representation to preserve its overall structure and unedited regions. Extensive experiments show that Hunyuan3D-Buffalo 1.0 achieves state-of-the-art or leading performance on text-to-3D generation and 3D editing benchmarks, while exhibiting strong understanding and part-generation capabilities. Our analysis further shows that both generation and understanding improve editing, demonstrating the effectiveness of unified 3D multimodal training. Project Page: https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0/