Boogu-Image-0.1:提升开源统一多模态理解与生成能力
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
July 14, 2026
作者: Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
cs.AI
摘要
我们推出Boogu-Image-0.1,一个开源的统一多模态理解与生成模型系列,包含Base、Turbo、Edit和Edit-Turbo四个变体。该系列在高质量文生图、快速推理、基于指令的图像编辑以及中英文双语文字渲染方面均展现出具有竞争力的性能。像Nano-Banana-Pro和GPT-Image-2这类闭源多模态系统,其强大性能源于系统级整合而非单一模型,但内部实现细节大多未公开。在本工作中,我们证明:即使在计算预算极为有限的条件下,通过针对性地改进模型理解能力、数据质量和训练流程,并配合基于代理的推理时扩展,也能显著提升生成与编辑性能。全面评估显示,Boogu-Image-0.1在各项标准基准上持续达到或超越其他开源模型,并取得接近领先闭源系统的结果。值得注意的是,这仅使用了2.0862亿张独特图像。基础模型的理论训练成本仅为约40万美元。我们分享了对更广泛研究社区有价值的实践讨论,并根据Apache 2.0协议开源模型权重、代码及配方,以推动统一多模态理解与生成的开源生态发展。代码地址:https://github.com/Boogu-Project/Boogu-Image。
English
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.