Boogu-Image-0.1:提升開源統一多模態理解與生成
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
July 14, 2026
作者: Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
cs.AI
摘要
我們推出了Boogu-Image-0.1,這是一個開源的統一多模態理解與生成模型系列,包含Base、Turbo、Edit和Edit-Turbo四種變體。它在高品質文字轉圖像生成、快速推理、基於指令的編輯以及雙語(中英文)文字渲染方面展現出具有競爭力的性能。像Nano-Banana-Pro和GPT-Image-2這類閉源多模態系統,是透過系統層級的整合而非單一模型達到強勁表現,但其內部做法大多未公開。在這項工作中,我們證明,針對模型理解能力、資料品質和訓練管線進行有針對性的改進,並結合智慧型推論時擴展(agentic inference-time scaling),即使在極為有限的計算預算下,也能顯著提升生成和編輯性能。全面評估結果顯示,Boogu-Image-0.1在標準基準測試中持續達到或超越其他開源模型,並取得接近領先閉源系統的成果。值得注意的是,這僅使用了2.0862億張獨特影像。基礎模型的理論訓練成本僅約40萬美元。我們分享這些實務討論,相信對更廣泛的研究社群具有價值,並在Apache 2.0許可證下釋出權重、程式碼和配方,以推動統一多模態理解與生成的開放生態系統。我們的程式碼可於此處取得:https://github.com/Boogu-Project/Boogu-Image。
English
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.