Boogu-Image-0.1: オープンソースの統一マルチモーダル理解と生成の促進

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation

July 14, 2026
著者: Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
cs.AI

要旨

我々は、オープンソースの統一マルチモーダル理解生成モデルファミリーであるBoogu-Image-0.1を紹介する。これはBase、Turbo、Edit、Edit-Turboの各バリアントから構成され、高品質なテキストから画像への生成、高速推論、指示ベース編集、およびバイリンガル(中国語-英語)テキストレンダリングにおいて競争力のある性能を発揮する。Nano-Banana-ProやGPT-Image-2のようなクローズドソースのマルチモーダルシステムは、単一モデルではなくシステムレベルの統合により強力な性能を達成しているが、その内部実践はほとんど公開されていない。本研究では、モデル理解、データ品質、トレーニングパイプラインへの的を絞った改善と、エージェント的推論時スケーリングを組み合わせることで、非常に制約の厳しい計算予算下でも生成・編集性能を大幅に向上できることを示す。包括的な評価により、Boogu-Image-0.1が標準ベンチマークにおいて他のオープンソースモデルと一貫して同等かそれ以上の性能を示し、主要なクローズドソースシステムに迫る結果を達成することが示された。特筆すべきは、これをわずか2億862万枚のユニークな画像で達成している点である。ベースモデルの理論上のトレーニングコストは約40万ドルに過ぎない。我々は、広範な研究コミュニティにとって価値があると考える実践的な議論を共有し、統一マルチモーダル理解生成のためのオープンエコシステムを推進するため、Apache 2.0のもとで重み、コード、レシピを公開する。コードはこちらで入手可能である:https://github.com/Boogu-Project/Boogu-Image
English
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.
PDF1072July 17, 2026