Boogu-Image-0.1: 오픈소스 통합 멀티모달 이해 및 생성 향상
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
July 14, 2026
저자: Guoxuan Chen, Chufeng Xiao, Haoran Yang, Siyue Xie, Binxiao Huang, Ming Zhang, Cheuk Him Chau, Xinyu Fu, Yingzhao Lian, Tom S. Y. Li, Jintao Lin, Bowen Dong, Zian Qian, Yuhao Liu, Yuxuan Hu, Weikang Shi, Bin Zou, Bowen Zheng, Haoxuan Che, Chang Chen, Yuyang He, Heyang Sun, Tianyu Huang, Chong Hou Choi, Cheng Gong, Han Shi, Haoli Bai, Xihui Liu, Hongsheng Li, Qifeng Chen, Chao Huang, Rui Liu, Chenyang Lei
cs.AI
초록
Boogu-Image-0.1을 소개합니다. 이는 Base, Turbo, Edit, Edit-Turbo 변형으로 구성된 오픈소스 통합 다중 모달 이해 및 생성 모델 제품군입니다. 고품질 텍스트-이미지 생성, 빠른 추론, 명령 기반 편집, 이중 언어(중국어-영어) 텍스트 렌더링에서 경쟁력 있는 성능을 제공합니다. Nano-Banana-Pro 및 GPT-Image-2와 같은 폐쇄형 다중 모달 시스템은 단일 모델이 아닌 시스템 수준 통합을 통해 강력한 성능을 달성하지만, 내부 관행은 대부분 공개되지 않은 상태입니다. 본 연구에서는 모델 이해도, 데이터 품질 및 훈련 파이프라인의 목표 지향적 개선과 에이전트적 추론 시간 확장을 결합하면 매우 제한된 연산 예산 하에서도 생성 및 편집 성능을 실질적으로 향상시킬 수 있음을 입증합니다. 종합 평가 결과, Boogu-Image-0.1은 표준 벤치마크에서 다른 오픈소스 모델과 일관되게 일치하거나 능가하며, 주요 폐쇄형 시스템에 근접한 결과를 달성합니다. 특히 이는 단 2억 862만 개의 고유 이미지만으로 이루어졌습니다. 기본 모델의 이론적 훈련 비용은 약 40만 달러에 불과합니다. 본 연구는 더 넓은 연구 커뮤니티에 가치가 있다고 판단되는 실용적 논의를 공유하며, 통합 다중 모달 이해 및 생성을 위한 개방형 생태계 발전을 위해 가중치, 코드 및 레시피를 Apache 2.0 라이선스로 공개합니다. 코드는 다음에서 확인할 수 있습니다: https://github.com/Boogu-Project/Boogu-Image.
English
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, and Edit-Turbo variants. It delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual (Chinese-English) text rendering. Closed-source multimodal systems like Nano-Banana-Pro and GPT-Image-2 achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. In this work, we demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets. Comprehensive evaluations show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks, and achieves results approaching leading closed-source systems. Notably, this is accomplished with only 208.62 million unique images. The base model's theoretical training cost is only approximately \$400K. We share practical discussions that we believe are valuable to the broader research community, and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. Our code is available here: https://github.com/Boogu-Project/Boogu-Image.