コーパスから共進化する能力へ:汎用画像生成のための能力中心のデータ設計
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
August 18, 2026
著者: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
cs.AI
要旨
大規模画像生成は、データ規模、品質、再バランス、再キャプション付けの進歩から恩恵を受けているが、従来のパイプラインは通常、タスク固有のデータセットを個別に最適化している。中心的な課題は、各タスク固有のコーパスをどのようにキュレートするかだけでなく、生成能力間の依存関係に従って異種の教師信号をどのように整理するかである。我々は、能力固有の教師信号構築と能力に整合したカリキュラムスケジューリングを結合する、能力駆動型データ基盤を提示する。その3つの専門化されつつ相互運用可能なデータエンジンは、テキスト-画像対応付け、画像間変換、画像-知識関連付けのための相補的な関係性教師信号を構築し、キャプション専門家はT2Iと編集の教師信号をタスクと粒度を横断して整合させる。多段階カリキュラムは、能力獲得の依存順序に沿って、タスク構成、視覚概念分布、データ品質、画像解像度を共進化させ、能力を考慮した評価は、ターゲットを絞った検索、エキスパート構築、ギャップを考慮したリサンプリングを通じてループを閉じる。大規模には、このフレームワークは4億4000万画像のT2Iコーパス、1億2000万の編集ペア、2700万以上の画像-エンティティペアをキュレートする。この基盤により、我々はマルチモーダル拡散モデルを2つのスケール(それぞれ30億および60億パラメータ)でスクラッチから訓練する。我々はCPI-Benchでの定量的評価と、多様なテキストから画像への変換および編集シナリオにわたる定性的評価を実施する。実験結果は、広範な視覚カバレッジ、多様なレンダリング、生成能力間の効果的な転移を示している。
English
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.