ChatPaper.aiChatPaper

從語料庫到共同演化能力:以能力為中心的通用圖像生成資料設計

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

August 18, 2026
作者: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
cs.AI

摘要

大規模影像生成已受益於資料規模、品質、重新平衡與重新生成描述等方面的進展,然而傳統流程通常孤立地最佳化特定任務的資料集。核心挑戰不僅在於如何整理每個任務特定的語料庫,還在於如何根據生成能力之間的依賴關係來組織異質監督。我們提出一個以能力為驅動的資料基礎設施,將特定能力的監督建構與對齊能力的課程排程相結合。其中的三個專業但可互操作的資料引擎,分別為文字-影像對應、影像間轉換與影像-知識關聯建構互補的關係監督;同時,描述專家使文字到影像(T2I)與編輯監督在不同任務與粒度上保持一致。一個多階段課程依循能力習得的依賴順序,共同演進任務組成、視覺概念分布、資料品質與影像解析度;能力感知評估則透過針對性檢索、專家建構與差距感知的重新取樣來閉環。在大規模應用中,該框架整理了包含 4.4 億張影像的 T2I 語料庫、1.2 億個編輯對,以及超過 2700 萬個影像-實體對。利用此基礎設施,我們從零開始訓練兩種規模的多模態擴散模型,參數量分別為 3B 與 6B。我們在 CPI-Bench 上進行定量評估,並在各種文字到影像與編輯場景中進行定性評估。實驗結果展現了廣泛的視覺覆蓋、多功能的渲染,以及跨生成能力的有效遷移。
English
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.