从语料库到协同演进的能力:面向通用图像生成的能力中心化数据设计
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
August 18, 2026
作者: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
cs.AI
摘要
大规模图像生成得益于数据规模、质量、再平衡和重新描述方面的进步,然而传统流程通常孤立地优化特定任务的数据集。核心挑战不仅在于如何整理每个任务专属的语料库,还在于如何根据生成能力之间的依赖关系来组织异构监督。我们提出了一种能力驱动的数据基础设施,将能力专属的监督构建与能力对齐的课程调度相结合。其三个专用但互操作的数据引擎为文本-图像基础、图像间变换和图像-知识关联构建互补的关系监督,而描述专家则跨任务和粒度对齐T2I与编辑监督。多阶段课程沿着能力习得的依赖顺序,联合演化任务组成、视觉概念分布、数据质量和图像分辨率;能力感知的评估则通过定向检索、专家构建和差距感知重采样来形成闭环。在大规模场景下,该框架构建了包含4.4亿图像的T2I语料库、1.2亿编辑对和超过2700万图像-实体对。我们利用这一基础设施,从零训练了两个规模的多模态扩散模型,参数量分别为3B和6B。我们在CPI-Bench上进行了定量评估,并在多种文本到图像和编辑场景中进行了定性评估。实验结果表明,该方法具有广泛的视觉覆盖、多样化的渲染能力,并能在生成能力之间有效迁移。
English
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.