말뭉치에서 공진화하는 능력으로: 범용 이미지 생성을 위한 능력 중심 데이터 설계
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
August 18, 2026
저자: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen
cs.AI
초록
대규모 이미지 생성은 데이터 규모, 품질, 재균형화, 재캡셔닝의 발전을 통해 이점을 얻었지만, 기존 파이프라인은 일반적으로 작업별 데이터셋을 개별적으로 최적화한다. 핵심 과제는 각 작업별 코퍼스를 큐레이션하는 방법뿐만 아니라, 생성 능력 간의 의존성에 따라 이질적 감독을 구성하는 방법에 있다. 우리는 능력별 감독 구성과 능력 정렬 커리큘럼 스케줄링을 결합한 능력 기반 데이터 인프라를 제시한다. 이 인프라의 세 가지 특화되었지만 상호 운용 가능한 데이터 엔진은 텍스트-이미지 그라운딩, 이미지 간 변환, 이미지-지식 연관을 위한 보완적 관계형 감독을 구축하며, 캡션 전문가들은 작업과 세분성에 걸쳐 T2I 및 편집 감독을 정렬한다. 다단계 커리큘럼은 능력 획득의 의존 순서에 따라 작업 구성, 시각 개념 분포, 데이터 품질, 이미지 해상도를 공동으로 진화시키며, 능력 인지 평가는 표적 검색, 전문가 구성, 격차 인지 리샘플링을 통해 루프를 닫는다. 대규모로, 이 프레임워크는 4억 4천만 개 이미지의 T2I 코퍼스, 1억 2천만 개의 편집 쌍, 2700만 개 이상의 이미지-엔티티 쌍을 큐레이션한다. 이 인프라를 사용하여 우리는 각각 3B 및 6B 크기의 두 가지 규모로 멀티모달 확산 모델을 처음부터 학습한다. 우리는 CPI-Bench에 대한 정량적 평가를 수행하고, 다양한 텍스트-이미지 및 편집 시나리오에 걸친 정성적 평가도 수행한다. 실험 결과는 광범위한 시각적 범위, 다재다능한 렌더링, 그리고 생성 능력 간의 효과적인 전이를 보여준다.
English
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a capability-driven data infrastructure that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.