ChatPaper.aiChatPaper

通过概念缩放与密集监督释放图像编辑的潜力

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

August 17, 2026
作者: Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang
cs.AI

摘要

现有的图像编辑框架主要遵循文本到图像扩散模型的训练范式。然而,将这一范式扩展到图像编辑时,暴露出两个固有的不一致性:具体而言,对编辑概念粒度的关注不足,以及稀疏监督信号导致的训练效率低下。为解决这些问题,我们建立了一个包含超过1000个细粒度编辑概念的综合层次化分类体系,并通过改进的合成框架构建了ConceptEdit-12M——一个包含1200万高质量编辑对的大规模数据集。这种基于库的方法有效纠正了生成数据的分布坍缩,同时保证了较高的数据保真度。此外,我们提出了一种密集监督训练策略,将多个互不干扰的概念合成为单个图像对。通过提供更丰富的学习信号,该策略显著提升了训练效率和整体模型性能。训练结果验证了我们的策略,显著优于先前工作。最后,我们提出了ConceptEdit-Bench,一个细粒度的评估套件,旨在广泛真实世界场景中诊断模型能力。
English
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.