透過概念縮放與密集監督釋放圖像編輯的潛力
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
August 17, 2026
作者: Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang
cs.AI
摘要
現有的影像編輯框架主要遵循文字到影像擴散模型的訓練範式。然而,將此範式擴展至影像編輯時,凸顯了兩個固有的差異:具體而言,一是對編輯概念粒度的關注不足,二是因稀疏監督訊號而導致的訓練效率低落。為了解決這些問題,我們建立了一個涵蓋超過 1,000 個細粒度編輯概念的全面層級式分類體系,並透過改良的合成框架構建了 ConceptEdit-12M,這是一個包含 1,200 萬對高品質編輯配對的大規模資料集。這種以概念庫驅動的方法有效矯正了生成資料的分佈崩潰問題,同時確保了高資料保真度。此外,我們提出了一種密集監督訓練策略,將多個互不干擾的概念合成至單一影像配對中。透過提供更豐富的學習訊號,此策略顯著提升了訓練效率與整體模型效能。訓練結果驗證了我們的策略,其表現顯著優於先前的研究。最後,我們提出了 ConceptEdit-Bench,這是一個旨在於廣泛真實世界場景中診斷模型能力的細粒度評估套件。
English
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.