ChatPaper.aiChatPaper

개념 스케일링과 밀집 감독을 통한 이미지 편집의 잠재력 발현

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

August 17, 2026
저자: Long Cui, Xiaoqian Liu, Qi Qin, Yi Xin, Tao Lin, Jianguo Li, Linfeng Zhang
cs.AI

초록

기존의 이미지 편집 프레임워크는 대부분 텍스트-이미지 확산 모델의 훈련 패러다임을 따른다. 그러나 이 패러다임을 이미지 편집으로 확장하면 두 가지 본질적인 불일치, 즉 편집 개념 세분성에 대한 주의 부족과 희소한 지도 신호로 인한 훈련 비효율성이 드러난다. 이러한 문제를 해결하기 위해, 우리는 1,000개 이상의 세분화된 편집 개념을 포함하는 포괄적인 계층적 분류 체계를 구축하고, 개선된 합성 프레임워크를 통해 1,200만 개의 고품질 편집 쌍으로 구성된 대규모 데이터셋 ConceptEdit-12M을 구축한다. 이 라이브러리 기반 접근법은 높은 데이터 충실도를 보장하면서 생성 데이터의 분포 붕괴를 효과적으로 교정한다. 또한, 우리는 단일 이미지 쌍에 서로 간섭하지 않는 여러 개념을 합성하는 밀집 지도 훈련 전략을 제안한다. 이 전략은 더 풍부한 학습 신호를 제공함으로써 훈련 효율과 전반적인 모델 성능을 모두 크게 향상시킨다. 훈련 결과는 우리의 전략을 검증하며, 이전 연구들을 크게 능가한다. 마지막으로, 우리는 다양한 실제 시나리오 전반에 걸쳐 모델의 능력을 진단하도록 설계된 세분화된 평가 스위트인 ConceptEdit-Bench를 제시한다.
English
Existing image editing frameworks predominantly follow the training paradigm of text-to-image diffusion models. However, extending this paradigm to image editing highlights two inherent discrepancies, specifically, the insufficient attention to edit concept granularity and the training inefficiency caused by sparse supervision signals. To address these issues, we establish a comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts and build ConceptEdit-12M, a massive dataset of 12 million high-quality editing pairs via an improved synthesis framework. This library-driven approach effectively rectifies the distribution collapse of generated data while ensuring high data fidelity. Furthermore, we propose a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs. By providing richer learning signals, this strategy significantly enhances both training efficiency and overall model performance. Training results validate our strategy, significantly outperforming prior works. Finally, we present ConceptEdit-Bench, a granular evaluation suite designed to diagnose model capabilities across a vast array of real-world scenarios.