ChatPaper.aiChatPaper

CPI-Bench: 실세계 이미지 편집을 위한 포괄적이고 실용적이며 지능적인 벤치마크

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

August 14, 2026
저자: Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen
cs.AI

초록

이미지 편집 모델의 급속한 발전과 다양한 영역에서의 광범위한 응용에 따라, 이러한 모델의 기능을 실제 세계 시나리오에 직접 배포해야 할 필요성이 점점 더 절실해지고 있다. 그러나 기존 벤치마크는 단순한 단일 이미지 작업에 국한되어 있어, 평가 범위가 제한적이고 다양한 모델 간의 성능을 효과적으로 변별하지 못한다. 결과적으로, 기존 벤치마크는 복잡한 다중 이미지 편집, 고도의 추론 능력을 요구하는 지시문, 그리고 실용적 배포 환경에서의 모델 성능을 신뢰성 있게 평가하지 못한다. 이러한 한계를 해결하기 위해, 우리는 실제 세계 이미지 편집을 위한 종합적이고 실용적이며 지능적인 벤치마크인 CPI-Bench를 제안한다. CPI-Bench는 세 가지 핵심 하위 집합으로 구성된다: 다양한 편집 작업을 포괄적으로 다루고 다중 이미지 편집 평가를 최초로 포함하는 CPI-General-Bench, 고빈도 실제 사용자 응용 시나리오에 초점을 맞춘 CPI-Practical-Bench, 그리고 고도의 추론 기반 편집 능력 평가에 전념하는 CPI-Intelligent-Bench이다. CPI-Bench에 기반한 주류 이미지 편집 모델의 평가 결과는 CPI-Bench가 모델 간 성능 변별력을 향상시킴을 보여준다. CPI-Bench는 일반 편집 능력, 실용적 배포 효율성, 그리고 고급 추론 기반 편집의 격차에 대한 종합적이고 신뢰할 수 있는 정량화를 제공하며, 향후 이미지 편집 모델 최적화에 귀중한 지침을 제시한다. 결정적으로, 우리의 순위 분석은 CPI-Bench가 Arena 이미지 편집 리더보드와 가장 높은 정렬도를 달성함을 보여주며, 이는 CPI-Bench가 인간 평가자의 선호도와 지각적 판단을 충실히 포착하여 실제 사용자 경험에 대한 견고한 대리 지표로 기능함을 시사한다.
English
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.