CPI-Bench:現実世界の画像編集のための包括的・実用的・インテリジェントなベンチマーク
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
August 14, 2026
著者: Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen
cs.AI
要旨
随着图像编辑模型的快速发展及其在各领域的广泛应用,将这些模型能力直接部署到真实世界场景中的需求日益迫切。然而,现有的基准测试仍局限于简单的单图像任务,存在覆盖维度有限、无法有效区分不同模型性能的问题。因此,它们无法可靠地评估模型在复杂多图像编辑、高要求推理指令以及实际部署场景中的表现。为了解决这些局限,我们提出了CPI-Bench,一个面向真实世界图像编辑的综合性、实用性与智能性基准测试。CPI-Bench包含三个核心子集:CPI-General-Bench,全面覆盖多样化编辑任务并率先引入多图像编辑评估;CPI-Practical-Bench,聚焦高频真实用户应用场景;以及CPI-Intelligent-Bench,专门评估高要求推理性编辑能力。基于CPI-Bench对主流图像编辑模型的评估结果表明,CPI-Bench增强了对不同模型性能的区分度。它为通用编辑能力、实际部署效果和高级推理性编辑之间的差距提供了全面且可靠的量化,为图像编辑模型未来的优化提供了宝贵的指导。关键在于,我们的排名分析显示,CPI-Bench与Arena Image Edit排行榜的对齐度最高,表明它忠实捕捉了人类评估者的偏好与感知判断,可作为真实用户体验的有力代理指标。
English
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.