ChatPaper.aiChatPaper

CPI-Bench:一个全面、实用且智能的真实世界图像编辑基准

CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

August 14, 2026
作者: Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen
cs.AI

摘要

随着图像编辑模型的快速发展及其在各领域的广泛应用,将这些模型能力直接部署到实际场景中的需求日益迫切。然而,现有基准仍然局限于简单的单图像任务,存在覆盖维度有限、无法有效区分不同模型性能等问题。因此,它们无法可靠地评估模型在复杂多图像编辑、高要求推理指令和实际部署场景中的性能。为解决上述局限,我们提出了CPI-Bench,一个面向真实世界图像编辑的综合、实用且智能的基准。CPI-Bench包含三个核心子集:CPI-General-Bench,全面覆盖多样化的编辑任务并首创性地纳入多图像编辑评估;CPI-Practical-Bench,聚焦于高频真实用户应用场景;以及CPI-Intelligent-Bench,专门评估高要求推理型编辑能力。基于CPI-Bench对主流图像编辑模型的评估结果表明,CPI-Bench增强了模型间的性能区分度。它为通用编辑能力、实际部署效果和高级推理型编辑之间的差距提供了全面可靠的量化,为图像编辑模型未来的优化提供了宝贵的指导。至关重要的是,我们的排名分析显示,CPI-Bench与Arena Image Edit排行榜实现了最高的一致性,表明其能够真实反映人类评估者的偏好和感知判断,可作为真实用户体验的可靠代理。
English
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.