CPI-Bench:一個全面、實用且智慧的真實世界影像編輯基準
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
August 14, 2026
作者: Qinye Zhou, Jun Zheng, Yongchao Du, Yuan Wang, Zhengrui Chen, Zuan Gao, Taihang Hu, Chao Lin, Yefeng Shen, Xingjian Wang, Zhao Wang, Zhengtao Wu, Xiaoli Xu, Zhengze Xu, Hao Yan, Denghui Yang, Yuhang Yu, Huayu Zhang, Mingzhou Zhang, Mengting Chen
cs.AI
摘要
隨著圖像編輯模型的快速發展及其在各個領域的廣泛應用,將這些模型能力直接部署到實際場景中的需求日益迫切。然而,現有的基準測試仍局限於簡單的單圖像任務,存在覆蓋維度有限、難以有效區分不同模型性能的問題。因此,它們無法可靠地評估模型在複雜多圖像編輯、高要求推理指令以及實際部署場景中的性能。為了解決這些局限性,我們提出了CPI-Bench,一個面向真實世界圖像編輯的綜合性、實用性與智能性基準。CPI-Bench包含三個核心子集:CPI-General-Bench,全面涵蓋多樣化的編輯任務,並首創性地納入多圖像編輯評估;CPI-Practical-Bench,專注於高頻率的真實用戶應用場景;以及CPI-Intelligent-Bench,致力於評估高要求推理型編輯的能力。基於CPI-Bench對主流圖像編輯模型的評估結果表明,CPI-Bench增強了模型之間的性能區分度。它為通用編輯能力、實際部署效果以及高級推理型編輯之間的差距提供了全面且可靠的量化,為圖像編輯模型未來的優化提供了寶貴的指導。關鍵的是,我們的排名分析顯示,CPI-Bench與Arena Image Edit排行榜的對齊程度最高,表明它能準確反映人類評估者的偏好與感知判斷,可作為真實用戶體驗的有力替代指標。
English
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.