CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing</p>\n","updatedAt":"2026-08-17T03:51:09.029Z","author":{"_id":"69e5c11e8605616ee4dcd92e","avatarUrl":"/avatars/edd17a0ed798a178df226e2ea25193a4.svg","fullname":"TaobaoTmall-AlgorithmProducts","name":"TaobaoTmall-AlgorithmProducts","type":"user","isPro":true,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":7,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7344925999641418},"editors":["TaobaoTmall-AlgorithmProducts"],"editorAvatarUrls":["/avatars/edd17a0ed798a178df226e2ea25193a4.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.14546","authors":[{"_id":"6a828385b601d59c65281470","name":"Qinye Zhou","hidden":false},{"_id":"6a828385b601d59c65281471","name":"Jun Zheng","hidden":false},{"_id":"6a828385b601d59c65281472","name":"Yongchao Du","hidden":false},{"_id":"6a828385b601d59c65281473","name":"Yuan Wang","hidden":false},{"_id":"6a828385b601d59c65281474","name":"Zhengrui Chen","hidden":false},{"_id":"6a828385b601d59c65281475","name":"Zuan Gao","hidden":false},{"_id":"6a828385b601d59c65281476","name":"Taihang Hu","hidden":false},{"_id":"6a828385b601d59c65281477","name":"Chao Lin","hidden":false},{"_id":"6a828385b601d59c65281478","name":"Yefeng Shen","hidden":false},{"_id":"6a828385b601d59c65281479","name":"Xingjian Wang","hidden":false},{"_id":"6a828385b601d59c6528147a","name":"Zhao Wang","hidden":false},{"_id":"6a828385b601d59c6528147b","name":"Zhengtao Wu","hidden":false},{"_id":"6a828385b601d59c6528147c","name":"Xiaoli Xu","hidden":false},{"_id":"6a828385b601d59c6528147d","name":"Zhengze Xu","hidden":false},{"_id":"6a828385b601d59c6528147e","name":"Hao Yan","hidden":false},{"_id":"6a828385b601d59c6528147f","name":"Denghui Yang","hidden":false},{"_id":"6a828385b601d59c65281480","name":"Yuhang Yu","hidden":false},{"_id":"6a828385b601d59c65281481","name":"Huayu Zhang","hidden":false},{"_id":"6a828385b601d59c65281482","name":"Mingzhou Zhang","hidden":false},{"_id":"6a828385b601d59c65281483","name":"Mengting Chen","hidden":false}],"publishedAt":"2026-08-14T00:00:00.000Z","submittedOnDailyAt":"2026-08-17T00:00:00.000Z","title":"CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing","submittedOnDailyBy":{"_id":"69e5c11e8605616ee4dcd92e","avatarUrl":"/avatars/edd17a0ed798a178df226e2ea25193a4.svg","isPro":true,"fullname":"TaobaoTmall-AlgorithmProducts","user":"TaobaoTmall-AlgorithmProducts","type":"user","name":"TaobaoTmall-AlgorithmProducts"},"summary":"With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.","upvotes":10,"discussionId":"6a828385b601d59c65281484","projectPage":"https://huggingface.co/datasets/TaobaoTmall-AlgorithmProducts/CPI-benchmark","githubRepo":"https://github.com/zqyzzz/CPI-benchmark","githubRepoAddedBy":"user","ai_summary":"CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.","ai_keywords":["CPI-Bench","multi-image editing","reasoning-based editing","image editing models","Arena Image Edit Leaderboard"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":2,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"6661ad9bc229e1da8e9b939d","avatarUrl":"/avatars/45d83ebdc1a2fc027309eb28f2584ef2.svg","isPro":false,"fullname":"chen","user":"mathildachen","type":"user"},{"_id":"65a909fe8581aad8c97a67d3","avatarUrl":"/avatars/96570e47117e957543d9f0fe5e1d9d57.svg","isPro":false,"fullname":"liutao","user":"byliutao","type":"user"},{"_id":"6470125bd742e9ef65222017","avatarUrl":"/avatars/6f1bfd673729f43f2a2a92e13c0287b8.svg","isPro":false,"fullname":"Zhengrui Chen","user":"zhengrchan","type":"user"},{"_id":"6381847a471a4550ff298c63","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6381847a471a4550ff298c63/RTKepvX67R6pLiiUidpUO.png","isPro":false,"fullname":"Jun","user":"zxbsmk","type":"user"},{"_id":"67594a15e561be24b25697c9","avatarUrl":"/avatars/73c8ab719c710903c10fd887ee0fcaca.svg","isPro":false,"fullname":"huayuzhang","user":"rainzhang666","type":"user"},{"_id":"634bec6ca5d4e109c6aa4895","avatarUrl":"/avatars/2f55327097b3782f6f4f59041f3b04b5.svg","isPro":false,"fullname":"hu","user":"taihang","type":"user"},{"_id":"650bbd1487dcda6616acb1e2","avatarUrl":"/avatars/925b83b15b239307b2cd444b8006a2ae.svg","isPro":false,"fullname":"lmnhsp","user":"lmnhsp","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"},{"_id":"647d92e09bb822b5cd3c7692","avatarUrl":"/avatars/73de68a74630667ea73e8dc86acd88ae.svg","isPro":false,"fullname":"Zhengtao Wu","user":"wuzht","type":"user"},{"_id":"69e5c11e8605616ee4dcd92e","avatarUrl":"/avatars/edd17a0ed798a178df226e2ea25193a4.svg","isPro":true,"fullname":"TaobaoTmall-AlgorithmProducts","user":"TaobaoTmall-AlgorithmProducts","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"68be41370a3fcebdcad6516a","name":"alibabagroup","fullname":"alibaba","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/68be3ab7e52df040b2cf80dc/li4G29u_EGswyTN1Sm_Kq.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.14546.md","query":{}}">
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
Abstract
CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.
With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.
Community
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.14546 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.14546 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.14546 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.