ChatPaper.aiChatPaper

VDiff-Bench:一個具挑戰性的細粒度影像差異辨識基準

VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification

September 5, 2026
作者: Yixin Wan, Tianle Zheng, Kai-Wei Chang
cs.AI

摘要

多模態大型語言模型(MLLMs)在視覺問答等一般視覺理解任務上表現強勁,卻常在基本的比較能力上受挫:辨識兩張相似影像之間有何變更。我們提出 VDiff-Bench,一個具挑戰性的選擇題基準,用於細粒度影像差異辨識。VDiff-Bench 包含 1,756 道基於影像對的四選一題目,涵蓋 10 個變更類別:位置、運動、區域影像色彩、整體影像色彩、出現/消失、雜訊/解析度、紋理、替換/尺寸、OCR/文字,以及光照。每道題對應兩張輸入影像,並有 4 個選項:真實差異、兩個困難負面描述,以及一個「無差異」干擾項。為使任務具挑戰性,我們特別策劃以真實答案為條件的負例,要求模型區分實際變更與鄰近的語意替代描述。使用 11 個最先進的開源與閉源 MLLMs 進行的實驗顯示,細粒度視覺比較仍然脆弱:模型在不同來源與變更類別上的表現不均,且在雜訊與紋理等細微低階變更上持續失敗。例如,三個 7-8B 規模的開源 MLLMs 在語意變更上得分 52.5-70.6%,但在雜訊與紋理等低階變更上僅有 8.7-33.3%,錯誤地假定兩張輸入影像之間沒有變更。令人驚訝的是,儘管其他閉源商業模型表現強勁,Grok 4.3 在辨識影像間雜訊與紋理差異時展現顯著的效能下降,明顯落後於 Kimi K2.5 與 K3 等大型開源模型。總體而言,VDiff-Bench 為評估 MLLMs 的比較式視覺理解提供了針對性的診斷工具,揭示標準單影像視覺語言任務未能捕捉的失敗。
English
Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often struggle with a basic comparative skill: identifying what has changed between two similar images. We introduce VDiff-Bench, a challenging multiple-choice benchmark for fine-grained Image Difference Identification. VDiff-Bench contains 1,756 four-way questions over image pairs and covers 10 change categories: position, motion, regional image color, overall image color, appearance/disappearance, noise/resolution, texture, substitution/size, OCR/text, and illumination. Each question corresponds to two image inputs with 4 choices: the true difference, two hard negative descriptions, and a "no difference" distractor. To make the task challenging, we specifically curate ground-truth-conditioned negatives that require models to distinguish the actual change from nearby semantic alternatives. Experiments with 11 state-of-the-art open- and closed-source MLLMs show that fine-grained visual comparison remains brittle: models exhibit uneven performance across sources and change categories, with persistent failures on subtle low-level changes like noises and textures. For instance, three 7-8B-scale open-source MLLMs score 52.5-70.6% on semantic changes but only 8.7-33.3% on low-level changes like noise and texture, falsely assuming no changes between two image inputs. Surprisingly, despite strong performance of other closed-source commercial models, Grok 4.3 demonstrate remarkable performance drop on identifying noise and texture differences between images, falling significantly behind large open-source models like Kimi K2.5 and K3. Overall, VDiff-Bench provides a targeted diagnostic for evaluating comparative visual understanding in MLLMs, exposing failures that are not captured by standard single-image vision-language tasks.