SnapBench:面向移动交互的即拍即问多模态检索评估基准
SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
August 30, 2026
作者: Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang
cs.AI
摘要
移动AI如同一位视觉向导,使用户能够拍摄物体照片并获取相关信息。拍图问答检索已成为移动AI最常见的入口之一,然而照片往往模糊不清,文本问题也可能简短或拼写有误。现有基准测试仅在干净输入上评估,或未在拍图问答检索中独立验证配对鲁棒性。为此,我们提出SnapBench——首个面向鲁棒拍图问答多模态检索的配对基准,涵盖1,145个查询和9,085个候选集项目,覆盖53种受控损坏条件并包含人工标注。我们评估了16种多模态检索器,涵盖双塔编码器和基于嵌入的VLM。结果表明,图像损坏会显著降低检索性能,而文本损坏主要影响纯文本检索,对联合检索的影响有限。干净图像上的纯图像检索往往优于联合检索,揭示了粗粒度文本的拖累效应以及噪声输入下跨模态回退机制的缺失。SnapBench为拍图问答场景下的鲁棒检索评估提供了受控测试平台。我们进一步提出MOOR(模态锚定、离群值感知、最优重加权)——一种简洁的自适应融合方法,强调拍图问答检索中可靠性感知模态校准的必要性。
English
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.