SnapBench:用於行動互動的即拍即問多模態檢索基準測試
SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
August 30, 2026
作者: Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang
cs.AI
摘要
行動AI扮演視覺神諭的角色,讓使用者拍下眼前事物並提出問題以獲取資訊。「拍問」(snap-and-ask)檢索已成為行動AI最常見的入口之一,然而照片往往模糊不清,文字提問也可能簡短或輸入有誤。現有基準測試僅在乾淨輸入上進行評估,或未能在拍問檢索中隔離配對穩健性的影響。為此,我們提出SnapBench——首個針對穩健拍問式多模態檢索的配對基準,涵蓋1,145個查詢、9,085個圖庫項目,並在53種受控損壞條件下提供人工註解。我們評估了16種多模態檢索器,包括雙塔編碼器與基於嵌入的視覺語言模型。結果顯示,影像損壞會大幅降低檢索效能,而文字損壞主要影響純文字檢索,對聯合檢索的影響有限。乾淨的純影像檢索表現往往優於聯合檢索,反映出粗略文字所造成的拖累,以及系統在雜訊輸入下缺乏跨模態回退機制的問題。SnapBench為拍問情境下的穩健檢索評估提供了受控測試環境。我們進一步提出MOOR(模態錨定、異常感知、最優重加權),一種簡單的自適應融合方法,突顯了拍問檢索中可靠性感知模態校正的必要性。
English
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.