SnapBench: モバイルインタラクションのためのSnap-and-Askマルチモーダル検索のベンチマーク評価
SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
August 30, 2026
著者: Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang
cs.AI
要旨
モバイルAIは視覚的なオラクルとして機能し、ユーザーが何かの写真を撮ってその情報を尋ねることを可能にする。スナップ&アスク検索は現在、モバイルAIの最も一般的な入り口の一つであるが、写真はしばしばぼやけており、テキストの質問は短かったり、打ち間違いを含んでいたりする。既存のベンチマークは、クリーンな入力のみを対象とするか、スナップ&アスク検索における画像・テキストペアのロバスト性を切り分けて評価していない。そこで本稿では、スナップ&アスクのマルチモーダル検索のロバスト性を評価する、画像・テキストのペアを扱う初のベンチマークであるSnapBenchを導入する。SnapBenchは、人手によるアノテーションを伴う1,145件のクエリと9,085件のギャラリー項目で構成され、53種類の制御された劣化条件を網羅している。我々は、デュアルタワー型エンコーダと埋め込みベースのVLMを含む16種類のマルチモーダル検索モデルを評価した。実験の結果、画像の劣化は検索性能を大幅に低下させる一方、テキストの劣化は主にテキストのみの検索に影響し、ジョイント検索(画像とテキストの併用検索)への影響は限定的であることが示された。また、クリーンな画像のみの検索がジョイント検索を上回る場合も多く、これはノイズを含む入力下では、粗いテキストが性能を引き下げ、クロスモーダルなフォールバックが働かないことを示している。SnapBenchは、スナップ&アスクのシナリオにおけるロバスト検索を評価するための制御されたテストベッドを提供する。さらに本稿では、MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting)とよぶ単純な適応的融合手法を提案し、スナップ&アスク検索における信頼性を考慮したモダリティキャリブレーションの必要性を強調する。
English
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.