ChatPaper.aiChatPaper

SnapBench: 모바일 상호작용을 위한 스냅-앤-애스크 멀티모달 검색의 벤치마킹

SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

August 30, 2026
저자: Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang
cs.AI

초록

모바일 AI는 시각적 오라클 역할을 하여, 사용자가 사물을 촬영하고 그에 관한 정보를 질의할 수 있게 해준다. 촬영-질의(snap-and-ask) 검색은 이제 모바일 AI의 가장 보편적인 진입점 중 하나가 되었지만, 촬영된 사진은 흔히 흐릿하고 텍스트 질문은 짧거나 오타가 포함되기 쉽다. 기존 벤치마크는 손상이 없는 깨끗한 입력만을 테스트하거나, snap-and-ask 검색에서 입력 쌍의 강건성(paired robustness)을 분리하여 검증하지 못한다. 따라서 우리는 강건한 snap-and-ask 멀티모달 검색을 위한 최초의 쌍(pair) 기반 벤치마크인 SnapBench를 제안한다. SnapBench는 53가지 통제된 손상 조건과 1,145개의 질의 및 9,085개의 갤러리 항목으로 구성되며, 인간 주석을 포함한다. 우리는 듀얼 타워 인코더와 임베딩 기반 VLM(비전-언어 모델)을 포함한 16개의 멀티모달 검색기를 평가했다. 실험 결과, 이미지 손상은 검색 성능을 크게 저하시키는 반면, 텍스트 손상은 주로 텍스트 전용 검색에만 영향을 미치고 결합 검색에는 제한적인 영향만 미치는 것으로 나타났다. 손상이 없는 이미지 전용 검색이 결합 검색보다 더 나은 성능을 보이는 경우도 많았는데, 이는 조악한(coarse) 텍스트가 성능을 끌어내리는 문제와 잡음 입력 상황에서 교차 모달 폴백이 부재함을 시사한다. SnapBench는 snap-and-ask 시나리오에서 강건한 검색을 평가하기 위한 통제된 테스트베드를 제공한다. 나아가 우리는 간단한 적응형 융합 기법인 MOOR(Modality-anchored, Outlier-aware, Optimal Reweighting)를 제안하며, snap-and-ask 검색에서 신뢰도를 고려한 모달리티 보정의 필요성을 강조한다.
English
Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.