SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions

Zirong Chen, Fuda Ye, Kuan Zhang, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Jin Ma, Yongqi Zhang

arXiv:2608.29607 · 2026-09-03 공개 · arXiv · PDF

multimodal-retrieval vlms mobile-ai snap-and-ask image-corruption text-corruption dual-tower-encoders moor

Abstract

Mobile AI acts as a visual oracle, empowering users to snap a picture of something and ask for information. Snap-and-ask retrieval is now one of the most common entry points for mobile AI, yet photos are often blurry, while text questions may be short or mistyped. Existing benchmarks only test on clean inputs or do not isolate paired robustness in snap-and-ask retrieval. Therefore, we introduce SnapBench, the first paired benchmark for robust snap-and-ask multimodal retrieval, spanning 1,145 queries, 9,085 gallery items under 53 controlled corruption conditions with human annotations. We evaluate 16 multimodal retrievers, covering dual-tower encoders and embedding-based VLMs. Results show that image corruptions substantially degrade retrieval, while text corruptions mainly affect text-only retrieval and have limited impact on joint retrieval. Clean image-only retrieval often outperforms joint retrieval, indicating the coarse-text drag and the lack of cross-modal fallback under noisy inputs. SnapBench provides a controlled testbed for evaluating robust retrieval in snap-and-ask scenarios. We further propose MOOR (Modality-anchored, Outlier-aware, Optimal Reweighting), a simple adaptive fusion approach, highlighting the need for reliability-aware modality calibration in snap-and-ask retrieval.

한국어 요약

한 줄 요약

SnapBench는 모바일 환경에서 사진과 질문을 결합한 다중모달 검색의 안정성을 평가하는 최초의 페어형 벤치마크이다.

핵심 기여도

핵심 아이디어

Snap-and-ask 검색은 사용자가 촬영한 사진과 짧은 질문을 기반으로 정보를 검색하는 모바일 상호작용이다. 기존 벤치마크는 깨끗한 입력만 테스트하거나, 페어형 안정성 평가를 고려하지 않았다. 따라서 저자들은 SnapBench를 제안하여, 동일한 검색 태스크에서 깨끗한 입력과 오염된 입력을 비교할 수 있는 구조를 설계했다. 이는 모델이 어떤 오염 조건에서 실패하는지 정확히 분석할 수 있게 한다. 또한, 질문이 너무 일반적일 경우(예: "What flower is this?") 여러 유사 항목의 점수가 높아지는 "coarse-text drag" 현상을 발견했으며, 이는 텍스트가 시각 신호보다 과도하게 가중치를 받기 때문으로 분석되었다.

기술적 접근법

주요 결과

의의 및 한계

SnapBench는 모바일 환경에서의 실제 오염 조건을 반영한 체계적인 평가 기반을 제공하며, 다중모달 검색 시 신뢰도 기반 모달 보정의 중요성을 강조한다. 그러나 모델 아키텍처 자체를 수정하지 않고 퓨전 방식만 변경한 MOOR은 근본적인 문제 해결은 아님. 또한, SnapBench는 특정 유형의 오염 조건만 포함하므로, 실제 모바일 환경의 모든 오염을 포괄하지는 못한다.

실용적 활용

SnapBench는 모바일 검색, AR/VR, 쇼핑 추천 등에서 사용자 촬영 사진과 짧은 질문을 기반으로 정보를 추출하는 시스템 개발에 활용 가능하다. MOOR은 모바일 환경에서 텍스트와 이미지 신호의 신뢰도를 동적으로 조정하는 데 유용한 기법으로, 실제 제품에 적용할 수 있는 가벼운 보완 솔루션이다.