SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Sirun Li, Minghao Liu, Ling Dai, Yong Li, Haoxin Lyu, Junting Zhou, Fan Zhang

arXiv:2608.04244 · 2026-08-06 공개 · arXiv · PDF

multimodal-llms mllm-evaluation scene-text text-vision localization-error dataset-variants conflict-resolution geolocation

Abstract

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.

한국어 요약

한 줄 요약

SIGNPOST-Bench는 텍스트-비전 충돌 해결 능력을 평가하는 5가지 조건의 대규모 벤치마크로, 20개 MLLM 모델을 평가하여 평균 위치 오차를 4.8배 증가시키는 텍스트 영향을 밝혀냈다.

핵심 기여도

핵심 아이디어

기존 벤치마크는 텍스트와 비전 정보가 충돌할 때 모델이 어떤 증거를 신뢰하는지 밝히지 못했다. SIGNPOST-Bench는 **지리적 위치 추정**을 기반으로, 텍스트와 비전 정보 간 충돌을 **연속적인 좌표 공간**에서 평가한다. 이는 위치 오차 감소와 충돌 텍스트에 의한 **목표 지점 향한 예측 이동**을 동시에 측정할 수 있게 한다.
**Synthetic, localized scene-text interventions**를 통해 텍스트만 변경하면서 비텍스트 요소는 보존함으로써, 텍스트의 영향을 분리해 측정한다. 이는 기존의 이산적 평가 방식과 달리, **충돌 텍스트가 예측에 미치는 방향성과 강도**를 정량적으로 분석할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

SIGNPOST-Bench는 텍스트-비전 충돌 해결 능력을 **연속적이고 재현 가능한 방식**으로 평가하는 첫 번째 벤치마크로, 모델이 충돌 텍스트를 **사용, 거부, 또는 따르는지**를 정량적으로 분석할 수 있다. 이는 MLLM의 신뢰성과 내성을 평가하는 데 중요한 도구가 될 수 있다.
하지만, **모델 내부의 결정 과정**이나 **추론 시 가중치 변화**는 측정하지 못하며, **추가적인 시각적 요소**나 **다국어 텍스트**에 대한 평가도 필요하다.

실용적 활용

SIGNPOST-Bench는 **지도, 로봇 네비게이션, 자율주행** 등에서 텍스트-비전 충돌이 발생하는 상황에서 모델 신뢰도를 평가하는 데 활용 가능하다. 또한, **MLLM의 내성 향상**을 위한 모델 개선 및 프롬프트 최적화 연구에 기초 자료로 사용될 수 있다.