vision-language multimodal dataset evaluation-protocol taxonomy contextual-entrainment dual-modality textual-entrainment
Abstract
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.
한국어 요약
한 줄 요약
ENTRAP-VL은 시각-언어 모델에서 문맥에 의한 출력 편향을 체계적으로 평가하기 위한 이중 모달성 탐사 도구이다.
핵심 기여도
- ENTRAP-VL은 1,500개의 수작업 데이터로 구성된, 8개 범주에 걸친 이중 문맥 조건 평가 데이터셋을 제시.
- 이중 문맥 조건: 텍스트-문맥 스트림(8개 조건), 시각-문맥 스트림(3개 조건)으로 구분.
- 기존 텍스트 중심의 'contextual entrainment' 개념을 시각-언어 모델로 확장하며, 'scene-relative veracity'라는 새로운 구분을 도입.
- 모델별 특정 성능 측정은 주장하지 않고, 평가 프로토콜과 세분화된 세분화 분류 체계를 제공.
핵심 아이디어
기존 언어 모델에서 관찰된 'contextual entrainment'는 모델이 입력 문맥에 따라 출력이 편향되는 현상으로, 이는 단순히 무관한 문맥의 존재에 의한 것이 아니라, 문맥 내 토큰의 반복에 기반한 확률적 경향성이다. 본 연구는 이러한 현상을 시각-언어 모델(VLM)로 확장할 때, 단일 모달성에서의 개념이 부족하다고 주장한다. VLM에서는 텍스트와 시각적 문맥이 독립적으로 편향을 유발할 수 있으며, 문맥이 실제 장면과 일치하지 않더라도 세계 지식 내에서 가능한 경우가 존재한다. 이를 'scene-relative veracity'로 명명하며, 이는 기존 텍스트 중심 연구에서는 다루어지지 않았던 새로운 차원이다.
기술적 접근법
- **ENTRAP-VL 데이터셋**: 1,500개의 수작업 데이터, 8개 범주로 구성, 두 축(문맥과 항목의 연관성, 진위 관계)으로 분류.
- **이중 스트림 구조**: 텍스트-문맥 스트림(8개 조건), 시각-문맥 스트림(3개 조건)으로 분리.
- **평가 프로토콜**: 각 조건별로 모델의 편향 경향성을 측정할 수 있도록 설계.
- **세분화된 세분화 분류**: 문맥의 진위, 연관성, 그리고 모달성(텍스트/시각)을 기준으로 구조화.
주요 결과
- ENTRAP-VL은 특정 모델의 성능을 측정하지 않으며, 평가 프로토콜과 세분화 분류 체계를 제공.
- 텍스트-문맥 스트림(8개 조건)과 시각-문맥 스트림(3개 조건)을 통해 모델의 이중 편향 경향성을 분석할 수 있음.
- 기존 텍스트 중심의 'contextual entrainment' 개념을 VLM으로 확장하며, 'scene-relative veracity'라는 새로운 구분을 도입.
의의 및 한계
ENTRAP-VL은 VLM에서 문맥에 의한 출력 편향을 체계적으로 평가할 수 있는 첫 번째 이중 모달성 도구로, 학술적으로는 VLM의 내재적 편향 메커니즘을 이해하는 데 기여한다. 실용적으로는 RAG(VQA)와 같은 시스템에서 부정확한 문맥이 모델 예측에 미치는 영향을 분석하는 데 활용될 수 있다. 그러나 본 연구는 특정 모델의 성능을 측정하지 않으며, 평가 프로토콜만 제공하므로, 모델별 실험 결과는 추후 연구에 맡겨야 한다는 한계가 있다.
실용적 활용
ENTRAP-VL은 시각-언어 모델이 사용되는 RAG(VQA), 이미지 기반 QA 시스템, 멀티모달 정보 추출 등에서 모델의 문맥 편향성을 평가하는 데 활용될 수 있다. 특히, 부정확하거나 유사한 문맥이 모델 예측에 미치는 영향을 분석하여 시스템의 신뢰성을 높이는 데 기여할 수 있다.