ChatPaper.aiChatPaper

ENTRAP-VL: 시각-언어 모델에서 이중 맥락적 동조를 위한 분류학적 프로브

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

July 22, 2026
저자: Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
cs.AI

초록

맥락적 동조화(contextual entrainment)는 모델이 입력의 보조 맥락이 해당 맥락의 관련성, 진실성, 또는 의미성과 무관하게 출력을 유도하는 경향을 의미한다. 최근 단일 양식 언어 모델에서 이 현상이 확인되고 메커니즘적으로 설명되었으나, 시각-언어 모델(VLM)에서의 발현 여부와 방식은 대체로 검토되지 않았으며, 현재 이 분야에는 이를 조사하기 위한 전용 도구가 부족하다. 본 연구는 VLM에서의 맥락적 동조화를 연구하는 것이 기존의 텍스트 전용 벤치마크를 다중 양식 환경으로 이식하는 것 이상을 요구한다는 입장을 취한다. 즉, 해당 항목(텍스트 스트림 내 묘사된 이미지, 시각 스트림 내 텍스트 질의)을 중심으로 구성된 조건들을 포함하는, 분류학적으로 구조화된 이중 양식 도구가 필요하다. 우리는 VLM으로의 전환이 점진적이 아닌 실질적이라고 주장한다. 이 전환은 동조화를 이중 현상으로 만들며, 텍스트 맥락과 시각 맥락에 의해 독립적으로 유도될 수 있게 하고, 이전 연구의 단일 양식 세계 지식 기반 공식화에는 존재하지 않는 진위 구분(묘사된 장면에 대해서는 거짓이지만 현실 세계에서는 가능한 맥락)을 열어준다. 이 입장을 구체화하고 실행 가능하게 만들기 위해, 본 연구는 ENTRAP-VL(시각-언어용 동조화 평가 프로브)을 소개한다. 이는 두 축(맥락과 항목 간의 연관성 및 진리와의 관계)으로 구성된 분류 체계에 따라 정리된 8개 범주의 1,500개 항목을 수동으로 선별한 데이터셋이며, 텍스트 동조화 스트림(8가지 맥락 조건)과 시각 동조화 스트림(3가지 맥락 조건)으로 나뉜다. 우리는 특정 모델에서의 동조화를 측정한다고 주장하지 않는다. 대신, 커뮤니티가 이 현상을 엄격히 조사할 수 있도록 도구와 이를 뒷받침하는 분류 체계, 그리고 이를 통해 가능해진 평가 프로토콜을 제공한다. 데이터셋과 관련 문서는 공개할 예정이다.
English
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.