ChatPaper.aiChatPaper

ENTRAP-VL:視覺語言模型中雙重語境誘發的類別探針

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

July 22, 2026
作者: Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
cs.AI

摘要

上下文引導(Contextual entrainment)指的是模型傾向於讓輸入中的輔助性上下文影響其輸出,而無關該上下文是否相關、真實甚至有意義。近期研究已在單模態語言模型中識別此現象並提出機制性解釋。然而,此現象在視覺語言模型(VLM)中是否及如何表現,在很大程度上仍未經檢驗,且該領域缺乏專為研究此現象設計的工具。我們認為,研究VLM中的上下文引導不僅需要將現有純文字基準移植到多模態環境中,更需建立一個基於分類結構的雙模態工具,其實驗條件需圍繞當前項目(文字串流中的描繪圖像、視覺串流中的文字查詢)構建。我們主張,過渡到VLM是實質性而非漸進性的進展——這使得引導成為雙重現象,可分別由文字上下文和視覺上下文獨立驅動,並開啟了前人所研究的單模態(僅基於世界知識)中不存在的真實性區分(即上下文對描繪場景為虛假但世界可能存在的情況)。為將此立場具體化且具可操作性,我們提出ENTRAP-VL(視覺與語言之引導評估探針),這是一個由人工校訂的資料集,包含1,500個項目,分屬八個類別,並依據橫跨兩軸的分類體系組織——即上下文與項目的關聯性及其與真實性的關係——再拆分為文字引導串流(八種上下文條件)與視覺引導串流(三種上下文條件)。我們不宣稱能測量任何特定模型的引導程度,而是提供這一工具、其背後的分類學依據以及所支援的評估協議,使學界能嚴謹地探究此現象。我們將公開釋出該資料集及其相關文件。
English
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.