ENTRAP-VL:一种用于视觉-语言模型中双重上下文诱导的分类学探针
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
July 22, 2026
作者: Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
cs.AI
摘要
语境牵制是指模型倾向于让输入中的辅助性语境引导其输出,无论该语境是否相关、真实甚至有意义。近期,这一现象已在单模态语言模型中被识别并获得了机制性解释。然而,在视觉-语言模型(VLM)中,该现象是否以及如何表现尚缺乏系统研究,且该领域缺少专门构建的用于研究此现象的工具。我们认为,研究VLM中的语境牵制不能仅将现有的纯文本基准迁移至多模态场景:它需要一套基于分类学结构、具备双模态特性的工具,且测试条件需围绕当前项目(文本流中的描绘图像、视觉流中的文本查询)构建。我们主张,转向VLM的研究是本质性的而非渐进性的。这使得牵制成为一种双重现象,可分别由文本语境和视觉语境独立驱动,并开辟了一个真实性区分维度(即对描绘场景为假但在现实中可能成立的语境),这在以往仅依赖世界知识的单模态研究中没有对应物。为具体化并实现这一主张,我们引入了ENTRAP-VL(面向视觉与语言的牵制评估探测工具),这是一个手工整理的数据集,包含8个类别共1500个项目,按照涵盖双轴(即语境与项目的关联性及其与真实性之间的关系)的分类体系组织,并划分为文本牵制流(8种语境条件)和视觉牵制流(3种语境条件)。我们并非宣称能测量特定模型中的牵制现象;而是提供该工具、支撑其设计的分类学体系以及该工具所支持的评估协议,以便学界能够严谨地研究这一现象。我们将公开发布该数据集及其文档。
English
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.