ENTRAP-VL: 視覚言語モデルにおける二重文脈的引き込みのための分類学的プローブ
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
July 22, 2026
著者: Karan Goyal, Afreen Hossain, Debojyoti Das, Vishal Bhutani
cs.AI
要旨
文脈的引き込みとは、モデルが入力中の補助的な文脈によって、その文脈が関連性があるか、真実か、あるいは意味があるかどうかにかかわらず、出力が引き寄せられる傾向である。近年、この現象は単一モダリティの言語モデルにおいて特定され、メカニズム的な説明が与えられている。対照的に、視覚言語モデル(VLM)において、それが現れるかどうか、またどのように現れるかはほとんど検討されておらず、その調査に特化した測定器が分野には欠けている。我々は、VLMにおける文脈的引き込みを研究するには、既存のテキストのみのベンチマークをマルチモーダル設定に移植する以上のものが必要であると考える。すなわち、分類学的に構造化された二重モダリティの測定器が必要であり、その条件は対象となる項目(テキストストリーム内の描写画像、ビジュアルストリーム内のテキストクエリ)を中心に構築されるべきである。我々は、VLMへの移行は漸進的なものではなく、実質的なものであると主張する。これにより、引き込みは二重の現象となり、テキスト文脈と視覚文脈によって独立して駆動可能となり、さらに、先行研究の単一モダリティで世界知識のみに基づく定式化には対応するものがない、真実性の区別(描写されたシーンでは偽であるが、世界では可能な文脈)が開かれる。この立場を具体的かつ実用的なものにするために、我々はENTRAP-VL(視覚と言語のための引き込み評価プローブ)を導入する。これは、手作業でキュレーションされた1,500項目からなるデータセットであり、8つのカテゴリにわたって、2つの軸(すなわち、項目に対する文脈の関連性と真実との関係)に沿った分類法で整理され、テキスト引き込みストリーム(8つの文脈条件)と視覚引き込みストリーム(3つの文脈条件)に分割されている。我々は特定のモデルにおける引き込みを測定すると主張するのではなく、測定器、それを動機づける分類法、およびそれが可能にする評価プロトコルを提供し、コミュニティがこの現象を厳密に調査できるようにする。データセットとそのドキュメントは公開する予定である。
English
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual entrainment in VLMs requires more than porting an existing text-only benchmark to the multimodal setting: it requires a taxonomically structured, dual-modality instrument whose conditions are constructed around the item at hand (the depicted image in the textual stream, the textual query in the visual stream). We argue that the move to VLMs is substantive rather than incremental. It makes entrainment a dual phenomenon, drivable independently by textual and by visual context, and it opens a veracity distinction (context that is false of the depicted scene yet possible in the world) that has no counterpart in the unimodal, world-knowledge-only formulation of prior work. To make this position concrete and actionable, we introduce ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes, i.e., the association of context with the item and its relationship to truth, and split into a textual-entrainment stream (eight context conditions) and a visual-entrainment stream (three context conditions). We do not claim to measure entrainment in any particular model; we provide the instrument, the taxonomy that motivates it, and the evaluation protocols it enables, so that the community can investigate the phenomenon rigorously. We will release the dataset and its documentation publicly.