聚集而非接納:注意力如何將潛在變數帶入可言語化的形式
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
August 15, 2026
作者: Parsa Mazaheri
cs.AI
摘要
語言模型以一種它們可報告的形式持有潛在量;當任務需要靈活重用某個量時,更多的該量以這種形式存在。導致表徵進入該形式的原因仍無定論,而「工作空間」一詞引出一個准入的敘事:一個決定何者進入的閘門。我們在開放權重模型上,利用雅可比透鏡,在一個五個分支共享相同上下文的基準上檢驗這個假設,結果在它預測有閘門之處並未發現閘門。需求會將一個概念的透鏡可見性提升至超過對所提供值施加算子所產生的可見性:在我們的主要檢查點上,百分位排名提升 +0.050 [+0.045, +0.057],且在我們測量的全部四個檢查點上均為正,儘管該分支的回答已達上限,而在該讀出下,準確度匹配的對照更強。同時,一個共享的線性映射可從每個分支(包括對照組)解碼該變量,其水平為選擇校正後下限的 6.4–9.0 倍。在被查詢位置產生後續可讀形式的,是中等深度窗口內由注意力中介的匯聚:將補丁深度與讀出深度分離後,在非飽和讀出下,該處的傳輸至少比任何更淺處高出 17 倍,而窗口內所有受測的多層感知機輸出均無正向貢獻。在飽和的百分位排名下,同一網格無法定位該窗口;這是該度量本身的事實。一個完全不需要該變量的分支,其匯聚量低至七分之一,因此該窗口具有需求特異性。該窗口有兩個經測量的邊界:下方是存活失敗,上方是破壞;在一個 64 層混合模型與來自另一家族的 62 層稠密模型中,它落在相同的相對深度。我們定位的是變量被安裝與被讀取的位置,而不是來自段落的路徑——該路徑不傳輸任何東西。但該讀出並非校準的使用度量:三個成分的讀出值彼此相差在 12% 以內,但它們對答案的實際作用卻相差 7.4 倍。
English
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.