汇集而非接纳:注意如何使潜变量进入可言语化的形式
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
August 15, 2026
作者: Parsa Mazaheri
cs.AI
摘要
语言模型以可报告的形式承载潜在量;当任务要求灵活重用时,该形式中存在的量会更多。是什么导致表征进入该形式仍属未知,而“工作空间”一词暗示了一种准入式的解释:一道决定何者进入的闸门。我们使用雅可比透镜在开放权重模型上对此加以检验,基准的五个分支共享相同的上下文,结果在它预测有闸门之处并未发现任何闸门。任务需求将概念的透镜可见性提高到对给定值应用算子所产生的水平之上:在我们主要的检查点上,百分位排名提升+0.050 [+0.045, +0.057];在我们测量的全部四个检查点上均为正,尽管该分支的回答已达到上限,而且在该读出下,准确率匹配的对比更强。同时,一个共享的线性映射能从每个分支(包括对照组)解码出该变量,其数值为选择校正后下限的6.4至9.0倍。在查询位置产生后续可读形式的原因,是中间深度窗口内由注意力介导的汇集:将补丁深度与读出深度分离后,在非饱和读出下,该处的传输至少比任何更浅处高17倍,而窗口内没有任何经测试的MLP输出作出正向贡献。在饱和百分位排名下,同一网格无法定位该窗口,这是该度量本身的事实。一个完全不需要该变量的分支,其汇集程度低至七分之一,因此该窗口是需求特定的。该窗口有两个经测量的边界:低于它则存续失败,高于它则被破坏;它在64层混合模型和来自另一个家族的62层稠密模型中,处于相同的相对深度。我们定位的是变量被安装和读取之处,而非来自文本段落的路径——该路径不传输任何内容。但该读出并非使用量的校准度量:三个组件使读数彼此相差在12%以内,但它们对答案的影响却相差7.4倍。
English
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.