集められるものであり、受け入れられるものではない:注意がどのようにして潜在変数を言語化可能な形態にもたらすのか
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
August 15, 2026
著者: Parsa Mazaheri
cs.AI
要旨
言語モデルは、報告可能な形で潜在量を保持しており、タスクがその量を柔軟に再利用することを要求する場合には、その形により多くの量が存在する。表現がその形に入る原因は未解明であり、「ワークスペース」という言葉は、何が入るかを決めるゲートという入場許可の物語を誘う。ヤコビアンレンズを用いて、五つのアームが同一の文脈を共有するベンチマーク上でオープンウェイトモデルをテストしたところ、ゲートが予測されるような場所にはゲートは見つからなかった。要求は、供給された値に演算子を適用した結果生じるものを超えて、概念のレンズ可視性を高める。主要チェックポイントでは百分位順位で+0.050 [+0.045, +0.057]であり、測定した四つすべてで正である。ただし、そのアームは天井効果を示し、精度を一致させた対照はその読み出しの下でより強い。同時に、一つの共有線形写像が、対照条件を含むすべてのアームから変数を、選択補正済み下限の6.4〜9.0倍で復号する。クエリされた位置における後期の読み出し可能な形を生み出すのは、中間深さのウィンドウ内での注意を介した収集である。パッチ深さと読み出し深さを分離すると、非飽和読み出しの下では、そこでの輸送がより浅い任意の場所よりも少なくとも17倍高くなるが、その内部でテストされたMLP出力のいずれも正に寄与しない。飽和する百分位順位の下では、同じグリッドはウィンドウを局在化しない。これはその測度についての事実である。変数を何にも必要としないアームでは、集中は7分の1であるため、ウィンドウは要求特異的である。そのウィンドウは、下側が生存の失敗、上側が破壊という二つの測定された縁を持ち、64層ハイブリッドモデルと別ファミリーの62層稠密モデルにおいて同じ相対深さに位置する。我々は変数がインストールされ読み出される場所を特定するのであって、パッセージからのルートではない。そのルートは何も輸送しない。しかし、読み出しは使用の較正された尺度ではない。三つの構成要素はそれを互いの12%以内に移動させ、答えに対する効果は7.4倍異なる。
English
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.