모아진 것이지 허용된 것이 아니다: 주의가 어떻게 잠재 변수를 언어화 가능한 형태로 이끄는가
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
August 15, 2026
저자: Parsa Mazaheri
cs.AI
초록
언어 모델은 보고할 수 있는 형태로 잠재 수량을 보유하며, 작업이 그 수량을 유연하게 재사용해야 할 때 그 형태에 더 많은 수량이 존재한다. 표현이 그 형태로 들어가게 하는 원인은 미해결 상태이며, 작업 공간(workspace)이라는 단어는 무엇이 들어갈지를 결정하는 게이트라는 진입 이야기를 암시한다. 다섯 개의 갈래가 동일한 문맥을 공유하는 벤치마크에 대해 오픈 가중치 모델에 Jacobian 렌즈를 적용해 테스트한 결과, 게이트가 예측되는 위치에서는 게이트를 발견하지 못했다. 수요는 공급된 값에 연산자를 적용함으로써 생성되는 것 이상으로 개념의 렌즈 가시성을 높인다. 기본 체크포인트에서 백분위 순위로 +0.050 [+0.045, +0.057]이며, 측정한 네 가지 모두에서 양성을 보인다. 다만 그 갈래는 상한에서 답하며, 정확도가 일치된 대비는 그 판독에서 더 강하다. 동시에 하나의 공유 선형 맵이 대조군을 포함한 모든 갈래에서 변수를 선택 보정 하한의 6.4–9.0배 수준으로 디코딩한다. 쿼리된 위치에서 이후 읽을 수 있는 형태를 생성하는 것은 중간 깊이 창 내부의 어텐션 매개 수집이다. 패치 깊이와 판독 깊이를 분리하면 비포화 판독에서 전달이 그 창에서 더 얕은 어느 위치보다 최소 17배 이상이며, 그 창 내부에서 테스트된 MLP 출력 중 긍정적으로 기여하는 것은 없다. 포화 백분위 순위에서는 동일한 격자가 그 창을 국소화하지 않는데, 이는 그 측정 방식에 관한 사실이다. 변수를 전혀 필요로 하지 않는 갈래는 7배 덜 집중시키므로 그 창은 수요 특이적이다. 그 창은 두 개의 측정된 경계를 가지며, 아래는 생존 실패, 위는 파괴로 나타난다. 또한 64층 하이브리드 모델과 다른 계열의 62층 밀집 모델에서 동일한 상대 깊이에 위치한다. 우리는 변수가 설치되고 읽히는 위치를 국소화하며, 통로로부터의 경로는 아무것도 운반하지 않으므로 국소화하지 않는다. 그러나 그 판독은 사용에 대한 보정된 측정은 아니다. 세 구성 요소가 판독값을 서로의 12% 이내로 이동시키지만, 답에 미치는 영향은 7.4배 차이가 난다.
English
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that decides what gets in. Testing it on open-weight models with Jacobian lenses, over a benchmark whose five arms share an identical context, we find no gate where it predicts one. Demand raises a concept's lens visibility beyond what applying an operator to a supplied value produces: +0.050 [+0.045, +0.057] in percentile rank on our primary checkpoint, positive on all four we measure, though that arm answers at ceiling and the accuracymatched contrast is stronger under that readout. At the same time one shared linear map decodes the variable from every arm, the control included, at 6.4-9.0x its selection-corrected floor. What produces the later readable form at the queried position is attention-mediated gathering inside a mid-depth window: separating patch depth from readout depth puts transport there at least 17x above anywhere shallower under non-saturating readouts, with no tested MLP output contributing positively inside it. Under the saturating percentile rank the same grid does not localise the window, which is a fact about that measure. An arm that needs the variable for nothing concentrates sevenfold less, so the window is demand-specific. That window has two measured edges, a survival failure below and destruction above, and it falls at the same fractional depth in a 64-layer hybrid and a 62-layer dense model from another family. We localise where the variable is installed and read, not the route from the passage, which transports nothing. But the readout is not a calibrated measure of use: three components move it to within 12% of one another and differ 7.4x in what they do to the answer.