희소 판독 프리즘: 토큰 대신 특성으로 로짓 렌즈 점수 설명하기
Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
September 1, 2026
저자: Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane
cs.AI
초록
언어 모델에서 다음 토큰 예측은 레이어를 거치며 발달한다. 렌즈 방법은 중간 은닉 상태를 토큰으로 디코딩함으로써 이 과정을 추적한다. 그러나 렌즈 판독은 은닉 상태와 이를 디코딩하는 데 사용되는 리드아웃(언임베딩 행렬)을 모두 반영한다. 많은 렌즈가 코퍼스에 피팅된다. 우리는 피팅 코퍼스만 다른 두 렌즈가 동일한 은닉 상태에 대해 서로 다른 토큰을 보고할 수 있음을 보이고, 이러한 의존성을 코퍼스 조건성(corpus conditionality)이라 부른다. 피팅 코퍼스와 무관하게 리드아웃 구조를 조사하기 위해, 희소 리드아웃 프리즘(Sparse Readout Prism, SRP)을 제안한다. SRP는 리드아웃의 가중치만을 사용해 리드아웃을 분해하고, 임의의 토큰 로짓 또는 로짓 차이를 희소 리드아웃 특성들의 기여도의 합으로 표현한다. 이를 통해 리드아웃 특성은 렌즈 판독의 새로운 분석 단위로 부각되며, 토큰 정체성에 가려질 수 있는 구조를 드러내고, 토큰·맥락·레이어·렌즈 간 비교를 가능하게 한다. 원래 리드아웃을 SRP의 희소 근사로 대체할 경우, 테스트된 로짓 차이는 리드아웃 행 간의 기하학적 관계에 기반해 구축된 여섯 가지 베이스라인 중 성능이 가장 좋은 베이스라인보다 8.9~17.3퍼센트 포인트 더 많이 재구성된다. 특성 절제 시 로짓 차이는 해당 특성의 SRP 기여도에 비례하여 이동한다. 판독되는 토큰은 피팅 코퍼스에 따라 달라지지만, 지배적인 리드아웃 특성은 안정적으로 유지된다. SRP는 구축 과정에서 코퍼스를 사용하지 않으므로, 렌즈 분석에 있어 피팅 코퍼스와 독립적인 통제 수단을 제공한다.
English
A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features. This reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses. Replacing the original readout with SRP's sparse approximation reconstructs 8.9-17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows. Ablating features shifts logit differences in proportion to their SRP contributions. Although token readings vary with the fitting corpus, the dominant readout feature remains stable. Because SRP uses no corpus in its construction, it provides a control independent of the fitting corpus for lens analyses.