일찍 인코딩되고 늦게 사용됨: 트랜스포머가 추론된 파트너의 전문성에 따라 행동하기 시작하는 지점
Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
September 7, 2026
저자: Mika Okamoto, Gabriele Sarti
cs.AI
초록
트랜스포머는 어떤 속성이 아직 출력에 영향을 미치지 않는 깊이에서 그 속성을 잔차 스트림 안에서 선형적으로 디코딩 가능하게 만들 수 있다. 정보를 읽을 수 있는 위치와 정보가 사용되는 위치 사이의 이러한 간극은 입력에 직접 명시된 속성들에 대해 입증되어 왔다. 우리는 이 간극이 대화 전반에 걸쳐 모델이 점진적으로 추론해야 하는 속성, 즉 대화 상대가 얼마나 전문가인지에 대해서도 성립하는지 묻는다. 네 가지 전문성 수준에서 모델이 연기한 페르소나들 간의 다중 턴 연구 계획 대화 코퍼스인 ExpertCollab을 사용하여, 우리는 대화 상대의 전문성이 초기 층에서 가장 잘 디코딩되며 네트워크 중간 지점 이전에 거의 우연 수준으로 떨어진다는 것을 발견한다. 반사실적 패칭은 최대 디코딩 가능성 층에 전문성 차이를 주입하는 것이 고정된 후기 층 리드아웃을 거의 변화시키지 않는 반면, 동일한 차이를 중간 지점 이후에 주입하면 거의 완전하게 전파되며, 이는 10배 이상의 분리를 보여준다. 내용을 일치시킨 무작위 통제와 프로브가 필요 없는 진단은 전환을 동일한 초기 층에 위치시키며, 정적으로 명시된 통제 속성은 전반에 걸쳐 디코딩 가능한 상태로 유지된다. 따라서 추론된 관계적 속성은 그것이 인과적으로 활성화되기 훨씬 전에 표상되며, 이는 대화 상대에 조건화된 행동을 읽어내거나 조향하려는 모든 시도가 개입해야 하는 위치를 제한한다. 우리는 합성 코퍼스에 대한 하나의 모델을 초기 시연으로 사용한다.
English
A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network. Counterfactual patching shows that injecting the expertise difference at the layer of peak decodability barely changes a fixed late-layer readout, whereas the same difference injected past the midpoint propagates almost completely, a separation of more than an order of magnitude. A content-matched random control and a probe-free diagnostic place the transition at the same early layer, and a statically specified control attribute stays decodable throughout. An inferred relational attribute is therefore represented well before it becomes causally active, which bounds where any attempt to read out or steer partner-conditioned behavior must intervene. We use one model on a synthetic corpus as an initial demonstration.