ChatPaper.aiChatPaper

「活性化オラクルが読まないことを学ぶとき:ファインチューニングされたオラクルにおける概念固有のブラインドスポット」

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

July 25, 2026
著者: Tobias Bersia, Tatiana Gaintseva
cs.AI

要旨

活性化オラクル(Activation Oracles; AOs)は、別のモデルの内部活性化に関する自然言語の質問に答えるように訓練された言語モデルである。これらは、モデルの状態から隠れた情報を読み取るための柔軟なインターフェースを提供する。特に、関連情報が内部では表現されているものの、表出行動には存在しない、または不完全である場合に有用である。しかしながら、AOはそれ自体が学習されたシステムであり、その回答は、表現された情報の中立的な読み出しではなく、訓練データ、目的関数、そして学習された報告行動によって形成される。我々はこれを、対象モデルが直接的な開示を避けつつ内部で隠れた概念を用いるようにファインチューニングされる、統制されたTaboo Word Guessing(タブーワード推測)設定において研究する。このような対象モデルに対して訓練されたAOが専門的な読み取り器になるとの期待に反し、ファインチューニングされたAOは概念特異的な反読み取り器(anti-reader)になり得ることが分かった。すなわち、それらは自身の訓練中に持続的に存在する概念を選択的に回復できない。この失敗は、対象モデルやオラクルの表現から概念が欠落していることだけでは説明できない。すなわち、対象概念はオラクル内部で依然として復号可能であり、LogitLensおよび層アブレーション分析は、この失敗がAOの読み出し経路において生じることを示している。我々の結果は、行動的漏洩、表現レベルの復号可能性、およびAO-verbalizability(AOによる言語化可能性)が互いに乖離し得ることを示しており、学習された解釈可能性インターフェースに対する信頼性の懸念を提起する。
English
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.