ChatPaper.aiChatPaper

當激活預言機學會不閱讀:微調預言機中的概念特定盲點

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

July 25, 2026
作者: Tobias Bersia, Tatiana Gaintseva
cs.AI

摘要

激活預言機(Activation Oracles, AOs)是語言模型,經訓練用來回答關於另一個模型內部激活值的自然語言問題。它們提供了一個靈活的介面,用於從模型狀態讀取隱藏資訊,尤其是在相關資訊已在內部表徵、但在可見行為中缺席或不完整的時候。然而,AOs 本身也是學習系統:它們的答案由訓練資料、目標以及習得的報告行為所形塑,而非對表徵資訊的中性讀出。我們在一個受控的「禁忌詞猜測」(Taboo Word Guessing)情境中研究此問題,其中受試模型經微調以在內部使用一個隱藏概念,同時避免直接揭露。與「在如此受試者上訓練的 AO 會成為專業讀取器」的預期相反,我們發現經微調的 AOs 可能變成概念特異性的反讀取器:它們選擇性地無法恢復在其自身訓練期間持續存在的概念。此失敗並不能簡單地用概念從受試模型或預言機表徵中消失來解釋:目標在預言機內部仍然可解碼,而 LogitLens 與層級消融分析顯示,失敗發生在 AO 的讀出路徑。我們的結果表明,行為洩漏、表徵層級的可解碼性以及 AO 的可語言化能力可能彼此脫節,這對學習式可解釋性介面的可靠性提出了擔憂。
English
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.