推論モデルのChain-of-Thought忠実性は、選好手がかりの提示位置と提示方法によって変化する
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
August 29, 2026
著者: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
cs.AI
要旨
Chain-of-thought(CoT)モニタリングは、推論トレースがモデルの回答を形成する情報を忠実に記録していると仮定する。既存の忠実性テストは多くの場合、明示的なバイアス手がかりをユーザーメッセージに配置するが、エージェントはツールの戻り値や生のアーティファクトを通じて選好に遭遇することがある。我々はFACE-Eval(Faithful Attribution of Cue Effects Evaluation:手がかり効果の忠実な帰属評価)を導入する。これは5,100サンプルの評価であり、手がかりの位置(ユーザーメッセージかツールの戻り値)と明示性(直接的要約か生のアーティファクト)を変化させる。我々は、手がかりに従う回答における言語化されたコミットメントと、すべての手がかり付きサンプルにおける言語化されない採用を測定する。我々は、8つのファミリーからの15のオープンウェイトモデルを評価する。総パラメータ数は4Bから1.60Tの範囲である。すべてのモデルにおいて、ユーザーメッセージの手がかりよりもツールの戻り値の手がかりに対する言語化されたコミットメントが低く、明示的な手がかりよりも暗黙的な手がかりに対する言語化されたコミットメントが低い。言語化されない採用は、15モデルすべてでツールの戻り値の手がかりに対して高く、30のモデル・チャネル比較のうち28で暗黙的な手がかりに対して高い。ソース帰属プロンプトは、7つのモデルでチャネルギャップを縮小する。これは、ユーザーチャネルの言語化されない採用を増やすことによる場合もあるが、モデルに推論が監視されると伝えることは、そのギャップを確実に埋めるわけではない。また、各ファミリーの最大モデルにおける選好採用を検出するために、2つのトランスクリプトモニタ(GPT-5.6-LunaとGPT-4o-mini)を使用する。32のモデル・チャネル・明示性セルにわたって、言語化されない採用が高いほど、両モニタの検出能力が低いことと関連している(ピアソンのr=-0.54およびr=-0.78)。これらの結果は、ここでテストされた単一呼び出し・プリフィルドツール設定の範囲において、選好情報がツールを通じて届く場合や生のアーティファクトから推測されなければならない場合に、CoTモニタリングの信頼性が低下する可能性があることを示唆している。
English
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.