ChatPaper.aiChatPaper

推理模型的思維鏈忠實度會因偏好提示的傳遞位置與方式而異

Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered

August 29, 2026
作者: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
cs.AI

摘要

思維鏈(CoT)監測假設推理軌跡能忠實記錄塑造模型答案的資訊。現有的忠實性測試常在用戶訊息中放置明確的偏見線索,而智能體可能透過工具返回或原始產物接觸到偏好。我們引入FACE-Eval(線索效應的忠實歸因評估),這是一個包含5,100個樣本的評估基準,其操縱線索位置(用戶訊息或工具返回)與明確性(直接摘要或原始產物)。我們衡量遵循線索之回答中的口頭化承諾,以及所有含線索樣本中的非口頭化採納。我們評估來自八個家族的15個開放權重模型,總參數量介於4B至1.60T之間。所有模型在工具返回線索下的口頭化承諾均低於用戶訊息線索,且在隱式線索下低於明確線索。非口頭化採納在全部15個模型上對工具返回線索較高,在30個模型-通道比較中的28個中對隱式線索較高。來源歸因提示在七個模型上縮小了通道差距,有時是透過增加用戶通道的非口頭化採納來實現;而告知模型其推理將受到監測則無法可靠地縮小這一差距。我們亦使用兩個轉錄監測器(GPT-5.6-Luna與GPT-4o-mini)來檢測每個家族中最大模型的偏好採納。在32個模型-通道-明確性單元中,較高的非口頭化採納與兩個監測器較低的檢測能力相關(分別為Pearson r=−0.54與r=−0.78)。這些結果表明,在本研究所測試的單次調用、預填充工具設置中,當偏好資訊經由工具到達或必須從原始產物中推斷時,思維鏈監測的可靠性可能較低。
English
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.