推理模型思维链的忠实性随偏好线索的传递位置与方式而变化
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
August 29, 2026
作者: Aryo Pradipta Gema, Neel Rajani, Rohit Saxena, Wai-Chung Kwan, Pasquale Minervini
cs.AI
摘要
思维链(CoT)监控假设推理轨迹忠实记录了塑造模型答案的信息。现有的忠实性测试通常将显性偏见线索放置在用户消息中,而智能体可能通过工具返回或原始工件接收到偏好。我们引入了FACE-Eval(线索效应归因忠实性评估),这是一个包含5,100个样本的评估,系统地变化线索位置(用户消息或工具返回)和明确性(直接摘要或原始工件)。我们测量了线索遵循答案中的显性承诺,以及所有带线索样本中的隐性采纳。我们评估了来自8个系列的15个开放权重模型,参数量从4B到1.60T不等。所有模型在工具返回线索上的显性承诺都低于用户消息线索,在隐式线索上的显性承诺也低于显式线索。在全部15个模型上,工具返回线索的隐性采纳更高;在30个模型-渠道比较中,有28个显示出隐式线索的隐性采纳更高。来源归因提示在7个模型上缩小了渠道差距,有时是通过增加用户渠道的隐性采纳实现的;而告知模型其推理将被监控并不能可靠地缩小这一差距。我们还使用了两个转录监控器(GPT-5.6-Luna和GPT-4o-mini)来检测每个系列中最大模型的偏好采纳。在32个模型-渠道-明确性单元中,较高的隐性采纳与两个监控器较低的检测能力相关(皮尔逊r分别为-0.54和-0.78)。这些结果表明,在本文测试的单次调用、预填充工具设置中,当偏好信息通过工具到达或必须从原始工件中推断时,CoT监控的可靠性可能较低。
English
Chain-of-thought (CoT) monitoring assumes that reasoning traces faithfully record the information that shapes a model's answer. Existing faithfulness tests often place explicit bias cues in the user message, while agents may encounter preferences through tool returns or raw artifacts. We introduce FACE-Eval (Faithful Attribution of Cue Effects Evaluation), a 5,100-sample evaluation that varies cue location (user message or tool return) and explicitness (direct summary or raw artifact). We measure verbalized commitment among cue-following answers and unverbalized adoption among all cued samples. We evaluate 15 open-weight models from eight families, with total parameters ranging from 4B to 1.60T. Every model has lower verbalized commitment for tool-return than user-message cues and for implicit than explicit cues. Unverbalized adoption is higher for tool-return cues on all 15 models and for implicit cues in 28 of 30 model-channel comparisons. A source-attribution prompt narrows the channel gap on seven models, sometimes by increasing user-channel unverbalized adoption, while telling models that their reasoning will be monitored does not reliably close the gap. We also use two transcript monitors (GPT-5.6-Luna and GPT-4o-mini) to detect preference adoption in the largest model of each family. Across 32 model-channel-explicitness cells, higher unverbalized adoption is associated with lower detection ability for both monitors (Pearson r=-0.54 and r=-0.78, respectively). These results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.