ChatPaper.aiChatPaper

提问还是回答:多轮健康错误信息干预的决策框架

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

August 22, 2026
作者: Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
cs.AI

摘要

在对话中纠正健康错误信息,需要的不仅仅是给出事实性反驳:用户所知道的、所相信的以及所需要听到的内容各不相同,因此有效的干预往往取决于首先提出恰当的澄清问题。然而,现有方法要么立即回应,要么不加区分地探询,将澄清视为要么没有必要、要么总是有益的。我们提出了奖励优化的探询-回应框架(RO-PnR),该框架能够学习何时提问是值得其成本的。在每个轮次中,RO-PnR 在进一步探询信息和给出最终纠正之间做选择,并由一个轮次级奖励引导,该奖励权衡探询的预期收益与其交互成本。为了捕捉用户异质性如何影响探询价值,我们为每个模拟用户建模,在健康素养和信念承诺维度上赋予其潜在状态。实验表明,RO-PnR 在三个健康错误信息数据集和三个基础模型上取得了最高的成本调整效用,且与始终探询的基线方法相比,使用的轮次减少了30%。
English
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.