ChatPaper.aiChatPaper

詢問或回答:多輪健康錯誤資訊介入之決策框架

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

August 22, 2026
作者: Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
cs.AI

摘要

修正健康錯誤資訊的對話不僅需要提出事實反駁:使用者所知、所信以及需要聽到的內容各不相同,因此有效的介入往往取決於先提出正確的澄清問題。然而,現有方法不是立即回應,就是不加以區別地探詢,要嘛認為澄清沒有必要,要嘛認為它永遠有益。我們提出「獎勵最佳化的探測與回應」(RO-PnR)框架,用以學習何時提問值得付出其成本。在每個回合,RO-PnR 會根據一個回合級獎勵,在探詢更多資訊與決定最終修正之間做出選擇;該獎勵權衡了探測的預期收益與其互動成本。為了捕捉使用者異質性如何影響探測價值,我們為每個模擬使用者建立一個包含健康素養與信念承諾的潛在狀態模型。實驗顯示,RO-PnR 在三個健康錯誤資訊資料集與三種基礎模型上達到最高的成本調整後效用,且比「總是探詢」的基線方法少用 30% 的回合。
English
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.