問うか答えるか:マルチターンの健康誤情報介入のための意思決定フレームワーク
Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention
August 22, 2026
著者: Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
cs.AI
要旨
対話における健康誤情報の訂正には、事実に基づく反論を示すだけでは不十分である。ユーザーごとに知っていることや信じていること、聞く必要があることは異なるため、効果的な介入はしばしば、最初に適切な明確化の質問を行うことに依存する。しかし既存手法は、即座に応答するか無差別に質問するかのどちらかであり、明確化を不要なものか常に有益なものかのいずれかとして扱っている。本稿では、質問がコストに見合う時期を学習するフレームワークであるReward-Optimized Probe-and-Respond(RO-PnR)を提案する。RO-PnRは各ターンにおいて、情報を得るための質問と最終的な訂正に踏み切ることのどちらを選ぶかを、質問の期待利得と対話コストを比較考量するターン単位の報酬に基づいて決定する。ユーザーの異質性が質問の価値にどう影響するかを捉えるため、各シミュレーションユーザーを、ヘルスリテラシーと信念の確信度に沿った潜在状態としてモデル化する。実験では、RO-PnRが3つの健康誤情報データセットと3つのベースモデルにわたり最高のコスト調整済み効用を達成し、常に質問するベースラインと比べてターン数を30%削減することを示した。
English
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.