ChatPaper.aiChatPaper

묻기 또는 답하기: 멀티턴 건강 오정보 개입을 위한 의사결정 프레임워크

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

August 22, 2026
저자: Xiaoying Song, Anirban Saha Anik, Jinyu Liu, Qitao Tan, Geng Yuan, Lingzi Hong
cs.AI

초록

대화에서 건강 오정보를 교정하는 일은 사실 기반의 반박을 제시하는 것 이상을 요구한다. 사용자는 아는 내용, 믿는 내용, 그리고 들어야 할 내용이 서로 다르기 때문에, 효과적인 개입은 종종 올바른 명확화 질문을 먼저 던지는 데 달려 있다. 그러나 기존 방법은 즉시 응답하거나 무차별적으로 탐색할 뿐, 명확화를 불필요하거나 항상 유익한 것으로 취급한다. 우리는 질문하는 것이 비용만한 가치가 있는 시점을 학습하는 보상 최적화 탐색-응답(RO-PnR) 프레임워크를 제안한다. 각 턴에서 RO-PnR은 추가 정보 탐색과 최종 교정 확정 사이를 선택하며, 탐색의 기대 이득을 상호작용 비용과 비교하는 턴 단위 보상에 따라 움직인다. 사용자 이질성이 탐색 가치에 미치는 영향을 포착하기 위해, 우리는 각 시뮬레이션 사용자를 건강 문해력과 신념 고착도를 축으로 하는 잠재 상태로 모델링한다. 실험 결과, RO-PnR은 세 가지 건강 오정보 데이터셋과 세 가지 기본 모델 전반에서 가장 높은 비용 조정 효용을 달성하며, 항상 탐색하는 기준선보다 30% 적은 턴을 사용한다.
English
Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.