ChatPaper.aiChatPaper

의도가 더 크게 말한다: 응답 모방을 넘어서는 제어 가능한 사용자 시뮬레이션

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

August 10, 2026
저자: Bo Wang, Ruixing Zhang, Yunqi Liu, Yang Zhang, Liangzhe Han, Tongyu Zhu, Leilei Sun
cs.AI

초록

사용자 시뮬레이터는 대화형 어시스턴트를 훈련하고 평가하기 위한 확장 가능한 환경으로 널리 사용된다. 다음 사용자 발화를 생성하는 것은 본질적으로 일대다(one-to-many) 문제이다. 동일한 프로필과 대화 맥락은 서로 다른 국소적 상호작용 의도를 가진 여러 그럴듯한 후속 발화를 지원할 수 있다. 따라서 유창한 응답이 수리(repair)가 아닌 수용(acceptance)과 같은 부적절한 의도로 대화를 진행시킬 수 있다. 우리의 핵심 통찰은 제어 가능한 사용자 시뮬레이션이 다음 사용자 발화가 실현해야 할 국소적 상호작용 의도(무엇을)와 그 의도가 언어로 표현되는 방식(어떻게)을 분리해야 한다는 것이다. 우리는 상호작용 의도를 각 턴의 명시적 지시문으로 노출하는 UserIDA(사용자 의도-지시 정렬, User Intent-Directive Alignment)를 제안한다. UserIDA는 여섯 가지 의도 인터페이스를 정의하고, 지도 미세조정을 통해 지시문 조건부 생성을 학습하며, 그룹 기반 강화학습에서 의도 보정 정책 최적화를 사용한다. 보상 함수는 혼합 그룹에서 의도를 위반한 후보가 의도를 준수한 대안보다 낮은 순위를 차지하도록 보장하면서 복합 응답 품질을 유지한다. LMSYS-USP에서 UserIDA는 86.6%의 의도 정확도를 달성하여 가장 강력한 전용 사용자 시뮬레이터 기준선보다 24.3% 포인트 높은 성능을 보이며, 의미적·문체적 유사도도 개선한다. 맥락 내 개입 실험에서 평가된 대화 상태의 91.7%에서 여섯 가지 목표 의도 중 최소 네 가지를 실현하며, 이는 가장 강력한 외부 기준선의 22.9%와 대비된다. 이러한 결과는 턴 단위 의도 제어가 사용자 시뮬레이션에서 응답 충실도와 상호보완적인 차원임을 입증한다.
English
User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6\% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7\% of evaluated dialogue states, compared with 22.9\% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.