ChatPaper.aiChatPaper

大语言模型在用户意图演变中迷失

LLMs Get Lost in Evolving User Intent

July 22, 2026
作者: Jihoon Tack, Philippe Laban, Jennifer Neville
cs.AI

摘要

随着大语言模型能力日益增强,它们越来越多地被部署为协作代理,通过迭代交互来承担用户委托的任务。然而,真正的交互本质上是动态的:用户很少一开始就明确表达意图,而是随着对话展开逐步披露、修正和重塑其意图。尽管如此,大语言模型目前仍主要在单轮、完全指定的设定下进行评估或训练,这留下了一个根本性问题:当用户意图在对话过程中演变时,大语言模型能在多大程度上追踪并响应这种变化?为研究这一问题,我们引入了一个框架,将静态的单轮任务转化为动态的多轮对话——在此过程中,用户的意图在各轮之间逐步揭示、修正,甚至中途转向——同时保留每项任务的原始评估协议,从而使现有基准测试无需额外标注即可作为受控实验平台。在多项任务中,我们发现一个一致的现象:在静态设定下表现优异的模型,在意图演变的设定中性能显著下降,且这一现象跨越不同模型家族。我们的研究结果揭示了一个根本差距:当前的大语言模型尚不能忠实追踪并响应用户不断演变的意图——这一能力在静态评估中不可见,但对未来的协作代理却至关重要。
English
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.