ChatPaper.aiChatPaper

持续技能基准:大语言模型智能体能否真正进化其能力?

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

August 4, 2026
作者: Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
cs.AI

摘要

现代智能体框架通过为大语言模型配备外部技能库来解决复杂任务。然而,这些系统能否有效进化其技能,以及由此产生的技能能否真正提升任务求解能力,目前仍不明确。为弥合这一差距,我们提出了ContinualSkillBench,一个面向上下文持续技能学习的动态评估框架。该框架涵盖五个代表性领域,每个领域包含100个相互关联的子任务,这些子任务按难度递增排列,并提供跨任务技能复用的机会。实验表明,顺序执行通常能提升性能,但提升幅度在不同模型和领域间存在显著差异。此外,上下文学习在平均水平上与显式技能维护表现相当,这表明性能提升主要源于对先前上下文和反馈的适应,而非仅依赖可复用的技能抽象。然而,显式技能在处理需要可复用流程或精确输出的任务时仍具有选择性优势。我们进一步发现,能力较弱的模型倾向于积累更大、更碎片化的任务特定技能集合。这些发现表明,当前的上下文技能进化机制能够支持持续适应,但在将经验一致地整合为稳健且可迁移的技能方面仍面临挑战。
English
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.