ContinualSkillBench: LLM 에이전트는 정말로 능력을 진화시킬 수 있는가?
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
August 4, 2026
저자: Tianyi Guan, Yiding Wang, Haotong Yang, Siyuan Cao, Shirui Liu, Yi Hu, Jiaqi Li, Muhan Zhang
cs.AI
초록
현대 에이전트 프레임워크는 대규모 언어 모델에 외부 기술 라이브러리를 장착하여 복잡한 작업을 해결한다. 그러나 이러한 시스템이 기술을 효과적으로 진화시킬 수 있는지, 그리고 그 결과로 생성된 기술이 작업 해결 능력을 실제로 향상시키는지는 여전히 불분명하다. 이러한 격차를 해소하기 위해 우리는 문맥 내 연속적 기술 학습을 위한 동적 평가 프레임워크인 ContinualSkillBench를 제안한다. 이 프레임워크는 5개의 대표적인 도메인을涵盖하며, 각 도메인은 난이도가 증가하는 순서와 교차 작업 기술 재사용 기회에 따라 배열된 100개의 상호 연결된 하위 작업을 포함한다. 실험 결과, 순차적 실행은 일반적으로 성능을 향상시키지만 그 이득은 모델과 도메인에 따라 상당히 달라진다. 또한 문맥 내 학습은 평균적으로 명시적 기술 유지 관리와 비슷한 성능을 보이는데, 이는 성능 향상의 상당 부분이 재사용 가능한 기술 추상화 자체보다는 이전 문맥과 피드백에 대한 적응에서 비롯됨을 시사한다. 그럼에도 명시적 기술은 재사용 가능한 절차나 정밀한 출력을 요구하는 작업에 대해 선택적 이점을 제공한다. 또한 성능이 낮은 모델일수록 더 크고 더 파편화된 작업별 기술 모음을 축적하는 경향이 있음을 발견했다. 이러한 발견은 현재의 문맥 내 기술 진화 메커니즘이 지속적 적응을 지원할 수 있지만, 경험을 견고하고 전이 가능한 기술로 일관되게 통합하는 데에는 여전히 어려움을 겪고 있음을 보여준다.
English
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.