PAST-Bench: 개인 에이전트에서의 재귀적 자기 개선의 기초 벤치마킹
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents
August 4, 2026
저자: Shuhan Xue, Zixin Ding, Yichen Shen, Yinjie Wang, Zhenfei Yin, Yingcheng Wu, Yuxin Chen, Mengdi Wang, Ling Yang
cs.AI
초록
재귀적 자기 개선은 에이전트가 축적된 경험을 더 나은 미래 행동으로 전환하도록 요구한다. 개인 AI 에이전트는 세션 간에 선호도, 작업 이력, 도구 루틴, 학습된 기술을 유지하므로 이러한 능력을 연구하기 위한 구체적인 환경을 제공한다. 그러나 유지된 경험이 실제로 시간이 지남에 따라 에이전트를 개선하는지는 체계적으로 테스트된 바 없다. 우리는 이 질문을 분리하기 위해 설계된 벤치마크인 PAST-Bench를 소개한다. 각 에이전트는 유지된 경험을 켜고 끄는 일치된 조건에서 새 세션 작업들의 순차적 시퀀스를 실행한다. 이 벤치마크는 기억, 절차적 재사용, 정보 수집, 갱신에 걸쳐 26개 시나리오와 204개 에피소드를 포함한다. 우리는 후속 작업에서의 개선과 그러한 개선이 의도된 저장·검색·갱신 경로를 따르는지 여부를 모두 보고한다. 일곱 개의 기본 모델과 네 개의 에이전트 프레임워크 전반에서 개선은 실재하지만 능력에 따라 고르지 않다. 동일한 주요 개선을 보이는 에이전트라도 그 개선이 의도된 경로의 증거에 의해 뒷받침되는지 여부는 크게 달라질 수 있다. 이러한 발견을 바탕으로 우리는 에이전트 루프의 여러 단계에 걸친 다섯 가지 표적 개입으로 Hermes를 확장한 Hermes+를 개발한다. Hermes+는 유지된 경험으로부터의 평균 개선을 높이고 더 명확한 경로 증거를 제공하며, 특히 오래된 상태를 교체해야 하는 작업에서 가장 큰 개선을 보인다. 다만 그 효과는 여전히 능력 및 모델 의존적이다. PAST-Bench와 Hermes+는 함께, 지속적 에이전트가 경험 보유에서 경험을 통한 체계적 개선으로 나아갈 수 있는 방식을 연구하기 위한 평가 및 진단 기반을 제공한다. 코드: https://github.com/Gen-Verse/PAST-Bench
English
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench