**PILOT in the Loop: 장기 지평선(long-horizon) 에이전트를 위한 실시간 자기 개선** **초록** 대규모 언어 모델(LLM) 기반 에이전트는 복잡한 추론 과제에서 뛰어난 성능을 보여 주었지만, 단계별 오류 누적과 긴 실행 궤적에 따른 지식 부족으로 인해 장기 지평선 작업에서 어려움을 겪는다. 본 논문에서는 반복 실행을 통해 에이전트를 동적으로 개선하기 위한 능동 학습 기반 프레임워크인 **PILOT**을 제안한다. PILOT은 (1) 불확실성이 높은 단계를 선택적으로 저장하는 **선별적 메모리**, (2) 실행 중 오류를 실시간으로 수정하는 **출력 보정**, (3) 작업 컨텍스트와 피드백을 통합하는 **맥락 적응**이라는 세 가지 핵심 메커니즘으로 구성된다. ALFWorld, WebShop, HotPotQA를 포함한 다양한 벤치마크에서 광범위한 실험을 수행한 결과, PILOT은 기존의 프롬프트 기반 방법 및 미세 조정 기반 방법보다 우수한 성능을 보였으며, 특히 긴 궤적을 요구하는 작업에서 성공률을 크게 향상시켰다. 또한 PILOT은 제한된 상호작용 예산에서도 안정적인 수렴을 보여주었으며, 오류 수정률과 지식 획득 효율성 측면에서도 우수한 성과를 입증하였다.
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
August 27, 2026
저자: Yang Xiao, Yusong Sun, Haoyi Wu, Wenyang Hui, Wen Da, Zhaokai Luo, Mu Chuan, Yao Hu, Wenjie Li, Chengyue Jiang
cs.AI
초록
장기 과제 에이전트 실행은 현재 실행과 향후 작업을 모두 개선할 수 있는 경험을 생성한다. 대부분의 자기 개선 방법은 실행이 종료된 후에만 이 경험을 처리하므로, 활성 실행을 전환하거나 그로부터 얻은 교훈을 즉시 적용하고 검증할 수 없다. 우리는 자기 개선이 대신 실시간으로 이루어져야 하며, 새롭게 생성되는 경험을 활용하여 활성 실행을 전환하고 지속적 하네스를 업데이트해야 한다고 주장한다. 기존 에이전트 아키텍처는 이러한 목표를 완전히 지원하지 못한다. 단일 에이전트 자기 교정은 작업 실행과 궤적 평가를 하나의 컨텍스트 내에서 결합하는 반면, 하위 에이전트 위임은 실행을 분리하지만 일반적으로 활성 하위 에이전트를 전환할 수 없다. 우리는 두 가지 결합 메커니즘을 통해 실시간 자기 개선을 구현하는 감독자-작업자 하네스인 PILOT을 제시한다: (1) 실시간 조향은 실행 중에 별도의 감독자가 활성 작업자를 전환하거나 중단할 수 있게 하며, (2) 실시간 자기 진화는 실행 중에 드러난 절차와 실패 모드를 재사용 가능한 기술과 메모리로 정제한다. 두 개의 고정 백본과 세 개의 벤치마크에 걸쳐 PILOT은 여섯 가지 구성 중 다섯 가지에서 1위를 차지한다. Terminal-Bench 2.0에서 PILOT은 비교 대상 하네스보다 최대 9.8퍼센트포인트 높은 성능을 보인다. 자기 개선 설정에서 PILOT은 GLM-5.1에서 14.6포인트, Kimi-K2.6에서 12.4포인트의 향상을 달성한다. 평균 출력 토큰은 각각 42.9%와 47.4% 감소하는 반면, 출력 토큰 백만 개당 성공적 평가 수는 각각 110.3%와 134.0% 증가한다.
English
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.