ChatPaper.aiChatPaper

스캐폴드로서의 환경: 장기 지평 과제에서 자기 진화 에이전트를 부트스트랩하기 위한 피드백 강화

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

September 8, 2026
저자: Hongbang Yuan, Zhuoran Jin, Yixin Cao
cs.AI

초록

대규모 언어 모델은 정적 추론에서 뛰어난 능숙도를 보이지만, 장기(long-horizon) 과업을 위해 강화학습(RL)을 통해 자율 에이전트로 훈련하는 것은 종종 심각한 보상 희소성으로 인해 저해된다. 기존의 감독 미세조정(SFT)을 통한 에이전트 측 워밍업이 이를 완화할 수 있지만, 이는 흔히 데이터 희소성과 제한된 탐색으로 인해 한계가 있다. 이를 해결하기 위해 우리는 피드백 강화 환경(Feedback-Enriched Environments, FEEs)을 구축함으로써 환경 측 적응으로의 패러다임 전환을 제안한다. 예비 연구를 통해 우리는 에피소드 내 탐색과 에피소드 간 진화 모두의 후반 단계에서 행동 안내에서 관측 풍부화로 전환함으로써 환경을 재구성하는 피드백 설계 전략을 정립한다. 다양한 Qwen3 모델 규모와 GRPO, GSPO, DAPO와 같은 RL 알고리즘을 사용한 SciWorld 및 BFCL 벤치마크에 대한 대규모 실험은 FEEs가 표준 설정보다 일관되게 성능 향상을 가져옴을 입증한다. 나아가 우리의 분석은 FEEs를 사용한 훈련이 (1) 엔트로피 변동성을 줄여 학습 동역학을 안정화하고, (2) 어려운 과업에서 능동적 상태 공간 탐색을 촉진하며, (3) 단순한 추론 시 사전(prior)으로 작용하기보다 환경적 안내가 정책 가중치로 내재화되도록 보장함을 보여주며, (4) 안정적 최적화를 위한 중요한 경계로서 그룹 내 피드백 일관성을 식별한다.
English
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional agent-side warming up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to environment-side adaptation by constructing Feedback-Enriched Environments (FEEs). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs (1) stabilizes training dynamics by reducing entropy volatility, (2) facilitates proactive state-space exploration in difficult tasks, (3) ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and (4) identifies intra-group feedback consistency as a critical boundary for stable optimization.