ChatPaper.aiChatPaper

足場としての環境:長期的タスクにおける自己進化型エージェントをブートストラップするためのフィードバックの豊富化

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

September 8, 2026
著者: Hongbang Yuan, Zhuoran Jin, Yixin Cao
cs.AI

要旨

大規模言語モデルは静的推論において顕著な高性能を示すが、長ホライズンタスクに向けて強化学習(RL)により自律エージェントとして訓練することは、しばしば深刻な報酬スパース性によって妨げられる。教師ありファインチューニング(SFT)による従来のエージェント側ウォームアップはこれを緩和し得るものの、データ不足と探索の制約によってしばしば限界がある。これに対処するため、我々はフィードバック豊富化環境(FEEs)を構築することによる、環境側適応へのパラダイムシフトを提案する。予備研究を通じて、エピソード内探索とエピソード間進化の双方の後半段階において、行動ガイダンスから観察の豊富化へ移行することにより環境を再定式化するフィードバック設計戦略を確立する。さまざまなQwen3モデル規模と、GRPO、GSPO、DAPOなどのRLアルゴリズムを用いたSciWorldおよびBFCLベンチマークでの大規模実験は、FEEsが標準設定に対して一貫して性能向上をもたらすことを示している。さらに、我々の分析は、FEEsを用いた訓練が (1) エントロピー変動を低減することで訓練ダイナミクスを安定化させ、(2) 困難なタスクにおける能動的な状態空間探索を促進し、(3) 単なる推論時事前として作用するのではなく、環境からのガイダンスを方策重みへ内在化させることを保証し、(4) グループ内フィードバック一貫性を安定な最適化のための重要な境界として特定することを明らかにする。
English
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional agent-side warming up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to environment-side adaptation by constructing Feedback-Enriched Environments (FEEs). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs (1) stabilizes training dynamics by reducing entropy volatility, (2) facilitates proactive state-space exploration in difficult tasks, (3) ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and (4) identifies intra-group feedback consistency as a critical boundary for stable optimization.