ChatPaper.aiChatPaper

環境作為支架:豐富回饋以啟動長程任務中的自我演化代理

Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks

September 8, 2026
作者: Hongbang Yuan, Zhuoran Jin, Yixin Cao
cs.AI

摘要

大型語言模型在靜態推理上展現卓越能力,然而透過強化學習(RL)將其訓練為自主代理以執行長程任務時,常受嚴重獎勵稀疏性阻礙。雖然透過監督式微調(SFT)進行傳統的代理端預熱可緩解此問題,但其常受資料稀缺與探索受限所限制。為解決此問題,我們提出轉向環境端適應的典範轉移,藉由建構回饋豐富化環境(FEEs)。透過先導研究,我們建立一套回饋設計策略,該策略藉由在回合內探索與回合間演化的後期階段,從行動引導轉向觀測豐富化,來重新構築環境。在 SciWorld 與 BFCL 基準測試上,使用多種 Qwen3 模型規模及 GRPO、GSPO、DAPO 等 RL 演算法進行的大規模實驗顯示,FEEs 相較於標準設定能持續提升效能。此外,我們的分析揭示,使用 FEEs 訓練可:(1) 藉由降低熵波動穩定訓練動態;(2) 在困難任務中促進主動的狀態空間探索;(3) 確保環境引導內化至策略權重,而非僅作為推論時先驗;以及 (4) 辨識出群組內回饋一致性是穩定最佳化的關鍵邊界。
English
Large Language Models demonstrate remarkable proficiency in static reasoning, yet training them as autonomous agents through Reinforcement Learning (RL) for long-horizon tasks is often hindered by severe reward sparsity. While conventional agent-side warming up via supervised fine-tuning (SFT) can alleviate this, it is frequently limited by data scarcity and constrained exploration. To address this, we propose a paradigm shift to environment-side adaptation by constructing Feedback-Enriched Environments (FEEs). Through a pilot study, we establish a feedback design strategy that reformulates environments by transitioning from action guidance to observation enrichment during the later stages of both intra-episode exploration and inter-episode evolution. Large-scale experiments on SciWorld and BFCL benchmarks using various Qwen3 model scales and RL algorithms such as GRPO, GSPO, and DAPO demonstrate that FEEs consistently yield performance improvements over standard settings. Furthermore, our analysis reveals that training with FEEs (1) stabilizes training dynamics by reducing entropy volatility, (2) facilitates proactive state-space exploration in difficult tasks, (3) ensures the internalization of environmental guidance into policy weights rather than acting as a mere inference-time prior, and (4) identifies intra-group feedback consistency as a critical boundary for stable optimization.