ChatPaper.aiChatPaper

LLM作为教练:面向不可验证任务的体验式学习

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

July 20, 2026
作者: Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei
cs.AI

摘要

面向开放式任务的强化学习将大语言模型的基于评分标准的评估压缩为标量奖励,丢弃了丰富的文本反馈,并将不同质量特征的响应混为一谈。我们提出经验学习,将大模型即评判者的反馈模型重新定位为大模型即教练。教练将每个在策略响应的评估提取为可迁移的经验知识,该知识用于条件化教师模型,并通过在策略上下文蒸馏被策略内化。与标量奖励相比,这种高带宽反馈通道提供密集监督,并保留高质量响应间的细粒度偏好。在两个策略族中,无论是使用策略自身反馈还是专有模型反馈,经验学习在保留和未见过的开放式任务上始终优于基于评分标准的强化学习。值得注意的是,经验学习在训练分布之外具有更好的泛化能力,并能缓解奖励攻讦。这些发现确立了经验知识作为对非可验证任务进行后训练的更丰富且更具泛化性的学习信号。
English
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.