ChatPaper.aiChatPaper

LLM-as-a-Coach: 검증 불가능한 과제를 위한 경험적 학습

LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks

July 20, 2026
저자: Tianzhu Ye, Li Dong, Guanheng Chen, He Zhu, Xun Wu, Shaohan Huang, Furu Wei
cs.AI

초록

개방형 작업에 대한 강화학습(RL)은 LLM의 루브릭 기반 평가를 스칼라 보상으로 압축하여 풍부한 텍스트 피드백을 무시하고 서로 다른 품질 프로필을 가진 응답들을 동일하게 취급한다. 본 논문에서는 경험 학습(Experiential Learning, EL)을 제안한다. 이는 LLM-심판(LLM-as-a-Judge) 방식의 피드백 모델을 LLM-코치(LLM-as-a-Coach)로 전환한다. 코치는 각 정책 기반 응답에 대한 평가를 전이 가능한 경험적 지식으로 추출하며, 이는 교사 모델을 조건화하고 정책이 정책 기반 문맥 증류(on-policy context distillation)를 통해 내면화한다. 스칼라 보상과 비교할 때, 이 더 높은 대역폭의 피드백 채널은 조밀한 감독을 제공하고 고품질 응답 간의 세분화된 선호도를 보존한다. 두 정책 계열에서, 정책 자체 또는 독점 모델의 피드백을 사용할 때, EL은 보류된 개방형 작업 및 보지 못한 개방형 작업에서 루브릭 기반 RL을 지속적으로 능가한다. 특히, EL은 학습 분포를 넘어 더 잘 일반화하며 보상 해킹을 완화한다. 이러한 발견은 경험적 지식이 비검증 가능한 작업의 후훈련(post-training)을 위한 더 풍부하고 일반화 가능한 학습 신호임을 입증한다.
English
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback channel provides dense supervision and preserves fine-grained preferences among high-quality responses. Across two policy families, with feedback from the policy itself or a proprietary model, EL consistently outperforms rubric-based RL on held-out and unseen open-ended tasks. Notably, EL generalizes better beyond the training distribution, and mitigates reward hacking. These findings establish experiential knowledge as a richer and more generalizable learning signal for post-training on non-verifiable tasks.