CAST: LLM 에이전트를 위한 턴 단위 교사로서의 게임 솔버
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
July 28, 2026
저자: Yu Wang, Yi-Kai Zhang, Wentao Shi, Ziang Ye, Yuchun Miao, Yueqing Sun, Qi Gu, Xunliang Cai, Lan-Zhe Guo, Han-Jia Ye, Fuli Feng
cs.AI
초록
장기 게임에서 행동하도록 대규모 언어 모델(LLM)을 훈련시키는 것은 범용 의사 결정을 향한 유망한 단계이지만, 검증 가능한 보상 기반 강화 학습(RLVR)은 어떤 결정이 성공을 결정짓는지에 대한 정보가 거의 없는 희소한 최종 보상에 의존합니다. 더 조밀한 프로세스 신호는 이러한 누락된 턴 단위 신용을 제공할 수 있지만, 기존 신호원은 저렴하면서도 정확하게 유지하기 어렵습니다. 우리는 게임 해결사의 상태 가치 변화가 특정 행동이 상태를 성공으로 진전시키는지 여부를 드러낸다는 점을 관찰했습니다. 이 통찰을 바탕으로, 우리는 이러한 가치 변화를 해결사 이점으로 변환하고 이를 턴 단위 신호로 RLVR에 주입하는 CAST(해결사 교사로부터의 신용 할당)를 제안합니다. 또한, 연성 최적 해결사 가정 하에서 해결사 이점을 최대화하는 것이 해결사로부터의 온-정책 증류와 동등하며, 교사 로짓 대신 스칼라 값만 필요함을 보여줍니다. Sokoban, Minesweeper, Rush Hour에서 CAST는 모든 게임에 대해 동일 도메인 및 미경험 난이도 평가에서 훈련된 모든 기준 모델을 능가했으며, ALFWorld와 WebShop에서 최고 평균 제로샷 성능을 달성했습니다. 우리의 코드는 https://github.com/Wloner0809/CAST에서 제공됩니다.
English
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.