ChatPaper.aiChatPaper

TREK: 탐구를 위한 정제, 개선을 위한 강화

TREK: Distill to Explore, Reinforce to Refine

July 6, 2026
저자: Yuanda Xu, Zhengze Zhou, Kayhan Behdin, Jelena Markovic-Voronov, Hejian Sang, Xiaomin Li, Wenhui Zhu, Xinchen Du, Aida Rahmattalabi, Ran He, Sen Na, Zhipeng Wang, Alborz Geramifard
cs.AI

초록

그룹 상대 정책 최적화(GRPO)는 현재 정책이 이미 유용한 추론 궤적을 샘플링할 때 효과적이지만, 올바른 해결 모드가 학생의 온-정책 지원 범위(on-policy support) 밖에 있는 어려운 프롬프트에서는 정체됩니다. 본 논문에서는 모방이 아닌 탐색 지원 확장을 위해 증류를 사용하는 간단한 단계적 절차인 TREK(Teacher-Routed Exploration via Forward KL)을 제안합니다. TREK의 주요 장점은 일반성입니다. 검증된 출력 궤적만 소비하기 때문에 외부 블랙박스 교사, 화이트박스 교사, 또는 추가 추론 시간 컨텍스트가 주어진 동일한 모델을 사용할 수 있으며, 교사 내부 정보를 사용할 수 없는 경우에도 어떤 어려운 프롬프트 샘플이 통합에 가장 가치 있는지 효율적으로 식별할 수 있습니다. TREK는 먼저 지원을 받지 않은 학생의 통과율이 매우 낮은 프롬프트를 식별하고, 제안 소스에 검증된 후보 해결책을 생성하도록 요청한 후, 현재 학생 가능도에 따라 상위 r개의 제안을 유지하고, 짧은 순방향 KL(forward-KL) 단계를 적용하여 이러한 검증된 모드를 학생의 지원 범위로 끌어들인 다음, 표준 온-정책 GRPO 정제로 돌아갑니다. 수학적 추론에서 DeepSeek-V4 제안을 사용한 TREK는 모든 테스트된 규모의 Qwen3 모델에 대해 AIME 2024 및 AIME 2025에서 성능을 향상시킵니다. Qwen3-8B의 경우 AIME 2025에서 36.9에서 40.3으로, AIME 2024에서 47.9에서 51.1로 향상되었으며(avg@16), 자체 컨텍스트 변형은 외부 교사 없이 38.5와 49.6에 도달합니다. 에이전트 과제에서 TREK는 ALFWorld 성공률을 75.8에서 82.8로, ScienceWorld 성공률을 12.5에서 26.7로 향상시킵니다. 특히 가장 어려운 과제 유형에서 TREK는 훈련 초기에 높은 성공률을 달성하는 반면, 지원되지 않은 GRPO는 유사한 수준에 도달하기 위해 훨씬 더 많은 최적화 단계가 필요합니다.
English
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support. We propose TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion. A key advantage of TREK is its generality: because it only consumes verified output trajectories, it can use an external black-box teacher, a white-box teacher, or the same model given additional inference-time context, and it can efficiently identify which hard-prompt samples are most worth consolidating even when teacher internals are unavailable. TREK first identifies prompts where the unaided student has very low pass rate, queries a proposal source to produce verified candidate solutions, keeps the top-r proposals ranked by current student likelihood, applies a short forward-KL phase to pull those verified modes into the student's support, and then returns to standard on-policy GRPO refinement. On mathematical reasoning, TREK with DeepSeek-V4 proposals improves Qwen3 models across all tested scales on AIME 2024 and AIME 2025; for Qwen3-8B, it improves AIME 2025 from 36.9 to 40.3 and AIME 2024 from 47.9 to 51.1 (avg@16), while the self-context variant reaches 38.5 and 49.6 without an external teacher. On agentic tasks, TREK raises ALFWorld success rate from 75.8 to 82.8 and ScienceWorld success rate from 12.5 to 26.7; notably, on the hardest task types, TREK achieves high success rates early in training while unaided GRPO requires substantially more optimization steps to reach comparable levels.