ChatPaper.aiChatPaper

TREK: Destilleren om te verkennen, Versterken om te verfijnen

TREK: Distill to Explore, Reinforce to Refine

July 6, 2026
Auteurs: Yuanda Xu, Zhengze Zhou, Kayhan Behdin, Jelena Markovic-Voronov, Hejian Sang, Xiaomin Li, Wenhui Zhu, Xinchen Du, Aida Rahmattalabi, Ran He, Sen Na, Zhipeng Wang, Alborz Geramifard
cs.AI

Samenvatting

Groepsrelatief Beleidsoptimalisatie (GRPO) is effectief wanneer het huidige beleid al bruikbare redeneertrajecten bemonstert, maar stagneert bij moeilijke prompts waarvan de correcte oplossingsmodi buiten de ondersteuning van het studentenbeleid liggen. Wij stellen TREK (Teacher-Routed Exploration via Forward KL) voor, een eenvoudige gefaseerde procedure die distillatie niet gebruikt voor imitatie, maar voor uitbreiding van de exploratieondersteuning. Een belangrijk voordeel van TREK is zijn algemeenheid: omdat het alleen geverifieerde uitvoertrajecten verbruikt, kan het een externe black-box-leraar, een white-box-leraar, of hetzelfde model met extra inferentiecontext gebruiken, en het kan efficiënt identificeren welke moeilijke-prompt-monsters het meest de moeite waard zijn om te consolideren, zelfs wanneer interne leraargegevens niet beschikbaar zijn. TREK identificeert eerst prompts waarbij de onbegeleide student een zeer laag slagingspercentage heeft, vraagt een voorstelbron om geverifieerde kandidaatoplossingen te produceren, behoudt de top-r voorstellen gerangschikt op huidige studentwaarschijnlijkheid, past een korte forward-KL-fase toe om die geverifieerde modi in de ondersteuning van de student te trekken, en keert dan terug naar standaard on-policy GRPO-verfijning. Bij wiskundig redeneren verbetert TREK met DeepSeek-V4-voorstellen Qwen3-modellen op alle geteste schalen op AIME 2024 en AIME 2025; voor Qwen3-8B verbetert het AIME 2025 van 36,9 naar 40,3 en AIME 2024 van 47,9 naar 51,1 (avg@16), terwijl de zelfcontextvariant 38,5 en 49,6 bereikt zonder externe leraar. Bij agentische taken verhoogt TREK het ALFWorld-succespercentage van 75,8 naar 82,8 en het ScienceWorld-succespercentage van 12,5 naar 26,7; opmerkelijk is dat TREK bij de moeilijkste taaktypen al vroeg in de training hoge succespercentages behaalt, terwijl onbegeleide GRPO aanzienlijk meer optimalisatiestappen nodig heeft om vergelijkbare niveaus te bereiken.
English
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls on hard prompts whose correct solution modes lie outside the student's on-policy support. We propose TREK (Teacher-Routed Exploration via Forward KL), a simple staged procedure that uses distillation not for imitation but for exploration support expansion. A key advantage of TREK is its generality: because it only consumes verified output trajectories, it can use an external black-box teacher, a white-box teacher, or the same model given additional inference-time context, and it can efficiently identify which hard-prompt samples are most worth consolidating even when teacher internals are unavailable. TREK first identifies prompts where the unaided student has very low pass rate, queries a proposal source to produce verified candidate solutions, keeps the top-r proposals ranked by current student likelihood, applies a short forward-KL phase to pull those verified modes into the student's support, and then returns to standard on-policy GRPO refinement. On mathematical reasoning, TREK with DeepSeek-V4 proposals improves Qwen3 models across all tested scales on AIME 2024 and AIME 2025; for Qwen3-8B, it improves AIME 2025 from 36.9 to 40.3 and AIME 2024 from 47.9 to 51.1 (avg@16), while the self-context variant reaches 38.5 and 49.6 without an external teacher. On agentic tasks, TREK raises ALFWorld success rate from 75.8 to 82.8 and ScienceWorld success rate from 12.5 to 26.7; notably, on the hardest task types, TREK achieves high success rates early in training while unaided GRPO requires substantially more optimization steps to reach comparable levels.