ChatPaper.aiChatPaper

UP: Onbegrensde Positieve Asymmetrische Optimalisatie voor het Doorbreken van het Exploratie-Stabiliteit Dilemma

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

July 8, 2026
Auteurs: Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin
cs.AI

Samenvatting

Reinforcement learning (RL) is de standaardparadigma geworden voor het verbeteren van de complexe redeneervaardigheden van grote taalmodellen (LLM's). Om steekproefefficiëntie te bereiken, vertrouwen moderne RL-frameworks op importance sampling (IS). Deze algoritmen kampen echter met een exploratie-stabiliteitsdilemma. Zuivere IS leidt vaak tot catastrofale trainingsinstabiliteit, terwijl standaard clippingmechanismen, die worden gebruikt om deze instabiliteit te beperken, het beleidsupdatebudget strikt inperken. Door het concept van Waarschijnlijkheidscapaciteit (Cap) te formaliseren, tonen we aan dat conservatief clippen de exploratie structureel belemmert door het updatebudget voor correcte maar onzekere redeneerpaden voortijdig af te kappen. Om deze beperkingen te doorbreken, stellen we Onbegrensde Positieve Asymmetrische Optimalisatie (UP) voor, een universele en plug-and-play-doelfunctie. UP herstructureert theoretisch het optimalisatieproces door het beleid via de stop-gradient-operator aan zijn huidige toestand te verankeren. Dit asymmetrische ontwerp ontsluit ongeclipte, stabiele gradiënten voor positieve voordelen om exploratie te maximaliseren, terwijl standaard clippbeveiligingen voor negatieve voordelen behouden blijven om trainingsinstabiliteit te voorkomen. Bovendien breidt onze formulering zich gemakkelijk uit over verschillende optimalisatiegranulariteiten, waaronder token-niveau (GRPO, DAPO) en sequentie-niveau (GSPO) frameworks. Uitgebreide experimenten tonen aan dat UP de exploratiecapaciteit vergroot en superieure redeneeraccuraatheid behaalt over diverse RL-algoritmen (DAPO, GSPO en GRPO), modelarchitecturen (Dicht, MoE en visie-taal) en trainingsmodaliteiten (taal en multimodaal), waarmee UP wordt gevalideerd als een werkelijk universele plug-and-play-verbetering voor op RL gebaseerde training.
English
Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs). To achieve sample efficiency, modern RL frameworks rely on importance sampling (IS). However, these algorithms suffer from an exploration-stability dilemma. Pure IS often leads to catastrophic training instability, while standard clipping mechanisms used to mitigate this instability strictly constrain the policy update budget. By formalizing the concept of Probability Capacity (Cap), we reveal that conservative clipping structurally stifles exploration by prematurely truncating the update budget for correct but low-confidence reasoning paths. To break free from these constraints, we propose Unbounded Positive Asymmetric Optimization (UP), a universal and plug-and-play objective. UP theoretically restructures the optimization process by anchoring the policy to its current state via the stop-gradient operator. This asymmetric design unleashes unclipped, stable gradients for positive advantages to maximize exploration, while maintaining standard clipping safeguards for negative advantages to prevent training instability. Furthermore, our formulation readily extends across different optimization granularities, including token-level (GRPO, DAPO) and sequence-level (GSPO) frameworks. Extensive experiments demonstrate that UP enhances exploration capacity and achieves superior reasoning accuracy across diverse RL algorithms (DAPO, GSPO, and GRPO), model architectures (Dense, MoE, and vision-language), and training modalities (language and multimodal), validating UP as a truly universal plug-and-play enhancement for RL-based training.