ChatPaper.aiChatPaper

TRIAGE: Rolgetypeerde Krediettoewijzing voor Agentisch Versterkingsleren

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

June 30, 2026
Auteurs: Yuanda Xu, Zhengze Zhou, Hejian Sang, Xiaomin Li, Jiaxin Zhang, Xinchen Du, Zhipeng Wang, Alborz Geramifard
cs.AI

Samenvatting

Agentische reinforcement learning vereist dat crediet wordt toegewezen aan omgevingsgerichte acties zoals zoekopdrachten, klikken, bewerkingen, navigatiecommando's en objectinteracties. Standaard GRPO gebruikt de uiteindelijke uitkomst van de verificateur als een uniform voordeel over alle actietokens. Dit uitkomstsignaal is nuttig maar structureel onvolledig: het bestraft nuttige exploratie in mislukte rollouts en versterkt overbodige of regressieve acties in succesvolle rollouts. Wij stellen TRIAGE voor, een rolgetypeerd credittoewijzingsraamwerk dat een semantische rollenas toevoegt aan uitkomstcrediet. Een gestructureerde beoordelaar classificeert elk segment als beslissende vooruitgang, nuttige exploratie, geen-vooruitgang infrastructuur of regressie, en een vaste rolgeconditioneerde regel kent deze labels toe aan begrensde segmentniveau procesbeloningen. Dit behoudt verificateursuitkomsten als de bron van optimalisatierichting, terwijl het de twee belangrijkste blinde vlekken van uitsluitend uitkomstcrediet corrigeert. Verder tonen wij aan dat rolgeconditioneerd crediet de optimale segmentniveau correctie is die uitdrukbaar is vanuit uitsluitend roletiketten – een projectie van het per-segment voordeelresidu op de rolvariabele – zodat de vaste rolconstanten de voordeelschatingsfout verminderen zodra de beoordelaar betrouwbaar is, en wij verbinden dit met beleidsgradiënten met lagere variantie. In ALFWorld, Search-QA en WebShop verbetert TRIAGE de succespercentages ten opzichte van GRPO voor twee beleidsmodellen en overtreft het zowel een scalaire door beoordelaar afgeleide procesbeloning als een uitkomstgesuperviseerde gedeelde backbone-waardebaseline. Ablatiestudies tonen aan dat de winst voortkomt uit roltypering in plaats van louter het toevoegen van dichte beloningen: betrouwbare detectie van regressie in succesvolle trajecten is de dominante bijdrager, terwijl exploratiecrediet een consistente secundaire winst oplevert; in voltooide ALFWorld- en WebShop-rollouts vermindert TRIAGE bovendien het aantal omgevingsgerichte beurten met respectievelijk 10,4% en 14,8% ten opzichte van GRPO.
English
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal is useful but structurally incomplete: it punishes useful exploration in failed rollouts and reinforces redundant or regressive actions in successful rollouts. We propose TRIAGE, a role-typed credit assignment framework that adds a semantic role axis to outcome credit. A structured judge classifies each segment as decisive progress, useful exploration, no-progress infrastructure, or regression, and a fixed role-conditioned rule maps these labels to bounded segment-level process rewards. This keeps verifier outcomes as the source of optimization direction while correcting the two main blind spots of outcome-only credit. We further show that role-conditioned credit is the optimal segment-level correction expressible from role labels alone -- a projection of the per-segment advantage residual onto the role variable -- so that the fixed role constants reduce advantage estimation error whenever the judge is reliable, and we connect this to lower-variance policy gradients. Across ALFWorld, Search-QA, and WebShop, TRIAGE improves success rates over GRPO for two policy models and outperforms both a scalar judge-derived process reward and an outcome-supervised shared-backbone value baseline. Ablations show that the gain comes from role typing rather than merely adding dense rewards: reliable detection of regression inside successful trajectories is the dominant contributor, while exploration credit provides a consistent secondary gain; on completed ALFWorld and WebShop rollouts, TRIAGE also reduces environment-facing turns by an additional 10.4% and 14.8% relative to GRPO.