DRACO: 장기 지평 에이전트 훈련을 위한 동적 루브릭 기반 세밀한 크레딧 할당
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
September 3, 2026
저자: Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk
cs.AI
초록
검증 가능한 보상으로부터의 강화 학습은 작업에 프로그래밍 방식의 검사기가 있을 때 효과적이지만, 대부분의 장기 지평 에이전트 도메인에는 그러한 검사기가 없습니다. 우리는 실측 성공 신호를 얻을 수 없는 결과 블라인드 설정을 다룹니다. 다중 기준 루브릭은 그러한 보상을 제공하는 널리 쓰이는 방법입니다. 루브릭은 궤적당 한 번만 채점되지만, 단일 스칼라 값은 수십 단계에 걸친 신호로는 부족합니다. 이에 우리는 신용 최적화를 위한 루브릭 기반 이점 분배(DRACO)를 제안합니다. DRACO는 훈련 중 정책의 변화하는 능력을 추적하도록 루브릭을 동적으로 생성하고, 완료된 궤적마다 해당 루브릭을 한 번 채점한 다음, 그 채점 결과를 주석이 달린 루브릭과 관련된 단계들에 재분배하여 GRPO에서 차별화된 단계별 이점을 생성합니다. 이 재분배는 닫힌 형태로 이루어지며, 학습된 기여도 할당 모듈을 도입하지 않습니다. AppWorld에서 DRACO는 검증기를 전혀 직접 사용하지 않음에도 불구하고 기본 모델보다 15.9점, 희소 실측 보상으로 훈련된 GRPO보다 5.3점 높은 성능을 얻습니다. 도메인 밖(out-of-domain) Tau-Bench에서는 최첨단 평가자 없이도 기본 모델보다 5.3점 높은 성능을 보여, 실측 보상 훈련 및 다른 루브릭 기반 훈련 설정을 모두 능가합니다. DRACO의 코드는 https://github.com/IBM/draco에서 확인할 수 있습니다.
English
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per completed trajectory, and redistributes that judgment over the steps responsible for annotated rubrics to produce differentiated per-step advantages in GRPO. The redistribution is closed-form and does not introduce any trained attribution module. On AppWorld, DRACO gains 15.9 points over the base model and 5.3 points over GRPO trained with a sparse ground-truth reward, despite not using any verifiers itself. On out-of-domain Tau-Bench, it gains 5.3 points over the base model even without a frontier judge, beating both ground-truth-reward training and other rubric-based training settings. The code for DRACO is available at https://github.com/IBM/draco.