ReflectRL: 골든 부정 궤적으로부터 반성적-직접적 추론을 통한 학습
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
August 4, 2026
저자: Jinhe Bi, Chennan Zhou, Zengjie Jin, Aniri, Shuo Lu, Wenke Huang, Hu Cao, Xun Xiao, Zhihong Zhu, Volker Tresp, Fei Shen, Yunpu Ma, Tat-Seng Chua
cs.AI
초록
온-폴리시 훈련은 대규모 언어 모델의 추론 능력을 향상시키기 위한 강력한 사후 훈련 패러다임으로 부상하였으며, 흔히 더 강력한 전문가 모델로부터 얻은 골든 궤적에 의해 강화된다. 그러나 전문가가 더 어려운 문제에 실패하는 경우, 기존의 궤적 기반 방법은 주요 감독 신호를 상실하며, 실패한 궤적은 대개 부정적 샘플로 간주되어 폐기된다. 우리는 이러한 실패, 즉 골든 네거티브 궤적(Golden Negative Trajectories)이 모방해야 할 시연이 아니라 반성해야 할 결함 있는 궤적으로 취급될 때 여전히 유용한 추론 신호를 제공할 수 있다고 주장한다. 우리는 반성 이점(Reflection Advantage)을 확인한다. 어려운 문제의 경우, 결함 있는 궤적을 반성하는 것이 처음부터 직접 문제를 해결하는 것보다 더 쉽고 효과적일 수 있다는 것이다. 이러한 동기에 기반하여, 우리는 온-폴리시 훈련 중에 골든 네거티브 궤적으로부터 학습하는 경량 플러그 앤 플레이 프레임워크인 ReflectRL을 제안한다. ReflectRL은 먼저 이 궤적들을 활용하여 반성적 추론(Reflective Reasoning)을 도출한 다음, 반성-직접 정책 전환(Reflective-to-Direct Policy Transition)을 적용하여 습득된 추론 행동을 직접 추론(Direct Reasoning)으로 전이한다. 9개 벤치마크, 4개 LLM 백본, 4개 온-폴리시 훈련 방법에 걸친 실험 결과는 ReflectRL이 최소한의 오버헤드로 추론 성능을 일관되게 향상시킨다는 것을 보여준다.
English
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories from stronger expert models. However, when the expert fails on harder problems, existing trajectory-guided methods lose their main source of supervision, and these failed trajectories are typically discarded as negative samples. We argue that such failures, which we call Golden Negative Trajectories, can still provide valuable reasoning signals when treated not as demonstrations to imitate, but as flawed trajectories to reflect upon. We identify a Reflection Advantage: for hard problems, reflecting on a flawed trajectory can be easier and more effective than solving the problem directly from scratch. Motivated by this, we propose ReflectRL, a lightweight plug-and-play framework that learns from Golden Negative Trajectories during on-policy training. ReflectRL first uses these trajectories to elicit Reflective Reasoning, then applies Reflective-to-Direct Policy Transition to transfer the acquired reasoning behavior back to Direct Reasoning. Experiments across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods show that ReflectRL consistently improves reasoning performance with minimal overhead.