Ring-Zero: 창발적 추론을 위한 조 개의 파라미터로의 Zero RL 스케일링
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
July 14, 2026
저자: Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou
cs.AI
초록
인간 주석 데이터 없이 검증 가능한 보상을 활용한 강화 학습, 흔히 제로 RL이라 불리는 접근법이 사고 사슬 추론을 이끌어내는 강력한 패러다임으로 부상했다. 그러나 계산적 제약으로 인해 기존 연구는 대부분 소규모 모델에 국한되어 있으며, 대규모에서의 훈련 동역학과 창발적 능력은 아직 탐구되지 않았다. 이 경계를 의미 있게 탐구하기 위해 우리는 모델로부터 고품질 추론 행동을 이끌어내는 것을 목표로 한다. 그러나 단순한 확장은 종종 가독성 저하, 토큰 중복, 적응적 추론 깊이의 부재로 이어짐을 발견했다. 이러한 문제를 해결하기 위해 우리는 클리핑된 중요도 샘플링, 훈련-추론 비율 보정, 혼합 정밀도 제어와 같은 알고리즘 및 시스템 최적화를 통합한 안정적이고 효율적인 훈련 파이프라인을 제시한다. 실험을 통해 확장의 "쓴 교훈"을 검증하는 세 가지 핵심 발견을 얻었다: (1) 1조 파라미터로의 확장이 샘플 효율성과 성능 상한을 크게 향상시킨다; (2) 훈련 과정이 초기 발견 단계와 이후 정밀화 단계를 순차적으로 거친다; (3) 모델이 의인화, 구조적 형식화, 자가 검증, 병렬 추론, 맥락 불안 등 고급 인지 행동을 자발적으로 발현하여 수작업 휴리스틱을 불필요하게 만든다. 7가지 수학 벤치마크에서 평가한 결과, Ring-2.5-1T-Zero는 경쟁력 있는 성능을 달성했다. 또한 최종 정답 정확성을 넘어 CoT 품질을 평가하기 위해, 이해 가능성, 재현 가능성, 효율성의 세 가지 차원에 걸친 구조화된 평가 프레임워크를 제안하며, 본 모델이 구조화되고 간결한 추론 흔적을 생성하는 데서 뚜렷한 이점을 보임을 입증한다. 관찰된 창발 현상을 공유함으로써, 특히 1조 규모에서의 확장 행동에 대한 더 깊은 통찰을 학계에 제공할 수 있기를 기대한다.
English
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.