Ring-Zero:創発的推論を実現する1兆パラメータへのZero RLのスケーリング
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
July 14, 2026
著者: Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou
cs.AI
要旨
人間による注釈データを必要としない検証可能な報酬を用いた強化学習(しばしばゼロRLと呼ばれる)は、思考連鎖推論を引き出すための強力なパラダイムとして登場した。しかし、計算リソースの制約から、既存の研究は主に小規模モデルに限られており、大規模スケールにおける学習ダイナミクスや創発的能力は未解明のままである。このフロンティアを有意義に探求するため、我々はモデルから高品質な推論行動を引き出すことを目指す。しかし、単純なスケーリングでは、可読性の低下、トークンの冗長性、適応的な推論の深さの欠如といった問題が生じることを発見した。これらの課題に対処するため、我々は安定かつ効率的な学習パイプラインを提案する。これは、クリップ付き重要度サンプリング、学習-推論比率補正、混合精度制御といったアルゴリズムおよびシステムの最適化を組み込んでいる。実験から得られた主な知見は3点であり、スケーリングの「苦い教訓」を裏付けるものである。(1) パラメータ1兆規模へのスケーリングは、サンプル効率と性能上限を大幅に向上させる。(2) 学習プロセスは、初期の発見フェーズとその後のシャープニングフェーズを順次経て進行する。(3) モデルは、擬人化、構造化フォーマット化、自己検証、並列推論、コンテクスト不安など、高度な認知行動を自発的に発達させ、手作りのヒューリスティックスを不要にする。7つの数学ベンチマークで評価した結果、Ring-2.5-1T-Zeroは競争力のある性能を達成した。さらに、最終的な正解の正確性を超えてCoTの品質を評価するため、理解可能性、再現性、効率性の3次元からなる構造化評価フレームワークを提案し、我々のモデルが構造化され簡潔な推論トレースの生成において明確な優位性を示すことを明らかにした。観測された創発現象を共有することで、特に1兆パラメータ規模におけるスケーリング行動に関する深い洞察をコミュニティに提供したい。
English
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.