Ring-Zero:將 Zero RL 擴展至一兆參數以實現湧現推理

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

July 14, 2026
作者: Xinyu Tang, Gangqiang Cao, Yurou Liu, Yuliang Zhan, Xiaochong Lan, Yifan Li, Yuchen Yan, Han Peng, Zican Dong, Zhenduo Zhang, Tianshu Wang, Xinyu Kong, Zujie Wen, Wayne Xin Zhao, Zhiqiang Zhang, Jun Zhou
cs.AI

摘要

強化學習結合可驗證獎勵而無需人工標註數據的方法(常被稱為零強化學習),已成為引發思維鏈推理的有效典範。然而,由於計算資源限制,現有研究大多侷限於小型模型,使得大規模下的訓練動態與湧現能力尚未被充分探索。為了有意義地探索此一前沿,我們旨在從模型中引發高品質的推理行為。然而我們發現,單純的規模擴張往往導致可讀性差、標記冗餘以及缺乏自適應推理深度等問題。為應對這些挑戰,我們提出了一套穩定且高效的訓練流程,整合了演算法與系統層面的優化,例如裁剪重要性取樣、訓練-推理比例校正以及混合精度控制。實驗提供了三項關鍵發現,驗證了規模擴張的「苦澀教訓」:(1) 將參數規模擴展至1兆,能顯著提升樣本效率與性能上限;(2) 訓練過程依序經歷初始的探索階段,隨後進入銳化階段;(3) 模型自發發展出先進的認知行為,包括擬人化、結構化格式、自我驗證、平行推理以及上下文焦慮,使得人工設計的啟發式規則變得不再必要。在七個數學基準測試中,Ring-2.5-1T-Zero達到了具競爭力的表現。此外,為在最終答案正確性之外評估思維鏈品質,我們提出了一套涵蓋三個維度的結構化評估框架:可理解性、可再現性與效率,而我們的模型在產生結構化且簡潔的推理軌跡方面展現出明顯優勢。透過分享我們觀察到的湧現現象,我們期望為學術界提供關於規模擴張行為(特別是在1兆參數規模下)更深入的洞見。
English
Reinforcement learning with verifiable rewards without human-annotated data, often referred to as zero RL, has emerged as a powerful paradigm for eliciting chain-of-thought reasoning. However, due to computational constraints, existing studies are largely restricted to small models, leaving the training dynamics and emergent capabilities at a large scale unexplored. To meaningfully explore this frontier, we aim to elicit high-quality reasoning behaviors from the model. However, we find that naive scaling often suffers from poor readability, token redundancy, and a lack of adaptive reasoning depth. To address these challenges, we present a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control. Our experiments offer three key findings that validate the "bitter lesson" of scaling: (1) scaling to 1T parameters significantly enhances sample efficiency and performance ceilings; (2) the training process progresses sequentially through an initial discovery phase followed by a sharpening phase; and (3) the model spontaneously develops advanced cognitive behaviors, including anthropomorphism, structured formatting, self-verification, parallel reasoning, and context anxiety, rendering hand-crafted heuristics redundant. Evaluated on seven mathematical benchmarks, Ring-2.5-1T-Zero achieves competitive performance. Additionally, to assess CoT quality beyond final-answer correctness, we propose a structured evaluation framework across three dimensions: comprehensibility, reproducibility, and efficiency, where our model demonstrates clear advantages in producing structured and concise reasoning traces. By sharing our observed emergent phenomena, we hope to provide the community with deeper insights into scaling behaviors, particularly at the 1-trillion scale.
PDF792July 17, 2026