超越人类监督的大规模推理模型扩展:通往超级智能的路径
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
August 31, 2026
作者: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
cs.AI
摘要
近期,大型推理模型(LRM)的进展表明,基于可验证奖励的强化学习(RLVR)能够显著提升数学和代码领域的推理能力,因为这些领域的结果可以自动检查。然而,将这一进展推广到开放式和智能体型任务仍然困难,原因在于可靠的奖励更难获得,且直接的人类监督无法跟上模型生成经验的规模与复杂性。本文研究当人类监督逐渐退出学习循环时,LRM如何持续改进。我们考察了这个问题的两个相互关联的维度。奖励轴追踪了从逐实例人工判断到可复用验证器、乃至无需人类反馈即可运作的奖励的发展历程。经验轴则考察学习如何从人工策划的任务与环境,演进到自生成课程、构建的环境以及自主共同进化。我们通过一个从L0到L4的五级阶梯连接这两个维度,该阶梯识别出学习过程中哪些部分仍然处于人类的持续控制之下。我们的分析进一步强调了日益自主的奖励和经验生成所带来的风险,包括奖励破解、反馈漂移、课程崩溃和环境错误。因此,我们还提供了围绕三个互补对象的评估:策略能力、反馈保真度和经验质量。这一分析为当前在超越人类监督条件下扩展LRM的方法,以及为迈向超级智能而开发自我维持学习系统所涉及的开放问题,提供了结构化的阐述。此外,我们维护了一个持续更新的GitHub仓库(https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision),以跟踪最新进展。
English
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.