ChatPaper.aiChatPaper

擴展大型推理模型以超越人類監督:通往超級智能之路

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

August 31, 2026
作者: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
cs.AI

摘要

近年來,大型推理模型(LRMs)的進展顯示,具可驗證獎勵的強化學習(RLVR)能大幅提升數學與程式編寫领域的推理能力,因為這些領域的結果可以自動檢查。然而,要將此進展延伸至開放式與代理型(agentic)任務仍具挑戰性,因為可靠的獎勵更難取得,而直接的人類監督也無法跟上模型生成經驗的規模與複雜度。本文探討當人類監督逐漸退出學習迴圈時,大型推理模型如何持續進步。我們檢視此問題的兩個相互關聯的面向。獎勵軸線追溯從單一實例的人類判斷,發展至可重複使用的驗證器(verifiers)與獎勵,甚至在沒有回饋的情況下也能運作。經驗軸線則探討學習如何從人類策劃的任務與環境,進展至自我生成的課程、建構環境,以及自主的共同演化。我們透過從L0到L4的五級階梯來連結這兩個面向,以辨識學習過程中哪些部分仍持續由人類控制。我們的分析進一步凸顯日益自主的獎勵與經驗生成所帶來的風險,包括獎勵駭取(reward hacking)、回饋漂移(feedback drift)、課程崩潰(curriculum collapse)與環境錯誤(environment errors)。因此,我們也圍繞三個互補的對象提供評估:政策能力(policy capability)、回饋忠實度(feedback fidelity)與經驗品質(experience quality)。此分析提供了當前超越人類監督擴展大型推理模型方法的結構化說明,以及發展朝向超級智能的自持續學習系統所涉及之開放問題。此外,我們維護一個持續更新的GitHub儲存庫(https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision),以追蹤最新進展。
English
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.