人間の監督を超えた大規模推論モデルのスケーリング:超知能への道
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
August 31, 2026
著者: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
cs.AI
要旨
大規模推論モデル(LRM)における近年の進展は、検証可能な報酬を用いた強化学習(RLVR)が、結果を自動的に検証できる数学やコードの領域において推論能力を大幅に向上させ得ることを示している。この進展をオープンエンドな課題やエージェント型タスクへ拡張することは、信頼性の高い報酬の獲得がより困難であり、また直接的な人間による監督がモデル生成経験の規模と複雑さに追従できないため、依然として難しい問題である。本論文は、人間による監督が学習ループから徐々に後退する中で、LRMがどのように改善を継続し得るかを研究する。我々はこの問題の相互に関連する二つの側面を考察する。報酬の軸は、インスタンス単位の人間による判定から、再利用可能な検証器や、人間のフィードバックなしでも機能する報酬への発展を追跡する。経験の軸は、人間がキュレーションしたタスクや環境から、自己生成型カリキュラム、構築された環境、自律的な共進化へと学習がどのように進展し得るかを考察する。我々はこれらの次元を、学習プロセスのどの部分が引き続き人間の制御下にあるかを特定するL0からL4までの五段階のラダーを通じて結び付ける。本分析はさらに、報酬ハッキング、フィードバックのドリフト、カリキュラムの崩壊、環境エラーなど、自律性を高めた報酬と経験生成によってもたらされるリスクを浮き彫りにする。その結果として、我々はポリシー能力、フィードバックの忠実性、経験の品質という三つの相補的な対象に関する評価も提供する。本分析は、人間の監督を超えたLRMのスケーリングに関する現在のアプローチと、超知能に向けた自己維持型学習システムの開発に関わる未解決問題について、構造化された説明を提供する。さらに、我々は最新の進展を追跡するために継続的に更新されるhttps://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHubリポジトリ}を維持している。
English
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.