인간 감독을 넘어선 대규모 추론 모델의 확장: 초지능을 향한 길
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
August 31, 2026
저자: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo
cs.AI
초록
대규모 추론 모델(LRM)의 최근 발전은 검증 가능한 보상을 통한 강화 학습(RLVR)이 결과를 자동으로 확인할 수 있는 수학 및 코드 영역에서 추론 능력을 실질적으로 향상시킬 수 있음을 보여주었다. 이러한 진전을 개방형 및 에이전트 기반 작업으로 확장하는 것은 여전히 어려운데, 이는 신뢰할 수 있는 보상을 확보하기가 더 까다롭고 직접적인 인간의 감독이 모델이 생성한 경험의 규모와 복잡성을 따라잡을 수 없기 때문이다. 본 논문은 인간의 감독이 학습 루프에서 점차 축소됨에 따라 LRM이 어떻게 지속적으로 개선될 수 있는지를 연구한다. 우리는 이 문제의 두 가지 상호 연결된 차원을 검토한다. 보상 축은 개별 사례에 대한 인간의 판단에서 재사용 가능한 검증기와 인간의 피드백 없이도 작동하는 보상으로의 발전 과정을 추적한다. 경험 축은 인간이 선별한 작업과 환경에서 자체 생성 커리큘럼, 구성된 환경, 자율적 공진화로 학습이 어떻게 발전할 수 있는지를 고찰한다. 우리는 학습 과정의 어느 부분이 지속적인 인간의 통제 하에 남아 있는지를 식별하는 L0부터 L4까지의 5단계 사다리를 통해 이러한 차원들을 연결한다. 우리의 분석은 또한 보상 해킹, 피드백 드리프트, 커리큘럼 붕괴, 환경 오류를 포함하여 점점 더 자율화되는 보상 및 경험 생성이 도입하는 위험을 강조한다. 결과적으로 우리는 정책 능력, 피드백 충실도, 경험 품질이라는 세 가지 상호 보완적 대상에 대한 평가도 제공한다. 이 분석은 인간의 감독을 넘어 LRM을 확장하기 위한 현재의 접근 방식과 초지능을 향한 자립적 학습 시스템 개발과 관련된 미해결 문제에 대한 체계적인 설명을 제공한다. 또한 우리는 최신 발전을 추적하기 위해 지속적으로 업데이트되는 GitHub 저장소(https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision)를 운영하고 있다.
English
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision{GitHub repository} to track the latest advances.