Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过从人工监督到自动验证奖励和自生成学习环境,解决大规模推理模型在复杂任务中持续改进的问题。
📝 Abstract
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.
Problem

Research questions and friction points this paper is trying to address.

large reasoning models
reinforcement learning with verifiable rewards
open-ended tasks
human supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement learning with verifiable rewards
autonomous co-evolution
self-generated curricula
Zhiqin Yang
Zhiqin Yang
The Chinese University of Hong Kong
Reasoning ModelsCollaborative Learning
Jingwen Fu
Jingwen Fu
Xi'an Jiaotong University
Computer Visionmachine learning
Y
Yuhan Liu
Xi’an Jiaotong University
Hengyu Liu
Hengyu Liu
The Chinese University of Hong Kong
Computer VisionMachine Learning
Y
Yonggang Zhang
The Hong Kong University of Science and Technology
K
Kainan Cao
The University of Hong Kong
Zizhuo Zhang
Zizhuo Zhang
Huazhong University of Science and Technology | Hong Kong Baptist University
LLM Post-TrainingMachine LearningRecommender SystemAI4Science
Chenxin Li
Chenxin Li
The Chinese University of Hong Kong
Multimodal LLMAgentWorld Model
Ruibin Yuan
Ruibin Yuan
HKUST
Artificial IntelligenceMusic GenerationMusic Information RetrievalComputer Music
Jiahao Pan
Jiahao Pan
Hong Kong University of Science and Technology
Speech ProcessingSpeech EnhancmentMusic Generation
J
Jiankai Sun
The Chinese University of Hong Kong
Zhenyuan Zhang
Zhenyuan Zhang
Professor, University of Electronic Science and Technology of China
Power System EquivalentSmart GridArc FlashPower Market
Yibo Li
Yibo Li
National University of Singapore
LLM
Yunlong Lin
Yunlong Lin
Xiamen University
MLLM AgentLarge (Foundation) ModelsInverse Problems
Jing Xiong
Jing Xiong
The University of Hong Kong
Natural Language ProcessingAutomated Theorem Proving
S
Sida Lin
The Hong Kong University of Science and Technology
B
Bo Han
Hong Kong Baptist University
Wei Xue
Wei Xue
Department of Applied Plant Science, Chonnam National University
Crop ecophysiology modellingclimate change
Y
Yike Guo
The Hong Kong University of Science and Technology