EdgeBench: 実世界環境からの学習のスケーリング則を解明する
EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
July 6, 2026
著者: Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi
cs.AI
要旨
事前学習スケーリング則は、データと計算量の増加に伴いモデルの能力が予測可能な形で向上することを示しています。しかし、デプロイ後に実環境から学習するプロセスについては、まだ理解がはるかに進んでいません。134の実環境タスクにおける約38,000時間のエージェントと環境とのインタラクションを分析した結果、我々は、環境学習中の全体的なパフォーマンスが、驚くべき精度(R^2 = 0.998)で対数シグモイドスケーリング則に従うという、我々の知る限り初めての証拠を発見しました。また、モデル世代を超えて、エージェントの学習速度が約3ヶ月ごとに倍増することも明らかになりました。この発見は、科学発見、ソフトウェアエンジニアリング、組合せ最適化、専門知識業務、形式数学、インタラクティブゲームにわたる、超長時間の134の実世界タスクスイートであるEdgeBenchに基づいています。各タスクは、豊富で多層的なフィードバックのもとで少なくとも12時間の連続エージェント動作を維持し、多大な専門家の労力をかけて構築されています。我々は、エージェントが実世界の経験からどのように学習するかという研究を加速するために、51のタスクと評価フレームワーク全体を公開します。
English
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.