ChatPaper.aiChatPaper

EdgeBench: 실세계 환경에서 학습의 스케일링 법칙 규명

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

July 6, 2026
저자: Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi
cs.AI

초록

사전 학습 스케일링 법칙은 모델의 성능이 데이터와 연산량에 따라 예측 가능하게 향상됨을 보여준다. 그러나 배포 후 실제 환경에서의 학습은 여전히 훨씬 덜 이해되고 있다. 134개의 실제 세계 과제에서 약 38,000시간의 에이전트-환경 상호작용을 분석한 결과, 우리가 아는 한, 환경 학습 중 전반적 성능이 놀라울 정도로 높은 정밀도(R² = 0.998)를 보이며 로그-시그모이드 스케일링 법칙을 따른다는 최초의 증거를 발견했다. 또한 모델 세대를 거치면서 에이전트 학습 속도가 약 3개월마다 두 배로 증가함을 발견했다. 이러한 발견은 EdgeBench에서 비롯되었는데, EdgeBench는 과학 발견, 소프트웨어 공학, 조합 최적화, 전문 지식 작업, 형식 수학, 대화형 게임에 걸친 초장기 지평을 가진 134개의 실제 세계 과제 모음이다. 각 과제는 풍부한 다단계 피드백 하에 최소 12시간의 연속적인 에이전트 작동을 유지하며, 상당한 전문가 노력을 통해 구축되었다. 우리는 51개의 과제와 전체 평가 프레임워크를 공개하여 에이전트가 실제 경험으로부터 학습하는 방식을 연구하는 것을 가속화하고자 한다.
English
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.