ChatPaper.aiChatPaper

EdgeBench: Revelando las leyes de escalamiento del aprendizaje en entornos del mundo real

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

July 6, 2026
Autores: Deyao Zhu, Xin Zhou, Shengling Qin, Xuekai Zhu, Hangliang Ding, Shu Zhong, Zixin Wen, Zhonglin Xie, Chenhui Gou, Linxuan Ren, Yueyang Wang, Junfeng Zhong, Rui Liu, Tian Gao, Yangguang Lin, Jingyuan Zhang, Maojia Song, Xuan Qi, Jinhong Wu, Chenyang Zhang, Yinzhu Piao, Ziru Niu, Hongbin Lin, Lingxiang Meng, Peng Tang, Chengyao Tang, Shanyu Wu, Huanyu Zheng, Yu Liu, Liya Zhu, He Wang, Ming Ding, Ziyu Wan, Hao Liu, Sibo Wang, Haotian Zhu, Xintian Zhang, Nan Chai, Yipeng Liu, Panhao Lai, Sihang Yuan, Zixin Su, Ge Zhang, Wangchunshu Zhou, Yantao Du, Wenhao Huang, Guang Shi
cs.AI

Resumen

Las leyes de escalado de preentrenamiento revelan que la capacidad del modelo mejora de manera predecible con los datos y el cómputo. Sin embargo, el aprendizaje en entornos del mundo real tras el despliegue sigue siendo mucho menos comprendido. Analizando aproximadamente 38 000 horas de interacción del agente con el entorno en 134 tareas del mundo real, encontramos, hasta donde sabemos, la primera evidencia de que el rendimiento general durante el aprendizaje ambiental sigue una ley de escalado log-sigmoidal con una precisión notablemente alta, alcanzando un R² = 0,998. A través de generaciones de modelos, también encontramos que la velocidad de aprendizaje del agente se duplica aproximadamente cada tres meses. Este descubrimiento proviene de EdgeBench, un conjunto de 134 tareas del mundo real con horizontes ultra largos, que abarcan descubrimiento científico, ingeniería de software, optimización combinatoria, trabajo profesional de conocimiento, matemáticas formales y juegos interactivos. Cada tarea sostiene al menos 12 horas de operación continua del agente bajo una retroalimentación rica y multinivel, y se construye mediante un esfuerzo experto sustancial. Publicamos 51 tareas y nuestro marco de evaluación completo para acelerar el estudio de cómo los agentes aprenden de la experiencia en el mundo real.
English
Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.