Skaling:Chinchilla 指数与 Kaplan 耦合的交汇
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
August 7, 2026
作者: Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz, Kartik Ahuja
cs.AI
摘要
神经缩放定律是语言模型发展的基础,然而标准公式在数据稀缺和过度训练的极端情况下,系统性地低估和高估了损失。这一失效的根源在于其潜在假设,即模型规模与训练数据对损失的影响是相互独立的。为解决此问题,我们引入了Skaling定律,一种广义函数形式,通过单一交互指数将模型能力与数据耦合起来。这一简单扩展在插值和外推两种机制下均将平均绝对百分比误差(MAPE)降低了1.5至3倍。当与仅限于低计算机制的稀疏网格策略相结合时,Skaling定律可实现准确的完整网格外推,所需计算量约为均匀扫描的十分之一。通过实现从小规模实验中进行可靠的性能预测,Skaling定律为下一代模型训练中的计算预算分配提供了更加稳健且资源高效的框架。
English
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.