スケーリング:Chinchillaの指数とKaplanの結合
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
August 7, 2026
著者: Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz, Kartik Ahuja
cs.AI
要旨
ニューラルスケーリング則は言語モデルの開発の基盤であるが、標準的な定式化は、データ不足の領域と過学習の極端な領域において損失を系統的に過小評価および過大評価する。この失敗は、モデルサイズと学習データが損失に独立に影響するという根底にある仮定に起因する。この問題に対処するため、我々は、モデルの容量とデータを単一の相互作用指数を通じて結合する一般化された関数形式であるSkaling則を導入する。この単純な拡張により、内挿および外挿の両領域において平均絶対パーセント誤差(MAPE)が1.5〜3倍減少する。低計算量領域に制限されたスパースグリッド戦略と組み合わせると、Skaling則は一様なスイープよりも約10分の1の計算量で正確なフルグリッド外挿を達成する。小規模実験からの信頼性の高い性能予測を可能にすることにより、Skaling則は、次世代モデルの学習における計算予算の配分のための、より堅牢でリソース効率の高い枠組みを提供する。
English
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.