ChatPaper.aiChatPaper

Skaling:Chinchilla 的指數遇上 Kaplan 的耦合

Skaling: Chinchilla's Exponents Meet Kaplan's Coupling

August 7, 2026
作者: Mathurin Videau, Badr Youbi-Idrissi, David Lopez-Paz, Kartik Ahuja
cs.AI

摘要

神經網絡縮放定律是語言模型發展的基礎,然而標準公式在資料匱乏與過度訓練的極端情境下,會系統性地低估及高估損失。此一缺陷源於其底層假設——模型規模與訓練資料對損失的影響彼此獨立。為了解決這個問題,我們提出 Skaling 定律,這是一個廣義函數形式,透過單一交互指數將模型容量與資料耦合。這個簡單的擴展在內插與外推兩種範圍內,均將平均絕對百分比誤差(MAPE)降低了 1.5 至 3 倍。當與僅限於低計算量範圍的稀疏網格策略搭配使用時,Skaling 定律能以約比均勻掃描少 10 倍的計算量,實現精確的全網格外推。透過從小型實驗中實現可靠的效能預測,Skaling 定律為下一代模型訓練中的計算預算分配提供了一個更穩健且資源高效的框架。
English
Neural scaling laws are foundational for language model development, yet standard formulations systematically under- and overestimate loss at data-scarce and overtraining extremes. This failure originates in the underlying assumption that model size and training data impact the loss independently. To address this, we introduce the Skaling law, a generalized functional form that couples model capacity and data through a single interaction exponent. This simple extension reduces the Mean Absolute Percentage Error (MAPE) by 1.5-3x across both interpolation and extrapolation regimes. When paired with a sparse grid strategy restricted to low-compute regimes, the Skaling law achieves accurate full-grid extrapolation using approximately 10x less compute than uniform sweeps. By enabling reliable performance prediction from small-scale experiments, the Skaling law provides a more robust and resource-efficient framework for allocating compute budgets in next-generation model training.