ChatPaper.aiChatPaper

基於超網絡的知識注入在大型語言模型中的縮放定律

Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models

July 21, 2026
作者: Nischay Dhankhar, Dos Baha, Abulhair Saparov
cs.AI

摘要

將事實知識可靠地且大規模地注入大型語言模型仍是一項開放性挑戰。超網絡為大規模知識注入提供了有前景的解決方案。儘管超網絡通常應用於測試時適應,但我們探討其在訓練時知識注入中的應用,即給定一個大規模事實語料庫,訓練一個超網絡來生成固定的LoRA適配器,當該適配器插入目標模型時,使模型能夠回答有關這些事實的問題。本研究探討超網絡是否可用於執行訓練時知識注入,以及此能力如何隨規模變化。超網絡的標度行為迄今仍未得到充分研究。我們的設計將超網絡的注入能力與目標模型的通用能力解耦,從而首次能夠對超網絡架構的標度定律進行嚴謹研究。我們描述了損失、推理準確率以及分佈外泛化如何隨超網絡深度、寬度及目標網絡大小而變化。我們構建了一個大規模數據集,名為MegaWikiQA,其中包含來自Wikidata5M示例構建的39個領域中的數千萬個多跳問答示例。我們的結果顯示:(i)基於超網絡的注入在所有架構維度上表現出廣泛的預測性冪律標度;以及(ii)超網絡能夠在不斷增大的規模下實現可靠的分佈外泛化,表明超網絡為LoRA微調和全微調等其他訓練時適應方法提供了有前景的替代方案,在所有分佈外評估中表現出更陡峭的標度指數。綜合這些結果,超網絡被確立為一種有原則且可擴展的訓練時適應基質,並提供了首個經驗基礎的標度定律,以指導超網絡在大型語言模型中進行事實推理。
English
Injecting factual knowledge into large language models (LLMs) reliably and at scale remains an open challenge. Hypernetworks provide a promising solution to large-scale knowledge injection. Although hypernetworks are typically applied for test-time adaptation, we explore their use in train-time knowledge injection, where, given a large corpus of facts, we train a hypernetwork to generate a fixed LoRA adapter that, when inserted into the target model, enable the model to answer questions about those facts. In this work, we investigate whether hypernetworks can be used to perform train-time knowledge injection and how this ability varies with scale. The scaling behavior of hypernetworks remains largely unstudied. Our design decouples the hypernetwork's injection capacity from the target model's general capability, enabling, for the first time, a rigorous study of scaling laws for hypernetwork architectures. We characterize how loss, reasoning accuracy, and out-of-distribution (OOD) generalization vary with hypernetwork depth, width, and target network size. We construct a large-scale dataset, called MegaWikiQA, containing tens of millions of multi-hop question-answer examples across 39 domains constructed from examples in Wikidata5M. Our results reveal: (i) hypernetwork-based injection exhibits broadly predictive power law scaling along all architecture axes; and (ii) hypernetworks are capable of reliable OOD generalization at increasing scales, suggesting that hypernetwork provides a promising alternative to other train-time adaptation methods such as LoRA finetuning and full fine-tuning, exhibiting steeper scaling exponents in all OOD evaluations. Together, these results establish hypernetworks as a principled and scalable substrate for train-time adaptation, and provide the first empirically grounded scaling laws to guide hypernetworks for factual reasoning in large language models.