ChatPaper.aiChatPaper

LittleLearner:教學控制下的知識曝露語言模型

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

August 13, 2026
作者: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
cs.AI

摘要

現代語言模型是在異質化的網路規模文本語料庫上訓練而成。因此,研究知識與技能的習得相當困難,因為模型先前接觸過的相關內容難以界定。為應對此挑戰,我們引入了LITTLECURRICULUM,這是一個精心策劃的880億token預訓練語料庫,專門針對美國小學教材內容設計,明確排除五年級以上所教授的概念、事實與詞彙。僅使用LITTLECURRICULUM從零開始訓練一個50億參數的大型語言模型,即得到LITTLELEARNER,此模型具備足以進行開放式評估的語言能力,但其知識與能力邊界清晰對應於可解釋的課程指引。我們釋出LITTLECURRICULUM與LITTLELEARNER,作為一個發展階段受限的沙盒環境,用以研究模型如何在明確界定的訓練範圍內習得、表徵及運用資料。我們在一系列初步實驗中展示了此沙盒的實用性,這些實驗探討透過後期訓練與情境內學習注入新知識的方法。這些方法使LITTLELEARNER能更有效地運用既有知識,但並未提升超出範圍的能力。我們的研究結果強調了此受控環境對於未來研究的重要價值。
English
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.