ChatPaper.aiChatPaper

LittleLearner:教学控制下知识暴露的语言模型

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

August 13, 2026
作者: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel
cs.AI

摘要

现代语言模型在异构的万维网规模文本语料库上进行训练。因此,研究知识与技能的习得颇具挑战性,因为难以刻画模型先前接触过的相关内容。为应对这一挑战,我们引入了LITTLECURRICULUM——一个精选的880亿词元预训练语料库,专为美国小学教材量身定制,明确排除了五年级以上教授的概念、事实和词汇。在LITTLECURRICULUM上从头训练一个拥有50亿参数的LLM,得到了LITTLELEARNER——一个具备足以进行开放式评估的语言能力,同时具有清晰的知识与能力边界的模型,这些边界可映射到可解释的课程大纲。我们发布LITTLECURRICULUM和LITTLELEARNER,作为一个发展受限的沙盒环境,用于研究模型如何在明确定义的训练范围内获取、表征和使用数据。我们通过一系列关于通过后训练和上下文学习注入新知识的初步实验,展示了该沙盒的实用性。这些方法使LITTLELEARNER能够更好地利用现有知识,但并未提升其范围之外的能力。我们的研究结果强调了这一受控环境对未来研究的重要价值。
English
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.