ChatPaper.aiChatPaper

人类认知与行为的小型基础模型

Small Foundation Models of Human Cognition and Behaviour

August 5, 2026
作者: Nick Oh, Fernand Gobet
cs.AI

摘要

基于人类行为数据微调的大型语言模型已成为通用认知代理,但所需规模以及这些模型是处理任务结构还是利用统计捷径,仍是开放性问题。我们在Psych-101数据集上训练了十四个模型,参数规模从1.35亿到140亿,涵盖四个架构家族;Psych-101包含来自160项实验的1070万个试次水平选择。在分布内,规模几乎无关紧要。这些模型落在一个狭窄区间内,仿佛触及天花板;0.6B到1B参数就足以在留出被试上匹配70B基线模型。在分布外,这一区间展开为明显更陡的规模梯度,较大的模型在新任务结构的泛化上具有明显优势。为确定这些模型使用哪些信息,我们进行了两项诊断。我们在27项实验中逐步剥离四个提示通道——任务指令、实验刺激、结果反馈和选择历史——并打乱试次顺序。屏蔽刺激和反馈的内容破坏了75.7%的学习信息,并将模型推至低于随机水平,这表明仅凭选择历史无法解释其表现。顺序打乱的结果显示,在试次独立的任务上模型表现不变,但在试次顺序由先前反应决定的任务上则表现出敏感性。因此,小型认知微调模型有望成为心理实验的噪声上限估计器,但其适用范围仍受限于训练中见过的范式。
English
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.