ChatPaper.aiChatPaper

小型人類認知與行為基礎模型

Small Foundation Models of Human Cognition and Behaviour

August 5, 2026
作者: Nick Oh, Fernand Gobet
cs.AI

摘要

在人類行為數據上微調的大型語言模型,已成為通用認知代理,但這所需的規模,以及這些模型究竟是處理任務結構還是利用統計捷徑,仍是未解的問題。我們在 Psych-101 資料集上訓練了十四個模型,參數量從 1.35 億到 140 億不等,涵蓋四個架構家族;該資料集包含來自 160 個實驗、共 1070 萬次試次級別的選擇。在分佈內,規模幾乎無關緊要。這些模型落在一個狹窄的區間內,彷彿觸及了天花板;0.6B 到 1B 的參數量即足以在留出參與者上匹配 70B 的基線。在分佈外,這個區間展開為明顯更陡峭的縮放梯度,較大的模型在泛化到新穎任務結構時明顯具有優勢。為了確定這些模型使用了哪些資訊,我們執行了兩項診斷。我們在 27 個實驗中逐步剝離四個提示通道——任務說明、實驗刺激、結果回饋與選擇歷史——並置換試次順序。遮蔽刺激與回饋的內容摧毀了 75.7% 的已學習資訊,並使模型低於隨機水準,證明僅憑選擇歷史並不能解釋其表現。置換則揭示了模型在具有獨立試次的任務上具有不變性,但在試次順序由先前反應決定的情況下則表現出敏感性。因此,小型的認知微調模型作為心理學實驗的雜訊上限估計器展現出潛力,儘管其適用範圍仍受限於訓練中所見的範式。
English
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.