人間の認知と行動の小型基盤モデル
Small Foundation Models of Human Cognition and Behaviour
August 5, 2026
著者: Nick Oh, Fernand Gobet
cs.AI
要旨
人間の行動データでファインチューニングされた大規模言語モデルは、汎用的な認知プロキシとして登場してきたが、それに必要なスケール、およびこれらのモデルが課題構造を処理しているのか、それとも統計的ショートカットを利用しているのかは、依然として未解決の問いである。我々は、160の実験における1070万件の試行レベルの選択からなるデータセットPsych-101を用いて、4つのアーキテクチャファミリーにわたる135M〜14Bパラメータの14モデルを訓練した。分布内では、スケールはほとんど影響しない。モデルは、あたかも天井に突き当たっているかのように狭い範囲に収まり、0.6B〜1Bパラメータで、未学習の被験者において70Bのベースラインに匹敵するのに十分である。分布外では、その範囲は著しく急峻なスケーリング勾配へと広がり、より大きなモデルが新奇な課題構造への汎化において明確に有利となる。これらのモデルがどの情報を用いているのかを特定するため、我々は2つの診断を実施した。27の実験にわたり、4つのプロンプトチャネル(課題教示、実験刺激、結果フィードバック、選択履歴)を段階的に除去し、試行順序を並べ替えた。刺激とフィードバックの内容をマスクすると、学習された情報の75.7%が破壊され、モデルは偶然水準以下に押し下げられた。これは、選択履歴だけではパフォーマンスを説明できないことを示している。並べ替えでは、独立な試行からなる課題では不変性が見られる一方、試行順序が以前の反応によって決定される課題では感受性が見られた。したがって、小規模な認知ファインチューニングモデルは、心理学実験のノイズ天井推定器として有望であるが、その適用範囲は訓練中に観察されたパラダイムに限定される。
English
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. In-distribution, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.