NeuroCogMapは大規模言語モデルの認知組織を明らかにする
NeuroCogMap Reveals Cognitive Organization of Large Language Models
July 1, 2026
著者: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI
要旨
複雑な認知機能が人工システム内でどのように組織化されているかを理解することは、大規模言語モデル(LLM)を解釈し、それらを生物学的認知と関連付ける上で中心的な課題である。しかしながら、LLMは広範な認知様行動を示すものの、その内部表現が行動、失敗、および人間の認知との関連を説明する再現可能な機能システムを形成しているかどうかは依然として不明である。ここでは、NeuroCogMapを提案する。これは認知神経科学に着想を得たフレームワークであり、LLMの内部特徴を機能的パーセルに整理し、それらを解釈可能な機能、認知能力、および認知階層へと結び付けるものである。これらのパーセルは、安定した意味的に一貫した組織を形成しており、モデル間で部分的に保存され、モデルの出力と機能的に結び付いている。この組織内では、幻覚、バイアス、拒否失敗、迎合(おべっか)などの主要なLLMの失敗は、表象システムおよび行動制御システムにおける明確な混乱に対応しており、メカニズムに基づく検出と標的を絞った介入のための内部シグネチャを生成する。モデルの行動を超えて、NeuroCogMapは自然な言語理解中のヒトの皮質応答の予測を改善し、特に高次連合野で最も強い対応を示す。認知レベルでは、その内部シグネチャは、人間の意思決定の古典的モデルの改良を導く潜在的な戦略を明らかにする。これらの知見は、NeuroCogMapを、人工システムにおける機能組織をマッピングし、この組織をヒトの皮質機能および認知行動に関連付けるためのシステムレベルのフレームワークとして確立するものである。
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.