ChatPaper.aiChatPaper

NeuroCogMap揭示大语言模型的认知组织

NeuroCogMap Reveals Cognitive Organization of Large Language Models

July 1, 2026
作者: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI

摘要

理解复杂认知功能在人工系统中的组织方式,对于解释大型语言模型(LLMs)并将其与生物认知相关联至关重要。然而,尽管LLMs展现出广泛的类认知行为,其内部表征是否构成可复现的功能系统,从而解释行为、失败模式及与人类认知的联系,仍不清楚。在此,我们提出NeuroCogMap——一个受认知神经科学启发的框架,将LLMs的内部特征组织成功能分区,并将其与可解释功能、认知能力及认知层级相联系。这些分区形成稳定且语义连贯的组织结构,部分跨模型保守,并与模型输出在功能上关联。在该组织内,LLMs的主要失败模式,包括幻觉、偏见、拒绝失败及谄媚,对应于表征与行为控制系统中的不同干扰,从而产生用于机制引导的检测与定向干预的内部特征。超越模型行为层面,NeuroCogMap提升了自然语言理解过程中人类皮层反应的预测能力,其中高阶联合皮层的对应关系最为显著。在认知层面,其内部特征揭示了潜在策略,指导了对人类决策经典模型的改进。综合而言,这些发现将NeuroCogMap确立为一种系统级框架,用于映射人工系统的功能组织,并将该组织与人类皮层功能和认知行为相关联。
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.