ChatPaper.aiChatPaper

NeuroCogMap 揭示大型語言模型的認知組織

NeuroCogMap Reveals Cognitive Organization of Large Language Models

July 1, 2026
作者: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI

摘要

理解人工系統中複雜認知功能的組織方式,是解讀大型語言模型(LLMs)並將其與生物認知連結的核心。然而,儘管LLMs展現出廣泛的認知行為,其內部表徵是否形成可再現的功能系統——足以解釋行為、失敗機制以及與人類認知的關聯——目前仍不明確。本文提出NeuroCogMap,一個受認知神經科學啟發的框架,將LLM的內部特徵組織為功能分區,並將其與可解釋功能、認知能力及認知層級結構相連結。這些分區形成穩定且語意一致的組織,在不同模型間部分保留,並與模型輸出在功能上相關聯。在此組織中,LLM的主要失誤——包括幻覺、偏見、拒絕失敗與諂媚——各自對應於表徵系統與行為控制系統的特定干擾,從而產生可供機制導向檢測與目標性干預的內部特徵。超越模型行為之外,NeuroCogMap能改善對自然語言理解過程中人類大腦皮質反應的預測,其中與高階聯合皮質的對應最為強烈。在認知層面上,其內部特徵揭露了潛在策略,可引導對人類決策經典模型的修正。綜合這些發現,NeuroCogMap建立了一套系統層級的架構,用以繪製人工系統的功能組織,並將此組織與人類大腦皮質功能及認知行為連結起來。
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.