NeuroCogMap이 대규모 언어 모델의 인지 조직을 밝히다
NeuroCogMap Reveals Cognitive Organization of Large Language Models
July 1, 2026
저자: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI
초록
복합 인지 기능이 인공 시스템 내에서 어떻게 조직화되는지 이해하는 것은 대규모 언어 모델(LLM)을 해석하고 이를 생물학적 인지와 연결 짓는 데 핵심적이다. 그러나 LLM이 광범위한 인지 유사 행동을 보임에도 불구하고, 내부 표상이 행동, 오류 및 인간 인지와의 연관성을 설명하는 재현 가능한 기능 시스템을 형성하는지는 여전히 불분명하다. 본 연구에서 우리는 NeuroCogMap을 제시한다. 이는 인지 신경과학에서 영감을 받은 프레임워크로, LLM의 내부 특징을 기능적 구획으로 조직하고 이를 해석 가능한 기능, 인지 능력 및 인지 계층 구조와 연결한다. 이러한 구획들은 안정적이고 의미론적으로 일관된 조직을 형성하며, 모델 간에 부분적으로 보존되고 모델 출력과 기능적으로 연결된다. 이 조직 내에서, 환각, 편향, 거부 실패 및 아첨 행동을 포함한 주요 LLM 오류는 표상 및 행동 제어 시스템에서 뚜렷한 교란에 해당하며, 이는 메커니즘 기반 탐지 및 표적 개입을 위한 내부 신호를 제공한다. 모델 행동을 넘어, NeuroCogMap은 자연스러운 언어 이해 동안 인간 피질 반응의 예측을 개선하며, 특히 고차 연합 피질에서 가장 강한 상응 관계를 보인다. 인지 수준에서, 그 내부 신호는 인간 의사 결정에 대한 고전적 모델의 개선을 안내하는 잠재적 전략을 드러낸다. 이러한 발견들은 종합적으로 NeuroCogMap을 인공 시스템의 기능적 조직을 매핑하고 이 조직을 인간 피질 기능 및 인지 행동과 연결 짓는 시스템 수준의 프레임워크로 확립한다.
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.