ChatPaper.aiChatPaper

NeuroCogMap onthult cognitieve organisatie van grote taalmodellen

NeuroCogMap Reveals Cognitive Organization of Large Language Models

July 1, 2026
Auteurs: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI

Samenvatting

Het begrijpen van hoe complexe cognitieve functies zijn georganiseerd binnen kunstmatige systemen is essentieel voor het interpreteren van grote taalmodellen (LLMs) en het relateren ervan aan biologische cognitie. Hoewel LLMs breed cognitief-achtig gedrag vertonen, blijft het onduidelijk of hun interne representaties reproduceerbare functionele systemen vormen die gedrag, falen en verbanden met menselijke cognitie verklaren. Hier presenteren we NeuroCogMap, een raamwerk geïnspireerd op de cognitieve neurowetenschappen dat interne kenmerken van LLMs in functionele parcelen indeelt en deze koppelt aan interpreteerbare functies, cognitieve vermogens en een cognitieve hiërarchie. Deze parcelen vormen een stabiele en semantisch coherente organisatie die gedeeltelijk behouden blijft over modellen heen en functioneel verbonden is met modeluitvoer. Binnen deze organisatie corresponderen belangrijke LLM-storingen, waaronder hallucinatie, bias, weigeringsfalen en vleierij, met specifieke verstoringen in representatie- en gedragscontrolesystemen, wat leidt tot interne signaturen voor mechanismegestuurde detectie en gerichte interventie. Naast modelgedrag verbetert NeuroCogMap de voorspelling van menselijke corticale responsen tijdens naturalistisch taalbegrip, met de sterkste overeenkomst in de hogere-orde-associatiecortex. Op cognitief niveau onthullen de interne signaturen latente strategieën die verfijningen van klassieke modellen van menselijke besluitvorming sturen. Samen vestigen deze bevindingen NeuroCogMap als een raamwerk op systeemniveau voor het in kaart brengen van functionele organisatie in kunstmatige systemen en voor het relateren van deze organisatie aan menselijke corticale functie en cognitief gedrag.
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.