NeuroCogMap Revela la Organización Cognitiva de los Grandes Modelos de Lenguaje
NeuroCogMap Reveals Cognitive Organization of Large Language Models
July 1, 2026
Autores: Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
cs.AI
Resumen
Comprender cómo se organizan las funciones cognitivas complejas dentro de los sistemas artificiales es fundamental para interpretar los modelos de lenguaje de gran escala (LLMs) y relacionarlos con la cognición biológica. Sin embargo, aunque los LLMs exhiben comportamientos ampliamente similares a los cognitivos, aún no está claro si sus representaciones internas forman sistemas funcionales reproducibles que expliquen el comportamiento, los fallos y los vínculos con la cognición humana. Aquí presentamos NeuroCogMap, un marco inspirado en la neurociencia cognitiva que organiza las características internas de los LLMs en paquetes funcionales y las vincula con funciones interpretables, capacidades cognitivas y una jerarquía cognitiva. Estos paquetes forman una organización estable y semánticamente coherente que se conserva parcialmente entre modelos y se vincula funcionalmente con las salidas del modelo. Dentro de esta organización, los principales fallos de los LLMs, incluyendo alucinación, sesgo, fallo de rechazo y sicofancia, corresponden a alteraciones distintivas en los sistemas de control representacional y conductual, generando firmas internas para la detección guiada por mecanismos y la intervención dirigida. Más allá del comportamiento del modelo, NeuroCogMap mejora la predicción de las respuestas corticales humanas durante la comprensión naturalista del lenguaje, con la correspondencia más fuerte en la corteza de asociación de alto orden. A nivel cognitivo, sus firmas internas exponen estrategias latentes que guían refinamientos de modelos clásicos de toma de decisiones humanas. En conjunto, estos hallazgos establecen a NeuroCogMap como un marco a nivel de sistema para mapear la organización funcional en sistemas artificiales y para relacionar esta organización con la función cortical humana y el comportamiento cognitivo.
English
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLMs) and relating them to biological cognition. Yet although LLMs exhibit broad cognitive-like behaviours, it remains unclear whether their internal representations form reproducible functional systems that explain behaviour, failure and links to human cognition. Here we present NeuroCogMap, a cognitive neuroscience-inspired framework that organizes internal features of LLMs into functional parcels and links them to interpretable functions, cognitive capabilities and a cognitive hierarchy. These parcels form a stable and semantically coherent organization that is partly conserved across models and functionally linked to model outputs. Within this organization, major LLM failures, including hallucination, bias, refusal failure and sycophancy, correspond to distinct disruptions in representational and behavioural-control systems, yielding internal signatures for mechanism-guided detection and targeted intervention. Beyond model behaviour, NeuroCogMap improves prediction of human cortical responses during naturalistic language comprehension, with the strongest correspondence in higher-order association cortex. At the cognitive level, its internal signatures expose latent strategies that guide refinements of classical models of human decision-making. Together, these findings establish NeuroCogMap as a system-level framework for mapping functional organization in artificial systems and for relating this organization to human cortical function and cognitive behaviour.