ChatPaper.aiChatPaper

Uma Única Camada para Explicá-las Todas: Compreendendo Ativações Massivas em Grandes Modelos de Linguagem

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models

May 8, 2026
Autores: Zeru Shi, Zhenting Wang, Fan Yang, Qifan Wang, Ruixiang Tang
cs.AI

Resumo

Investigamos as origens das ativações massivas em modelos de linguagem de grande porte (LLMs) e identificamos uma camada específica denominada Camada de Emergência Massiva (CEM), que é consistentemente observada em diferentes famílias de modelos, onde as ativações massivas emergem pela primeira vez e subsequentemente se propagam para camadas mais profundas por meio de conexões residuais. Mostramos que, dentro da camada CEM, tanto os parâmetros da RMSNorm quanto os da FFN contribuem conjuntamente para o surgimento das ativações massivas. Uma vez formada, a representação do token de ativação massiva permanece amplamente invariante entre as camadas, reduzindo a diversidade das representações ocultas passadas ao módulo de atenção. Motivados por essa limitação, propomos um método simples e eficaz para reduzir a rigidez do token de ativação massiva. Nossa abordagem melhora consistentemente o desempenho dos LLMs em diversas tarefas, incluindo seguimento de instruções e raciocínio matemático, tanto em contextos livres de treinamento quanto de ajuste fino. Além disso, mostramos que nosso método mitiga os sumidouros de atenção ao enfraquecer seletivamente sua influência, elucidando sua origem no nível de estado oculto e lançando nova luz sobre estratégias de mitigação baseadas em princípios.
English
We investigate the origins of massive activations in large language models (LLMs) and identify a specific layer named the Massive Emergence Layer (ME Layer), that is consistently observed across model families, where massive activations first emerge and subsequently propagate to deeper layers through residual connections. We show that, within the ME Layer both the RMSNorm and the FFN parameters jointly contribute to the emergence of massive activations. Once formed, the massive activation token representation remains largely invariant across layers, reducing the diversity of hidden representations passed to the attention module. Motivated by this limitation, we propose a simple and effective method to reduce the rigidity of the massive activation token. Our approach consistently improves LLM performance across multiple tasks, including instruction following and math reasoning, in both training free and fine tuning settings. Moreover, we show that our method mitigates attention sinks by selectively weakening their influence, elucidating their origin at the hidden state level and shedding new light on principled mitigation strategies.