MISA: Mistura de Indexador com Atenção Esparsa para Inferência em LLM de Contexto Longo
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
May 8, 2026
Autores: Ruijie Zhou, Fanxu Meng, Yufei Xu, Tongxuan Liu, Guangming Lu, Muhan Zhang, Wenjie Pei
cs.AI
Resumo
O DeepSeek Sparse Attention (DSA) estabelece o estado da arte para atenção esparsa em tempo de inferência de alta granularidade ao introduzir um indexador aprendido por token que pontua cada token de prefixo e seleciona os mais relevantes para a atenção principal. Para manter a expressividade, o indexador utiliza múltiplas cabeças de consulta (por exemplo, 64 no DeepSeek-V3.2) que compartilham o mesmo conjunto de tokens selecionados; esse design multicabeça é justamente o que torna o indexador o principal custo em contextos longos. Propomos o MISA (Mixture of Indexer Sparse Attention), uma substituição direta para o indexador DSA que trata suas cabeças de indexação como um conjunto de especialistas mistos (mixture-of-experts). Um roteador leve utiliza estatísticas em nível de bloco para selecionar um subconjunto dependente da consulta, com apenas algumas cabeças ativas, e apenas essas cabeças executam a pesada pontuação em nível de token. Isso preserva a diversidade do pool original do indexador enquanto reduz o custo por consulta, passando de pontuar cada token de prefixo com todas as cabeças para pontuar apenas com um punhado de cabeças roteadas, além de um termo de roteamento insignificante calculado em um pequeno conjunto de chaves agrupadas. Introduzimos ainda uma variante hierárquica do MISA que utiliza a passagem roteada para manter um conjunto de candidatos ampliado e, em seguida, reclassifica-o com o indexador DSA original para recuperar quase exatamente os tokens selecionados finais. Com apenas oito cabeças ativas e sem treinamento adicional, o MISA iguala o desempenho do indexador denso DSA no LongBench tanto para o DeepSeek-V3.2 quanto para o GLM-5, enquanto opera com oito e quatro vezes menos cabeças de indexação, respectivamente, e supera o HISA em média. Ele também preserva mapas de calor totalmente verdes no teste "Needle-in-a-Haystack" em contextos de até 128 mil tokens e recupera mais de 92% dos tokens selecionados pelo indexador DSA por camada. Nosso kernel TileLang proporciona um ganho de velocidade aproximadamente 3,82 vezes maior em relação ao kernel original do indexador DSA em uma única GPU NVIDIA H200.
English
DeepSeek Sparse Attention (DSA) sets the state of the art for fine-grained inference-time sparse attention by introducing a learned token-wise indexer that scores every prefix token and selects the most relevant ones for the main attention. To remain expressive, the indexer uses many query heads (for example, 64 on DeepSeek-V3.2) that share the same selected token set; this multi-head design is precisely what makes the indexer the dominant cost on long contexts. We propose MISA (Mixture of Indexer Sparse Attention), a drop-in replacement for the DSA indexer that treats its indexer heads as a pool of mixture-of-experts. A lightweight router uses cheap block-level statistics to pick a query-dependent subset of only a few active heads, and only those heads run the heavy token-level scoring. This preserves the diversity of the original indexer pool while reducing the per-query cost from scoring every prefix token with every head to scoring it with only a handful of routed heads, plus a negligible router term computed on a small set of pooled keys. We further introduce a hierarchical variant of MISA that uses the routed pass to keep an enlarged candidate set and then re-ranks it with the original DSA indexer to recover the final selected tokens almost exactly. With only eight active heads and no additional training, MISA matches the dense DSA indexer on LongBench across DeepSeek-V3.2 and GLM-5 while running with eight and four times fewer indexer heads respectively, and outperforms HISA on average. It also preserves fully green Needle-in-a-Haystack heatmaps up to a 128K-token context and recovers more than 92% of the tokens selected by the DSA indexer per layer. Our TileLang kernel delivers roughly a 3.82 times speedup over DSA's original indexer kernel on a single NVIDIA H200 GPU.