게이티드 델타넷(Gated DeltaNet)이 4비트 양자화에서 생존하는 이유: 하이브리드 27B LLM의 순환 절반에 대한 NVFP4 W4A4
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
September 3, 2026
저자: Sergii Kozyrev, Davyd Maiboroda
cs.AI
초록
하이브리드 LLM은 소프트맥스 어텐션과 Gated DeltaNet(GDN) 같은 선형 어텐션 계층들을 결합하며, 이 계층들의 순환 상태는 고정된 크기로 문맥을 요약한다. Qwen3.8-27B(48개 GDN 계층, 16개 어텐션 계층)에 대한 초기 커뮤니티의 4비트 양자화는 GDN 블록, 특히 그 decay 및 write-strength 게이트를 8비트 또는 16비트 정밀도로 유지했다. 순환 구조에서 발생하는 오류가 긴 문맥에 걸쳐 누적될 것이라는 직관 때문이다. 우리는 Minima를 구축하여 이 직관을 시험한다. Minima는 GDN을 포함한 총 496개의 선형 계층에 NVFP4 W4A4를 적용한 모델이다. 4K/32K 문맥에서의 perplexity, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, 64K까지의 RULER 검색 태스크 전반에 걸쳐, Minima는 시드 노이즈 범위 내에서 BF16과 일치하며(5개 과제 평균 -0.52), 비교 대상 중 가장 작은 용량(17.5 GiB)과 가장 빠른 프리필(+14~19%)을 제공한다. 또한 32K perplexity 격차는 위치가 증가함에 따라 줄어든다. 네 부분으로 구성된 메커니즘 연구가 그 이유를 설명한다. (i) NVFP4의 16-요소 블록 스케일링은 잔차 스트림의 극단적 이상치를 국소화하여 계층 역할에 따른 활성화 오류를 평준화한다. (ii) 취약할 것으로 예상되었던 게이트 투영은 실제로는 민감도가 가장 낮다. softplus/exponential 및 sigmoid 파라미터화가 GEMM 오류의 약 11%를 출력 오류의 약 2%로 압축하기 때문이다. (iii) 델타 규칙 순환은 주입된 노이즈를 32K 토큰에 걸쳐 평평한 고원 수준으로 유지하며 수백 스텝 내에 상태 임펄스를 잊는다. 각 쓰기(write) 연산이 현재 키 방향을 따라 상태를 덮어쓰기 때문이다. (iv) 토큰별 양자화 비용은 누적되지 않고 문맥이 길어질수록 희석된다. 또한 우리는 모듈별로 보정된 NVFP4 체크포인트를 해당 모듈들을 단일 GEMM으로 융합하는 커널로 서빙할 때 발생하는 전역 스케일 불일치를 해결하고, 보정된 FP8 KV-캐시 스케일은 성능 손실이 없음을 보인다. 그 결과는 실용적인 방법, 즉 모든 것을 양자화하고 KV 스케일을 포함하는 방식과, 하이브리드 LLM의 순환 절반이 양자화하기 쉬운 절반인 이유에 대한 메커니즘적 설명으로 요약된다. 체크포인트: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4
English
Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by building Minima: NVFP4 W4A4 on all 496 linear layers, GDN included. Across perplexity at 4K/32K, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, and RULER retrieval to 64K, Minima matches BF16 within seed noise (5-task average -0.52) while being the smallest (17.5 GiB) and fastest-prefill (+14-19%) recipe we compare, and its 32K perplexity gap shrinks with position. A four-part mechanism study explains why: (i) NVFP4's 16-element block scaling localizes the residual stream's extreme outliers, equalizing activation error across layer roles; (ii) the supposedly fragile gate projections are the least sensitive -- softplus/exponential and sigmoid parameterizations compress ~11% GEMM error to ~2% output error; (iii) the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens and forgets a state impulse within hundreds of steps, because each write overwrites the state along the current key direction; (iv) the per-token quantization cost washes out with context instead of compounding. We also repair a global-scale mismatch that arises when per-module-calibrated NVFP4 checkpoints are served by kernels that fuse those modules into one GEMM, and show calibrated FP8 KV-cache scales are performance-free. The result: a practical recipe -- quantize everything, ship KV scales -- and a mechanistic account of why the recurrent half of a hybrid LLM is the easy half to quantize. Checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4