Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Sergii Kozyrev, Davyd Maiboroda

arXiv:2609.04098 · 2026-09-04 공개 · arXiv · PDF

long-context kv-cache delta-rule nvfp4 quantization-error gated-deltanet w4a4-quantization mmlu-pro

Abstract

Hybrid LLMs pair softmax attention with linear-attention layers such as Gated DeltaNet (GDN), whose recurrent state summarizes the context in fixed size. Early community 4-bit quantizations of Qwen3.8-27B (48 GDN layers, 16 attention layers) left the GDN block in 8- or 16-bit precision -- especially its decay and write-strength gates -- on the intuition that errors in a recurrence accumulate over long contexts. We test that intuition by building Minima: NVFP4 W4A4 on all 496 linear layers, GDN included. Across perplexity at 4K/32K, MMLU-Pro, GSM8K, AIME'25, GPQA-Diamond, LiveCodeBench, and RULER retrieval to 64K, Minima matches BF16 within seed noise (5-task average -0.52) while being the smallest (17.5 GiB) and fastest-prefill (+14-19%) recipe we compare, and its 32K perplexity gap shrinks with position. A four-part mechanism study explains why: (i) NVFP4's 16-element block scaling localizes the residual stream's extreme outliers, equalizing activation error across layer roles; (ii) the supposedly fragile gate projections are the least sensitive -- softplus/exponential and sigmoid parameterizations compress ~11% GEMM error to ~2% output error; (iii) the delta-rule recurrence holds injected noise at a flat plateau over 32K tokens and forgets a state impulse within hundreds of steps, because each write overwrites the state along the current key direction; (iv) the per-token quantization cost washes out with context instead of compounding. We also repair a global-scale mismatch that arises when per-module-calibrated NVFP4 checkpoints are served by kernels that fuse those modules into one GEMM, and show calibrated FP8 KV-cache scales are performance-free. The result: a practical recipe -- quantize everything, ship KV scales -- and a mechanistic account of why the recurrent half of a hybrid LLM is the easy half to quantize. Checkpoint: https://huggingface.co/minima-ai/mnma_qwen3.8_27b_nvfp4

한국어 요약

한 줄 요약

Gated DeltaNet(GDN) 기반의 Qwen3.8-27B 모델을 NVFP4 W4A4로 전체 496개의 선형층을 4비트로 양자화하여 BF16 수준 성능을 달성한 연구.

핵심 기여도

핵심 아이디어

Gated DeltaNet(GDN)은 고정 크기의 재귀 상태로 컨텍스트를 요약하는 선형 어텐션 레이어로, 기존 연구에서는 재귀 오류가 누적될 수 있다고 가정해 8비트 이상으로 보호했다. 그러나 본 연구는 NVFP4 W4A4 양자화를 통해 GDN 전체를 4비트로 처리해도 성능 저하 없이 작동함을 입증했다.

GDN의 재귀 상태 업데이트는 $ S_t = \alpha_t S_{t-1} + \beta_t k_t (v_t - S_{t-1}^\top k_t)^\top $ 형태로, 각 단계에서 상태가 덮어쓰여 오류가 누적되지 않도록 설계되어 있다. 특히, gate parameterization(softplus/exponential, sigmoid)가 GEMM 오류를 약 11%에서 2%로 압축하며, delta-rule이 오류를 효과적으로 제거한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용