DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs

Haokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui, Yingtao Zhang, Linzhan Mou, Linqi Song, Zhenan Sun, Ying Wei

arXiv:2406.01721 · 2026-07-27 공개 · arXiv · PDF

llm-quantization model-quantization low-bit-quantization permutation-transformation activation-smoothing block-wise-rotation outlier-mitigation rotation-transformation

Abstract

Quantization of large language models (LLMs) faces significant challenges, particularly due to the presence of outlier activations that impede efficient low-bit representation. Traditional approaches predominantly address Normal Outliers, which are activations across all tokens with relatively large magnitudes. However, these methods struggle with smoothing Massive Outliers that display significantly larger values, which leads to significant performance degradation in low-bit quantization. In this paper, we introduce DuQuant, a novel approach that utilizes rotation and permutation transformations to more effectively mitigate both massive and normal outliers. First, DuQuant starts by constructing the rotation matrix, using specific outlier dimensions as prior knowledge, to redistribute outliers to adjacent channels by block-wise rotation. Second, We further employ a zigzag permutation to balance the distribution of outliers across blocks, thereby reducing block-wise variance. A subsequent rotation further smooths the activation landscape, enhancing model performance. DuQuant simplifies the quantization process and excels in managing outliers, outperforming the state-of-the-art baselines across various sizes and types of LLMs on multiple tasks, even with 4-bit weight-activation quantization. Our code is available at https://github.com/Hsu1023/DuQuant.

한국어 요약

한 줄 요약

DuQuant은 회전 및 순열 변환을 활용해 4비트 정량화에서도 LLM 성능을 크게 향상시키는 새로운 정량화 방법이다.

핵심 기여도

핵심 아이디어

DuQuant은 기존 정량화 방법이 **Massive Outliers**를 효과적으로 처리하지 못하는 문제를 해결하기 위해 **회전**(rotation)과 **zigzag 순열**(permutation) 변환을 결합한 새로운 접근법을 제안한다. 이는 **outlier의 분포를 채널 간 균일하게 재배치**함으로써 정량화 시 발생하는 정확도 손실을 줄인다.

기존 방법은 **Normal Outliers**를 다루는 데는 효과적이지만, **Massive Outliers**는 값이 매우 크고 토큰 내에서 **제한적으로 나타나는 특성** 때문에 기존의 스무딩 기법이 실패한다. DuQuant은 **outlier 차원을 사전 지식으로 활용한 회전 행렬**(greedy 알고리즘 기반)을 통해 **블록 단위로 outlier를 이웃 채널로 재분배**하고, **zigzag 순열**을 통해 **블록 간 분포 불균형을 완화**한다. 이후 추가 회전을 통해 **활성화 레이어의 스무딩**을 수행함으로써 **정량화 정확도를 향상**시킨다.

기술적 접근법

주요 결과

의의 및 한계

DuQuant은 **4-bit 정량화 환경에서도 뛰어난 성능**을 보이며, **LLM의 에지 기기 배포 가능성**을 높인다. 특히, **Massive Outliers 처리**라는 기존 정량화의 핵심 문제를 해결함으로써, **정량화 과정을 단순화**하고 **성능 저하를 최소화**하는 데 기여한다.

그러나, **Phi2-2.8B 모델**에서는 **쿼리-밸류 행렬 곱셈 불안정성**으로 인해 **정량화 과정에서 오버플로우 발생**하며, **FP 모델 수준의 성능은 달성하지 못**한다. 또한, **DuQuant은 학습 기반 파라미터 조정 없이 calibration 데이터만 사용**하므로, **학습 가능한 가중치 클리핑**(learnable weight clipping)을 적용하지 않아 일부 모델에서 한계가 있을 수 있다.

실용적 활용

DuQuant은 **엣지 기기**(Edge Devices)에서 **LLM 배포**를 가능하게 하며, 특히 **메모리 및 연산 자원이 제한된 환경**에서 유용하다. **LLaMA, Vicuna, Mistral, Phi2 등 다양한 LLM 아키텍처**에 적용 가능하며, **4-bit 정량화 기반의 모바일, 클라우드, IoT 애플리케이션**에 활용할 수 있다.