llm model-compression post-training-quantization llama2 optimal-splitting high-accuracy-inference time-efficiency binarization
Abstract
Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive computation and memory requirements. However, existing quantization techniques fall short of maintaining LLM performance under ultra-low bit-widths. In response to this challenge, we present BiLLM, a groundbreaking 1-bit post-training quantization scheme tailored for pretrained LLMs. Based on the weight distribution of LLMs, BiLLM first identifies and structurally selects salient weights, and minimizes the compression loss through an effective binary residual approximation strategy. Moreover, considering the bell-shaped distribution of the non-salient weights, we propose an optimal splitting search to group and binarize them accurately. BiLLM achieving for the first time high-accuracy inference (e.g. 8.41 perplexity on LLaMA2-70B) with only 1.08-bit weights across various LLMs families and evaluation metrics, outperforms SOTA quantization methods of LLM by significant margins. Moreover, BiLLM enables the binarization process of the LLM with 7 billion weights within 0.5 hours on a single GPU, demonstrating satisfactory time efficiency. Our code is available at https://github.com/Aaronhuang-778/BiLLM.
한국어 요약
한 줄 요약
BiLLM은 1.08-bit 평균 가중치로 LLaMA2-70B 모델에서 8.41의 퍼플렉시티를 달성한 1-bit 후학습 양자화 기법이다.
핵심 기여도
- 구조적 중요 가중치를 Hessian 기반 메트릭으로 선택하고, 잔차 근사법(residual approximation)을 통해 정밀도 손실 최소화.
- 비중요 가중치에 대해 최적 분할 검색(optimal splitting search)을 도입하여 정확한 이진화 수행.
- LLaMA2-70B 모델에서 1.08-bit 평균 가중치로 8.41의 퍼플렉시티를 달성하며, 기존 최고 성능을 49.4% 이상 개선.
- 70억 개의 가중치를 가진 LLM을 단일 GPU에서 0.5시간 이내에 이진화 가능.
핵심 아이디어
BiLLM은 LLM의 가중치 분포 특성을 기반으로, 중요 가중치와 비중요 가중치를 구분하여 각각 다른 이진화 전략을 적용한다. 중요 가중치는 Hessian 행렬을 기반으로 선정하며, 잔차 근사법을 통해 정밀도 유지와 저장 공간 절약의 균형을 맞춘다. 비중요 가중치는 종벨형 분포를 고려하여 최적 분할 검색을 통해 정확한 이진화를 수행한다. 이는 기존 1-bit 양자화 기법이 정밀도를 유지하지 못하는 문제를 해결하기 위한 핵심 아이디어이다.
기술적 접근법
- **구조적 중요 가중치 선택**: Hessian 기반 메트릭을 사용하여 중요 가중치를 선정.
- **잔차 근사법**: 중요 가중치에 대해 잔차를 근사하여 정밀도 손실 최소화.
- **최적 분할 검색**: 비중요 가중치의 종벨형 분포를 고려하여 최적의 분할점을 찾고, 그룹별로 이진화.
- **블록 기반 오류 보상**: 기존 방법과 동일하게 블록 단위로 오류 보상 수행.
- **하드웨어 효율성**: 70억 개의 가중치를 가진 모델을 0.5시간 이내에 단일 GPU에서 처리 가능.
주요 결과
- **LLaMA2-70B**: 1.08-bit 평균 가중치로 8.41 퍼플렉시티 (FP16 OPT-66B 대비 +9.34 개선).
- **OPT-30B**: 1.1-bit 평균 가중치로 PB-LLM(1.7-bit) 대비 49.4% 정밀도 개선.
- **Vicuna-7B/13B**: 1.08-bit 평균 가중치로 GPTQ 대비 49.4%~77.0% 성능 향상.
- **메모리 절약**: BiLLM은 2-bit GPTQ 대비 69.9%의 메모리 점유율을 기록.
의의 및 한계
BiLLM은 LLM의 1-bit 후학습 양자화 기술에서 기존 한계를 돌파한 첫 사례로, 1.08-bit 평균 가중치로 높은 정밀도를 유지하는 것을 입증하였다. 이는 LLM의 에지 기기 및 자원 제한 환경에서의 실용적 배포 가능성을 높인다. 그러나 현재는 주로 오픈소스 LLM에만 적용되었으며, 상용 모델에서의 성능 검증이 필요하다. 또한, GEMM 연산의 이진화 구현이 어려운 점은 실제 하드웨어 적용 시 한계로 작용할 수 있다.
실용적 활용
BiLLM은 GPU 메모리 절약과 빠른 양자화 시간을 통해 모바일 기기, IoT, 클라우드 엣지 서버 등 자원 제한 환경에서 LLM 배포를 가능하게 한다. 특히, Vicuna와 같은 조정된 지시 모델에도 적용 가능하여, 다양한 응용 시나리오에서 활용성이 높다.