Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, Zhifang Sui

arXiv:2401.07851 · 2026-07-27 공개 · arXiv · PDF

llm-inference speculative-decoding autoregressive-decoding latency-reduction decoding-strategies token-drafting draft-verification llm-efficiency

Abstract

To mitigate the high inference latency stemming from autoregressive decoding in Large Language Models (LLMs), Speculative Decoding has emerged as a novel decoding paradigm for LLM inference. In each decoding step, this method first drafts several future tokens efficiently and then verifies them in parallel. Unlike autoregressive decoding, Speculative Decoding facilitates the simultaneous decoding of multiple tokens per step, thereby accelerating inference. This paper presents a comprehensive overview and analysis of this promising decoding paradigm. We begin by providing a formal definition and formulation of Speculative Decoding. Then, we organize in-depth discussions on its key facets, such as drafter selection and verification strategies. Furthermore, we present a comparative analysis of leading methods under third-party testing environments. We aim for this work to serve as a catalyst for further research on Speculative Decoding, ultimately contributing to more efficient LLM inference.

한국어 요약

한 줄 요약

스펙ulative 디코딩을 통해 대규모 언어 모델의 추론 효율성을 향상시키는 방법을 체계적으로 조사하고 분석한다.

핵심 기여도

핵심 아이디어

스펙ulative 디코딩은 기존의 자동회귀 디코딩 방식에서 토큰을 하나씩 생성하는 대신, 여러 토큰을 먼저 예측(드래프트)하고 병렬로 검증함으로써 추론 속도를 향상시키는 새로운 패러다임이다. 이는 컴퓨터 아키텍처에서 사용되는 "스펙ulative execution" 개념을 언어 모델 추론에 적용한 것이다.

이 접근법은 두 가지 핵심 관찰에 기반한다: 1) 많은 토큰은 작은 모델로도 예측 가능하며, 2) 언어 모델 추론의 주요 병목은 메모리 대역폭이지 연산 자체가 아니다. 따라서, 스펙ulative 디코딩은 메모리 접근을 줄이고, 병렬 검증을 통해 추론 효율성을 높인다.

기술적 접근법

주요 결과

의의 및 한계

스펙ulative 디코딩은 대규모 언어 모델의 추론 속도를 향상시키는 데 기여하며, 실시간 응답이 필요한 서비스에 유용하다. 또한, 메모리 효율성과 병렬 처리 가능성을 제시하며, 추론 인프라 최적화에 기여할 수 있다.

하지만, 드래프트 생성 모델의 정확도와 검증 전략의 선택은 추론 품질에 영향을 미치므로, 정밀도와 속도의 균형을 유지하는 것이 중요하다. 또한, 기존 연구는 다양한 환경에서 평가되어 있어, 표준화된 벤치마크(Spec-Bench)의 필요성이 강조된다.

실용적 활용

스펙ulative 디코딩은 실시간 챗봇, 고객 지원 시스템, 대화형 인터페이스 등에서 추론 속도 향상에 활용될 수 있다. 또한, 클라우드 기반 언어 모델 서비스에서 인프라 비용 절감과 응답 시간 단축에 기여할 수 있다.