SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

Chung-En Ho, Weiyu Sun, Cheng-Jhih Shih, He Li, Yong Liu, Yingyan Celine Lin

arXiv:2610.04875 · 2026-10-11 공개 · arXiv · PDF

speculative-decoding throughput-optimization diffusion-language-models triton-kernel sparse-execution hidden-state-reuse multi-branch-redundancy residual-gating

Abstract

Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.

한국어 요약

한 줄 요약

SpecFold는 다중 분기 추측 디코딩에서 계산 중복성을 활용해 디퓨전 언어 모델의 처리 속도를 최대 1.99배 향상시키는 알고리즘-시스템 협력 설계이다.

핵심 기여도

핵심 아이디어

기존 디퓨전 언어 모델(DLLM)의 디코딩 가속화 연구는 주로 **시간적 중복**(temporal redundancy)을 활용했으나, 본 연구는 **분기 간 계산 중복**(multi-branch redundancy)이라는 새로운 축을 제시한다. 추측 디코딩 과정에서, **draft branches는 부모 분기에서 대부분의 토큰을 상속**받고, **추가로 언마스크된 위치만 다름**으로 인해, **은닉 상태가 분기 간 매우 유사**하다는 점을 관찰했다.

이를 기반으로, **SpecFold**는 토큰 수준에서 **잔여 게이트**(residual gate)를 적용해 부모 계산을 선택적으로 재사용하고, **folded attention**과 **FFN 재사용**을 통해 중복 계산을 제거한다. 특히, **bidirectional attention** 구조를 고려해 부모의 attention state를 재사용하면서 자식의 query, key, value로 보정하는 방식을 제안한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용