On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu
arXiv:2608.30320 · 2026-09-01 공개 · arXiv · PDF
model-architecture parameter-efficiency training-stability muon-optimizer gated-delta-net pretraining-benchmarks sparse-mixture-of-experts n-gram-embedding
Abstract
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.
한국어 요약
한 줄 요약
Qwen3.8-Flash-Next는 125B 파라미터를 가진 스파스 믹스-오브-익스퍼트 모델로, 효율성과 훈련 안정성을 동시에 향상시킨다.
핵심 기여도
- Gated DeltaNet (GDN)과 Qwen Sparse Attention (QSA)를 결합한 토큰 믹싱 구조로, 1/3의 활성화 파라미터와 1/9의 FLOPs로 기존 397B-A17B 모델과 유사한 성능 달성.
- Gated Residual (GR)을 도입하여 잔류 스트림을 네 가지 브랜치로 확장하고, 훈련 안정성 향상.
- Muon 최적화기와 함께 학습률과 배치 크기 최적화, 배치 크기 워밍업 불필요화.
- n-gram 임베딩 레이어를 사용해 파라미터 확장 없이 성능 향상.
핵심 아이디어
Qwen3.8-Flash-Next는 훈련 효율성과 모델 성능, 훈련 안정성을 동시에 고려한 아키텍처 설계를 목표로 한다. 기존의 full-attention 기반 모델은 파라미터와 FLOPs가 많아 비효율적이라는 문제를 해결하기 위해 GDN과 QSA를 결합한 토큰 믹싱 방식을 도입했다. GDN은 선형 비용으로 prefix를 고정 크기의 상태로 압축하고, QSA는 마이크로 블록 단위로 컨텍스트를 평가하여 인덱싱 비용을 줄인다. GR은 잔류 스트림을 네 가지 브랜치로 확장하고, 요소별 게이트를 통해 안정성을 유지한다. 이는 단순히 파라미터 수를 줄이는 것 이상의 구조적 혁신을 의미한다.
기술적 접근법
- **토큰 믹싱**: GDN과 QSA의 하이브리드 구조. 4개 레이어 중 1개는 full-attention, 나머지는 GDN 또는 QSA.
- **잔류 스트림**: 네 가지 브랜치로 확장된 GR 구조. 요소별 게이트를 통해 각 브랜치의 기여도 조절.
- **n-gram 임베딩**: 51B 파라미터를 가진 레이어로, 호스트 메모리에서 사전 로드하여 추가 파라미터 확장.
- **최적화기**: Muon 최적화기 사용. 학습률과 배치 크기 최적화, 배치 크기 워밍업 불필요.
- **하이퍼파라미터**: 학습률과 배치 크기 상향 조정으로 훈련 안정성 향상.
주요 결과
- **14개 훈련 벤치마크**: Qwen3.8-Flash-Next는 기존 397B-A17B 모델보다 8개에서 우수, 나머지 6개에서 최대 2.6점 이하로 뒤처짐.
- **활성화 파라미터**: 1/3, 훈련 토큰 1/3, FLOPs 1/9로 동일 수준 성능 유지.
- **QSA 평가**: 8개 벤치마크 중 7개에서 full-attention 기반 모델과 동일 또는 우수한 성능.
- **RULER 벤치마크**: 512K 이상에서 QSA가 full-attention 대비 93.00 (기존 90.08)로 성능 향상.
- **MRCR 벤치마크**: 512K에서 40.53 (기존 30.66), 1M에서 26.44 (기존 20.71)로 향상.
의의 및 한계
Qwen3.8-Flash-Next는 훈련 효율성과 모델 성능, 훈련 안정성을 동시에 고려한 설계를 통해 기존 대형 모델의 한계를 극복한다. 특히, GDN과 QSA의 결합은 파라미터와 FLOPs를 줄이면서도 성능을 유지하는 데 성공했다. GR 구조는 훈련 안정성을 향상시키고, Muon 최적화기는 학습률과 배치 크기 최적화를 통해 훈련 과정을 안정화한다. 그러나, n-gram 임베딩의 확장은 호스트 메모리 의존성이 높아 확장성에 한계가 있을 수 있다. 또한, QSA는 긴 컨텍스트에서 우수하지만, 단기 기억 처리 능력은 full-attention과 비교해 다소 제한될 수 있다.
실용적 활용
Qwen3.8-Flash-Next는 대규모 언어 모델이 필요한 산업 분야, 특히 효율성과 비용 절감이 중요한 클라우드 인프라, 대화형 AI, 코드 생성, 다국어 지원 등에 적용 가능하다. 특히, GSA와 GR 구조는 고성능 컴퓨팅 자원이 제한된 환경에서도 안정적으로 작동하며, Muon 최적화기는 대규모 훈련 과정에서도 안정적인 수렴을 보장한다.