On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability

Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu

arXiv:2608.30320 · 2026-09-01 공개 · arXiv · PDF

model-architecture parameter-efficiency training-stability muon-optimizer gated-delta-net pretraining-benchmarks sparse-mixture-of-experts n-gram-embedding

Abstract

We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.

한국어 요약

한 줄 요약

Qwen3.8-Flash-Next는 125B 파라미터를 가진 스파스 믹스-오브-익스퍼트 모델로, 효율성과 훈련 안정성을 동시에 향상시킨다.

핵심 기여도

핵심 아이디어

Qwen3.8-Flash-Next는 훈련 효율성과 모델 성능, 훈련 안정성을 동시에 고려한 아키텍처 설계를 목표로 한다. 기존의 full-attention 기반 모델은 파라미터와 FLOPs가 많아 비효율적이라는 문제를 해결하기 위해 GDN과 QSA를 결합한 토큰 믹싱 방식을 도입했다. GDN은 선형 비용으로 prefix를 고정 크기의 상태로 압축하고, QSA는 마이크로 블록 단위로 컨텍스트를 평가하여 인덱싱 비용을 줄인다. GR은 잔류 스트림을 네 가지 브랜치로 확장하고, 요소별 게이트를 통해 안정성을 유지한다. 이는 단순히 파라미터 수를 줄이는 것 이상의 구조적 혁신을 의미한다.

기술적 접근법

주요 결과

의의 및 한계

Qwen3.8-Flash-Next는 훈련 효율성과 모델 성능, 훈련 안정성을 동시에 고려한 설계를 통해 기존 대형 모델의 한계를 극복한다. 특히, GDN과 QSA의 결합은 파라미터와 FLOPs를 줄이면서도 성능을 유지하는 데 성공했다. GR 구조는 훈련 안정성을 향상시키고, Muon 최적화기는 학습률과 배치 크기 최적화를 통해 훈련 과정을 안정화한다. 그러나, n-gram 임베딩의 확장은 호스트 메모리 의존성이 높아 확장성에 한계가 있을 수 있다. 또한, QSA는 긴 컨텍스트에서 우수하지만, 단기 기억 처리 능력은 full-attention과 비교해 다소 제한될 수 있다.

실용적 활용

Qwen3.8-Flash-Next는 대규모 언어 모델이 필요한 산업 분야, 특히 효율성과 비용 절감이 중요한 클라우드 인프라, 대화형 AI, 코드 생성, 다국어 지원 등에 적용 가능하다. 특히, GSA와 GR 구조는 고성능 컴퓨팅 자원이 제한된 환경에서도 안정적으로 작동하며, Muon 최적화기는 대규모 훈련 과정에서도 안정적인 수렴을 보장한다.