Qwen3.8-Next 아키텍처 설계에 관하여: 평가, 효율성, 그리고 학습 안정성
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
August 31, 2026
저자: Zihan Qiu, Zekun Wang, Xiao Li, Yanpeng Li, Yang Xu, Yixuan Wang, Huaqing Zhang, Rui Men, Bochao Mao, Chengruidong Zhang, Fan Zhou, Hao Luo, Haofeng Huang, Haoran Lian, Haoyan Huang, Hongqing Chen, Jianwei Zhang, Jing Xu, Junjie Wang, Langshi Chen, Liangyu Wang, Linlang Jiang, Man Yuan, Minmin Sun, Peng Jin, Siqi Zhang, Siyu Wang, Xingzhang Ren, Yakai Wang, Yi Zhang, Yiming Dong, Yizhong Cao, Yubo Ma, Yunfei Mao, Bo Zheng, Dayiheng Liu
cs.AI
초록
본 논문은 Qwen3.8-Flash-Next의 아키텍처와 절제(ablation) 실험을 기술한다. 이 모델은 125B 파라미터를 가진 희소 혼합 전문가(sparse mixture-of-experts) 모델로, 토큰당 6B 파라미터가 활성화되며, 추가로 51B 파라미터 규모의 n-gram 임베딩 테이블을 가속기 밖에 보관한다. 열네 개의 사전 학습 벤치마크에서 이 모델은 397B-A17B 전작보다 여덟 개에서 앞서고, 나머지에서는 최대 2.6포인트 뒤처진다. 이는 활성화 파라미터가 1/3, 학습 토큰 수가 1/3, 학습 FLOPs가 약 1/9인 조건에서 달성한 결과다. 토큰 혼합은 Gated DeltaNet(GDN)과 전역 어텐션을 레이어별로 혼합한 방식을 사용하며, 네 개 레이어마다 하나의 전체 어텐션 레이어를 둔다. 지속 사전 학습(continued pretraining) 단계에서는 이 전체 어텐션 레이어를 Qwen 스파스 어텐션(Qwen Sparse Attention, QSA)으로 대체한다. QSA는 압축된 경량 인덱서를 사용해 마이크로 블록 단위로 컨텍스트를 점수화한다. 잔차 스트림은 네 개의 분기로 확장되고 요소별 게이트를 통해 읽히는데, 이 설계를 Gated Residual(GR)이라고 부른다. 백본 외부에는 n-gram 임베딩 레이어 하나를 두어 용량을 추가하며, 해당 테이블은 호스트 메모리에서 프리페치된다. 모든 후보 변경 사항을 세 가지 축으로 평가한다: 손실 및 다운스트림 벤치마크, 학습·프리필·디코딩에서의 변경 비용, 그리고 최적 하이퍼파라미터와 학습 안정성에 미치는 영향. 손실과 다운스트림 정확도는 항상 함께 움직이지는 않는다. n-gram 어휘를 확장하면 손실은 단조롭게 감소하지만 다운스트림 정확도는 포화된다. 이 아키텍처와 Muon 옵티마이저를 함께 사용하면 최적 학습률과 배치 크기가 상향 이동하고, 배치 크기 워밍업이 불필요해지며, 스트레스 테스트에서 안정성이 크게 개선된다. 손실, 벤치마크, 효율성, 안정성은 하나의 설계 문제를 이룬다. 이들을 함께 해결하면 동시에 더 효율적이고, 더 우수하며, 더 안정적인 방법론이 도출된다.
English
We describe the architecture and ablations of Qwen3.8-Flash-Next, a sparse mixture-of-experts model with 125B parameters, 6B activated per token, and additional 51B parameters of n-gram embedding tables held off the accelerator. On fourteen pre-training benchmarks the model leads the 397B-A17B predecessor on eight and trails it on the rest by at most 2.6 points, at 1/3 the activated parameters, 1/3 the training tokens, and roughly 1/9 the training FLOPs. Token mixing uses a layer-wise hybrid of Gated DeltaNet (GDN) and global attention, with one full-attention layer in every four; at continued-pretraining time those full-attention layers are replaced by Qwen Sparse Attention (QSA), which scores context at micro-block granularity with a compressed lightweight indexer. The residual stream is widened to four branches and read through an elementwise gate, a design we call the Gated Residual (GR). Capacity is added outside the backbone by a single n-gram embedding layer whose tables are prefetched from host memory. We evaluate every candidate change along three axes: loss together with downstream benchmarks; the cost of the change in training, prefill and decode; and its effect on the optimal hyperparameters and training stability. Loss and downstream accuracy do not always move together: enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates. The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests. Loss, benchmarks, efficiency and stability form one design problem. Solved jointly, they yield a recipe that is simultaneously more efficient, more capable and more stable.