母題三:技術報告
Motif 3: Technical Report
August 10, 2026
作者: Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu
cs.AI
摘要
我們介紹 Motif 3,這是一個僅解碼器的混合專家語言模型,擁有3140億總參數,每個 token 激活132億參數。每個稀疏 MoE 層包含384個路由專家,每個 token 選取其中8個。這種細粒度稀疏性在限制計算量的同時,提供了大量的專家容量。Motif 3 以分組差分潛在注意力(GDLA)為核心,該機制將分組差分注意力與多頭潛在注意力的壓縮鍵值表示相結合。此架構進一步整合了改良的流形約束超連接、專家特定 PolyNorm 激活,以及多 token 預測,以提升最佳化穩定性、專家專業化與推理效率。我們在約12.5兆個 token 上對 Motif 3 進行預訓練,涵蓋網頁文件、STEM、程式碼、數學、多語言內容與領域專業語料庫。專家平衡與數值穩定技術支援大規模穩定訓練,而選擇性 MXFP8 計算與通訊、記憶體高效融合核心,以及視窗感知的上下文並行則使訓練可支援高達256K token 的上下文長度。我們的後訓練流程結合了通用監督式微調、六個以強化學習訓練的專家教師、一個以監督式微調訓練的軟體工程教師,以及多教師同策略蒸餾。由此產生的統一模型整合了推理、程式設計、工具使用、專業工作、長上下文理解、校準式棄答與指令遵循等互補能力。在廣泛的評測套件中,Motif 3 與領先的開放權重模型相比,展現出具競爭力的表現,包括在長程代理任務、數學推理、科學知識,以及對幻覺敏感的評測中取得強勁成果。
English
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.