Motif 3: 기술 보고서
Motif 3: Technical Report
August 10, 2026
저자: Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu
cs.AI
초록
우리는 총 3140억 개의 파라미터와 토큰당 132억 개의 활성화 파라미터를 가진 디코더 전용 Mixture-of-Experts 언어 모델인 Motif 3를 소개한다. 각 희소 MoE 레이어는 384개의 라우팅 전문가를 포함하며, 토큰당 8개가 선택된다. 이러한 세분화된 희소성은 연산을 제한하면서도 상당한 전문가 용량을 제공한다. Motif 3는 그룹 차등 잠재 어텐션(GDLA)을 중심으로 구축되었으며, GDLA는 그룹 차등 어텐션과 멀티헤드 잠재 어텐션의 압축된 키-값 표현을 통합한다. 이 아키텍처는 또한 최적화 안정성, 전문가 전문화, 추론 효율성을 개선하기 위해 수정된 매니폴드 제약 하이퍼 연결, 전문가별 PolyNorm 활성화, 다중 토큰 예측을 통합한다. 우리는 웹 문서, STEM, 코드, 수학, 다국어 콘텐츠, 도메인 특화 말뭉치에 걸친 약 12.5조 개의 토큰으로 Motif 3를 사전 학습한다. 전문가 균형 및 수치 안정화 기법은 대규모에서의 안정적인 학습을 지원하며, 선택적 MXFP8 연산 및 통신, 메모리 효율적인 융합 커널, 윈도우 인지 컨텍스트 병렬 처리는 최대 256K 토큰의 컨텍스트 길이로 학습을 가능하게 한다.
우리의 후학습 파이프라인은 일반 지도 미세 조정, 강화 학습으로 학습된 6개의 전문 교사 모델, 지도 미세 조정으로 학습된 소프트웨어 엔지니어링 교사 모델, 다중 교사 온-정책 증류를 결합한다. 그 결과 얻어진 통합 모델은 추론, 코딩, 도구 사용, 전문 작업, 장문맥 이해, 보정된 기권, 지시 수행의 상호 보완적 능력을 결집한다. 광범위한 평가 스위트에서 Motif 3는 주요 오픈 가중치 모델 대비 경쟁력 있는 성능을 입증하며, 특히 장기 지평 에이전트 작업, 수학적 추론, 과학적 지식, 환각에 민감한 평가에서 강력한 결과를 나타낸다.
English
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.