Motif 3: Technical Report

Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu

arXiv:2608.09119 · 2026-08-11 공개 · arXiv · PDF

long-context code-generation mathematical-reasoning mixture-of-experts model-distillation multi-token-prediction decoder-only sparse-activation

Abstract

We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.

한국어 요약

한 줄 요약

Motif 3는 3140억 파라미터를 가진 MoE 언어 모델로, GDLA와 다양한 최적화 기법을 통해 다중 분야에서 뛰어난 성능을 보인다.

핵심 기여도

핵심 아이디어

Motif 3는 기존 MoE 아키텍처의 계산 비용 문제를 해결하기 위해 GDLA라는 새로운 어텐션 메커니즘을 도입한다. GDLA는 Differential Attention과 Multi-head Latent Attention의 장점을 결합하여 키-밸류 캐시 요구량을 줄이며, 표현력 있는 어텐션 동작을 유지한다. 또한, 384개의 전문가 중 8개만 선택하는 스파스 MoE 구조를 통해 전문가 계산 부담을 줄이고, Expert-Specific PolyNorm 활성화 함수와 Multi-Token Prediction을 통해 최적화 안정성과 추론 효율성을 동시에 향상시킨다. 이는 기존 MoE 모델이 전문가 과부하나 특화 붕괴를 겪는 문제를 완화하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

Motif 3는 대규모 MoE 모델의 안정적 학습과 효율적 추론을 가능하게 하는 GDLA와 MOPD 기반 통합 학습 파이프라인을 제시하며, 기존 오픈 모델 대비 뛰어난 성능을 보인다. 특히, 256K 토큰 길이의 컨텍스트 처리와 전문가 균형 조절 기법은 대규모 모델의 학습 안정성 향상에 기여한다. 그러나 SciCode와 CritPt에서 상대적으로 낮은 성능을 보이는 점은 과학적 코드 생성 및 전문 과학 추론 분야에서의 개선 필요성을 시사한다. 또한, 모델의 복잡성은 추론 및 학습에 높은 하드웨어 요구를 동반한다.

실용적 활용

Motif 3는 긴 컨텍스트 이해, 도구 사용, 전문 작업, 수학 및 과학 추론 등이 필요한 산업 현장에서 활용 가능하다. 특히, IT, 금융, 법률 분야의 복잡한 작업 자동화 및 대화형 시스템 구축에 적합하며, 학술 연구에서의 과학적 추론 및 코드 생성에도 유용할 수 있다.