Motif 3: Technical Report
Junghwan Lim, Joon Son Chung, Sungmin Lee, Wai Ting Cheung, Gihun Cho, Minsu Ha, Sangho Kang, Beomgyu Kim, Dongseok Kim, Jangwoong Kim, Taehyun Kim, Taewhan Kim, Jeesoo Lee, Jeongdoo Lee, Junhyeok Lee, Dongpin Oh, Hyeyeon Cho, Dahye Choi, Jaeheui Her, Hanbin Jung, Changjin Kang, Minjae Kim, Youngrok Kim, Hyukjin Kweon, Hongjoo Lee, Yeongjae Park, Bokki Ryu
arXiv:2608.09119 · 2026-08-11 공개 · arXiv · PDF
long-context code-generation mathematical-reasoning mixture-of-experts model-distillation multi-token-prediction decoder-only sparse-activation
Abstract
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per token. Each sparse MoE layer contains 384 routed experts, with eight selected per token. This fine-grained sparsity provides substantial expert capacity while limiting computation. Motif 3 is built around Grouped Differential Latent Attention (GDLA), which integrates grouped differential attention with the compressed key-value representation of Multi-head Latent Attention. The architecture further incorporates modified manifold-constrained hyper-connections, Expert Specific PolyNorm activations, and multi-token prediction to improve optimization stability, expert specialization, and inference efficiency. We pretrain Motif 3 on approximately 12.5 trillion tokens spanning web documents, STEM, code, mathematics, multilingual content, and domain-specialized corpora. Expert-balancing and numerical-stabilization techniques support stable training at scale, while selective MXFP8 computation and communication, memory-efficient fused kernels, and window-aware context parallelism enable training with context lengths up to 256K tokens. Our post-training pipeline combines general supervised fine-tuning, six specialist teachers trained with reinforcement learning, a software-engineering teacher trained with supervised fine-tuning, and Multi-teacher On-Policy Distillation. The resulting unified model consolidates complementary capabilities in reasoning, coding, tool use, professional work, long-context understanding, calibrated abstention, and instruction following. Across a broad evaluation suite, Motif 3 demonstrates competitive performance against leading open weight models, including strong results on long-horizon agentic tasks, mathematical reasoning, scientific knowledge, and hallucination-sensitive evaluation.
한국어 요약
한 줄 요약
Motif 3는 3140억 파라미터를 가진 MoE 언어 모델로, GDLA와 다양한 최적화 기법을 통해 다중 분야에서 뛰어난 성능을 보인다.
핵심 기여도
- 3140억 파라미터, 1320억 토큰당 활성화 파라미터를 가진 MoE 언어 모델 Motif 3 제안.
- 384개의 전문가 중 8개만 선택하는 GDLA 기반의 고밀도 스파스 아키텍처 도입.
- 12.5트릴리언 토큰으로 사전 학습하며, 256K 토큰 컨텍스트 길이를 지원하는 최적화 기법 제시.
- MOPD를 활용한 7개 전문가 교사 모델의 통합 학습 파이프라인 구축.
핵심 아이디어
Motif 3는 기존 MoE 아키텍처의 계산 비용 문제를 해결하기 위해 GDLA라는 새로운 어텐션 메커니즘을 도입한다. GDLA는 Differential Attention과 Multi-head Latent Attention의 장점을 결합하여 키-밸류 캐시 요구량을 줄이며, 표현력 있는 어텐션 동작을 유지한다. 또한, 384개의 전문가 중 8개만 선택하는 스파스 MoE 구조를 통해 전문가 계산 부담을 줄이고, Expert-Specific PolyNorm 활성화 함수와 Multi-Token Prediction을 통해 최적화 안정성과 추론 효율성을 동시에 향상시킨다. 이는 기존 MoE 모델이 전문가 과부하나 특화 붕괴를 겪는 문제를 완화하는 데 기여한다.
기술적 접근법
- **모델 구성**: 53개의 Transformer 레이어, 4,096 차원의 모델 차원.
- **MoE 구조**: 384개의 전문가, 8개 선택 (13.2억 파라미터/토큰).
- **GDLA**: 80개의 쿼리 헤드, 16개의 KV 헤드, 1,024/512 차원의 저랭크 표현.
- **mHC**: 매니폴드 제약 하이퍼커넥션을 통해 잔차 활성화 이상치를 제어.
- **트레이닝 최적화**: MXFP8, 윈도우 인식 컨텍스트 병렬 처리, 메모리 효율 퓨즈 커널 적용.
- **데이터**: 12.5트릴리언 토큰, STEM, 코드, 수학, 다국어, 도메인 전문 데이터 포함.
- **포스트-트레이닝**: SFT, 6개의 RL 교사, 1개의 SFT 교사, MOPD를 통한 통합 학습.
주요 결과
- **Agentic Task**: τ³-Banking 35.3, Terminal-Bench 2.1 74.9, τ²-Bench Telecom 94.7.
- **Professional Task**: GDPval-AA v2 38.7, ITBench-AA 51.5 (최고).
- **코드 및 수학**: SWE-bench Verified 76.2, SciCode 40.6.
- **논리 추론**: IMO-AnswerBench 83.2, GPQA Diamond 83.4.
- **Hallucination 제어**: AA-Omniscience 정확도 30.1, non-hallucination 71.6 (상위권).
- **긴 컨텍스트**: AA-LCR 72.3, IFBench 78.2.
의의 및 한계
Motif 3는 대규모 MoE 모델의 안정적 학습과 효율적 추론을 가능하게 하는 GDLA와 MOPD 기반 통합 학습 파이프라인을 제시하며, 기존 오픈 모델 대비 뛰어난 성능을 보인다. 특히, 256K 토큰 길이의 컨텍스트 처리와 전문가 균형 조절 기법은 대규모 모델의 학습 안정성 향상에 기여한다. 그러나 SciCode와 CritPt에서 상대적으로 낮은 성능을 보이는 점은 과학적 코드 생성 및 전문 과학 추론 분야에서의 개선 필요성을 시사한다. 또한, 모델의 복잡성은 추론 및 학습에 높은 하드웨어 요구를 동반한다.
실용적 활용
Motif 3는 긴 컨텍스트 이해, 도구 사용, 전문 작업, 수학 및 과학 추론 등이 필요한 산업 현장에서 활용 가능하다. 특히, IT, 금융, 법률 분야의 복잡한 작업 자동화 및 대화형 시스템 구축에 적합하며, 학술 연구에서의 과학적 추론 및 코드 생성에도 유용할 수 있다.