SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu, Wenhao Huang, Shen Yan, Jian Li

arXiv:2609.01343 · 2026-09-02 공개 · arXiv · PDF

code-generation in-context-learning scaling-laws parameter-efficiency looped-transformers depth-reuse attention-sink compute-matching

Abstract

Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

한국어 요약

한 줄 요약

SMELT는 계산 예산을 일치시키면서 반복된 중간 레이어를 통해 Transformer 성능을 6.8–18.0% 개선하는 MoE 기반 설계법이다.

핵심 기여도

핵심 아이디어

기존 연구는 반복된 Transformer 레이어가 성능을 향상시킨다고 주장했으나, 추가 FLOPs와 파라미터 증가를 구분하지 못했다. SMELT는 Mixture-of-Experts(MoE) 구조를 활용해 FLOPs, 파라미터, KV 캐시 세 가지 예산을 일치시키며 반복 설계를 연구한다. 핵심 아이디어는 중간 레이어만 반복하고, 두 번 반복하는 것이 성능 향상에 기여한다는 점이다. 이는 반복이 단순히 더 많은 계산을 하게 하는 것이 아니라, 정보를 재사용하고 집중도를 높이는 방식으로 작동한다는 통찰을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

SMELT는 반복 설계가 단순히 추가 계산이 아닌, 구조적 개선을 가능하게 함을 보여준다. 특히, 반복이 attention 메커니즘 내에서 정보 집중도를 높이는 방식으로 작동한다는 점에서 학술적 의의가 있다. 그러나 연구는 200M 규모에서 설계 실험을 수행했기 때문에, 더 큰 모델에서는 최적 반복 수나 범위가 달라질 수 있다. 또한, 반복이 하드웨어 효율성에 미치는 영향은 아직 명시되지 않았다.

실용적 활용

SMELT는 컴퓨트 예산이 제한된 상황에서 더 높은 성능을 추구하는 대형 언어 모델 개발에 적용 가능하다. 특히, 코드 생성, 긴 텍스트 처리, in-context 학습 등 구조화된 데이터 처리에 유용하다.