Loop the Loopies!

Zitian Gao, Yilong Chen, Yihao Xiao, Xinyu Yang, Ran Tao, Joey Zhou, Bryan Dai

arXiv:2607.16051 · 2026-07-20 공개 · arXiv · PDF

mixture-of-experts post-training parameter-efficiency ablation-study looped-transformer large-model reasoning-abilities imo

Abstract

We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6Bparameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N-fold increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. Our novel post-training pipeline equips Loopie with strong reasoning abilities. At the 2025 IMO and IPhO, Loopie achieves gold-medal performance without tools.

한국어 요약

한 줄 요약

Loopie는 기존 루프형 트랜스포머의 한계를 극복한 고성능 MoE 모델이다.

핵심 기여도

핵심 아이디어

기존 루프형 트랜스포머는 동일 컴퓨트 예산 내에서 파라미터 수를 늘리는 것이 루프 수를 늘리는 것보다 더 효과적이라는 문제를 안고 있었다. Loopie는 이 문제를 해결하기 위해 Mixture-of-Experts(MoE) 구조를 도입하고, 20B-파라미터 모델과 6B-파라미터 모델로 시리즈를 구성함. 이는 루프 수 대신 파라미터 수를 늘리는 전략과 비교해 동일한 컴퓨트 예산 내에서 더 나은 성능을 보장한다. 또한, Loopie는 추론 능력을 강화하기 위해 새로운 포스트-트레이닝 파이프라인을 도입함.

기술적 접근법

주요 결과

의의 및 한계

Loopie는 루프형 트랜스포머의 성능 한계를 극복한 첫 번째 MoE 기반 모델로, 대규모 모델의 효율적 설계에 기여함. 또한, 포스트-트레이닝 파이프라인을 통해 추론 능력을 강화한 점이 학술적·실용적 가치를 높임. 그러나 Loopie의 구체적인 학습 데이터셋이나 하이퍼파라미터는 명시되지 않아 재현성 측면에서 한계가 있음.

실용적 활용

Loopie는 과학·수학 분야의 고난도 문제 해결, 대규모 언어 모델의 효율적 설계, 컴퓨트 예산이 제한된 환경에서의 모델 최적화에 활용될 수 있음.