SMELT: 연산량-일치 MoE 루프드 트랜스포머를 위한 스케일링 법칙
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
September 1, 2026
저자: Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu, Wenhao Huang, Shen Yan, Jian Li
cs.AI
초록
루프를 적용한 트랜스포머(Looped Transformers)는 공유된 레이어 블록을 반복함으로써 유효 깊이를 증가시키지만, 대부분의 평가는 모델 크기를 고정한 상태에서 이루어져 구조적 이점과 추가 FLOPs를 혼동한다. 우리는 토큰당 FLOPs, 총 비임베딩 파라미터 수, KV 캐시를 긴밀히 일치시키면서 MoE(Mixture-of-Experts) 트랜스포머에서의 루프(looping)를 연구한다. 일련의 어블레이션을 통해 우리는 SMELT(Sparse MoE Transformer, middle layers Loop Twice)라는 방법을 도출했다. 이는 전체 레이어 중 가운데 절반을 두 번 반복하면서, 세 가지 예산 모두에서 루프를 적용하지 않은 기준(Baseline) 모델과 일치한다. 우리는 SMELT를 비임베딩 파라미터 최대 540억(54B) 개까지 네 가지 규모로 확장하고, 각 아키텍처에 대해 별도의 Chinchilla식 스케일링 법칙을 적합시켰다. SMELT의 손실은 연산량이 증가함에 따라 더 빠르게 감소하여, 계산 최적 경계에서 훈련 FLOPs의 6.8~18.0%를 절약한다. 이러한 이점은 검증 손실이 예측하는 수준을 넘어 다운스트림 벤치마크로 전이되며, 코드 작업에서 가장 크고 샘플 길이와 맥락 내 예제 수가 증가함에 따라 더 커진다. 메커니즘 분석은 두 번째 통과가 어텐션 싱크(attention sink)를 줄이고 어텐션 확률 질량을 내용 관련 토큰으로 재분배함을 보여주는데, 이러한 귀납적 편향이 관찰된 성능 향상의 바탕이 될 수 있다. 이러한 결과는 루프가 예산을 일치시킨 조건에서도 트랜스포머를 개선할 수 있음을 보여주며, 깊이 재사용을 측정 가능한 성과로 전환하는 실용적인 방법을 제공한다.
English
Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.