ChatPaper.aiChatPaper

기하학적 상보성을 넘어서: 스파스 혼합 전문가 라우팅에서의 일관된 중첩

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

July 30, 2026
저자: Huiyuan Tian, Bonan Xu, Shijian Li
cs.AI

초록

희소 혼합 전문가(MoE) 언어 모델은 각 토큰을 여러 전문가로 라우팅하며, 이는 그 이점에 대한 기하학적 설명을 시사한다: 함께 선택된 전문가들은 서로 다른 표현 방향을 기여해야 한다는 것이다. 기존 증거는 경로 일관성, 후보 품질, 후보-문맥 상호작용을 종종 혼동한다. 우리는 전문가 부분공간 분리 지수(ESSI), 대응 경로 잔차, 접두사 통제 2×2 요인 설계를 사용하여 이러한 양들을 구분하며, 고정 경로 개입과 통제된 Top-k 연구를 통해 기능적 가치를 평가한다. 세 가지 쌍별 대비가 결과를 정리한다. 첫째, 여섯 가지 MoE 아키텍처 전반에서 전문가 부분공간은 상당히 중첩되지만, 실제 경로는 대응 대안보다 토큰 표현을 더 잘 설명한다. 둘째, OLMoE, Mixtral, DeepSeek의 39개 요인 셀 전반에서 선택된 후보는 모든 셀에서 가장 강한 미선택 경쟁 후보보다 잔차 표현을 더 많이 설명하지만, 실제 접두사는 이 이점을 전반적으로 좁힌다: 모든 상호작용이 음수이며 모든 95% 신뢰 구간이 0보다 아래에 위치한다. 셋째, 이러한 기하학적 축소는 기능적 중복을 의미하지 않는다: 이후 전문가를 추가하는 것은 39개의 고정 경로 비교 중 24개에서 다음 토큰 예측을 개선하는 반면, 나머지 15개의 추정치는 결정적이지 않다; 통제된 훈련 연구는 또한 세 시드 모두에서 Top-2가 Top-1보다 유리함을 보여준다. 우리는 이러한 결합 패턴을 정합적 중첩(coherent overlap)이라고 부른다: 라우팅은 공유된 기하학적 이웃에서 토큰 관련 전문가를 선택하지만, 유용한 다중 전문가 계산은 서로소 선형 포괄(disjoint linear coverage) 없이도 지속된다. 이러한 양들을 분리함으로써 기하학적 유사성만으로는 중복성이나 가지치기 가치를 결정할 수 없는 이유가 명확해진다.
English
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled 2times2 factorial; frozen-route interventions and a controlled Top-k study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.