超越幾何互補性:稀疏混合專家路由中的連貫重疊
Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing
July 30, 2026
作者: Huiyuan Tian, Bonan Xu, Shijian Li
cs.AI
摘要
稀疏混合專家(MoE)語言模型將每個 token 路由至多個專家,這暗示其效益可由幾何學觀點解釋:共同被選中的專家應貢獻不同的表徵方向。既有證據往往混淆了路由一致性、候選品質,以及候選與上下文之間的交互作用。我們使用專家子空間分離指數(ESSI)、匹配路由殘差與前綴控制的 2×2 因子設計區分這些量;並以凍結路由干預與受控 Top-k 研究評估其功能價值。三組配對對比組織了研究結果。首先,在六種 MoE 架構中,專家子空間大幅重疊,但實際路由對 token 表徵的解釋力仍優於匹配的替代方案。其次,在 OLMoE、Mixtral 與 DeepSeek 的 39 個因子設計單元中,被選中的候選者對殘差表徵的解釋力皆高於最強的未選中競爭者;然而,實際前綴在每個單元中都縮小了此優勢:所有交互作用皆為負向,且每個 95% 信賴區間都低於零。第三,此幾何縮窄並不意味著功能性冗餘:在 39 個凍結路由比較中,有 24 個顯示添加後續專家能改善下一個 token 預測,其餘 15 個估計則不具結論性;一項受控訓練研究也在全部三個隨機種子中偏好 Top-2 勝於 Top-1。我們將此聯合模式稱為「一致性重疊」:路由從共享的幾何鄰域中選取與 token 相關的專家,而實用的多專家計算仍然存在,無需不相交的線性覆蓋。區分這些量有助於釐清為何僅憑幾何相似性無法決定冗餘性或剪枝價值。
English
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled 2times2 factorial; frozen-route interventions and a controlled Top-k study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.