ChatPaper.aiChatPaper

幾何的相補性を超えて:スパース混合エキスパート・ルーティングにおけるコヒーレントな重複

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

July 30, 2026
著者: Huiyuan Tian, Bonan Xu, Shijian Li
cs.AI

要旨

スパース混合エキスパート(MoE)言語モデルは、各トークンを複数のエキスパートにルーティングする。このことは、その利点に関する幾何学的説明を示唆する。すなわち、同時選択されたエキスパートは、互いに異なる表現方向に寄与するはずである。既存の証拠は、ルートの一貫性、候補の品質、および候補と文脈の交互作用をしばしば混同している。我々は、エキスパート部分空間分離指標(ESSI)、マッチングしたルート残差、およびプレフィックスを制御した2×2要因計画を用いて、これらの量を区別する。ルート固定介入と制御されたTop-k研究は、機能的価値を評価する。3つの対応のある対比が所見を整理する。第一に、6つのMoEアーキテクチャにわたって、エキスパート部分空間はかなり重複しているが、実際のルートは、マッチングした代替ルートよりもトークン表現をよく説明する。第二に、OLMoE、Mixtral、DeepSeekの39の要因セルすべてにおいて、選択された候補は、すべてのセルで最も強い非選択の対立候補よりも残差表現を多く説明するが、実際のプレフィックスはこの優位性を全体にわたって縮小する。すなわち、すべての交互作用は負であり、すべての95%信頼区間はゼロを下回る。第三に、この幾何学的縮小は機能的冗長性を意味しない。39のルート固定比較のうち24で、後続のエキスパートの追加が次トークン予測を改善し、残りの15の推定は決定的でない。さらに、制御された訓練研究も、3つのシードすべてでTop-1よりもTop-2を支持する。我々は、この結合パターンを整合的重複(coherent overlap)と呼ぶ。ルーティングは、共有された幾何学的近傍からトークン関連のエキスパートを選択する一方で、有用な多エキスパート計算は、非交差の線形被覆なしに持続する。これらの量を分離することで、幾何学的類似性だけでは冗長性や枝刈り価値を決定できない理由が明確になる。
English
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled 2times2 factorial; frozen-route interventions and a controlled Top-k study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.