ChatPaper.aiChatPaper

超越几何互补性:稀疏混合专家路由中的相干重叠

Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routing

July 30, 2026
作者: Huiyuan Tian, Bonan Xu, Shijian Li
cs.AI

摘要

稀疏混合专家(MoE)语言模型将每个token路由到多个专家,这提示了一个关于其优势的几何解释:共同被选中的专家应当贡献不同的表征方向。已有证据往往将路由一致性、候选质量和候选-上下文交互混为一谈。我们通过专家子空间分离指数(ESSI)、匹配路由残差和前缀控制的2×2因子设计来区分这些量;并通过冻结路由干预和受控Top-k研究来评估其功能价值。三组配对对比构成了研究发现的核心。 第一,在六种MoE架构中,专家子空间存在实质性重叠,但实际路由对token表征的解释优于匹配的替代方案。第二,在OLMoE、Mixtral和DeepSeek的39个因子设计单元中,所选候选在每个单元中对残差表征的解释均强于最强的未选中竞争者,然而实际前缀在所有单元中均收窄了这一优势:所有交互作用均为负值,且每个95%置信区间均低于零。第三,这种几何上的收窄并不意味着功能冗余:在39次冻结路由比较中,增加后续专家改善了其中24次的下一个token预测,其余15次估计不具结论性;一项受控训练研究在三个随机种子中也一致倾向于Top-2优于Top-1。我们将这一联合模式称为一致性重叠:路由从共享的几何邻域中选择与token相关的专家,同时,多专家计算的有用性在不具有不相交线性覆盖的情况下得以持续。区分这些量有助于阐明为何几何相似性本身无法决定冗余性或剪枝价值。
English
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction. We distinguish these quantities using an Expert Subspace Separation Index (ESSI), matched-route residuals, and a prefix-controlled 2times2 factorial; frozen-route interventions and a controlled Top-k study assess functional value. Three paired contrasts organize the findings. First, across six MoE architectures, expert subspaces overlap substantially, yet actual routes explain token representations better than matched alternatives. Second, across the 39 factorial cells in OLMoE, Mixtral, and DeepSeek, the selected candidate explains more of the residual representation than the strongest unselected rival in every cell, yet the actual prefix narrows this advantage throughout: all interactions are negative, and every 95% confidence interval lies below zero. Third, this geometric narrowing does not imply functional redundancy: adding later experts improves next-token prediction in 24 of 39 frozen-route comparisons, while the other 15 estimates are inconclusive; a controlled training study also favors Top-2 over Top-1 in all three seeds. We call this joint pattern coherent overlap: routing selects token-relevant experts from a shared geometric neighborhood, while useful multi-expert computation persists without disjoint linear coverage. Separating these quantities clarifies why geometric similarity alone cannot determine redundancy or pruning value.