面向探索、净化与模型合并的频谱重连
Spectral Rewiring for Exploration, Purification, and Model Merging
July 3, 2026
作者: Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI
摘要
强化学习已成为大型语言模型的标准后训练方法,但密集的全参数更新带来了两个部署相关的瓶颈:一是推理性能受到抑制,通常表现为测试时缩放出现过早饱和;二是在通过多领域训练或模型合并整合多种能力时产生干扰。我们证明,这些更新的推理有效成分主要集中于基模型的谱空间中,由此提出了子空间对齐重连(SAR)——一种事后编辑方法,它在保留谱核心的同时移除正交分量。因此,SAR能够维持推理增益,并过滤掉那些抑制性能或放大跨领域干扰的残余更新方向。在多个模型家族和规模上,SAR仅使用约0.58%的总参数即可提取紧凑的推理核心:它保留了超过99%的后训练性能,改进了数学推理中的高k探索,并在内部模型上通过改进七个开放基准中的六个泛化至智能体编码任务。SAR还能净化混合领域训练更新,在保持数学推理和指令跟随能力的同时释放被抑制的编码能力。此外,它实现了跨专家模型合并,产生的跨领域泛化性能超越了以往的合并基线,甚至超过了最佳的单领域专家。总体而言,SAR表明,从参数几何中提取推理有效更新可以作为无需训练的机制,用于提升推理性能与多领域表现。
English
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilities through multi-domain training or model merging. We show that the reasoning-effective component of these updates is largely concentrated in the base model's spectral space, motivating Subspace-Aligned Rewiring (SAR), a post-hoc editing method that retains this spectral core while removing orthogonal components. SAR therefore preserves reasoning gains and filters residual update directions that suppress performance or amplify cross-domain interference. Across several model families and scales, SAR extracts compact reasoning cores using as little as approximately 0.58% of total parameters: it preserves over 99% of post-training performance and improves high-k exploration in mathematical reasoning, and generalizes to agentic coding by improving six of seven open benchmarks on an in-house model. SAR also purifies mixed-domain training updates by releasing suppressed coding capability while maintaining math reasoning and instruction following. It further enables model merging across experts, yielding cross-domain generalization that surpasses previous merging baselines and even the best single-domain experts. Overall, SAR shows that extracting reasoning-effective updates from parameter geometry can serve as a training-free mechanism to improve reasoning and multi-domain performance.