탐색, 정제 및 모델 병합을 위한 스펙트럴 재배선
Spectral Rewiring for Exploration, Purification, and Model Merging
July 3, 2026
저자: Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
cs.AI
초록
강화 학습은 대규모 언어 모델의 표준적인 사후 훈련 방법이 되었으나, 조밀한 전체 파라미터 업데이트는 두 가지 배포 관련 병목 현상을 초래한다: 흔히 테스트 시간 스케일링의 조기 포화로 나타나는 억제된 추론 성능, 그리고 다중 도메인 훈련이나 모델 병합을 통해 여러 능력을 통합할 때 발생하는 간섭이다. 본 연구는 이러한 업데이트 중 추론에 효과적인 요소가 주로 기저 모델의 스펙트럼 공간에 집중되어 있음을 보이며, 이를 바탕으로 부분공간 정렬 재배선(SAR)을 제안한다. SAR은 이 스펙트럼 코어를 유지하면서 직교 성분을 제거하는 사후 편집 방법이다. 따라서 SAR은 추론 성능 향상을 보존하고, 성능을 억제하거나 교차 도메인 간섭을 증폭시키는 잔여 업데이트 방향을 걸러낸다. 여러 모델 계열과 규모에 걸쳐 SAR은 전체 파라미터의 약 0.58%만으로 소형 추론 코어를 추출한다: 사후 훈련 성능의 99% 이상을 유지하면서 수학적 추론에서 고-k 탐색을 개선하며, 자체 모델의 7가지 에이전트 코딩 공개 벤치마크 중 6가지에서 성능을 향상시켜 코딩으로 일반화된다. 또한 SAR은 혼합 도메인 훈련 업데이트를 정제하여, 수학 추론 및 지시 따르기 능력을 유지하면서 억제된 코딩 능력을 해제한다. 나아가 전문가 간 모델 병합을 가능하게 하여, 이전 병합 기준선과 최고의 단일 도메인 전문가조차 능가하는 교차 도메인 일반화를 달성한다. 전반적으로, SAR은 파라미터 기하학에서 추론 효과적인 업데이트를 추출하는 것이 추론 및 다중 도메인 성능을 개선하는 훈련 없는 메커니즘으로 사용될 수 있음을 보여준다.
English
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilities through multi-domain training or model merging. We show that the reasoning-effective component of these updates is largely concentrated in the base model's spectral space, motivating Subspace-Aligned Rewiring (SAR), a post-hoc editing method that retains this spectral core while removing orthogonal components. SAR therefore preserves reasoning gains and filters residual update directions that suppress performance or amplify cross-domain interference. Across several model families and scales, SAR extracts compact reasoning cores using as little as approximately 0.58% of total parameters: it preserves over 99% of post-training performance and improves high-k exploration in mathematical reasoning, and generalizes to agentic coding by improving six of seven open benchmarks on an in-house model. SAR also purifies mixed-domain training updates by releasing suppressed coding capability while maintaining math reasoning and instruction following. It further enables model merging across experts, yielding cross-domain generalization that surpasses previous merging baselines and even the best single-domain experts. Overall, SAR shows that extracting reasoning-effective updates from parameter geometry can serve as a training-free mechanism to improve reasoning and multi-domain performance.