ISO: RLVR-네이티브 최적화 스택
ISO: An RLVR-Native Optimization Stack
July 21, 2026
저자: Hanqing Zhu, Wenyan Cong, Zhizhou Sha, Sagnik Mukherjee, Xinyuan Song, David González-Martínez, Xiaoxia Wu, Yuandong Tian, Shiwei Liu, David Z. Pan, Zhangyang "Atlas" Wang
cs.AI
초록
검증 가능한 보상을 통한 강화 학습(RLVR)은 언어 모델의 추론 능력을 빠르게 향상시키고 있지만, 보상 피드백을 가중치 공간 업데이트로 변환하는 최적화 계층에 대한 이해는 여전히 부족하다. 우리의 이전 분석(Zhu et al., 2025)을 바탕으로, 우리는 모델 가중치의 특이 구조를 통해 이 누락된 계층을 연구하고 스펙트럼 상속을 확인한다: RLVR은 관련 입력 및 출력 특이 프레임의 변화를 통해 새로운 행동을 획득하면서 기본 모델의 가중치 스펙트럼을 재사용할 수 있다.
우리는 스펙트럼 상속을 등스펙트럼 최적화(ISO)로 구현한다. 이는 상호 보완적인 오프라인 및 온라인 구현체를 갖춘 RLVR 고유의 고정 스펙트럼 최적화 프레임워크이다. 오프라인에서 ISO-Merger는 공유 기본 모델을 가진 전문가들의 프레임 변화를 단일 고정 스펙트럼 모델로 결합하며, 병합 후 데이터, 롤아웃, 그래디언트 업데이트 또는 자기 정책 증류(OPD)가 필요하지 않다. 이는 상호 보완적인 전문가 능력을 복원하고 비교된 데이터 없는 병합 방법 중 가장 강력한 종합 성능을 달성한다. 온라인에서 ISO-Optimizer는 기본 스펙트럼을 고정한 상태에서 프레임 변수에 AdamW 및 Muon을 포함한 선택된 기본 최적화기를 적용한다. 1.5B에서 8B 파라미터 범위의 추론 및 코딩 작업에서 ISO-Optimizer는 보고된 실행에서 정확도를 향상시키고 훨씬 적은 학습 단계로 일치하는 점수에 도달한다. Qwen3-8B-Base에서 AdamW는 270 학습 단계 후에 종합 정확도 0.495에 도달한다. ISO-AdamW는 단 100 학습 단계 후에 동일한 정확도에 도달하고 210 학습 단계 후에는 0.509로 더 향상된다. 종합적으로, ISO는 RLVR의 누락된 최적화 계층에 대한 구체적인 답을 제시한다: 사전 학습 최적화를 통째로 상속하는 대신, 보상 기반 적응의 구조에 맞게 사후 학습을 설계하라: 스펙트럼을 상속하고 프레임을 최적화하라.
English
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames.
We operationalize spectral inheritance as Isospectral Optimization (ISO), an RLVR-native, fixed-spectrum optimization framework with complementary offline and online instantiations. Offline, ISO-Merger combines the frame changes of shared-base specialists into a single fixed-spectrum model, requiring no post-merge data, rollouts, gradient updates, or on-policy distillation (OPD). It recovers complementary specialist capabilities and achieves the strongest aggregate performance among the compared data-free merging methods. Online, ISO-Optimizer applies a chosen base optimizer, including AdamW and Muon, to the frame variables while keeping the base spectra fixed. Across reasoning and coding tasks ranging from 1.5B to 8B parameters, ISO-Optimizer improves accuracy in the reported runs and reaches matched scores with substantially fewer training steps. On Qwen3-8B-Base, AdamW reaches an aggregate accuracy of 0.495 after 270 training steps. ISO-AdamW reaches the same accuracy after only 100 training steps and improves further to 0.509 after 210 training steps. Together, ISO offers a concrete answer to RLVR's missing optimization layer: rather than inheriting pre-training optimization wholesale, design post-training around the structure of reward-driven adaptation: inherit the spectrum, optimize the frames.