LLM 추론을 위한 진화 전략 이해: GRPO보다 더 넓은 추론 커버리지
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
August 27, 2026
저자: Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
cs.AI
초록
진화 전략(ES)은 최근 LLM 추론을 위한 메모리 효율적인 사후 훈련 패러다임으로 부상했다. 그러나 ES의 최적화 동작은 충분히 연구되지 않아, 주류 사후 훈련 패러다임(예: 그룹 상대 정책 최적화(GRPO))과 비교하여 ES의 장점 범위를 정의하기 어렵다. 본 논문은 ES의 동역학과 메커니즘을 체계적으로 조사함으로써, 먼저 ES가 GRPO보다 성능상 이점을 가짐을 이론적·경험적으로 보여주며, ES가 더 넓은 추론 범위를 유도하여 사전 훈련된 LLM의 추론 능력을 더 잘 활용할 수 있음을 입증한다. 이론적으로, ES 모집단 전체의 검증기 투영 젠슨-샤논 다양성이 더 높은 Pass@K 성능에 기여함을 보인다. 경험적으로, 엔트로피 붕괴를 보이는 GRPO와 달리, ES는 Pass@1을 향상시키면서도 GRPO보다 높은 Pass@K를 달성한다. 또한 GRPO의 Pass@1 강점과 ES의 Pass@K 이점을 결합하는 순차적 GRPO-ES 훈련 전략을 개발한다. 둘째, 상당한 전체 모델 파라미터 드리프트에도 불구하고 ES의 작업 성능 향상은 더 큰 크기의 업데이트로 구성된 희소 부분집합에 의해서만 기여됨을 발견한다. 이러한 기능적 희소성은 큰 파라미터 이동이 광범위한 기능 변화를 의미할 필요가 없음을 시사하며, 홀드아웃 평가는 또한 이것이 반드시 파괴적 망각으로 이어지지 않음을 보여준다. 마지막으로, 하이퍼파라미터 설계가 ES의 효과성에 어떻게 영향을 미치는지 연구하여, ES가 더 큰 LLM에서 더 작은 모집단 크기를 요구함을 입증한다. 이러한 발견은 ES를 덜 효과적이고 메모리 효율적인 GRPO의 대안이 아닌, 별개의 추론 사후 훈련 패러다임으로 자리매김하게 한다.
English
Evolution Strategies (ES) have recently emerged as a memory-efficient post-training paradigm for LLM reasoning. However, the optimization behavior of ES remains understudied, making it hard to define its advantage scope compared to mainstream post-training paradigms (e.g., Group Relative Policy Optimization (GRPO)). By systematically investigating ES dynamics and mechanisms, this paper first identifies a performance advantage of ES over GRPO, theoretically and empirically showing that ES can lead to broader reasoning coverage, thereby better exploiting the reasoning capabilities of pretrained LLMs. Theoretically, we show that verifier-projected Jensen-Shannon diversity across the ES population is helpful to higher Pass@K performances. Empirically, unlike GRPO, which exhibits entropy collapse, ES improves Pass@1 while attaining higher Pass@K than GRPO. We further develop a sequential GRPO-ES training strategy that combines GRPO's strength in Pass@1 with ES's gains in Pass@K. Second, we find that despite substantial whole-model parameter drift, the task-performance gains of ES are only contributed to a sparse subset of larger-magnitude updates. This functional sparsity suggests that large parameter movement need not imply widespread functional change, and held-out evaluations further show that it does not necessarily lead to catastrophic forgetting. Finally, we study how hyperparameter design affects the effectiveness of ES, demonstrating that ES requires a smaller population size in a larger LLM. These findings position ES as a distinct reasoning post-training paradigm rather than a less effective, memory-efficient alternative to GRPO.