DecoEvo: 텍스트 공간에서 솔버와 루브릭 생성기 기술의 점수 분리 공진화
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
July 28, 2026
저자: Jiangwang Chen, Zixin Song, Junlin Liu, Shuaiyu Zhou, Haiyan Wu, Haihan Shi, Chenxi Zhou, Hanqing Li, Xiao Yang, Da Zhu, Guanjun Jiang, Hai Wan, Xibin Zhao
cs.AI
초록
텍스트 공간 최적화는 모델 가중치 대신 외부 자연어 아티팩트를 편집하여 대규모 언어 모델(LLM)을 적응시키므로, 최적화된 아티팩트는 검사 가능한 상태로 유지되며 모델을 블랙박스로 취급할 수 있다. 그러나 기존의 대부분 텍스트 공간 방법은 평가를 고정된 상태로 유지한다. 개방형 과제에서는 이로 인해 병목 현상이 발생할 수 있다. 즉, 해결사가 평가 기준이 측정하는 기준을 개선한 후에는 누락된 차원이 최적화 신호에 보이지 않게 된다. 단순히 평가 기준을 진화시키는 것도 현재 해결사의 점수에 기반하여 업데이트가 선택될 때 신뢰할 수 없는데, 이는 평가 기준을 충족시키기 쉽게 만들어 걸보기에는 진전이 있는 것처럼 보일 수 있기 때문이다. 우리는 DecoEvo(분리 공진화, Decoupled Co-Evolution)를 도입한다. 이는 분리된 목표 하에서 해결사 기술과 평가 기준 생성 기술을 공진화시키며, 최적화 과정에서 황금 평가 기준을 사용하지 않는다. 해결사 기술은 기준 수준의 피드백을 통해 업데이트되고, 평가 기준 생성 기술은 해결사 종합 점수와 무관한 요구 사항 범위와 응답 변별력에 대한 보완적 감사를 통해 수정된다. 이러한 분리는 생성자 업데이트를 새로 드러난 해결사 약점에 집중시켜, 해결사가 이미 충족하는 기준에 대한 반복적 강조를 줄인다. 각 벤치마크의 공식 평가에서 DecoEvo는 다섯 개 벤치마크 및 세 개의 LLM 백본에서 모든 비교 방법을 능가하며, 다섯 벤치마크 평균에서 SkillOpt 대비 2.8~5.0%의 상대적 성능 향상을 보인다.
English
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also unreliable when updates are selected by the current solver's score, because apparent progress can come from making the rubric easier to satisfy. We introduce DecoEvo (Decoupled Co-Evolution), which co-evolves a solver skill and a rubric-generator skill under decoupled objectives without using gold rubrics during optimization. The solver skill is updated using criterion-level feedback, while the rubric-generator skill is revised through complementary audits of requirement coverage and response discrimination that are independent of aggregate solver score. This separation focuses generator updates on newly exposed solver weaknesses, reducing repeated emphasis on criteria the solver already satisfies. Under each benchmark's official evaluation, DecoEvo outperforms all compared methods across five benchmarks and three LLM backbones, yielding 2.8--5.0\% relative gains over SkillOpt in the five-benchmark average.