J-Zero: 제로 데이터에서의 통합 도전자-솔버-심판 공진화
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
August 27, 2026
저자: Gyouk Chu, Myeongho Jeon, Eunho Yang
cs.AI
초록
자기진화 언어 모델은 최근 인간 감독 비용을 절감할 수 있는 이점으로 초지능을 향한 유망한 경로로 부상하고 있다. 검증 가능한 영역에서는 상당한 진전이 이루어졌지만, 검증 불가능한 영역에서의 자기진화는 여전히 탐구가 크게 부족하다. 우리는 두 영역 모두에서 자기 개선을 지원하는 통합 도전자-해결자-판정자 공진화 프레임워크인 J-Zero(제로 데이터 기반 판정자 공동 적응)를 제안한다. 도전자와 해결자는 적대적 상호작용을 통해 공진화한다. 도전자는 점점 더 어려운 작업을 생성하고, 해결자는 이에 대한 더 높은 품질의 응답을 생성하는 방법을 학습한다. 동시에 판정자는 각 응답이 생성된 방식에서 순서가 사전에 알려진 선호 쌍, 즉 해결자의 답변이 도전자의 답변보다, 분해·재결합된 답변이 원샷 답변보다 우선하는 선호 쌍을 사용하여 공동 적응한다(판정자 자신의 점수에 기반하지 않음). J-Zero는 검증 가능한 영역에서 평균 4.2점, 검증 불가능한 영역에서 평균 8.0점으로 기준선을 능가하며, 기준선은 두 번의 반복 후 성능이 저하되는 반면 J-Zero는 최소 10번의 반복을 통해 계속 개선된다.
English
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.