ChatPaper.aiChatPaper

J-Zero:零数据驱动的挑战者-求解器-评判器统一协同进化

J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

August 27, 2026
作者: Gyouk Chu, Myeongho Jeon, Eunho Yang
cs.AI

摘要

自进化语言模型近来已成为迈向超级智能的一条有前景的路径,其优势在于降低人类监督的成本。尽管在可验证领域已取得可观进展,但在不可验证领域中的自进化研究仍相对不足。我们提出了从零数据协同适应的评判者框架(J-Zero),这是一个统一的挑战者—求解者—评判者协同进化框架,可同时支持两个领域的自我改进。挑战者与求解者通过对抗性交互协同进化:挑战者生成难度递增的任务,而求解者学习为其产出更高质量的答案。与此同时,评判者利用偏好对进行协同适应,这些偏好对的顺序并非来自评判者自身的评分,而是根据各答案的生成方式预先确定,即求解者的答案优于挑战者的答案,其经分解重组后的答案优于其一次性生成的答案。J-Zero在可验证领域平均领先基线4.2分,在不可验证领域领先8.0分,并且可持续改进至少十次迭代,而基线方法在两次迭代后即出现性能退化。
English
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.