J-Zero:挑戰者—求解器—裁判的統一共同演化(從零資料出發)
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
August 27, 2026
作者: Gyouk Chu, Myeongho Jeon, Eunho Yang
cs.AI
摘要
自我演化語言模型近來已成為通往超級智能的一條極具前景的途徑,其優勢在於能降低人工監督的成本。儘管在可驗證領域已取得相當進展,自我演化在不可驗證領域中的探索仍明顯不足。我們提出了「零資料裁判共同適應」(Judge co-adaptation from Zero data, J-Zero)框架,這是一個統一的「挑戰者—求解者—裁判」共同演化架構,可同時支援兩大領域的自我改進。挑戰者與求解者透過對抗性互動共同演化:挑戰者生成難度漸增的任務,而求解者則學習產出更高品質的回應。與此同時,裁判透過偏好對進行共同適應,這些偏好對的排序在生成階段即已事先確定,係根據各回應的產生方式而定,亦即求解者的答案優於挑戰者的答案,以及求解者分解再重組後的答案優於其單次生成的答案,而非依賴裁判自身的評分。J-Zero在可驗證領域平均領先基線模型4.2分,在不可驗證領域平均領先8.0分,且至少可持續改善十次迭代,而基線模型在兩次迭代後即開始衰退。
English
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.