J-Zero: ゼロデータからの統合的挑戦者-解法者-判定者の共進化
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
August 27, 2026
著者: Gyouk Chu, Myeongho Jeon, Eunho Yang
cs.AI
要旨
近年、自己進化型言語モデルは、人間の監督コストの削減という利点を備えた、スーパーインテリジェンスへの有望な道筋として登場している。検証可能な領域ではかなりの進歩が見られる一方で、検証不可能な領域における自己進化は、依然として十分に探究されていない。本稿では、ゼロデータからの判定器共適応(J-Zero)を提案する。これは、挑戦者(Challenger)・解決者(Solver)・判定器(Judge)を統合した共進化フレームワークであり、両領域にわたる自己改善を支援する。挑戦者と解決者は、対抗的相互作用を通じて共進化する。すなわち、挑戦者がますます困難なタスクを生成し、解決者がそれらに対するより高品質な応答を生成することを学習する。並行して、判定器は、判定器自身のスコアではなく、各応答の生成方法から順序が事前に分かっている選好ペア、すなわち、解決者の回答を挑戦者の回答よりも優先し、解決者の分解・再結合された回答をそのワンショット回答よりも優先するペアを用いて共適応する。J-Zeroは、検証可能領域で平均4.2ポイント、検証不可能領域で平均8.0ポイント、ベースラインを上回り、少なくとも10回の反復を通じて改善を続ける。一方、ベースラインは2回の反復後に性能が低下する。
English
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains substantially less explored. We propose Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge co-evolution framework that supports self-improvement across both domains. The Challenger and Solver co-evolve through an adversarial interaction: the Challenger generates increasingly difficult tasks, while the Solver learns to produce higher-quality responses to them. In parallel, the Judge co-adapts using preference pairs whose ordering is known in advance from how each response was produced, i.e., the Solver's answer over the Challenger's, and its decomposed-and-recombined answer over its one-shot answer, rather than from the Judge's own scores. J-Zero outperforms the baselines by an average of 4.2 points on verifiable and 8.0 points on unverifiable domains, and continues to improve through at least ten iterations, whereas the baselines degrade after two.