ChatPaper.aiChatPaper

DiagEvo: 계층적 오류 메모리를 통한 진단 기반 자가 진화

DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory

September 1, 2026
저자: Xincheng Wei, Yifan Ding, Yoshua Li, Dongsheng Ma, Rongxiang Weng, Xunliang Cai, Wenjian Ding, Yao Zhang
cs.AI

초록

자기 대결(self-play)은 언어 모델 자기 진화를 위한 효과적인 패러다임이지만, 안내(guidance)가 없다면 해결기(solver)의 성능은 라운드가 거듭될수록 정체되거나 하락할 수 있다. 비안내(unguided) 방식들은 문제 생성을 난이도, 학습 가능성, 또는 다양성과 같은 신호를 사용해 유도한다. 이러한 신호들은 문제를 도전적이고 다양하게 유지하지만, 이후 라운드가 해결되지 않은 어떤 추론 약점을 중점적으로 겨냥해야 하는지는 명시하지 않는다. 안내(guided) 방식들은 인간 예시, 문서 말뭉치, 또는 지정된 난이도 목표를 포함한 외부 태스크 자원에서 방향을 얻으므로, 자기 대결 루프 외부에서 공급되는 태스크 정보에 의존한다. 우리는 필요한 방향이 해결기 자신의 실패 이력으로부터 도출될 수 있음을 보여준다. 우리는 DiagEvo를 제안한다. DiagEvo의 진단기(diagnostician)는 이 이력에서 반복되는 오류 원인을 추출하여 계층적 오류 원인 메모리에 저장한다. 메모리는 관련 원인들을 기술 노드(skill node) 아래에 그룹화하고, 표적 문제에 대한 자기 일관성(self-consistency)에 따라 각각을 활성(Active) 또는 숙달(Mastered) 상태로 추적한다. 도전기(challenger)는 이러한 상태들과 재발 횟수를 사용하여 원인 표적 생성을 자유 탐색과 균형 있게 배분한다. 이중 신뢰도 필터링은 최빈 해결기 답변이 명확한 투표 우위를 보일 때만 중간 난이도 문제를 유지한다. DiagEvo는 외부 태스크 자원 없이 자기 대결 과정에서 생성된 정보로부터 커리큘럼을 도출한다. 기본 4B 진단기를 사용할 때, DiagEvo는 Qwen3-4B, Qwen3-8B, OctoThinker-8B의 세 해결기 각각에 대해 9개 벤치마크 전체의 평균 정확도에서 모든 기준선을 능가한다. Qwen3-8B의 경우, 5개 수학적 추론 벤치마크에 걸쳐 평균 정확도 72.3%를 달성하여 R-Zero보다 4.5%포인트 높다. 또한 9개 벤치마크 전체 평균 정확도는 57.4%로 DARC보다 1.1%포인트 높다. 절제 실험은 계층적 오류 원인 메모리와 이중 신뢰도 필터링이 모두 이러한 성과 향상에 기여함을 보여준다.
English
Self-play is an effective paradigm for language-model self-evolution, but without guidance, solver performance can plateau or decline across rounds. Unguided methods steer question generation with signals such as difficulty, learnability, or diversity. These signals keep questions challenging and varied but do not specify which unresolved reasoning weaknesses later rounds should target. Guided methods obtain direction from external task resources, including human examples, document corpora, or specified difficulty targets, and therefore rely on task information supplied outside the self-play loop. We show that the needed direction can instead be derived from the solver's own failure history. We introduce DiagEvo, whose diagnostician extracts recurring error causes from this history and stores them in a hierarchical error-cause memory. The memory groups related causes under skill nodes and tracks each as Active or Mastered according to self-consistency on targeted questions. The challenger uses these states and recurrence counts to balance cause-targeted generation with free exploration. Double-confidence filtering retains intermediate-difficulty questions only when the most common solver answer has a clear vote lead. DiagEvo derives its curriculum from information produced during self-play, without external task resources. With the default 4B diagnostician, DiagEvo outperforms every baseline in mean accuracy across all nine benchmarks for each of the three solvers: Qwen3-4B, Qwen3-8B, and OctoThinker-8B. On Qwen3-8B, it reaches 72.3% mean accuracy across five mathematical reasoning benchmarks, 4.5 percentage points above R-Zero. Its mean accuracy across all nine benchmarks is 57.4%, 1.1 percentage points above DARC. Ablations show that the hierarchical error-cause memory and double-confidence filtering both contribute to these gains.