ChatPaper.aiChatPaper

CalibForge: 학습 가능한 종단 작업 확장을 위한 적대적 해석기 보정

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

August 6, 2026
저자: Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
cs.AI

초록

터미널 에이전트를 훈련하려면 실행 가능하고 검증 가능한 태스크가 필요하며, 이는 단순히 해결 가능할 뿐만 아니라 학습에 적절한 도전 수준을 갖추어야 한다. 실행 가능한 검증은 실현 가능성을 확립하지만, 주어진 솔버 설정에서 태스크가 어떻게 작동하는지는 드러내지 않는다. 본 논문에서는 검증된 솔버 동작을 활용하여 적대적 솔버 캘리브레이션을 통해 후보 태스크를 수정하는 자율 터미널 태스크 합성 시스템인 CalibForge를 제시한다. 다중 솔버 캘리브레이션은 이질적 솔버 풀 내에서의 불일치를 목표로 하는 반면, 대조적 솔버 캘리브레이션은 지정된 강-통과/약-실패 관계를 목표로 하며, 두 전략 모두 입증된 해결 가능성에 기반한 솔버 상대적 학습 가능 영역을 작동화한다. CalibForge를 사용하여 5,431개의 캘리브레이션된 터미널 태스크를 구축하였다. 절제 실험 결과, 두 전략 모두 작성 및 검증 단독이나 일반적인 단일 솔버 피드백보다 더 효과적인 지도를 제공하는 것으로 나타났다. 전체 컬렉션으로 훈련된 모델은 Terminal-Bench 2.0에서 32.58%와 47.57%를 달성하였다. 해당 기본 모델 대비 가장 큰 개선 폭은 Terminal-Bench 2.0에서 24.71퍼센트 포인트, SWE-bench Pro에서 27.68포인트, Doc2Repo에서 30.04포인트에 달한다. 종합적으로, 이러한 결과는 효과적이고 전이 가능한 에이전트 훈련 데이터를 구축하기 위한 실용적 목표로서 솔버 상대적 학습 가능성을 뒷받침한다.
English
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.