ChatPaper.aiChatPaper

CalibForge:面向可学习终端任务扩展的对抗式求解器校准

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

August 6, 2026
作者: Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
cs.AI

摘要

训练终端智能体需要可执行且可验证的任务,这些任务不仅要可求解,还必须具备适合学习的适当挑战性。可执行性验证确立了可行性,却无法揭示任务相对于给定求解器配置的行为表现。本文提出CalibForge——一个自主式终端任务合成系统,通过利用经过验证的求解器行为,以对抗性求解器校准方式修订候选任务。多求解器校准针对异构求解器池中的分歧,而对比性求解器校准则针对指定的强通过/弱失败关系;两者均将锚定于已验证可解性的求解器相对可学习区间加以操作化。借助CalibForge,我们构建了5,431个经校准的终端任务。消融实验表明,相较于仅依赖编写与验证或普通单求解器反馈,这两种策略均能提供更有效的监督信号。在完整数据集上训练的模型在Terminal-Bench 2.0上分别达到32.58%和47.57%的成绩。相对对应基础模型的最大改进幅度在Terminal-Bench 2.0上达到24.71个百分点,在SWE-bench Pro上达到27.68个百分点,在Doc2Repo上达到30.04个百分点。综上,这些结果支持将求解器相对可学习性作为构建有效且可迁移的智能体训练数据的实用目标。
English
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.