CalibForge:面向可學習終端任務擴展的對抗式求解器校準
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks
August 6, 2026
作者: Fanzhe Meng, Guoxin Chen, Jiale Zhao, Shuang Sun, Zhiyu Lin, Wayne Xin Zhao, Ruihua Song, Ji-Rong Wen, Kai Jia
cs.AI
摘要
訓練終端代理需要可執行且可驗證的任務,這些任務不僅要可解,還須具有適當的學習挑戰性。可執行驗證能確立任務的可行性,但無法揭示任務在特定求解器設定下的行為表現。本文提出 CalibForge,這是一個自主式終端任務合成系統,利用已驗證的求解器行為,透過對抗式求解器校準來修訂候選任務。多求解器校準針對異質求解器池內部的分歧,而對比式求解器校準則針對指定的強通過/弱失敗關係;兩者皆將求解器相對的可學習區間具體化,並以已證實的可解性為錨點。利用 CalibForge,我們建構了 5,431 個經校準的終端任務。消融實驗顯示,這兩種策略所提供的監督訊號,均優於僅依賴任務撰寫與驗證,或一般的單一求解器回饋。在完整資料集上訓練的模型,於 Terminal-Bench 2.0 達到 32.58% 與 47.57% 的成績。與對應的基礎模型相比,最大提升幅度在 Terminal-Bench 2.0 上達 24.71 個百分點、在 SWE-bench Pro 上達 27.68 個百分點、在 Doc2Repo 上達 30.04 個百分點。綜合而言,這些結果支持將求解器相對可學習性視為建構有效且可遷移的代理訓練資料之實用目標。
English
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a given solver setting. In this paper, we present CalibForge, an autonomous terminal-task synthesis system that uses verified solver behavior to revise candidate tasks through adversarial solver calibration. Multi-solver calibration targets disagreement within a heterogeneous solver pool, whereas contrastive solver calibration targets a designated strong-pass/weak-fail relation; both operationalize a solver-relative learnable zone anchored in demonstrated solvability. Using CalibForge, we construct 5,431 calibrated terminal tasks. Our ablations show that both strategies yield more effective supervision than authoring and validation alone or ordinary single-solver feedback. Models trained on the full collection achieve 32.58% and 47.57% on Terminal-Bench 2.0. The largest improvements over the corresponding base model reach 24.71 percentage points on Terminal-Bench 2.0, 27.68 points on SWE-bench Pro, and 30.04 points on Doc2Repo. Together, these results support solver-relative learnability as a practical target for constructing effective and transferable agent training data.