JIT-Agent:ジャストインタイム・ハーネス進化によるハーネス知能のスケーリング
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution
August 26, 2026
著者: Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang, Xiaobin Hu, Qibing Ren, Wangchunshu Zhou, Shuicheng Yan
cs.AI
要旨
エージェントの能力は、モデルだけでは決まらない。メモリ管理、計画戦略、アクションプロトコル、ツール/スキルのオーケストレーションを含むエージェントハーネスは、その基盤となるモデルの貢献を凌駕し得る。しかし、ハーネスの設計は依然として手作業に依存し、タスク固有であり、根本的にスケーラブルではない。我々は、任意の既製エージェント型LLMに対して、タスク適応型のエージェントハーネスをその場で合成するよう訓練されたハーネス知能モデルJIT-Agentを提案する。我々はエージェントハーネスを、固定の4モジュールプロトコルに従う構成可能かつ機械生成可能な成果物として形式化し、JIT-Agentに、目の前のタスクに応じたハーネスのカスタマイズ、安定かつ信頼性の高い実行のためのハーネスの修復、そして拡大し続ける過去のハーネス構成のアーカイブから性能シグナルを蒸留することによる自己進化を学習させる。ハーネスヘルパーとしてJIT-Agentを備えたDeepSeek-V4-Flashは、DeepSearchQA (+9.1) とOdysseyBench (+4.3) でGPT-5.6を上回り、既に強力なGLM-5.2は最大+20.2ポイントの向上を達成した。統制評価を通して、JIT-Agentが生成したハーネスは、OpenCodeやClaude Codeなどの成熟したエージェントランタイムと性能面で競合し、DeepSeek V4、Mimo-V2.5、Qwen3.6のマルチスケールモデルファミリーを一貫して改善する。我々の知る限り、JIT-Agentはジャストインタイム・ハーネス生成専用に構築された初のモデルであり、ハーネス知能を、モデルスケーリングと直交する訓練可能かつ転移可能で複利的なエージェント能力の次元として確立する。
English
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.