ChatPaper.aiChatPaper

JIT-Agent: 적시 하네스 진화를 통한 하네스 지능 확장

JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

August 26, 2026
저자: Guibin Zhang, Leo Lu, Fangzhou Xie, Kang Zhu, Junhao Wang, Zhifei Xie, Zhaochen Yu, Zihang Liu, Zhongxiang Sun, Qiankun Li, Yue Liao, Heng Chang, Xiaobin Hu, Qibing Ren, Wangchunshu Zhou, Shuicheng Yan
cs.AI

초록

에이전트 능력은 모델 단독으로 결정되지 않는다. 메모리 관리, 계획 전략, 행동 프로토콜, 도구/스킬 오케스트레이션을 포괄하는 에이전트 하네스는 기저 기반 모델의 기여도를 압도할 수 있다. 그럼에도 하네스 설계는 여전히 수동적이고 작업별로 특화되어 있으며 근본적으로 확장이 불가능하다. 우리는 임의의 기성 에이전트형 LLM을 위해 작업 적응형 에이전트 하네스를 즉석에서 합성하도록 훈련된 하네스 지능 모델인 JIT-Agent를 제시한다. 에이전트 하네스를 고정된 4모듈 프로토콜에 의해 규율되는 합성 가능하고 기계 생성 가능한 산출물로 정형화하고, JIT-Agent가 주어진 작업에 맞춰 하네스를 맞춤화하고, 안정적이고 신뢰할 수 있는 실행을 위해 하네스를 수리하며, 이전 하네스 구성들의 확장 아카이브에서 성능 신호를 증류하여 자기 진화하도록 훈련한다. JIT-Agent를 하네스 헬퍼로 장착한 DeepSeek-V4-Flash는 DeepSearchQA(+9.1)와 OdysseyBench(+4.3)에서 GPT-5.6을 능가하며, 이미 강력한 GLM-5.2는 최대 +20.2포인트를 추가로 획득한다. 통제된 평가 전반에서 JIT-Agent가 생성한 하네스는 OpenCode 및 Claude Code와 같은 성숙한 에이전트 런타임과 성능 경쟁력을 갖추고, DeepSeek V4, Mimo-V2.5, Qwen3.6의 다중 규모 모델 계열을 일관되게 개선한다. 우리가 아는 한, JIT-Agent는 적시 하네스 생성을 위해 특별히 설계된 최초의 모델로서, 모델 확장과 직교하는 에이전트 능력의 훈련 가능하고 전이 가능하며 누적 가능한 차원으로 하네스 지능을 확립한다.
English
Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model. Yet harness design remains manual, task-specific, and fundamentally unscalable. We present JIT-Agent, a harness intelligence model trained to synthesize task-adaptive agent harnesses on the fly for arbitrary off-the-shelf agentic LLMs. We formalize the agent harness as a composable, machine-generatable artifact governed by a fixed four-module protocol, and train JIT-Agent to customize harnesses for a given task at hand, repair harnesses for stable and reliable execution, and self-evolve by distilling performance signals from an expanding archive of prior harness configurations. Equipped with JIT-Agent as a harness helper, DeepSeek-V4-Flash surpasses GPT-5.6 on DeepSearchQA (+9.1) and OdysseyBench (+4.3), while the already strong GLM-5.2 gains up to +20.2 points. Across controlled evaluations, JIT-Agent-generated harnesses are performance-competitive with mature agent runtimes such as OpenCode and Claude Code and consistently improve multi-scale model families of DeepSeek V4, Mimo-V2.5, and Qwen3.6. To our knowledge, JIT-Agent is the first model purpose-built for just-in-time harness generation, establishing harness intelligence as a trainable, transferable, and compounding dimension of agent capability orthogonal to model scaling.