ClawGym II: 에이전트 하네스에서의 블랙박스 RL 탐구
ClawGym II: Exploring Black-Box RL on Agent Harness
August 17, 2026
저자: Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
cs.AI
초록
에이전트 하네스는 에이전트와 환경 간의 상호작용을 조정함으로써 장기 지평 작업에서의 성능을 크게 향상시켜 왔다. 그러나 복잡한 하네스를 통한 강화 학습은 여전히 대체로 탐구되지 않은 상태이며, 이러한 훈련을 장기 지평 에이전트 작업으로 확장하는 데는 근본적인 도전 과제가 존재한다. 본 연구에서는 복잡한 하네스를 통한 일반 에이전트의 안정적이고 확장 가능한 최적화를 위한 통합 블랙박스 RL 프레임워크를 제시한다. 구체적으로, 먼저 대규모 동시 롤아웃을 위해 임시 샌드박스 내에 작업 환경과 하네스를 격리하는 샌드박스 기반 실행 인프라를 구축한다. 그런 다음 정책 최적화를 불투명한 하네스 실행으로부터 분리하고, 모델 경계에 서빙 프록시를 배치하여 모델 호출을 포착한다. 다중 턴 궤적을 재구성하고 훈련 효율성을 개선하기 위해, 포착된 호출을 접두사 트리로 구성하고 비평가 기반 PPO와 비평가 없는 GRPO를 모두 적응시켜 복원된 트리 구조에 대해 최적화한다. 한편, 최적화 과정 전반에 걸쳐 훈련-추론 일관성을 유지한다. 마지막으로, 단일 모델이 이종 하네스에 의해 공동 최적화될 수 있도록 하는 믹스 하네스 훈련을 도입한다. Qwen3-30A3B를 사용할 때, 블랙박스 RL은 OpenClaw와 Claude Code를 통해 ClawGym-Bench에서 Pass@1을 각각 9.98점과 14.81점 향상시키며 200-400회의 최적화 단계 동안 안정성을 유지한다. 또한, 이 프레임워크는 JobBench와 OfficeQA와 같은 더 어려운 작업에서도 일관된 성능 향상을 보여준다. 전반적으로, 본 프레임워크는 이종 실행 시스템 전반에 걸친 통합 훈련을 지원하며, 블랙박스 하네스를 통한 일반 에이전트의 효과적이고 안정적이며 확장 가능한 최적화를 가능하게 한다.
English
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.