NeoHorse-1: 라우팅 하네스를 활용한 에이전트형 사후 학습을 통한 재귀적 자기 개선을 향하여
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
September 8, 2026
저자: NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo, Kai Han, Hailin Hu, Zihan Jiang, Xiang Kuang, Boxun Li, Yulong Li, Zehua Pei, Yuchuan Tian, Jiamin Wang, Yu Wang, Yunhe Wang, Yihong Wu, Haiyang Xu, Shuo Zhang, Hang Zhou, Siyang Cheng, Jiayu Fan, Wei He, Qingrui Jiao, Hongguang Li, Zhiyuan Li, Runke Liu, Xi Liu, Xinchen Liu, Sinno Jialin Pan, Yi Ren, Liuyang Song, Chenyu Wang, Bei Yu, Quanlu Zhang, Xiangyu Zhang, Mengyu Zheng, Yingjie Zong
cs.AI
초록
재귀적 자기 개선(RSI)은 AI 시스템이 자신의 역량을 관찰하고 그 증거를 다음 학습 라운드로 전환하는 구체적인 메커니즘을 필요로 한다. 우리는 에이전트 후속 학습을 통해 이 경로를 탐구하기 위해 개발된 에이전트 네이티브 모델 계열인 NeoHorse-1을 제시한다. 우리의 시스템은 지능형 라우팅을 갖춘 이종 모델 풀을 결합하여 각 사용자 턴마다 예측된 역량 수요, 선택된 서비스 계층, 후속 상호작용을 기록한다. 이 기록들은 인터리브된 추론, 도구 호출, 하네스 컨텍스트를 보존하는 학습 예시로 변환되며, 구조 검증, 6차원 의미 평가, 서브씬 수준 라벨링을 통해 승인된다. 라우팅 신호는 지도 미세 조정을 3단계 커리큘럼으로 조직하고, 동일한 진행 과정에서 교사가 학생 생성 응답을 감독하는 라우팅 유도 온폴리시 증류로 확장된다. 이어서 역량 유도 할당은 평가 피드백을 다음 학습 혼합으로 변환함으로써, 시스템이 수행하도록 배우는 것이 다음에 무엇으로부터 배울지를 형성하는 평가-선택-업데이트 루프를 닫는다. 하네스 기반 에이전트, 도구 사용, 코딩, 지시 수행을 포괄하는 11개 벤치마크 전반에서, 후속 학습은 4B에서 매크로 평균을 58.94에서 64.87로, 9B에서 65.60에서 69.04로 상승시켜, 후속 학습된 4B 모델과 9B 기본 모델 간의 총체적 격차를 상당히 좁힌다. NeoHorse-1은 이러한 피드백 기반 과정의 초기 프로토타입과 연속적인 반복에 걸친 하네스 매개 RSI를 향한 경로를 제공한다.
English
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.