NeoHorse-1:透過帶有路由框架的代理式後訓練邁向遞迴自我改進
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
September 8, 2026
作者: NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo, Kai Han, Hailin Hu, Zihan Jiang, Xiang Kuang, Boxun Li, Yulong Li, Zehua Pei, Yuchuan Tian, Jiamin Wang, Yu Wang, Yunhe Wang, Yihong Wu, Haiyang Xu, Shuo Zhang, Hang Zhou, Siyang Cheng, Jiayu Fan, Wei He, Qingrui Jiao, Hongguang Li, Zhiyuan Li, Runke Liu, Xi Liu, Xinchen Liu, Sinno Jialin Pan, Yi Ren, Liuyang Song, Chenyu Wang, Bei Yu, Quanlu Zhang, Xiangyu Zhang, Mengyu Zheng, Yingjie Zong
cs.AI
摘要
遞迴自我改進(RSI)需要一個具體機制,使 AI 系統得以觀察自身能力,並將該證據轉化為下一輪學習。我們提出 NeoHorse-1,一個代理原生模型系列,其開發目的是透過代理式後訓練探索此路徑。我們的系統結合異質模型池與智慧路由,記錄每個使用者回合的預測能力需求、所選服務層級及後續互動。這些記錄會轉換為訓練樣本,保留交錯推理、工具呼叫與測試框架脈絡,並透過結構驗證、六維語意評估與子場景層級標註予以納入。路由訊號將監督式微調組織為三階段課程,並延伸至路由引導的同策略蒸餾,其中教師在相同進程下監督學生生成的回應。能力引導分配接著將評估回饋轉換為下一輪訓練混合資料,形成評估-選擇-更新循環,其中系統學會做什麼,將形塑它下一步從何學習。在涵蓋基於測試框架的代理、工具使用、程式設計與指令遵循的十一項基準測試上,後訓練在 4B 時將巨觀平均從 58.94 提升至 64.87,在 9B 時從 65.60 提升至 69.04,大幅縮小後訓練 4B 模型與 9B 基礎模型之間的整體差距。NeoHorse-1 提供了此回饋驅動過程的初始原型,以及一條邁向跨連續迭代、由測試框架中介的遞迴自我改進的路徑。
English
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.