ChatPaper.aiChatPaper

NeoHorse-1: ルーティングハーネスを用いたエージェント型ポストトレーニングによる再帰的自己改善に向けて

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

September 8, 2026
著者: NeoHorse Team, Guoliang Cao, Guohao Dai, Tianyu Guo, Kai Han, Hailin Hu, Zihan Jiang, Xiang Kuang, Boxun Li, Yulong Li, Zehua Pei, Yuchuan Tian, Jiamin Wang, Yu Wang, Yunhe Wang, Yihong Wu, Haiyang Xu, Shuo Zhang, Hang Zhou, Siyang Cheng, Jiayu Fan, Wei He, Qingrui Jiao, Hongguang Li, Zhiyuan Li, Runke Liu, Xi Liu, Xinchen Liu, Sinno Jialin Pan, Yi Ren, Liuyang Song, Chenyu Wang, Bei Yu, Quanlu Zhang, Xiangyu Zhang, Mengyu Zheng, Yingjie Zong
cs.AI

要旨

再帰的自己改善(RSI)には、AIシステムが自身の能力を観察し、その証拠を次の学習ラウンドへ変換する具体的なメカニズムが必要である。我々はNeoHorse-1を提示する。これは、エージェント的ポストトレーニングを通じてこの道を探求するために開発されたエージェントネイティブモデル群である。本システムは、知的ルーティングを備えた異種モデルプールを組み合わせ、各ユーザーターンについて、予測された能力需要、選択されたサービス層、および後続のインタラクションを記録する。これらの記録は、インターリーブされた推論、ツール呼び出し、およびハーネス文脈を保持する学習例へ変換され、構造検証、6次元意味評価、サブシーン単位のラベリングを通じて採用される。ルーティング信号は、教師あり微調整を3段階のカリキュラムに編成し、ルーティング誘導オンポリシー蒸留へと拡張する。そこでは、教師が同じ進行の下で学生生成応答を監督する。能力誘導型の割り当ては、評価フィードバックを次の学習混合へ変換し、評価-選択-更新ループを閉じる。このループでは、システムが何をすることを学ぶかが、次に何から学ぶかを形作る。ハーネスベースのエージェント、ツール使用、コーディング、指示追従を対象とする11のベンチマークにわたって、ポストトレーニングはマクロ平均を4Bで58.94から64.87へ、9Bで65.60から69.04へ引き上げ、ポストトレーニング済み4Bモデルと9Bベースモデルの間の総合ギャップを大幅に狭める。NeoHorse-1は、このフィードバック駆動プロセスの初期プロトタイプと、連続する反復にわたるハーネス媒介RSIへの道筋を提供する。
English
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.