Long-Horizon-Terminal-Bench:密な報酬に基づく評価を用いた長期視野終端タスクにおけるエージェントの限界の検証
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
July 9, 2026
著者: Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang
cs.AI
要旨
AIエージェントは、短く明確に定義されたタスクを自律的に完了できるようになった。しかし、既存のターミナルベンチマークの多くは、数分で完了し、最終的な結果のみで評価される単純な問題に焦点を当てている。この設定では、中間的な進捗や部分的な解決策が見落とされ、疎な報酬信号とエージェント能力の不完全な評価しか得られない。我々は、実験の再現、ソフトウェア工学、マルチモーダル解析、インタラクティブゲーム、科学計算など、9つのカテゴリにわたる46の長時域タスクからなるターミナルベンチマーク「Long-Horizon-Terminal-Bench」を導入する。各タスクは、参照解やシミュレーションエンジンを備えたTerminal-Benchスタイルの設定に従うが、さらに細分化された段階的な評価用サブタスクに分解される。この設計により、密な中間報酬と部分点が可能となり、エージェントが最終目標に到達するかどうかだけでなく、オープンエンドなワークフロー上でどこまで進捗するかを評価で捉えられる。Long-Horizon-Terminal-Benchのタスクは、典型的には数百のエピソードと数分から数時間の実行時間を必要とし、一回限りの問題解決ではなく、長期的な計画、長いコンテキストの管理、反復的なデバッグに重点を置いている。我々は15の最先端モデルを評価し、エージェントがタスクあたり平均990万トークンを消費し、1回の実行あたり約231エピソード、85.3分の実行時間を要することを発見した。これは、Long-Horizon-Terminal-Benchが従来のターミナルベースのベンチマークよりも要求が厳しいことを示している。最も性能の高いテスト済みモデルでも、部分報酬の閾値0.95でpass@1が15.2%、完全報酬の閾値1.0で10.9%にとどまり、両閾値におけるモデル全体の平均合格率はそれぞれ4.3%と1.7%であった。これらの結果は、改善の余地が大きいことを示している。さらに、失敗モードとエラーパターンを分析し、長時域ターミナルエージェントの今後の進歩を支援するために、Long-Horizon-Terminal-Benchを公開する。
English
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.