Long-Horizon-Terminal-Bench: 밀집 보상 기반 평가를 통한 장기 종단 과제에서의 에이전트 한계 테스트

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

July 9, 2026
저자: Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang
cs.AI

초록

AI 에이전트는 이제 짧고 명확하게 정의된 작업을 자율적으로 완료할 수 있게 되었다. 그러나 기존의 터미널 벤치마크는 대부분 수 분 이내에 종료되며 최종 결과만으로 평가되는 단순한 문제에 초점을 맞추고 있다. 이러한 구성은 중간 진행 상황과 부분 해결책을 간과하여 희소한 보상 신호와 불완전한 에이전트 능력 평가를 초래한다. 본 연구에서는 실험 재현, 소프트웨어 공학, 멀티모달 분석, 대화형 게임, 과학 계산 등 9개 범주에 걸친 46개의 장기 지평(long-horizon) 과제로 구성된 터미널 벤치마크인 Long-Horizon-Terminal-Bench를 제안한다. 각 과제는 참조 솔루션 또는 시뮬레이션 엔진을 갖춘 Terminal-Bench 스타일의 설정을 따르면서도, 세분화된 단계별 평가 과제로 분해된다. 이러한 설계는 조밀한 중간 보상과 부분 점수를 가능하게 하여, 에이전트가 최종 목표에 도달했는지 여부뿐만 아니라 개방형 워크플로우에서 얼마나 진행되었는지도 평가할 수 있게 한다. Long-Horizon-Terminal-Bench의 과제는 일반적으로 수백 회의 에피소드와 수 분에서 수 시간의 실행 시간을 필요로 하며, 일회성 문제 해결보다는 장기 지평 계획, 장기 맥락 관리, 반복적 디버깅을 강조한다. 15개의 최첨단 모델을 평가한 결과, 에이전트는 과제당 평균 9.9M 토큰을 소비하고, 실행당 약 231회의 에피소드와 85.3분의 실행 시간을 기록하여, Long-Horizon-Terminal-Bench가 기존의 터미널 기반 벤치마크보다 더 높은 요구 조건을 제시함을 확인했다. 가장 강력한 평가 모델조차 부분 보상 임계값 0.95에서 15.2%, 완전 보상 임계값 1.0에서 10.9%의 pass@1을 기록했으며, 전체 모델의 평균 통과율은 각각 4.3%와 1.7%에 불과했다. 이러한 결과는 개선의 여지가 크다는 것을 보여준다. 또한 실패 모드와 오류 패턴을 분석하고, Long-Horizon-Terminal-Bench를 공개하여 장기 지평 터미널 에이전트의 향후 발전을 지원하고자 한다.
English
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.
PDF461July 14, 2026