StartupBench: 시장 검증된 엔드투엔드 워크플로우에서 범용 에이전트 벤치마킹
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
August 18, 2026
저자: Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
cs.AI
초록
최근 대규모 언어 모델(LLM)과 에이전트의 발전은 AI 시스템이 복잡한 작업을 수행하는 능력을 실질적으로 향상시켰다. 그러나 기존 벤치마크는 대부분 연구자가 선별한 작업에 의존하고 있어, 이러한 발전이 실제 사용자가 AI 시스템에 요구하는 작업으로까지 확장되는지는 불확실하다. 우리는 시장에서 검증된 AI 스타트업 제품에 기반한 엔드투엔드(E2E) 에이전트 벤치마크인 StartupBench를 소개한다. 유용한 에이전트 능력에 대한 사전 정의된 가정에서 작업을 도출하는 대신, 실제 채택이 입증된 AI 제품과 해당 제품의 작업 흐름 및 사용자를 체계적으로 분석하여 다양한 전문 분야에서 AI에 대한 실질적 수요가 확립된 실제 작업을 식별한다. 우리는 이러한 작업 흐름을 완전한 산출물 중심의 작업으로 변환하고, 복잡한 요구사항을 포착하는 세분화된 평가 루브릭으로 평가한다. 통합 에이전트 하네스 환경에서 평가된 대표적 모델들 중 최강의 모델조차도 많은 작업에서 상당한 부분적 진전을 보였음에도 불구하고 StartupBench의 약 30%만 성공적으로 완료했다. 추가 분석은 복잡한 지시 따르기와 도메인별 전문성이 주요 실패 원인임을 식별한다. 우리의 결과는 많은 시장 검증된 작업 흐름이 현재 범용 에이전트의 신뢰할 수 있는 능력을 여전히 넘어서는 영역에 있음을 보여주며, 실제 사용자 작업의 엔드투엔드 완료를 향한 진전을 측정하는 경험적 지표로서 StartupBench를 확립한다.
English
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce StartupBench, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.