StartupBench:面向市场验证的端到端工作流的通用型代理基准测试
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
August 18, 2026
作者: Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
cs.AI
摘要
近年来,大语言模型(LLMs)与智能体(agents)的快速发展显著提升了AI系统执行复杂任务的能力。然而,现有基准测试大多依赖研究者自行选取的任务,这使得我们难以确定这种进步是否真正延伸至现实用户实际期望AI系统完成的工作。为此,我们提出了StartupBench——一个基于市场验证的AI初创公司产品的端到端(E2E)智能体基准测试。与从预设的智能体功能假设出发定义任务不同,我们系统性地研究那些已被证实获得用户采用(demonstrated adoption)的AI产品及其产品工作流和用户群体,从而在多个专业领域中识别出AI已具备实际需求(practical demand)的现实任务。我们将这些工作流转化为以完整交付物为导向的任务,并通过捕捉其复杂需求的细粒度评分标准(fine-grained rubrics)进行评估。在统一智能体测试框架(unified agent harness)下对代表性模型进行评估后发现,即便最强的模型也只能成功完成StartupBench中约30%的任务,尽管在许多任务上已取得实质性部分进展。进一步的分析表明,复杂指令遵循(complex instruction following)与领域专业知识(domain-specific expertise)是主要的失败来源。我们的研究结果表明,许多经市场验证的工作流仍然超出了当前通用智能体的可靠能力范围,这使StartupBench成为衡量面向现实用户任务端到端完成进展的实证基准。
English
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce StartupBench, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.