ChatPaper.aiChatPaper

StartupBench:於市場驗證之端到端工作流程上評測通用型代理

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

August 18, 2026
作者: Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
cs.AI

摘要

大型語言模型(LLMs)與代理的最新進展大幅提升了人工智慧系統執行複雜任務的能力。然而,現有基準測試大多依賴研究人員選定的任務,使得此類進展是否能延伸至真實世界使用者實際要求人工智慧系統執行的工作,仍屬未知。我們提出 StartupBench,這是一個以市場驗證之人工智慧新創產品為基礎的端對端代理基準測試。我們並非基於對代理有用能力的預設假設來定義任務,而是系統性地研究已被證實獲得採用的 AI 產品、其產品工作流程與使用者,以識別 AI 已在多種專業領域中展現實際需求的真實世界任務。我們將這些工作流程轉化為完整的、以交付成果為導向的任務,並以能夠捕捉其複雜需求的細粒度評分標準進行評估。在統一的代理測試框架下評估多個具代表性的模型時,即使最強的模型,也僅能成功完成約 30% 的 StartupBench 任務,儘管在許多任務上已取得顯著的部分進展。進一步的分析指出,複雜的指令遵循與特定領域專業知識等方面是主要的失敗來源。我們的結果顯示,許多經過市場驗證的工作流程仍超出當前通用代理的可靠能力範圍,這使 StartupBench 成為衡量朝向端對端完成真實世界使用者任務進展的實證指標。
English
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce StartupBench, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.