StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建基于市场验证的AI创业产品的工作流基准StartupBench,评估通用智能体在真实世界任务中的表现,揭示现有模型在复杂指令执行和领域特定知识上的不足。
📝 Abstract
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Agents
Real-World Tasks
Market-Validated Workflows
End-to-End Completions
Innovation

Methods, ideas, or system contributions that make the work stand out.

StartupBench
Market-Validated Workflows
End-to-End Tasks
Complex Instruction Following
Domain-Specific Expertise
🔎 Similar Papers
No similar papers found.
L
Liya Zhu
ByteDance Seed
X
Xin Ma
Nanjing University
T
Tao Liu
M-A-P
H
Haodong Wang
TokenWave.AI
G
Ge Zhang
ByteDance Seed
J
Jingzhe Ding
Nanjing University
Q
Qingshui Gu
M-A-P
Y
Yongjie Zhong
TokenWave.AI
Jinxiang Meng
Jinxiang Meng
Nanjing University of Posts and Telecommunications
LLM AgentTable ReasoningTool Use
Y
Yuan Gao
Nanjing University
Y
Yunqiu Zhou
M-A-P
H
Hao Zhu
TokenWave.AI
J
Jifeng He
ByteDance Seed
Y
Yongzhi Liao
Nanjing University
X
Xinyi Zhang
M-A-P
C
Chaoxin Li
TokenWave.AI
Y
Yi Zhu
ByteDance Seed
X
Xi Lin
Nanjing University
D
Duju Zeng
M-A-P
X
Xiang Gao
TokenWave.AI
W
Wen Zhang
ByteDance Seed
Y
Yunyang Wang
Nanjing University
D
Duo Wang
M-A-P
Huan Zhou
Huan Zhou
Northwestern Polytechnical University
Mobile Edge ComputingFederated LearningMobile Social NetworksVANETsData Offloading
Z
Zuo Wang
ByteDance Seed