VibeLifeBench:你的生活智能体能在生活世界中主动且持久地行动吗?
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
August 11, 2026
作者: Xiaohongshu Inc
cs.AI
摘要
大语言模型(LLM)智能体越来越多地被部署为个人助理。然而,现有评估大多使用静态环境中的简短、自包含的请求。日常生活辅助则不同:一项任务会持续数周而非数分钟;在智能体未被提示时,世界也在不断变化;许多约束从未被明确说明。一个仅仅回答眼前请求的智能体在此类任务中将会失败。相反,所需要的是一个保持主动性和一致性的智能体:它自行决定何时行动、何时提问、何时保持沉默;它注意到无人宣告的变化;它从第一天到最后一天都维持同一个连贯的计划。目前没有任何基准测试能衡量这一点。我们提出了 VibeLifeBench,这是一个包含200个长期任务、横跨十个日常生活领域的基准测试。每个任务都是一个脚本化的多周时间线,发生在包含22个模拟服务的模拟世界中。这个世界按自己的时钟推进,而且其许多变化都是静默的,因此只有重新检查世界的智能体才能发现它们。每个任务都通过细粒度、加权的检查来评分,这些检查只读取智能体实际留下的痕迹,涵盖最终状态、行动的及时性以及是否遵守了隐含约束。我们评估了七个前沿模型。它们的得分都很低,这显示了当前智能体与现实生活辅助之间的差距。我们将开源所有任务、环境和评估框架。
English
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.