ChatPaper.aiChatPaper

VibeLifeBench:你的生活智能體能否在生活世界中主動且持續地運作?

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

August 11, 2026
作者: Xiaohongshu Inc
cs.AI

摘要

大型語言模型(LLM)智能體日益被部署為個人助理。然而,現有的評測大多使用靜態環境中的簡短、自包含請求。日常生活輔助則不同:一項任務持續數週而非數分鐘;當智能體未被提示時,世界持續變化;許多限制條件從不直接言明。只會回應眼前請求的智能體,必將在這種任務上失敗。反之,所需要的是一個能保持主動性與一致性的智能體:它自行決定何時行動、何時詢問、何時保持沉默;它能察覺無人宣告的變化;它讓同一個計畫從第一天到最後一天都保持一致。目前沒有任何評測基準衡量這類能力。我們提出 VibeLifeBench,這是一個涵蓋十個日常生活領域、共 200 個長時程任務的評測基準。每個任務都是一條腳本化的多週時間線,位於一個由 22 個模擬服務組成的模擬世界中。世界依自身時鐘推進,且其中許多變化是靜默的,因此只有會重新檢視世界的智能體才能發現這些變化。每個任務都以細粒度、加權的檢查來評分,這些檢查只讀取智能體實際留下的痕跡,涵蓋最終狀態、行動的及時性,以及是否遵守了隱含約束。我們評估了七個前沿模型。所有模型得分都很低,這顯示當前智能體距離真實生活輔助仍有很大差距。我們將開源所有任務、環境與評測框架。
English
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.