ChatPaper.aiChatPaper

VibeLifeBench:あなたのライフエージェントは、生きた世界でプロアクティブかつ粘り強く行動できるか?

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

August 11, 2026
著者: Xiaohongshu Inc
cs.AI

要旨

大規模言語モデル(LLM)エージェントは、個人アシスタントとしてますます広く展開されている。しかし、既存の評価は、そのほとんどが静的な環境における短く自己完結的なリクエストに基づいている。日常生活の支援はこれとは異なる。タスクは数分ではなく数週間にわたって継続する。エージェントがプロンプトを受けていない間に、世界は変化し続ける。多くの制約は、決して明示的に述べられることはない。目前のリクエストに単に応答するだけのエージェントは、そのようなタスクでは失敗するだろう。代わりに必要とされるのは、積極的(プロアクティブ)でありながら一貫性を保つエージェントである。すなわち、いつ行動し、いつ質問し、いつ沈黙するかを自ら決定し、誰も知らせてくれない変化に気づき、初日から最終日まで単一の計画を首尾一貫して維持するエージェントである。現在のベンチマークではこれを測定するものはない。我々は、10の日常生活領域にわたる200の長期的タスクからなるベンチマーク、VibeLifeBenchを紹介する。各タスクは、22の模擬サービスからなるシミュレーション世界における、スクリプト化された複数週間のタイムラインである。世界は独自の時計で進行し、その変化の多くは静かであるため、世界を再調査するエージェントだけがそれらを発見できる。すべてのタスクは、エージェントが実際に残した出力のみを読み取る、きめ細かい重み付きチェックによって採点される。このチェックは、最終状態、アクションの適時性、および暗黙の制約を順守したかどうかを網羅する。我々は7つの最先端モデルを評価した。そのすべてが低スコアであり、現在のエージェントが実際の生活の支援からいかに遠いかを示している。我々は、すべてのタスク、環境、および評価フレームワークをオープンソースとして公開する予定である。
English
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.