ChatPaper.aiChatPaper

VibeLifeBench: 당신의 라이프 에이전트는 살아있는 세계에서 능동적이고 지속적으로 행동할 수 있는가?

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

August 11, 2026
저자: Xiaohongshu Inc
cs.AI

초록

대규모 언어 모델(LLM) 에이전트는 점점 더 개인 비서로 배포되고 있다. 그러나 기존 평가들은 대부분 정적 환경에서 짧고 독립적인 요청을 사용한다. 일상생활 지원은 상황이 다르다. 작업은 몇 분이 아니라 몇 주에 걸쳐 진행된다. 에이전트가 프롬프트를 받지 않는 동안에도 세계는 계속 변화한다. 많은 제약 조건은 명시적으로 언급되지 않는다. 눈앞의 요청에 단순히 응답하는 에이전트는 그러한 과제에서 실패할 것이다. 대신 필요한 것은 능동적이고 일관성 있는 에이전트다. 언제 행동하고, 언제 질문하고, 언제 침묵할지를 스스로 결정한다. 아무도 알리지 않은 변화를 감지한다. 첫날부터 마지막 날까지 하나의 계획을 일관되게 유지한다. 현재 어떤 벤치마크도 이를 측정하지 않는다. 우리는 10개의 일상생활 영역에 걸친 200개의 장기 과제로 구성된 벤치마크인 VibeLifeBench를 소개한다. 각 과제는 22개의 모의 서비스로 구성된 시뮬레이션 세계에서의 스크립트 기반 다주간 타임라인이다. 세계는 자체 시계에 따라 진행되며, 많은 변화는 알림 없이 발생하므로 세계를 다시 점검하는 에이전트만이 이를 발견할 수 있다. 각 과제는 에이전트가 실제로 남긴 것만을 확인하는 세분화된 가중치 기반 검사로 평가되며, 최종 상태, 행동의 적시성, 암묵적 제약 조건의 준수 여부를 포함한다. 우리는 7개의 프론티어 모델을 평가했다. 모든 모델이 낮은 점수를 기록했으며, 이는 현재 에이전트들이 실제 생활을 지원하는 데 얼마나 부족한지를 보여준다. 우리는 모든 과제, 환경, 평가 프레임워크를 오픈소스로 공개할 예정이다.
English
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.