VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Xiaohongshu Inc

arXiv:2608.10875 · 2026-08-12 공개 · arXiv · PDF

llm-agents long-horizon agent-evaluation evaluation-framework proactive-agents simulated-environments life-assistance vibelifebench

Abstract

Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.

한국어 요약

한 줄 요약

VibeLifeBench는 일상 생활 도메인에서 주도적이고 지속적인 대리인의 능력을 평가하기 위한 200개의 장기 태스크 기반 벤치마크이다.

핵심 기여도

핵심 아이디어

VibeLifeBench는 기존 벤치마크가 간단하고 정적 환경에서 단기 작업을 중심으로 평가하는 반면, 실제 일상 생활에서는 주도적이고 지속적인 행동이 필요하다는 점을 반영하여 설계되었다. 예를 들어, 여행 계획을 수주일부터 종료일까지 일관되게 유지하면서 예약 변경, 날씨 변화, 예산 제약 등을 자동으로 감지하고 조정해야 한다. 이는 단순히 명령을 수행하는 대리인을 넘어, 환경 변화를 자발적으로 인식하고 계획을 조정하는 능력을 평가하는 데 목적이 있다. VibeLifeBench는 이러한 능력을 세 가지 핵심 속성—주도성(proactivity), 동적 세계 적응(living-world adaptation), 장기 일관성(long-horizon coherence)—으로 정의하고, 이를 구체적인 시뮬레이션 태스크와 점검 기준으로 측정한다.

기술적 접근법

주요 결과

의의 및 한계

VibeLifeBench는 일상 생활 도메인에서의 대리인 평가 기준을 명확히 정의하고, 기존 벤치마크가 간과한 주도성, 동적 세계 적응, 장기 일관성을 측정하는 첫 시도이다. 특히, 22개의 모의 서비스와 200개의 장기 태스크를 통해 실제 생활 환경을 유사하게 재현하고, 12,261개의 세부 점검 항목을 통해 정밀한 평가가 가능하다는 점에서 학술적 의의가 크다. 그러나 일부 태스크는 실제 세계보다 단순화되어 있을 수 있으며, 사용자와의 상호작용을 고려한 평가가 부족한 점이 한계로 지적된다.

실용적 활용

VibeLifeBench는 가정용 AI 개인 비서, 건강 관리 대리인, 재정 계획 도구 등 일상 생활 지원 시스템 개발에 활용될 수 있다. 특히, 장기 계획을 유지하면서 환경 변화를 자동으로 반영하는 능력을 가진 대리인의 개발을 촉진할 수 있다.