Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym

Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang, Dongyeop Kang

arXiv:2609.37267 · 2026-10-06 공개 · arXiv · PDF

proactive-agents llm-harness compute-allocation simulation-evaluation proactivity-gym user-trust llm-judgment task-capability

Abstract

Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a design space organized around five dimensions: task scope, anticipation horizon, activation trigger, processing timing, and intervention depth, and specify the situation and system modeling needed to support its choices, including user and environment representations, backbone LLMs, and agent harnesses. Lastly, we propose PROACTIVITY-GYM, a simulation-based evaluation testbed including multi-day scenarios, stateful environments, and persona-conditioned simulated users that can evaluate the consequences of proactive assistance across interactions. Evaluations across 23 model-harness configurations uncover substantial performance gaps across 3T and reveal that LLM judges often conflate task capability and trust. A human study with 30 participants demonstrates the importance of the joint 3T optimization: participants show sharp trust declines after intervention misalignment despite correct outcomes, and prefer sleep-time assistance, even when imperfect, to preserve ongoing focus. Together, these findings support designing and evaluating proactive agents through the joint consideration of useful work, compute allocation, and evolving user trust.

한국어 요약

한 줄 요약

이 연구는 3T(작업 능력, 시간 할당, 신뢰) 원칙을 기반으로 한 프로액티브 LLM 에이전트의 설계, 구현, 평가 기초를 제시하며, Proactivity-Gym이라는 시뮬레이션 평가 테스트베드를 소개한다.

핵심 기여도

핵심 아이디어

기존 연구는 사용자 요청에 반응하는 **반응형**(reactive) 에이전트에 집중했으나, 이 연구는 사용자가 요청하기 전에 유용한 작업을 수행하는 **프로액티브**(proactive) 에이전트의 설계 기초를 제시한다. 핵심 아이디어는 3T(3T: Task Capability, Temporal Allocation, Trust) 원칙을 기반으로, 작업 능력, 시간 할당, 신뢰를 **동시에 고려**하여 에이전트의 도움이 사용자에게 **실질적 가치**를 제공하도록 하는 것이다.

예를 들어, 에이전트가 올바르게 작업을 수행하더라도, **적절하지 않은 시점**(processing timing)에 개입하거나, **과도한 개입 깊이**(intervention depth)로 인해 사용자의 신뢰(TR)가 급격히 하락할 수 있다. 이는 단순히 작업 성공률(Task Capability, TC)을 높이는 것만으로는 충분하지 않음을 보여준다. 연구는 이러한 3T 요소가 **동시 최적화**되어야 하며, 이를 위해 Proactivity-Gym이라는 시뮬레이션 기반 평가 환경을 제안한다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 프로액티브 에이전트 설계에서 **작업 능력**(TC)만이 아닌, **시간 할당**(TA)과 **사용자 신뢰**(TR)를 함께 고려해야 한다는 점을 강조하며, 3T 최적화의 중요성을 입증한다. Proactivity-Gym은 시뮬레이션 기반 평가 환경으로, 에이전트의 **다양한 상호작용**을 고려한 평가가 가능하다는 점에서 학술적·실용적 가치가 있다.

그러나, 현재의 테스트베드는 **단순화된 자원 제약**과 **사용자 모델**을 기반으로 하며, **장기적 현실 세계 평가**는 아직 이루어지지 않았다. 또한, 3T 간의 **동시 최적화 알고리즘**은 제시되지 않았으며, 향후 연구 주제로 제시된다.

실용적 활용

이 연구는 개인 AI 에이전트가 사용자의 **수면 시간**이나 **공백 시간**을 활용해 작업을 수행하는 데 적용 가능하다. 예를 들어, OpenClaw, Claude Code, Codex와 같은 허네스를 사용한 에이전트는 사용자의 일정과 자원을 고려해 **적절한 시점에 도움을 제공**할 수 있다. 특히, 사용자의 **집중도 유지**와 **신뢰 구축**이 중요한 업무 자동화, 개인 비서, 교육 도우미 등에 활용 가능하다.