GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

Geyi Yang, Zikun Qu, Xiang Li, Zhiyong Wang, Min Zhang, Shipei Zeng, Zhongxiang Dai

arXiv:2610.00948 · 2026-10-07 공개 · arXiv · PDF

gui-agents qwen3-vl harness-optimization osworld-verified meta-harness gpt-5 evidence-driven windowsagentarena

Abstract

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

한국어 요약

한 줄 요약

GUI-HARVEST는 GUI 모델의 실행 환경을 자동 최적화하여 동결된 백본 모델 기반의 자가 개선 GUI 에이전트를 구현하는 시스템이다.

핵심 기여도

핵심 아이디어

GUI 에이전트의 실행 환경(런타임 허네스)은 관찰 구성, 동작 실행, 검증, 복구, 종료를 제어하는 핵심 요소이다. 기존 허네스 최적화 방법은 GUI 특성에 맞지 않아, **GUI-HARVEST**는 세 가지 주요 문제를 해결한다:
1. 모델 의도와 실제 시각적 효과 간의 불일치를 진단하는 것.
2. 실행 결과의 변동성으로 인한 실패 원인을 정확히 파악하는 것.
3. 특정 태스크에서의 실패를 일반화하여 재사용 가능한 허네스 수정으로 전환하는 것.

GUI-HARVEST는 **Evidence Analyst**, **Cross-task Clusterer**, **Harness Engineer**, **Validator** 모듈을 통해 실행 증거를 수집하고, 이를 기반으로 허네스를 자동 수정한다. 특히, **before-and-after screenshots**를 사용하여 모델의 의도와 실제 동작 간의 차이를 명확히 파악하며, **within-task comparisons**를 통해 실행 변동성을 분석한다.

기술적 접근법

주요 결과

의의 및 한계

GUI-HARVEST는 GUI 에이전트의 실행 증거를 기반으로 허네스를 자동 최적화함으로써, 동결된 모델에서도 지속적인 성능 향상을 가능하게 한다. 특히, **multimodal execution evidence**를 활용한 진단과 수정은 기존 텍스트 기반 허네스 최적화 방법과 차별화된다.

하지만, **GUI-HARVEST**는 특정 GUI 환경(예: OSWorld, WindowsAgentArena)에서 실험되었으며, 다른 환경으로의 일반화 가능성은 추가 연구가 필요하다. 또한, **source-code edits**는 **bounded**로 제한되어 있어, 복잡한 허네스 구조에 적용 시 한계가 있을 수 있다.

실용적 활용

GUI-HARVEST는 GUI 기반 자동화 시스템, 특히 **OSWorld**, **WindowsAgentArena**와 유사한 환경에서 자가 개선 에이전트를 구축하는 데 유용하다. 또한, **Qwen3-VL-32B-Instruct**, **GPT-5** 등 대형 모델의 실행 환경 최적화에도 활용 가능하다.