On the Effects of Data Scale on UI Control Agents

Wei Li, Will Bishop, Alice Li, Christopher Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, O. Riva

arXiv:2406.03679 · 2026-07-27 공개 · arXiv · PDF

llm-finetuning task-complexity ui-control-agents android-control-dataset high-level-tasks low-level-tasks out-of-domain-performance data-scale

Abstract

Autonomous agents that control computer interfaces to accomplish human tasks are emerging. Leveraging LLMs to power such agents has been of special interest, but unless fine-tuned on human-collected task demonstrations, performance is still relatively low. In this work we study whether fine-tuning alone is a viable approach for building real-world computer control agents. In particularly, we investigate how performance measured on both high and low-level tasks in domain and out of domain scales as more training data is collected. To this end we collect and release a new dataset, AndroidControl, consisting of 15,283 demonstrations of everyday tasks with Android apps. Compared to existing datasets, each AndroidControl task instance includes both high and low-level human-generated instructions, allowing us to explore the level of task complexity an agent can handle. Moreover, AndroidControl is the most diverse computer control dataset to date, including 14,548 unique tasks over 833 Android apps, thus allowing us to conduct in-depth analysis of the model performance in and out of the domain of the training data. Using the dataset, we find that when tested in domain fine-tuned models outperform zero and few-shot baselines and scale in such a way that robust performance might feasibly be obtained simply by collecting more data. Out of domain, performance scales significantly more slowly and suggests that in particular for high-level tasks, fine-tuning on more data alone may be insufficient for achieving robust out-of-domain performance.

한국어 요약

한 줄 요약

UI 컨트롤 에이전트의 성능이 데이터 규모에 어떻게 영향을 받는지 AndroidControl 데이터셋을 통해 분석한 연구.

핵심 기여도

핵심 아이디어

UI 컨트롤 에이전트는 인간의 작업을 수행하기 위해 디바이스 화면을 관찰하고 클릭, 타이핑 등의 행동을 생성해야 한다. 본 연구는 미세조정(fine-tuning)만으로도 실제 UI 컨트롤 에이전트를 구축할 수 있는지 탐구하며, 특히 데이터 규모가 성능에 미치는 영향을 분석한다. 기존 연구는 주로 고레벨 추론(예: 작업 계획)과 저레벨 행동(예: 클릭)을 별개로 다루었지만, 본 연구는 AndroidControl 데이터셋을 통해 둘을 동시에 평가할 수 있는 구조를 제공한다. 이는 고레벨 작업과 저레벨 작업의 복잡성 차이를 명확히 분석하는 데 기여한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용