llm-finetuning task-complexity ui-control-agents android-control-dataset high-level-tasks low-level-tasks out-of-domain-performance data-scale
Abstract
Autonomous agents that control computer interfaces to accomplish human tasks are emerging. Leveraging LLMs to power such agents has been of special interest, but unless fine-tuned on human-collected task demonstrations, performance is still relatively low. In this work we study whether fine-tuning alone is a viable approach for building real-world computer control agents. In particularly, we investigate how performance measured on both high and low-level tasks in domain and out of domain scales as more training data is collected. To this end we collect and release a new dataset, AndroidControl, consisting of 15,283 demonstrations of everyday tasks with Android apps. Compared to existing datasets, each AndroidControl task instance includes both high and low-level human-generated instructions, allowing us to explore the level of task complexity an agent can handle. Moreover, AndroidControl is the most diverse computer control dataset to date, including 14,548 unique tasks over 833 Android apps, thus allowing us to conduct in-depth analysis of the model performance in and out of the domain of the training data. Using the dataset, we find that when tested in domain fine-tuned models outperform zero and few-shot baselines and scale in such a way that robust performance might feasibly be obtained simply by collecting more data. Out of domain, performance scales significantly more slowly and suggests that in particular for high-level tasks, fine-tuning on more data alone may be insufficient for achieving robust out-of-domain performance.
한국어 요약
한 줄 요약
UI 컨트롤 에이전트의 성능이 데이터 규모에 어떻게 영향을 받는지 AndroidControl 데이터셋을 통해 분석한 연구.
핵심 기여도
- AndroidControl이라는 새로운 UI 컨트롤 데이터셋을 수집·공개 (15,283개의 데모, 833개 앱, 14,548개의 고유 작업 포함).
- LoRA 미세조정 모델의 성능이 도메인 내에서는 데이터 증가에 따라 상승하지만, 도메인 외에서는 1~2개 수준의 데이터 증가가 필요하다는 분석.
- 고레벨 작업(5단계)의 경우 도메인 내에서 2M 에피소드가 95% 완료율 달성에 필요하다는 추정.
- 기존 제로샷·피샷 기반 모델 대비 미세조정 모델이 도메인 내 성능에서 우수함을 보임.
핵심 아이디어
UI 컨트롤 에이전트는 인간의 작업을 수행하기 위해 디바이스 화면을 관찰하고 클릭, 타이핑 등의 행동을 생성해야 한다. 본 연구는 미세조정(fine-tuning)만으로도 실제 UI 컨트롤 에이전트를 구축할 수 있는지 탐구하며, 특히 데이터 규모가 성능에 미치는 영향을 분석한다. 기존 연구는 주로 고레벨 추론(예: 작업 계획)과 저레벨 행동(예: 클릭)을 별개로 다루었지만, 본 연구는 AndroidControl 데이터셋을 통해 둘을 동시에 평가할 수 있는 구조를 제공한다. 이는 고레벨 작업과 저레벨 작업의 복잡성 차이를 명확히 분석하는 데 기여한다.
기술적 접근법
- **데이터셋**: AndroidControl (15,283개 데모, 833개 앱, 14,548개 고유 작업 포함).
- **모델**: LoRA 기반 미세조정 모델 사용.
- **평가 지표**: 에피소드 완료율(episode completion rate)을 주요 평가 지표로 사용.
- **비교 대상**: 제로샷, 피샷 기반 모델과 비교.
- **데이터 분할**: 도메인 내/외 작업을 구분하여 성능 분석.
- **하이퍼파라미터**: 명시되지 않음.
주요 결과
- **도메인 내 성능**: 미세조정 모델이 제로샷/피샷 기반 모델 대비 우수한 성능을 보임.
- **저레벨 작업**: 1M 에피소드로 95% 완료율 달성 가능.
- **고레벨 작업**: 2M 에피소드로 95% 완료율 달성 가능.
- **도메인 외 성능**: 고레벨 작업의 경우 150M 에피소드가 필요하며, 이는 도메인 내 대비 1~2개 수준의 데이터 증가를 요구함.
- **결론**: 미세조정은 도메인 내에서는 효과적이지만, 도메인 외에서는 데이터 증가만으로는 한계가 있음.
의의 및 한계
- **의의**: AndroidControl은 가장 다양한 UI 컨트롤 데이터셋으로, 도메인 내/외 성능을 비교할 수 있는 구조를 제공.
- **한계**: 분석은 단일 모델(LoRA 기반)에 기반하며, 다른 모델 아키텍처나 학습 전략에 대한 일반화는 제한적.
- **추가 연구 필요**: 도메인 외 성능 향상을 위한 새로운 학습 전략(예: 멀티도메인 학습, 강화학습)이 필요함.
실용적 활용
- 모바일 앱 자동화, 고객 지원 자동화, 테스트 자동화 등에서 UI 컨트롤 에이전트를 적용 가능.
- AndroidControl 데이터셋은 연구자들이 UI 컨트롤 에이전트의 성능을 정량적으로 평가하는 데 유용한 자원.
- 도메인 내 성능 향상을 위한 데이터 수집 전략 수립에 활용 가능.