Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, E. Schoop, Floris Weers, Amanda Swearngin, Jeffrey Nichols, Yinfei Yang, Zhe Gan

arXiv:2404.05719 · 2026-07-27 공개 · arXiv · PDF

vision-language instruction-following multimodal-llms grounding aspect-ratio ferret-ui mobile-ui icon-recognition

Abstract

Recent advancements in multimodal large language models (MLLMs) have been noteworthy, yet, these general-domain MLLMs often fall short in their ability to comprehend and interact effectively with user interface (UI) screens. In this paper, we present Ferret-UI, a new MLLM tailored for enhanced understanding of mobile UI screens, equipped with referring, grounding, and reasoning capabilities. Given that UI screens typically exhibit a more elongated aspect ratio and contain smaller objects of interest (e.g., icons, texts) than natural images, we incorporate"any resolution"on top of Ferret to magnify details and leverage enhanced visual features. Specifically, each screen is divided into 2 sub-images based on the original aspect ratio (i.e., horizontal division for portrait screens and vertical division for landscape screens). Both sub-images are encoded separately before being sent to LLMs. We meticulously gather training samples from an extensive range of elementary UI tasks, such as icon recognition, find text, and widget listing. These samples are formatted for instruction-following with region annotations to facilitate precise referring and grounding. To augment the model's reasoning ability, we further compile a dataset for advanced tasks, including detailed description, perception/interaction conversations, and function inference. After training on the curated datasets, Ferret-UI exhibits outstanding comprehension of UI screens and the capability to execute open-ended instructions. For model evaluation, we establish a comprehensive benchmark encompassing all the aforementioned tasks. Ferret-UI excels not only beyond most open-source UI MLLMs, but also surpasses GPT-4V on all the elementary UI tasks.

한국어 요약

한 줄 요약

Ferret-UI는 모바일 UI 화면을 정밀하게 이해하고 지시에 따라 행동할 수 있는 다중모달 대형 언어 모델로, GPT-4V보다 기본 UI 작업에서 우수한 성능을 보인다.

핵심 기여도

핵심 아이디어

기존의 다중모달 대형 언어 모델(MLLMs)은 자연 이미지에 최적화되어 있지만, 모바일 UI 화면은 더 긴 종횡비와 더 작은 UI 요소(아이콘, 텍스트)를 포함하므로 일반적인 모델로는 정확한 이해가 어려웠다. Ferret-UI는 UI 화면의 이러한 특성을 고려해 "any resolution" 기법을 도입하여 화면을 2개의 하위 이미지로 분할하고, 각각을 별도로 인코딩함으로써 세부 정보를 유지한다. 이는 Ferret 모델의 기존 구조를 기반으로 하되, UI 화면에 특화된 공간 인식 능력을 강화한 핵심 아이디어이다. 또한, 정밀한 지시 수행과 추론 능력을 키우기 위해 GPT-4를 활용한 고급 작업 데이터셋을 추가로 구축했다.

기술적 접근법

주요 결과

의의 및 한계

Ferret-UI는 모바일 UI 화면의 특성(긴 종횡비, 작은 요소)을 고려한 첫 번째 UI 중심 MLLM으로, UI 자동화, 접근성 향상, 앱 테스팅 등 다양한 분야에 기여할 수 있다. 또한, GPT-4V보다 기본 작업에서 우수한 성능을 보이는 점에서 실용적 가치가 높다. 그러나 고급 작업에서는 여전히 Fuyu나 CogAgent와 비교해 개선이 필요한 부분이 있으며, 일부 작업에서 라벨 노이즈로 인해 성능이 저하될 수 있다는 한계가 있다.

실용적 활용

Ferret-UI는 모바일 앱 자동화, 접근성 도구 개발, UI/UX 테스트, 사용자 행동 분석 등 다양한 산업 및 연구 분야에서 활용 가능하다. 특히, 시각 정보를 기반으로 정밀한 UI 요소를 식별하고 지시에 따라 행동하는 능력은 앱 개발 및 테스트 프로세스를 효율화하는 데 기여할 수 있다.