UI-Venus-2 Technical Report

Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou

arXiv:2609.00028 · 2026-09-02 공개 · arXiv · PDF

foundation-models rl-training gui-agents multimodal verification task-automation safety-mechanisms closed-loop-reasoning

Abstract

Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

한국어 요약

한 줄 요약

UI-Venus-2는 모바일, 웹, 데스크톱 환경에서 작동하는 다목적 GUI 에이전트로, 170개 이상의 앱과 강력한 검증 메커니즘을 통해 실용성을 확보한다.

핵심 기여도

핵심 아이디어

기존 GUI 에이전트는 주로 벤치마크 중심의 제한된 환경에서 작동하며, 실제 세계 적용 시 환경 커버리지, 태스크 구성, 보상 검증의 문제를 겪는다. UI-Venus-2는 이 세 가지 차원을 **동시 확장**하는 전략을 채택한다. 특히, **기능 기반 인스트럭션 생성 파이프라인**을 통해 실제 앱 기능에 부합하는 태스크를 대규모로 생성하고, **트레이스 레벨 및 샘플 레벨 검증**을 도입하여 정확한 보상 신호를 확보한다. 또한, CAPTCHA 처리를 포함한 **종단 간 자율 작동**을 지원함으로써 데이터 확장을 촉진한다.

기술적 접근법

주요 결과

의의 및 한계

UI-Venus-2는 기존 GUI 에이전트의 벤치마크 중심 접근을 넘어, 실제 세계에서 작동 가능한 **일반적이고 검증 가능한 에이전트**로의 진화를 제시한다. 특히, **다중 플랫폼**(모바일, 웹, 데스크톱)에서의 일반화 능력과 **강력한 검증 메커니즘**은 실용적 가치를 높인다. 그러나, **모든 앱 및 웹사이트를 커버하는 것은 여전히 어려움**이 있으며, **복잡한 인터페이스 변화에 대한 적응력**은 추가 연구가 필요하다.

실용적 활용

UI-Venus-2는 **자동화된 디지털 작업**(예: 로그인, 등록, 탐색)을 필요로 하는 산업 분야에서 활용 가능하며, **연구자들이 안전하고 신뢰할 수 있는 GUI 에이전트를 개발하는 기초 플랫폼**으로도 사용될 수 있다. 특히, **다국어 및 다플랫폼 환경**에서의 자동화 요구가 높은 기업 및 연구소에 유용하다.