ChatPaper.aiChatPaper

UI-Venus-2 기술 보고서

UI-Venus-2 Technical Report

August 27, 2026
저자: Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
cs.AI

초록

멀티모달 GUI 에이전트는 디지털 작업 자동화를 위한 유망한 패러다임으로 부상했지만, 벤치마크 중심 모델에서 신뢰할 수 있는 실세계 애플리케이션으로의 전환은 제한된 환경 커버리지, 취약한 작업 구성, 그리고 불신뢰성 높은 보상 검증으로 인해 여전히 어려운 과제로 남아 있다. 본 연구에서는 모바일, 웹, 데스크톱 환경 전반에서 통합된 폐루프 추론-행동 프레임워크를 통해 작동하는 범용 기반 GUI 에이전트인 UI-Venus-2를 제시한다. 실질적 배포로의 격차를 해소하기 위해, 우리는 세 가지 핵심 차원을 공동으로 확장한다: (1) 환경(Environments): 170개 이상의 다국어 모바일 앱과 네이티브 데스크톱 운영체제로 커버리지 확장; (2) 작업(Tasks): 기능 기반 명령 생성을 위한 심층 조사(deep-research) 파이프라인 도입; (3) 검증(Verification): 훈련을 위한 신뢰할 수 있는 강화학습 신호를 보장하기 위해 시각적 키포인트와 다중 모델 투표를 활용한 추적 수준(trace-level) 및 샘플 수준(sample-level) 평가자 채택. 또한, 중대한 결과를 초래하는 행동의 통제된 실행을 보장하기 위해 안전 인식 메커니즘을 통합한다. UI-Venus-2는 유능하고 효율적이며 오픈소스인 기반을 제공함으로써, 실세계 애플리케이션을 위한 보다 일반화 가능하고 검증 가능하며 자기반성적 에이전트로의 발전을 촉진한다.
English
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.