ChatPaper.aiChatPaper

UI-Venus-2 技术报告

UI-Venus-2 Technical Report

August 27, 2026
作者: Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
cs.AI

摘要

多模态GUI智能体已成为数字任务自动化的前沿范式,然而,从以基准测试为导向的模型迈向可靠的现实世界应用仍面临诸多挑战,其根源在于环境覆盖有限、任务构造脆弱以及奖励验证不可靠。为此,我们提出UI-Venus-2,一种通用型基础GUI智能体,能够通过统一的闭环推理—行动框架在移动、网页与桌面环境中运行。为弥合实际部署之间的鸿沟,我们对三个关键维度进行了协同扩展:(1) 环境维度,将覆盖范围扩展至170余个多语言移动应用及原生桌面操作系统;(2) 任务维度,引入深度研究流水线,实现基于功能的指令生成;(3) 验证维度,采用轨迹级与样本级评估器,并结合视觉关键点与多模型投票机制,从而为训练提供可靠的强化学习信号。此外,我们还整合了安全感知机制,确保关键操作得到受控执行。凭借自身能力强、效率高且开源的基座特性,UI-Venus-2推动了该领域向着更可泛化、更可验证、更可自省的现实世界应用智能体方向发展。
English
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.