ChatPaper.aiChatPaper

UI-Venus-2 技術報告

UI-Venus-2 Technical Report

August 27, 2026
作者: Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
cs.AI

摘要

多模態GUI代理已成為數位任務自動化的一個前景可期的典範,然而,由於環境覆蓋範圍有限、任務建構脆弱以及獎勵驗證不可靠,從以基準為導向的模型過渡到可靠的實際應用仍然充滿挑戰。在本研究中,我們提出了 UI-Venus-2,這是一個通用型基礎GUI代理,旨在透過統一的閉環推理-行動框架,在行動裝置、網頁和桌面環境中運行。為縮小通往實際部署的差距,我們聯合擴展了三個關鍵維度:(1) 環境——將覆蓋範圍擴展至逾170個多語言行動應用程式及原生桌面作業系統;(2) 任務——採用深度研究管線進行以功能為根基的指令生成;(3) 驗證——採用具有視覺關鍵點與多模型投票的軌跡層級與樣本層級評估器,以確保為訓練提供可靠的強化學習訊號。此外,我們整合了安全感知機制,以確保高影響力行動的受控執行。透過提供一個功能強大、高效且開源的基礎,UI-Venus-2 推動該領域朝向更可泛化、可驗證且具自我反思能力的代理邁進,以因應實際世界的應用需求。
English
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification. In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework. To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training. Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.