Qwen-UI-Agent 技術報告:邁向下一代以真實世界為中心的基礎 GUI 代理
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
July 30, 2026
作者: Hanzhang Zhou, Panrong Tong, Xu Zhang, Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang, Long Li, Long Chen, Lei Wang, Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi
cs.AI
摘要
GUI代理有望成為現有數位裝置上的通用執行器。為推動其邁向真實世界應用,我們設想代理能在真實裝置上可靠運作、跨平台執行工作流程、結合GUI互動與CLI執行、完成長時程任務、主動發起有用服務,並以最少的人力投入自主提升其能力。在此願景指引下,我們提出Qwen-UI-Agent,一個以真實世界為核心的基礎GUI代理,涵蓋行動裝置、電腦操作、網頁及DeepSearch等環境。Qwen-UI-Agent結合多樣化的沙箱環境與大規模真實裝置行動執行環境。其統一動作空間將GUI操作與CLI執行交錯結合,並在單一模型回合中產生批次動作。自動研究式資料飛輪利用代理來建構任務與環境、診斷失敗原因,並規劃後續迭代。線上強化學習支援超過100回合的軌跡訓練,並以超過10,000個並行環境加速資料蒐集。輕量化的中介層支援行動裝置與電腦之間的主動服務發起及有狀態工作流程。
在一系列廣泛的評測中,Qwen-UI-Agent在行動裝置使用基準上達到最先進效能,同時在電腦及瀏覽器操作任務上展現與前沿模型(包括Opus 4.8、Gemini 3.1 Pro與GPT-5.6 Sol)相當的競爭力。在行動裝置使用方面,其在MobileWorld上達到82.1%、在MobileWorld-Real上達到92.2%、在AndroidDaily上達到97.5%。在電腦操作方面,其在OSWorld-Verified上達到79.5%,並在OSWorld-v2上取得40.0%的部分進展得分。在瀏覽器操作與GUI定位方面,其分別在WebArena上達到73.6%,以及在ScreenSpot-Pro上達到81.5%。
English
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer.
Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.