ChatPaper.aiChatPaper

Qwen-UIエージェント技術報告書:次世代の実世界中心基盤GUIエージェントを目指して

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

July 30, 2026
著者: Hanzhang Zhou, Panrong Tong, Xu Zhang, Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang, Long Li, Long Chen, Lei Wang, Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi
cs.AI

要旨

GUIエージェントは、既存のデジタルデバイスに対する汎用実行器となる可能性を秘めている。実世界での利用に向けて、我々は、実デバイス上で信頼性高く動作し、プラットフォーム横断的にワークフローを実行し、GUI操作とCLI実行を組み合わせ、長期的なタスクを完了し、有用なサービスを能動的に開始し、最小限の人間の労力で能力を自律的に向上させるエージェントを構想している。このビジョンに基づき、我々は、モバイル、コンピュータ利用、ウェブ、DeepSearch環境にわたる実世界中心の基盤GUIエージェントであるQwen-UI-Agentを提案する。Qwen-UI-Agentは、多様なサンドボックス環境と大規模な実機モバイルランタイムを組み合わせる。その統一されたアクション空間は、GUI操作とCLI実行を織り交ぜ、単一のモデルターンでバッチ化されたアクションを生成する。AutoResearch方式のデータフライホイールは、エージェントを使用してタスクと環境を構築し、障害を診断し、その後の反復を計画する。オンライン強化学習は、100ターンを超える軌跡でのトレーニングをサポートし、10,000を超える並行環境がロールアウトを高速化する。軽量なハーネス層は、モバイルとコンピュータにわたる能動的なサービス開始とステートフルなワークフローをサポートする。 広範な評価スイートにわたり、Qwen-UI-Agentは、モバイル利用ベンチマークで最先端の性能を達成しつつ、Opus 4.8、Gemini 3.1 Pro、GPT-5.6 Solなどのフロンティアモデルに対して、コンピュータ利用およびブラウザ利用タスクで競争力のある性能を発揮する。モバイル利用では、MobileWorldで82.1%、MobileWorld-Realで92.2%、AndroidDailyで97.5%を達成する。コンピュータ利用では、OSWorld-Verifiedで79.5%、OSWorld-v2で部分進捗スコア40.0%を達成する。ブラウザ利用とGUIグラウンディングでは、それぞれWebArenaで73.6%、ScreenSpot-Proで81.5%を達成する。
English
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.