ChatPaper.aiChatPaper

Qwen-UI-Agent 기술 보고서: 차세대 실세계 중심 파운데이션 GUI 에이전트를 향하여

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

July 30, 2026
저자: Hanzhang Zhou, Panrong Tong, Xu Zhang, Quyu Kong, Chenglin Cai, Tianyu Xia, Gongjie Zhang, Jianan Zhang, Long Li, Long Chen, Lei Wang, Gaole Dai, Pengxiang Li, Liangyu Chen, Yue Wang, Steven Hoi
cs.AI

초록

GUI 에이전트는 기존 디지털 기기 위에서 범용 실행기가 될 잠재력을 지니고 있다. 실제 세계에서의 사용으로 발전시키기 위해, 우리는 실제 기기에서 안정적으로 작동하고, 플랫폼 간 워크플로우를 실행하며, GUI 상호작용과 CLI 실행을 결합하고, 장기 지평 작업을 완료하며, 유용한 서비스를 선제적으로 시작하고, 최소한의 인간 노력으로 능력을 자율적으로 개선하는 에이전트를 구상한다. 이러한 비전에 따라, 우리는 모바일, 컴퓨터 사용, 웹 및 DeepSearch 환경을 포괄하는 실제 세계 중심의 기반 GUI 에이전트인 Qwen-UI-Agent를 제시한다. Qwen-UI-Agent는 다양한 샌드박스 환경과 대규모 실제 기기 모바일 런타임을 결합한다. 통합된 행동 공간은 GUI 작업과 CLI 실행을 교차시키며, 단일 모델 턴에서 배치된 행동을 생성한다. AutoResearch 방식의 데이터 플라이휠은 에이전트를 사용하여 작업과 환경을 구성하고, 실패를 진단하며, 후속 반복을 계획한다. 온라인 강화학습은 100턴을 초과하는 궤적에 대한 훈련을 지원하며, 10,000개 이상의 동시 환경이 롤아웃을 가속화한다. 경량 하네스 계층은 모바일과 컴퓨터 전반에서 선제적 서비스 시작과 상태 유지 워크플로우를 지원한다. 광범위한 평가 제품군 전반에 걸쳐, Qwen-UI-Agent는 모바일 사용 벤치마크에서 최첨단 성능을 달성하는 동시에 Opus 4.8, Gemini 3.1 Pro 및 GPT-5.6 Sol을 포함한 최첨단 모델들과 비교해 컴퓨터 및 브라우저 사용 작업에서 경쟁력 있는 성능을 제공한다. 모바일 사용에서는 MobileWorld에서 82.1%, MobileWorld-Real에서 92.2%, AndroidDaily에서 97.5%를 달성한다. 컴퓨터 사용에서는 OSWorld-Verified에서 79.5%, OSWorld-v2에서 40.0%의 부분 진행 점수를 달성한다. 브라우저 사용 및 GUI 그라운딩에서는 각각 WebArena에서 73.6%, ScreenSpot-Pro에서 81.5%를 달성한다.
English
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.