ChatPaper.aiChatPaper

UI-Mate:透過上下文示範推進開放權重基礎GUI智能體

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

August 16, 2026
作者: Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng
cs.AI

摘要

基礎GUI代理可以自動化複雜的數位任務,但部署受到稀缺且有偏誤的訓練資料、模糊的提示以及不可靠的執行所阻礙。常規工作流程依賴於使用者特定的工具和默契慣例,因此未明說的指令可能在多次執行之間產生任意差異。我們提出UI-Mate,一個基礎GUI代理,整合了環境接地的訓練堆疊與上下文內的示範學習。UI-Mate有三項貢獻:可擴展的環境接地訓練堆疊:一個閉環資料引擎,透過統一的任務驗證器套件,在大規模並行環境中自動化任務生成、環境建構、展開、過濾、能力平衡、監督式微調(SFT)和線上強化學習。上下文內的示範學習:一種將多模態示範轉換為彈性子任務層級流程、遵循相關示範步驟並從即時介面重新規劃的機制。OSWorkerBench基準與洞見:一個橫跨41個應用程式、包含100個長時程辦公室任務的基準,支援僅指令和示範引導的評估。其示範資源將一個33任務的自我示範設定(由同一目標的成功強代理展開軌跡所建構)與一個45任務的變體示範設定(由相關但不完全相同任務的人類錄製所建構)加以區分。實驗顯示,UI-Mate-27B在一般電腦使用基準上樹立了新的開放權重最先進水平,在OSWorld-Verified上得分77.0%,在WindowsAgentArena上得分66.2%。在OSWorkerBench上,它達到41.0%的嚴格成功率和76.9%的進展,比其Qwen3.6-27B基礎模型高出17.7和24.5個百分點。在33任務的自我示範子集中,一個示範將嚴格成功率從17.2%提升至35.4%,進展從67.9%提升至81.1%,大幅提高了長時程可靠性。專案頁面:https://ui-mate.github.io。
English
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.