UI-Mate: インコンテキストデモンストレーションによるオープンウェイト基盤GUIエージェントの高度化
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
August 16, 2026
著者: Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng
cs.AI
要旨
基盤GUIエージェントは複雑なデジタルタスクを自動化できるが、その展開は、不足し偏ったトレーニングデータ、曖昧なプロンプト、そして信頼性の低い実行によって妨げられている。日常的なワークフローはユーザー固有のツールや暗黙の慣習に依存しているため、明示されていない指示は実行ごとに恣意的なばらつきを生む可能性がある。我々は、環境に基づくトレーニングスタックと文脈内デモンストレーション学習を統合した基盤GUIエージェントであるUI-Mateを提案する。UI-Mateは以下の3つの貢献を行う。
スケーラブルな環境基盤型トレーニングスタック: 閉ループのデータエンジンは、統合タスク検証器バンドルを通じて、大規模並列環境全体でタスク生成、環境構築、ロールアウト、フィルタリング、能力バランス調整、SFT、およびオンラインRLを自動化する。
文脈内デモンストレーション学習: マルチモーダルなデモンストレーションを柔軟なサブタスクレベルのワークフローに変換し、関連する実演ステップに従い、ライブインターフェースから再計画するメカニズム。
OSWorkerBenchベンチマークと知見: 41のアプリケーションにわたる100の長期的なオフィスタスクからなるベンチマークであり、指示のみによる評価とデモンストレーション誘導による評価の両方をサポートする。そのデモンストレーションリソースは、同じターゲットに対する強力なエージェントの成功ロールアウトから構築された33タスクの自己デモ設定と、関連するが同一ではないタスクの人間による記録から構築された45タスクの変種デモ設定を分離する。
実験により、UI-Mate-27Bは一般的なコンピュータ操作ベンチマークにおいて、オープンウェイトの新たな最高水準(SOTA)を達成し、OSWorld-Verifiedで77.0%、WindowsAgentArenaで66.2%を記録することを示す。OSWorkerBenchでは、厳密成功率41.0%、進捗率76.9%を達成し、ベースモデルであるQwen3.6-27Bをそれぞれ17.7ポイントおよび24.5ポイント上回る。33タスクの自己デモサブセットでは、単一のデモンストレーションにより厳密成功率が17.2%から35.4%へ、進捗率が67.9%から81.1%へ向上し、長期的な信頼性が大幅に改善される。
プロジェクトページ: https://ui-mate.github.io
English
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.