UI-Mate: 인컨텍스트 시연을 통한 오픈 가중치 기반 GUI 에이전트의 고도화
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
August 16, 2026
저자: Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng
cs.AI
초록
기반 GUI 에이전트는 복잡한 디지털 작업을 자동화할 수 있지만, 배포는 희소하고 편향된 학습 데이터, 모호한 프롬프트, 불안정한 실행으로 인해 저해된다. 일상적인 워크플로는 사용자별 도구와 암묵적 관례에 의존하므로, 명시되지 않은 지시는 실행 간에 임의적인 변동을 초래할 수 있다. 우리는 환경 기반 학습 스택과 맥락 내 데모 학습을 통합한 기반 GUI 에이전트인 UI-Mate를 제시한다. UI-Mate는 세 가지 기여를 한다. 확장 가능한 환경 기반 학습 스택: 폐루프 데이터 엔진은 통합 작업-검증기 번들을 통해 대규모 병렬 환경에서 작업 생성, 환경 구축, 롤아웃, 필터링, 능력 균형 조정, SFT, 온라인 RL을 자동화한다. 맥락 내 데모 학습: 다중 모달 데모를 유연한 하위 작업 수준 워크플로로 변환하고, 관련 데모 단계를 따르며, 실시간 인터페이스에서 재계획하는 메커니즘. OSWorkerBench 벤치마크 및 통찰: 41개 애플리케이션에 걸친 100개의 장기 지평 사무 작업 벤치마크로, 지시문 전용 및 데모 유도 평가를 지원한다. 데모 자원은 동일한 목표에 대한 성공적인 강력한 에이전트 롤아웃으로 구축된 33개 작업의 자체 데모 설정과, 관련되지만 동일하지 않은 작업의 인간 녹화로 구축된 45개 작업의 변형 데모 설정으로 구분된다. 실험 결과 UI-Mate-27B는 일반 컴퓨터 사용 벤치마크에서 새로운 공개 가중치 최고 수준을 달성하여, OSWorld-Verified에서 77.0%, WindowsAgentArena에서 66.2%를 기록했다. OSWorkerBench에서는 엄격 성공 41.0%, 진행률 76.9%를 달성하여 기본 모델인 Qwen3.6-27B보다 각각 17.7포인트와 24.5포인트 앞섰다. 33개 작업의 자체 데모 하위 집합에서 데모 하나를 제공하면 엄격 성공이 17.2%에서 35.4%로, 진행률이 67.9%에서 81.1%로 향상되어 장기 지평 신뢰성을 크게 개선한다. 프로젝트 페이지: https://ui-mate.github.io.
English
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce arbitrary variations across runs. We present UI-Mate, a foundation GUI agent that integrates an environment-grounded training stack with in-context demonstration learning. UI-Mate makes three contributions: A Scalable Environment-Grounded Training Stack: A closed-loop data engine automates task generation, environment construction, rollout, filtering, capability balancing, SFT, and online RL across massively parallel environments via unified task-verifier bundles. In-Context Demonstration Learning: A mechanism that transforms multimodal demonstrations into flexible subtask-level workflows, follows relevant demonstrated steps, and re-plans from the live interface. OSWorkerBench Benchmark and Insights: A benchmark of 100 long-horizon office tasks across 41 applications that supports instruction-only and demonstration-guided evaluation. Its demonstration resources separate a 33-task self-demo setting, built from successful strong-agent rollouts of the same targets, from a 45-task variant-demo setting, built from human recordings of related but non-identical tasks. Experiments show that UI-Mate-27B sets a new open-weight state of the art on general computer-use benchmarks, scoring 77.0% on OSWorld-Verified and 66.2% on WindowsAgentArena. On OSWorkerBench, it reaches 41.0% strict success and 76.9% progress, outperforming its Qwen3.6-27B base by 17.7 and 24.5 points. On the 33-task self-demo subset, one demonstration raises strict success from 17.2% to 35.4% and progress from 67.9% to 81.1%, substantially improving long-horizon reliability. Project page: https://ui-mate.github.io.