Xiaomi-GUI-0 Technisch Rapport
Xiaomi-GUI-0 Technical Report
June 30, 2026
Auteurs: Wanxia Cao, Chengzhen Duan, Pei Fu, Pengzhi Gao, Niu Lian, Fazhan Liu, Hui Liu, Heng Qu, Qinzhuo Wu, Zhehao Yu, Tongbo Chen, Shiqi Cui, Anan Du, Shukai Jia, Yuanfa Li, Yike Liu, Wenchao Lu, Haoyuan Sun, Jiatong Sun, Cheng Tan, Yajie Wang, Changqiao Wu, Tao Xiong, Jiahui Yang, Yuxuan Yuan, Ruoceng Zhang, Shaojie Zhang, Jian Zhu, Jian Luan, Cong Zou
cs.AI
Samenvatting
Grafische gebruikersinterface (GUI)-agenten bouwen voort op visie-taalmodelen om eindgebruikerstaken end-to-end te voltooien in echte toepassingen via interfaceacties zoals tikken, vegen, tekstinvoer en navigatie. Bestaande GUI-agenten worden echter grotendeels getraind en geëvalueerd op offlinetrajecten, gesimuleerde omgevingen en gestandaardiseerde benchmarks. Deze verschillen aanzienlijk van echte toepassingen wat betreft interfacelayout, interactielogica en de verdeling van abnormale toestanden, en kunnen de uitvoeringsstabiliteit in realistisch gebruik niet getrouw karakteriseren, waarbij accountstatussen, machtigingsdialogen, betalingsauthenticatie en risicobeheer voortdurend de toestandsverdeling hervormen en een aanhoudende kloof openen tussen benchmarkscores en echte bruikbaarheid. Om deze kloof te dichten, stellen wij Xiaomi-GUI-0 voor, een native multimodale GUI-agent voor echte mobiele omgevingen, getraind en geëvalueerd binnen een gesloten lus met echte apparaten. De kern is een echte-apparaat-dominante hybride infrastructuur, waarbij fysieke apparaten de primaire uitvoeringsomgeving zijn en sandboxen ondersteunende hulp bieden, zodat gegevensverzameling, training, rollout en evaluatie een uitvoeringsverdeling delen die dicht bij de daadwerkelijke implementatie ligt. We construeren multi-source trainingsgegevens die hoogfrequente hoofdtaken, hoog-generaliserende gegevens voor long-tail-intenties en capaciteitsverbeteringsgegevens voor reflectie en geheugen omvatten, en introduceren een foutgestuurd gegevensvliegwiel dat faaltrajecten omzet in gecorrigeerde acties, reflectieve uitleg en hersteldemonstraties. Het model wordt getraind via een progressieve driefasige pijplijn van gesuperviseerde finetuning, stapsgewijs bekrachtigingsleren en agentisch bekrachtigingsleren. Geëvalueerd op openbare benchmarks en ons eigen RealMobile, behaalt Xiaomi-GUI-0 72,0% succes op RealMobile en 78,9% op AndroidWorld, terwijl de uitvoeringsstabiliteit en herkenning van abnormale toestanden in realistische taken aanzienlijk verbeteren.
English
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigation. However, existing GUI agents are trained and evaluated largely on offline trajectories, simulated environments, and standardized benchmarks. These differ substantially from real applications in interface layout, interaction logic, and abnormal-state distribution, and cannot faithfully characterize execution stability in real-world use, where account states, permission dialogs, payment authentication, and risk control continually reshape the state distribution and open a persistent gap between benchmark scores and real usability. To close this gap, we propose Xiaomi-GUI-0, a native multimodal GUI agent for real mobile environments, trained and evaluated within a real-device closed loop. At its core is a real-device-dominant hybrid infrastructure, where physical devices are the primary execution environment and sandboxes provide auxiliary support, so that data collection, training, rollout, and evaluation share an execution distribution close to real deployment. We construct multi-source training data spanning high-frequency head tasks, high-generalization data for long-tail intents, and capability-enhancement data for reflection and memory, and introduce an error-driven data flywheel that turns failure trajectories into corrected actions, reflective explanations, and recovery demonstrations. The model is trained through a progressive three-stage pipeline of supervised fine-tuning, step-level reinforcement learning, and agentic reinforcement learning. Evaluated on public benchmarks and our in-house RealMobile, Xiaomi-GUI-0 achieves 72.0% success on RealMobile and 78.9% on AndroidWorld, while substantially improving execution stability and abnormal-state recognition in real-world tasks.