ChatPaper.aiChatPaper

KnowAct-GUIClaw: Diep kennen, perfect handelen – Persoonlijke GUI-assistent met zelf-evoluerend geheugen en vaardigheid

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

July 15, 2026
Auteurs: Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang
cs.AI

Samenvatting

OpenClaw is uitgegroeid tot een toonaangevend agentframework voor de automatisering van complexe taken, maar kent onvoldoende ondersteuning voor cross-platform GUI-interactie en een goed opgebouwd zelfevolutiemechanisme. Deze tekortkomingen beperken de aanpassing aan diverse apparaatecosystemen en voorkomen prestatieverbeteringen door continu leren van uitvoeringservaring. Om deze problemen op te lossen, stellen we het Know Deeply, Act Perfectly-paradigma voor persoonlijke assistenten voor, dat stelt dat opgebouwde gebruikersinteractie- en taakuitvoeringservaring de uitvoeringsnauwkeurigheid en -efficiëntie direct verbetert, waarbij cognitief begrip en operationele uitvoering worden verenigd. Op basis van dit paradigma introduceren we KnowAct-GUIClaw, een nieuw Know-Route-Act-Reflect-framework dat is ontworpen om de GUI-manipulatietekorten van OpenClaw aan te pakken en de beperkingen op het gebied van cross-platform en recursieve zelfverbetering te doorbreken. Ten eerste maakt de host-agent gebruik van opgebouwde interactie-ervaring en taakrelevante kennis voor de ontleding en toewijzing van langetermijntaken (Know). Ten tweede, een inplugbare GUI-subagent met een ervaringstoewijsbaar geheugensysteem (Know) en een zelfevoluerende vaardigheidsbibliotheek (Act), waardoor naadloze cross-platformmigratie en snelle padintegratie mogelijk worden. Met name dit framework slaat continu gebruikersprofielen en feedback op om de nauwkeurigheid van taakontleding en toolaanroepen te verbeteren. Uitgebreide experimenten op Android, iOS, HarmonyOS en Windows tonen aan dat KnowAct-GUIClaw superieure efficiëntie, nauwkeurigheid en cross-platform aanpasbaarheid bereikt. Met name de GUIClaw met opensource Kimi-2.6-modellen behaalt de beste prestaties (64,1%) op de long-horizon MobileWorld-benchmark, waarbij alle agentische frameworks en closed-source agentische modellen, zoals Seed-2.0-Pro en GPT-5.5, worden verslagen. Bovendien zijn de kennisrijke geheugen- en uitvoeringsvaardigheden die door ons framework worden ondersteund, overdraagbaar naar diverse basismodellen, met een verbetering van 8,5% met Kimi-2.6.
English
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-running experience directly improve execution accuracy and efficiency, unifying cognitive comprehension and operational execution. Based on this paradigm, we introduce KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework designed to address OpenClaw's GUI manipulation deficits and break through its cross-platform and recursive self-improvement constraints. First, the host agent leverages accumulated interaction experience and task-relevant knowledge for long-horizon task decomposition and allocation (Know). Second, a pluggable GUI subagent with an experience-attributable memory system (Know) and self-evolving skill library (Act), enabling seamless cross-platform migration and fast-path integration. Especially, this framework continuously stores user profiles and feedback to improve the accuracy of task decomposition and tool calls. Extensive experiments across Android, iOS, HarmonyOS and Windows show that KnowAct-GUIClaw achieves superior efficiency, accuracy and cross-platform adaptability. Especially, the GUIClaw with open-source Kimi-2.6 models achieves the best performance (64.1%) on the long-horizon MobileWorld benchmark, beating all agentical frameworks and closed-source agentical models, e.g., Seed-2.0-Pro and GPT-5.5. Additionally, the knowledgeable memory and execution skills supported by our framework are transferable across diverse base models, improving by 8.5% with Kimi-2.6.