KnowAct-GUIClaw:深度認知,完美行動——具備自我進化記憶與技能的個人化圖形介面助手

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

July 15, 2026
作者: Yunxin Li, Jinchao Li, Shibo Su, Zhenran Xu, Chenrui Zhao, Tongshu Bian, Xiaoman Liang, Meishan Zhang, Baotian Hu, Min Zhang
cs.AI

摘要

OpenClaw 已成為複雜任務自動化的領先代理框架,但其跨平台 GUI 互動支援不足,且缺乏完善的自我演化機制。這些缺陷限制了它對多樣化裝置生態系統的適應能力,也使其無法透過持續學習執行經驗來提升效能。為解決這些問題,我們提出「深知而篤行」的個人助理範式,主張累積使用者互動與任務執行經驗,能直接提升執行準確性與效率,從而融合認知理解與操作執行。基於此範式,我們推出 KnowAct-GUIClaw——一套全新的「知-路-行-思」框架,旨在彌補 OpenClaw 在 GUI 操作上的不足,並突破其跨平台與遞迴自我改進的限制。首先,主代理利用累積的互動經驗與任務相關知識,進行長程任務分解與分配(知)。其次,一個可插拔的 GUI 子代理配備具經驗歸因能力的記憶系統(知)與自我演化技能庫(行),實現無縫跨平台遷移與快速路徑整合。特別是,此框架持續儲存使用者設定檔與回饋,以提升任務分解與工具呼叫的準確性。在 Android、iOS、HarmonyOS 與 Windows 上的大量實驗顯示,KnowAct-GUIClaw 在效率、準確性與跨平台適應性方面表現卓越。尤為突出的是,基於開源 Kimi-2.6 模型的 GUIClaw 在長程 MobileWorld 基準測試中達到最佳效能(64.1%),超越所有代理框架與封閉源代理模型(如 Seed-2.0-Pro 與 GPT-5.5)。此外,我們框架所支援的知識性記憶與執行技能可跨多種基礎模型遷移,在 Kimi-2.6 上提升 8.5%。
English
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws limit its adaptation to diverse device ecosystems and prevent performance improvements through continuous learning from execution experience. To resolve these issues, we propose the Know Deeply, Act Perfectly paradigm for personal assistants, which holds that accumulated user interaction and task-running experience directly improve execution accuracy and efficiency, unifying cognitive comprehension and operational execution. Based on this paradigm, we introduce KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework designed to address OpenClaw's GUI manipulation deficits and break through its cross-platform and recursive self-improvement constraints. First, the host agent leverages accumulated interaction experience and task-relevant knowledge for long-horizon task decomposition and allocation (Know). Second, a pluggable GUI subagent with an experience-attributable memory system (Know) and self-evolving skill library (Act), enabling seamless cross-platform migration and fast-path integration. Especially, this framework continuously stores user profiles and feedback to improve the accuracy of task decomposition and tool calls. Extensive experiments across Android, iOS, HarmonyOS and Windows show that KnowAct-GUIClaw achieves superior efficiency, accuracy and cross-platform adaptability. Especially, the GUIClaw with open-source Kimi-2.6 models achieves the best performance (64.1%) on the long-horizon MobileWorld benchmark, beating all agentical frameworks and closed-source agentical models, e.g., Seed-2.0-Pro and GPT-5.5. Additionally, the knowledgeable memory and execution skills supported by our framework are transferable across diverse base models, improving by 8.5% with Kimi-2.6.
PDF441July 17, 2026