ChatPaper.aiChatPaper

機器人操作中視覺-語言-行動模型的雙重潛在記憶

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

July 8, 2026
作者: Hongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan
cs.AI

摘要

主流视觉-语言-动作(VLA)模型基于马尔可夫假设,主要从当前观测预测动作,因此在处理长时域、时间依赖型任务时表现不佳。现有的记忆增强型VLA方法要么扩展观测窗口,要么从记忆库中检索历史信息作为辅助策略上下文。然而,这些方法将记忆置于VLA推理的原生潜嵌入空间之外,阻碍了历史经验与多模态推理及动作生成过程的无缝交织。为此,我们提出LaMem-VLA——一种原生潜记忆框架,通过将历史经验重构为潜记忆标记,并将其直接融入VLA推理流程。其核心包含四个协同组件:(i)策展器,将历史经验组织为互补的短期与长期记忆库;(ii)搜索器,基于多模态认知对两个记忆库进行查询以检索上下文相关证据;(iii)压缩器,将检索到的证据重构为紧凑的短期与长期潜记忆标记;(iv)编织器,将这些记忆标记与当前观测和指令注入到一个连续的嵌入序列中。通过在连续潜空间中完成历史经验的表征、检索与消耗,LaMem-VLA使记忆能够直接参与VLA推理,并在有限上下文约束下引导动作生成。在SimplerEnv和LIBERO上的大量实验证明了LaMem-VLA的优越性。
English
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.