ChatPaper.aiChatPaper

ロボット操作のための視覚言語行動モデルにおける二重潜在記憶

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

July 8, 2026
著者: Hongyu Qu, Jianzhe Gao, Xiaobin Hu, Shaohuan Yang, Xinlei Yu, Rui Yan, Wenguan Wang, Xiangbo Shu, Shuicheng Yan
cs.AI

要旨

主流の視覚・言語・行動(VLA)モデルは、マルコフ性仮定のもとで主に現在の観測から行動を予測するため、長期的かつ時間依存的なタスクに苦戦している。既存のメモリ拡張型VLAは、観測ウィンドウを拡大するか、メモリバンクから履歴を補助的なポリシー側のコンテキストとして取得する。しかし、それらはメモリをVLA推論の本来の潜在埋め込み空間の外に置いたままであり、履歴経験がマルチモーダル推論や行動生成と流動的に統合されるのを妨げている。そこで我々はLaMem-VLAを提案する。これは、履歴経験を潜在メモリトークンに再構築し、VLA推論に直接織り込む潜在メモリネイティブフレームワークである。その中核として、LaMem-VLAは4つの連携コンポーネントを導入する:(i) 履歴経験を補完的な短期・長期メモリ保管庫に整理するキュレーター、(ii) マルチモーダル認知を用いて両方の保管庫をクエリし、コンテキストに関連する証拠を取得するシーカー、(iii) 取得した証拠をコンパクトな短期・長期潜在メモリトークンに再構築するコンデンサー、(iv) これらのメモリトークンを現在の観測と指示とともに1つの連続的な埋め込みシーケンスに注入するウィーバー。履歴経験を同一の連続潜在空間で表現・取得・消費することにより、LaMem-VLAはメモリが直接VLA推論に参加し、制限されたコンテキスト下で行動生成を導くことを可能にする。SimplerEnvおよびLIBEROにおける広範な実験により、我々のLaMem-VLAの優位性が実証された。
English
Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally dependent tasks. Existing memory-augmented VLAs either expand the observation window or retrieve history from the memory bank as auxiliary policy-side context. However, they leave memory outside the native latent embedding space of VLA reasoning, preventing historical experience from being fluidly interleaved with multimodal reasoning and action formation. To this end, we introduce LaMem-VLA, a latent-memory-native framework that reconstructs historical experience into latent memory tokens and directly interweaves them with VLA reasoning. At its core, LaMem-VLA introduces four coordinated components: (i) a curator that organizes historical experience into two complementary short-term and long-term memory vaults; (ii) a seeker that queries both vaults using the multimodal cognition to retrieve context-relevant evidence; (iii) a condenser that reconstructs the retrieved evidence into compact short-term and long-term latent memory tokens; and (iv) a weaver that injects these memory tokens with the current observation and instruction into one continuous embedding sequence. By representing, retrieving, and consuming historical experience entirely in the same continuous latent space, LaMem-VLA enables memory to directly participate in VLA reasoning and guide action generation under a bounded context. Extensive experiments on SimplerEnv and LIBERO demonstrate the superiority of our LaMem-VLA.