ChatPaper.aiChatPaper

Enfold:將世界模型想像摺疊進預測性表徵,以實現超高效率具身控制

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

August 6, 2026
作者: Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan
cs.AI

摘要

世界生成模型通常通過它們所產生的內容來使用:渲染的未來、視頻條件化的動作,或由昂貴的生成分支計算出的潛在上下文。我們認為,它們更具可重複使用價值的資產是構建未來的計算過程。當生成器將受損的未來轉化為連貫的軌跡時,其中間狀態在不同抽象層次上組織外觀、空間佈局和交互。這種未來生成計算能否被內化為僅從當前狀態推斷出的表徵?我們提出 Enfold,它將此計算轉移到一種從當前視覺上下文和語言指令預測的表徵中。在訓練期間,生成器處理觀察到的未來時暴露出的多層級狀態,監督一個僅使用當前輸入的編碼器。學習到的表徵被反饋以條件化未來生成,並由任務頭讀取,而不允許任務梯度重塑編碼器。在部署時,動作預測不再執行生成器。在 LIBERO、RoboTwin2.0 和真實機器人任務上,Enfold 支持強控制,同時相對於 Fast--WAM 將動作延遲降低了 3.7 倍,而 Enfold-Flash 達到了 10.1 倍。表徵分析表明,它抑制了干擾性變異,並優先捕捉在更長時域中出現的變化。當當前場景被人為干預改變時,生成的續接和執行的動作都會適應,這與固定軌跡重放不一致。這些結果將世界生成器重新定義為預測控制表徵的來源:如果其內部結構可以被摺疊進當前狀態,那麼其未來無需在每一步中被實體化。
English
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by 3.7times relative to Fast--WAM, Enfold-Flash reaches 10.1times. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.