Enfold:将世界模型想象折叠为预测表征以实现超高效具身控制
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
August 6, 2026
作者: Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan
cs.AI
摘要
世界生成模型通常通过其产出来使用:一段渲染的未来、一个以视频为条件的动作,或由昂贵生成分支计算出的潜在上下文。我们认为,它们更具可复用性的资产是构造未来的计算过程。当生成器将受损的未来转化为连贯轨迹时,其中间状态在不同抽象层级上组织外观、空间布局与交互。这种未来生成计算能否被内化为一种仅从当下推断出的表征?我们提出 Enfold,它将该计算转移为从当前视觉上下文和语言指令预测得到的表征。训练期间,生成器处理观测到的未来时所暴露的多层级状态,监督一个仅基于当前的编码器。学习到的表征被反馈以条件化未来生成,并被任务头读取,同时不允许任务梯度重塑编码器。部署时,动作预测不再执行生成器。在 LIBERO、RoboTwin2.0 和真实机器人任务中,Enfold 在实现强控制的同时,将动作延迟相对于 Fast--WAM 降低了 3.7 倍;Enfold-Flash 则达到 10.1 倍。表征分析表明,它抑制了干扰性变化,并优先捕捉在更长时域中涌现的变化。当当前场景因人为干预而改变时,生成的后续内容与执行的动作都会相应调整,这与固定轨迹回放不一致。这些结果将世界生成器重新定义为预测控制表征的来源:如果其内部结构能够被融入当下,则无需在每一步都将未来具体化。
English
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by 3.7times relative to Fast--WAM, Enfold-Flash reaches 10.1times. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.