Enfold:超効率的身体化制御のための予測的表現への世界モデル想像の折り込み
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
August 6, 2026
著者: Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan
cs.AI
要旨
世界生成モデルは通常、それらが生成するものを通じて使用される。すなわち、レンダリングされた未来、ビデオ条件付きの行動、あるいはコストのかかる生成ブランチによって計算される潜在コンテキストである。我々は、それらのより再利用可能な資産は未来を構築する計算であると主張する。生成器が破損した未来を一貫した軌道へと変換する際、その中間状態は外観、空間的レイアウト、相互作用を抽象度のレベルにわたって整理する。この未来生成計算は、現在のみから推論される表現に内部化できるのだろうか。我々は、この計算を現在の視覚的文脈と言語指示から予測される表現へ転送するEnfoldを提案する。訓練中は、生成器が観測された未来を処理する際に露出される多レベルの状態が、現在のみを入力とするエンコーダを監視する。学習された表現は未来生成を条件付けるためにフィードバックされ、タスクの勾配がエンコーダを再形成するのを許すことなくタスクヘッドによって読み取られる。デプロイ時には、行動予測はもはや生成器を実行しない。LIBERO、RoboTwin2.0、および実ロボットタスクにおいて、Enfoldは強力な制御を実現しつつ、行動レイテンシをFast-WAM比で3.7倍削減し、Enfold-Flashは10.1倍の削減を達成する。表現分析は、それが無関係な変動を抑制し、より長い時間的範囲にわたって現れる変化を優先的に捕捉することを示している。現在のシーンが人間の介入によって変更されると、生成された継続と実行された行動の両方が適応する。これは固定された軌道のリプレイとは矛盾する。これらの結果は、世界生成器を予測的制御表現の源泉として再解釈するものである。すなわち、その未来は、その内部構造が現在に折り畳まれ得るのであれば、すべてのステップで具現化される必要はない。
English
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by 3.7times relative to Fast--WAM, Enfold-Flash reaches 10.1times. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.