潛在世界模型中的決策指標對齊:MPC規劃的診斷方法與行動條件化目標
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
August 19, 2026
作者: Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
cs.AI
摘要
JEPA 風格的潛在世界模型可以將到目標潛在表徵的歐氏距離用作模型預測控制(MPC)的成本。然而,對任務變數的強大解碼並不能保證此特定成本能依據真實任務進展對候選動作序列進行排序。我們將後者性質稱為決策-度量對齊。我們引入了 Plan-Real Spearman,用於衡量隨機計畫上潛在與真實排序之間的一致性;以及 CEM-stage Spearman,用於衡量當交叉熵法(CEM)搜尋集中其提議時,同一種一致性的表現。我們分析了潛在距離在何種充分條件下能保持真實成本排序,並指出編碼器失真、終端展開誤差與候選邊際為主要控制因素。在觀察到的經驗對齊差距引導下,DA-LeWM 以逆向動力學與示範條件化目標-動作頭增強了 LeWM。在我們所有的實驗中,DA-LeWM 加速了收斂,並在線上成功率上高於 LeWM,同時探測分數保持相似。這些結果表明,動作條件化目標改善了基於歐氏成本與 CEM 的潛在 MPC 所使用的幾何結構。
English
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property decision-metric alignment. We introduce Plan-Real Spearman, which measures latent--real rank agreement on random plans, and CEM-stage Spearman, which measures the same agreement as cross-entropy-method (CEM) search concentrates its proposal. We analyze sufficient conditions under which latent distance preserves real-cost rankings, identifying encoder distortion, terminal rollout error, and candidate margins as the controlling quantities. Guided by the observed empirical alignment gap, DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads. Across all our experiments, DA-LeWM accelerates convergence and achieves higher online success than LeWM, while probe scores remain similar. These results show that action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC.