ChatPaper.aiChatPaper

潜空间世界模型中的决策度量对齐:用于模型预测控制规划的诊断方法与动作条件化目标

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

August 19, 2026
作者: Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
cs.AI

摘要

JEPA式潜在世界模型可以使用到目标潜在状态的欧氏距离作为模型预测控制(MPC)的成本函数。然而,对任务变量的强解码并不能保证该特定成本能够依据真实任务进展对候选动作序列进行排序。我们将后一种性质称为决策度量对齐。我们引入了计划-真实斯皮尔曼系数(Plan-Real Spearman),用于衡量随机计划上潜在排序与真实排序的一致性,以及CEM阶段斯皮尔曼系数(CEM-stage Spearman),用于衡量当交叉熵方法(CEM)搜索集中其提议分布时二者的排序一致性。我们分析了潜在距离保持真实成本排序的充分条件,识别出编码器失真、终端展开误差和候选裕度是控制量。在观测到的经验对齐差距的指导下,DA-LeWM在LeWM的基础上增加了逆动力学和演示条件化目标-动作头。在所有实验中,DA-LeWM相较于LeWM加速了收敛并取得了更高的在线成功率,同时探针得分保持相近。这些结果表明,动作条件化目标改善了基于欧氏成本、CEM的潜在MPC所使用的几何结构。
English
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property decision-metric alignment. We introduce Plan-Real Spearman, which measures latent--real rank agreement on random plans, and CEM-stage Spearman, which measures the same agreement as cross-entropy-method (CEM) search concentrates its proposal. We analyze sufficient conditions under which latent distance preserves real-cost rankings, identifying encoder distortion, terminal rollout error, and candidate margins as the controlling quantities. Guided by the observed empirical alignment gap, DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads. Across all our experiments, DA-LeWM accelerates convergence and achieves higher online success than LeWM, while probe scores remain similar. These results show that action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC.