잠재 세계 모델에서의 의사결정 지표 정렬: MPC 계획을 위한 진단 및 행동 조건화 목적 함수
Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
August 19, 2026
저자: Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
cs.AI
초록
JEPA 스타일 잠재 세계 모델은 모델 예측 제어(MPC)의 비용으로 목표 잠재 표현까지의 유클리드 거리를 사용할 수 있다. 그러나 과제 변수의 강력한 디코딩은 이 특정 비용이 후보 행동 시퀀스를 실제 과제 진행도에 따라 순위화한다는 것을 보장하지 않는다. 우리는 후자의 속성을 결정-지표 정렬(decision-metric alignment)이라고 부른다. 우리는 무작위 플랜에 대한 잠재-실제 순위 일치도를 측정하는 Plan-Real 스피어만과, 교차 엔트로피 방법(CEM) 탐색이 제안 분포를 집중시킬 때 동일한 일치도를 측정하는 CEM-stage 스피어만을 도입한다. 우리는 잠재 거리가 실제 비용 순위를 보존하는 충분 조건을 분석하고, 제어 변수로 인코더 왜곡, 종단 롤아웃 오차, 후보 마진을 식별한다. 관찰된 경험적 정렬 격차에 기반하여, DA-LeWM은 역동역학 및 시연-조건화 목표-행동 헤드를 통해 LeWM을 확장한다. 모든 실험에서 DA-LeWM은 수렴을 가속화하고 LeWM보다 더 높은 온라인 성공률을 달성하며, 프로브 점수는 유사하게 유지된다. 이 결과들은 행동-조건화 목적 함수가 유클리드 비용 기반 CEM 잠재 MPC가 사용하는 기하 구조를 개선함을 보여준다.
English
JEPA-style latent world models can use Euclidean distance to a goal latent as the cost for model-predictive control (MPC). Strong decoding of task variables, however, does not guarantee that this particular cost ranks candidate action sequences by real task progress. We call the latter property decision-metric alignment. We introduce Plan-Real Spearman, which measures latent--real rank agreement on random plans, and CEM-stage Spearman, which measures the same agreement as cross-entropy-method (CEM) search concentrates its proposal. We analyze sufficient conditions under which latent distance preserves real-cost rankings, identifying encoder distortion, terminal rollout error, and candidate margins as the controlling quantities. Guided by the observed empirical alignment gap, DA-LeWM augments LeWM with inverse-dynamics and demonstration-conditioned goal-action heads. Across all our experiments, DA-LeWM accelerates convergence and achieves higher online success than LeWM, while probe scores remain similar. These results show that action-conditioned objectives improve the geometry used by Euclidean-cost, CEM-based latent MPC.