QQWorld:以分位數-分位數匹配進行世界模型正則化
QQWorld: Quantile-Quantile Matching for World Model Regularization
July 30, 2026
作者: Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu
cs.AI
摘要
潛在世界模型透過在緊湊的表示空間中預測未來狀態來實現高效規劃,但其效能關鍵地取決於所學習到的潛在分布的品質。LeWorldModel(LeWM)使用 Epps-Pulley(EP)目標函數,以各向同性高斯分布為目標對其潛在變數進行正則化。我們證明,EP 的修正梯度對於孤立的尾部樣本會迅速消失,導致重尾偏差無法獲得充分控制。為了解決此限制,我們提出 QQWorld,以分位數-分位數(QQ)匹配目標函數取代 EP,直接將投影後的潛在樣本與秩匹配的高斯分位數對齊,從而在尾部維持有效的修正梯度。我們進一步開發了跨批次 QQ,利用先前批次的解耦樣本擴大有效排序池,並刻畫其偏差-方差權衡。在四個控制環境中,QQWorld 有效提升了 LeWM 的平均規劃成功率,同時持續表現出更佳的高斯對齊與更薄的潛在尾部。
English
Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.