QQWorld: ワールドモデル正則化のための分位点-分位点マッチング
QQWorld: Quantile-Quantile Matching for World Model Regularization
July 30, 2026
著者: Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu
cs.AI
要旨
潜在世界モデルは、コンパクトな表現空間における将来状態の予測を通じて効率的な計画を可能にするが、その性能は学習された潜在分布の品質に決定的に依存する。LeWorldModel (LeWM) は、Epps-Pulley (EP) 目的関数を用いてその潜在表現を等方ガウス分布に正則化する。我々は、EPの修正勾配が孤立した裾のサンプルに対して急速に消失し、重尾のずれが十分に制御されないまま残ることを示す。この制限に対処するため、我々はQQWorldを提案する。これはEPを分位数-分位数マッチング目的関数に置き換え、射影された潜在サンプルを順位整合されたガウス分位数に直接整列させることで、裾における効果的な修正勾配を維持するものである。さらに、前のバッチから切り離されたサンプルを用いて実効的な順位プールを拡大するクロスバッチQQを開発し、そのバイアス-バリアンスのトレードオフを特徴付ける。4つの制御環境において、QQWorldはLeWMの平均計画成功率を効果的に向上させ、一貫してより良いガウス整合性とより薄い潜在裾を達成する。
English
Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.