ChatPaper.aiChatPaper

QQWorld: 세계 모델 정규화를 위한 분위수-분위수 매칭

QQWorld: Quantile-Quantile Matching for World Model Regularization

July 30, 2026
저자: Zhoushun Yu, Xiaoyu Hu, Xiangyu Xu
cs.AI

초록

잠재 세계 모델은 컴팩트 표현 공간에서 미래 상태를 예측함으로써 효율적인 계획을 가능하게 하지만, 그 성능은 학습된 잠재 분포의 품질에 결정적으로 의존한다. LeWorldModel(LeWM)은 Epps-Pulley(EP) 목적 함수를 사용하여 잠재 변수를 등방성 가우시안 분포로 정규화한다. 우리는 EP의 보정 기울기가 고립된 꼬리 표본에 대해 급격히 소멸하여 두꺼운 꼬리 편차가 충분히 제어되지 않은 채 남게 됨을 보인다. 이러한 한계를 해결하기 위해, 우리는 EP를 분위수-분위수 정합 목적 함수로 대체하는 QQWorld를 제안한다. QQWorld는 투영된 잠재 표본을 순위 정합된 가우시안 분위수와 직접 정렬함으로써 꼬리 영역에서 효과적인 보정 기울기를 유지한다. 또한, 우리는 이전 배치에서 분리된 표본을 사용하여 유효 순위 풀을 확장하는 크로스 배치 QQ를 개발하고, 그 편향-분산 트레이드오프를 규명한다. 네 가지 제어 환경에서 QQWorld는 LeWM의 평균 계획 성공률을 효과적으로 개선하며, 동시에 더 나은 가우시안 정합과 더 얇은 잠재 꼬리를 일관되게 산출한다.
English
Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.