FactorJEPA:将整体式未来预测分解为布局-智能体-交互通道,以应对全球南方拥挤混乱的城市环境
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
August 2, 2026
作者: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
cs.AI
摘要
世界模型因其捕捉和预测物理世界结构与动态的能力而受到广泛关注。在这一新兴研究领域中,联合嵌入预测架构(JEPA)提供了一条尤为引人注目的方向。
我们研究了一个很大程度上未被探索的场景:全球南方人口稠密、拥挤且混乱的城市环境,我们称之为DENSEWORLD。与主导现有评估的低密度、车道结构化场景不同,这些场景呈现出软性空间边界、极端的智能体异质性、持续遮挡以及混行交通下快速的社会性协商。我们为该场景引入了首个大规模数据集:涵盖22个城市的1,000小时行车、步行和航拍视频。现有JEPA方案难以在异质性和部分可观测性条件下保留密集交互动态。
我们提出FactorJEPA,将世界结构作为一等预测原语。它并非将未来编码为单一整体的潜在表示,而是通过可见性门控和分离子空间组合布局、实体与交互,以保留仅部分可观测的智能体并抑制跨因子捷径。FactorJEPA改善了(i)未来潜在表示的准确性(未来帧L1)、(ii)干预敏感预测(因果L1)和(iii)对视觉证据减少的鲁棒性(掩码率斜率),同时展现出(iv)可复现的运动-信息权衡(运动余弦)。方法排名在2B和1B V-JEPA 2.1骨干网络上一致复现,ρ = 0.895至0.978。
我们公开发布了DENSEWORLD-115k数据集(https://huggingface.co/datasets/anonymousML123/denseworld-115k)和经surgery训练的FactorJEPA检查点(https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder)。
English
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction.
We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability.
We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978.
We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).