FactorJEPA:將整體未來表徵分解為佈局-智能體-交互通道,應用於擁擠且混亂的全球南方城市世界
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
August 2, 2026
作者: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
cs.AI
摘要
世界模型因其捕捉與預測物理世界結構與動態的能力而備受關注。在此新興領域中,聯合嵌入預測架構(JEPA)提供了一個特別引人注目的方向。
我們研究了一個 largely unexplored 的領域:人口稠密、擁擠且混亂的全球南方城市環境,我們稱之為 DENSEWORLD。與現有評估中主導的低密度、車道結構化場景不同,這些場景展現出軟性空間邊界、極端的代理異質性、持續性遮擋,以及混合交通下的快速社會協商。我們引入了該領域第一個大規模資料集:涵蓋 22 個城市的 1,000 小時行車、步行及空拍影片。現有的 JEPA 公式在異質性與部分可觀測性條件下,難以保留密集的互動動態。
我們提出 FactorJEPA,將世界結構作為第一級預測原語。與其將未來編碼為單一潛在表示,它透過可視性閘門與分離子空間來組合佈局、實體與互動,以保留部分可觀測的代理並阻止跨因子捷徑。FactorJEPA 改善了 (i) 未來潛在表示準確度(Future-frame L1)、(ii) 干預敏感預測(Causal L1),以及 (iii) 對減少視覺證據的穩健性(Mask-ratio slope),同時展現 (iv) 一個可重現的運動資訊權衡(Motion cosine)。方法排名在 2B 與 1B V-JEPA 2.1 骨幹上重現,rho 值介於 0.895 至 0.978 之間。
我們公開釋出 DENSEWORLD-115k 資料集(https://huggingface.co/datasets/anonymousML123/denseworld-115k)以及經手術訓練的 FactorJEPA 檢查點(https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder)。
English
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction.
We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability.
We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978.
We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).