ChatPaper.aiChatPaper

FactorJEPA: 混雑かつ混沌としたグローバル・サウスの都市世界における、モノリシックな未来予測のレイアウト・エージェント・インタラクションの各チャネルへの因子分解

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

August 2, 2026
著者: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
cs.AI

要旨

世界モデルは、物理世界の構造とダイナミクスを捕捉・予測する能力により、大きな注目を集めている。この新たな領域において、Joint Embedding Predictive Architectures(JEPA)は特に有力な方向性を提供する。我々は、ほとんど未探索の領域を研究する。すなわち、人口が多く、混雑し、混沌としたグローバル・サウスの都市環境であり、これをDENSEWORLDと呼ぶ。既存の評価を支配する低密度でレーン構造化された設定とは異なり、これらのシーンは、曖昧な空間的境界、極端なエージェントの不均一性、持続的な遮蔽、および混合交通下での迅速な社会的交渉を示す。我々は、この領域のための最初の大規模データセット、すなわち22都市にわたる1,000時間の走行映像、歩行映像、および航空映像を導入する。既存のJEPA定式化は、不均一性と部分的可観測性の下で、密集した相互作用ダイナミクスを保持することが困難である。我々は、世界構造を第一級の予測プリミティブとするFactorJEPAを導入する。将来をモノリシックな潜在表現にエンコードするのではなく、可視性ゲートと分離された部分空間を用いて、部分的に観測されたエージェントを保持し、因子間のショートカットを防ぎながら、レイアウト、エンティティ、および相互作用を構成する。FactorJEPAは、(i) 将来潜在精度(Future-frame L1)、(ii) 介入に敏感な予測(Causal L1)、(iii) 視覚的証拠の減少に対する頑健性(Mask-ratio slope)を改善し、その一方で、(iv) 再現可能な運動情報のトレードオフ(Motion cosine)を明らかにする。手法のランキングは、2Bおよび1BのV-JEPA 2.1バックボーン間で再現され、ρ = 0.895〜0.978である。我々は、DENSEWORLD-115kデータセット(https://huggingface.co/datasets/anonymousML123/denseworld-115k)と、サージェリー訓練済みのFactorJEPAチェックポイント(https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder)を公開する。
English
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).