ChatPaper.aiChatPaper

FactorJEPA: 혼잡하고 혼란스러운 글로벌 사우스 도시 세계를 위한 모놀리식 미래의 레이아웃-에이전트-상호작용 채널 분해

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

August 2, 2026
저자: Kapil Wanaskar, Gaytri Jena, Aman Chadha, Vinija Jain, Vasu Sharma, Amitava Das
cs.AI

초록

세계 모델은 물리적 세계의 구조와 역학을 포착하고 예측하는 능력으로 큰 주목을 받아왔다. 이러한 새로운 연구 지형에서 공동 임베딩 예측 구조(JEPA)는 특히 매력적인 방향을 제시한다. 우리는 DENSEWORLD라고 부르는, 인구가 많고 혼잡하며 혼란스러운 글로벌 사우스(Global South) 도시 환경이라는 거의 탐구되지 않은 영역을 연구한다. 기존 평가를 주도하는 저밀도·차선 구조 환경과 달리, 이러한 장면들은 불분명한 공간 경계, 극단적인 에이전트 이질성, 지속적인 가림, 그리고 혼합 교통에서의 빠른 사회적 협상을 보인다. 우리는 이 영역에 대한 최초의 대규모 데이터셋인 22개 도시에 걸친 1,000시간 분량의 주행·보행·항공 영상을 소개한다. 기존 JEPA 방식들은 이질성과 부분 관측 가능성 하에서 밀집된 상호작용 역학을 보존하는 데 어려움을 겪는다. 우리는 세계 구조를 일급 예측 기본 요소로 만드는 FactorJEPA를 제안한다. 이는 미래를 단일(모놀리식) 잠재 표현으로 인코딩하는 대신, 가시성 게이트와 분리된 부분공간을 사용하여 부분적으로 관측된 에이전트를 보존하고 교차 요인 지름길을 억제함으로써 레이아웃, 개체, 그리고 상호작용을 구성한다. FactorJEPA는 (i) 미래 잠재 표현 정확도(Future-frame L1), (ii) 개입 민감 예측(Causal L1), (iii) 감소된 시각적 증거에 대한 강건성(Mask-ratio slope)을 개선하며, 동시에 (iv) 재현 가능한 모션-정보 트레이드오프(Motion cosine)를 드러낸다. 방법 순위는 2B 및 1B V-JEPA 2.1 백본에서도 재현되며, rho = 0.895~0.978이다. 우리는 DENSEWORLD-115k 데이터셋(https://huggingface.co/datasets/anonymousML123/denseworld-115k)과 surgery 방식으로 훈련된 FactorJEPA 체크포인트(https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder)를 공개한다.
English
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly compelling direction. We study a largely unexplored regime: populous, crowded, and chaotic Global South urban environments, which we call DENSEWORLD. Unlike the lower-density, lane-structured settings that dominate existing evaluations, these scenes exhibit soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. We introduce the first large-scale dataset for this regime: 1,000 hours of drive-through, walk-through, and aerial video across 22 cities. Existing JEPA formulations struggle to preserve dense interaction dynamics under heterogeneity and partial observability. We introduce FactorJEPA, which makes world structure a first-class predictive primitive. Rather than encoding the future in a monolithic latent, it composes layout, entities, and interactions, using a visibility gate and separated subspaces to preserve partially observed agents and discourage cross-factor shortcuts. FactorJEPA improves (i) future-latent accuracy (Future-frame L1), (ii) intervention-sensitive prediction (Causal L1), and (iii) robustness to reduced visual evidence (Mask-ratio slope), while exposing (iv) a reproducible motion-information trade-off (Motion cosine). Method rankings replicate across 2B and 1B V-JEPA 2.1 backbones, with rho = 0.895 to 0.978. We publicly release the DENSEWORLD-115k dataset (https://huggingface.co/datasets/anonymousML123/denseworld-115k) and the surgery-trained FactorJEPA checkpoints (https://huggingface.co/datasets/anonymousML123/factorjepa-outputs/tree/main/outputs/full/vjepa_2_1_vitg_1B/train/m09c_surgery_3stage_DI_diheavy_encoder).