하나의 미래, 모든 로봇: 분산형 JEPA를 이용한 레이블 효율적 집단 상태 예측
One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
July 30, 2026
저자: Alan-Barsag Gazzaev, Alexey Garvilov, Sergey Muravyov
cs.AI
초록
군집의 모든 로봇이 국소 관측과 대역폭 제한 메시지만으로 동일한 미래 집단 상태를 예측할 수 있을까? 우리는 이를 분산 공유 상태 예측으로 정식화하고, 각 로봇의 출력이 하나의 공통 미래 토큰 필드를 나타내는 순환 결합-임베딩 예측 구조인 CS-JEPA(Collective-State JEPA)를 소개한다. 배포 시 각 로봇은 16프레임 국소 이력과 방향 간선당 하나의 64-부동소수점 순환 메시지를 사용하며, 전역 풀링, 대상 인코더, 에피소드 시계 또는 기록된 미래 행동은 사용하지 않는다. 하위 과제 집단 레이블 없이 사전학습한 후, 고정된 표현을 전역 레이블이 있는 6, 12, 24개 에피소드에 적합시킨 릿지 프로브로 평가한다. 동일한 수신기 앵커와 배포 용량을 가지지만 9,607개의 학습 전용 추가 파라미터를 사용하는 원시 미래 재구성과 비교하여, 사전 등록된 5-시드 후속 연구는 최대 108대 로봇의 분포 내, 링, 상호 kNN, 미지의 크기 계열에서 예측 오차 및 로봇 간 일치도 레이블-예산 AUC를 개선한다. 모든 효과는 5/5 외부 시드에서 CS-JEPA에 유리하다. 별도의 봉인된 8-시드 후속 연구에서는 일치된 행동 조건 예측기가 수신기-국소 예측 표현을 생성하기 전에 각 후보 4단계 계획을 입력받는다. CS-JEPA는 분기-가치 MSE를 45.5% 줄이고 문맥 내 후보 점수 Pearson 상관관계를 0.1291만큼 개선하며, 두 효과 모두 미지의 N=32를 포함한 8/8 시드에서 유리하다. 이 결과는 공통 미래 JEPA 목표가 토폴로지 및 크기 변화 하에서 분산 군집 예측을 위한 레이블 효율적인 기본 요소임을 지지하며, 계획 관련 가치 추정에 대한 추가 증거를 제공한다.
English
Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture whose output at every robot represents one common future token field. At deployment, each robot uses a 16-frame local history and one 64-float recurrent message per directed edge; there is no global pooling, target encoder, episode clock, or recorded future action. After pretraining without downstream collective labels, frozen representations are evaluated with ridge probes fitted on 6, 12, or 24 globally labeled episodes. Against raw-future reconstruction with the same receiver anchor and deployment capacity but 9,607 additional training-only parameters, a prospectively registered five-seed follow-up improves prediction-error and inter-robot-agreement label-budget AUC on in-distribution, ring, mutual-kNN, and unseen-size families up to 108 robots. Every effect favors CS-JEPA in 5/5 outer seeds. In a separate sealed eight-seed follow-up, matched action-conditioned predictors receive each candidate four-step plan before producing receiver-local predictive representations. CS-JEPA reduces branch-value MSE by 45.5% and improves within-context candidate-score Pearson correlation by 0.1291, with both effects favorable in 8/8 seeds, including at unseen N=32. These results support common-future JEPA targets as a label-efficient primitive for decentralized swarm prediction under topology and size shift, with additional evidence of planning-relevant value estimation.