ChatPaper.aiChatPaper

一個未來,每個機器人:基於去中心化JEPA的標籤高效集體狀態預測

One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA

July 30, 2026
作者: Alan-Barsag Gazzaev, Alexey Garvilov, Sergey Muravyov
cs.AI

摘要

群體中的每台機器人能否僅憑局部觀測和頻寬受限的訊息,預測相同的未來集體狀態?我們將此問題形式化為分散式共享狀態預測,並提出集體狀態JEPA(CS-JEPA),這是一種遞迴聯合嵌入預測架構,其在每台機器人上的輸出代表一個共同的未來token場。在部署時,每台機器人使用16幀的局部歷史,以及每條有向邊一個64浮點數的遞迴訊息;沒有全域池化、目標編碼器、回合時鐘或記錄的未來動作。在未使用下游集體標籤進行預訓練後,以凍結表徵透過在6、12或24個全域標記回合上擬合的嶺迴歸探針進行評估。相對於具有相同接收端錨點和部署容量、但多出9,607個僅訓練用參數的原始未來重建方法,一項前瞻性註冊的五種子後續實驗,在最多108台機器人的分佈內、環狀、相互kNN及未見規模系列中,改善了預測誤差和機器人間一致性標籤預算AUC。所有效應在5/5外部種子中均有利於CS-JEPA。在一項獨立的密封八種子後續實驗中,匹配的動作條件預測器在產生接收端局部預測表徵之前,會接收每個候選的四步計畫。CS-JEPA將分支價值均方誤差降低了45.5%,並將上下文內候選分數的皮爾森相關提高了0.1291,兩項效應在8/8種子中均有利,包括在未見的N=32時。這些結果支持將共同未來JEPA目標作為在拓撲和規模變化下進行分散式群體預測的標籤高效基本單元,並提供了與規劃相關的價值估計的額外證據。
English
Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture whose output at every robot represents one common future token field. At deployment, each robot uses a 16-frame local history and one 64-float recurrent message per directed edge; there is no global pooling, target encoder, episode clock, or recorded future action. After pretraining without downstream collective labels, frozen representations are evaluated with ridge probes fitted on 6, 12, or 24 globally labeled episodes. Against raw-future reconstruction with the same receiver anchor and deployment capacity but 9,607 additional training-only parameters, a prospectively registered five-seed follow-up improves prediction-error and inter-robot-agreement label-budget AUC on in-distribution, ring, mutual-kNN, and unseen-size families up to 108 robots. Every effect favors CS-JEPA in 5/5 outer seeds. In a separate sealed eight-seed follow-up, matched action-conditioned predictors receive each candidate four-step plan before producing receiver-local predictive representations. CS-JEPA reduces branch-value MSE by 45.5% and improves within-context candidate-score Pearson correlation by 0.1291, with both effects favorable in 8/8 seeds, including at unseen N=32. These results support common-future JEPA targets as a label-efficient primitive for decentralized swarm prediction under topology and size shift, with additional evidence of planning-relevant value estimation.