一个未来,所有机器人:基于去中心化JEPA的标签高效集体状态预测
One Future, Every Robot: Label-Efficient Collective-State Prediction with Decentralized JEPA
July 30, 2026
作者: Alan-Barsag Gazzaev, Alexey Garvilov, Sergey Muravyov
cs.AI
摘要
群体中的每个机器人能否仅凭局部观测和带宽受限的消息预测相同的未来集体状态?我们将此形式化为去中心化共享状态预测,并引入集体状态JEPA(CS-JEPA)——一种循环联合嵌入预测架构,其在每个机器人处的输出表示一个共同的未来token场。在部署时,每个机器人使用16帧局部历史以及每条有向边一条由64个浮点数组成的循环消息;没有全局池化、目标编码器、回合时钟或已记录的未来动作。在无下游集体标签的预训练之后,使用在6、12或24个全局标记回合上拟合的岭回归探针评估冻结表示。与具有相同接收器锚点和部署容量但额外拥有9,607个仅训练时参数的原始未来重建相比,一项前瞻性注册的五种子后续实验在分布内、环形、相互kNN及未见规模族(最多108个机器人)上,改善了预测误差和机器人间一致性的标签预算AUC。所有效应在5/5个外层种子中均有利于CS-JEPA。在一项独立的密封八种子后续实验中,匹配的动作条件预测器在生成接收器局部的预测表示之前接收每个候选四步计划。CS-JEPA将分支值均方误差降低了45.5%,并将上下文内候选得分的皮尔逊相关性提高了0.1291;两项效应在8/8个种子中均有利,包括在未见过的N=32时亦然。这些结果支持将共同未来JEPA目标作为拓扑和规模变化下去中心化群体预测的标签高效基元,并提供了与规划相关的价值估计的额外证据。
English
Can every robot in a swarm predict the same future collective state from only local observations and bandwidth-limited messages? We formulate this as decentralized shared-state prediction and introduce Collective-State JEPA (CS-JEPA), a recurrent joint-embedding predictive architecture whose output at every robot represents one common future token field. At deployment, each robot uses a 16-frame local history and one 64-float recurrent message per directed edge; there is no global pooling, target encoder, episode clock, or recorded future action. After pretraining without downstream collective labels, frozen representations are evaluated with ridge probes fitted on 6, 12, or 24 globally labeled episodes. Against raw-future reconstruction with the same receiver anchor and deployment capacity but 9,607 additional training-only parameters, a prospectively registered five-seed follow-up improves prediction-error and inter-robot-agreement label-budget AUC on in-distribution, ring, mutual-kNN, and unseen-size families up to 108 robots. Every effect favors CS-JEPA in 5/5 outer seeds. In a separate sealed eight-seed follow-up, matched action-conditioned predictors receive each candidate four-step plan before producing receiver-local predictive representations. CS-JEPA reduces branch-value MSE by 45.5% and improves within-context candidate-score Pearson correlation by 0.1291, with both effects favorable in 8/8 seeds, including at unseen N=32. These results support common-future JEPA targets as a label-efficient primitive for decentralized swarm prediction under topology and size shift, with additional evidence of planning-relevant value estimation.