ChatPaper.aiChatPaper

SyncWorld:視覺校準使世界模型得以作為零樣本模擬器

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

September 8, 2026
作者: Yuncong Yang, Zhengtao Han, Furkan Ozyurt, Zeyuan Yang, Han Yang, Junyi Cao, Haoyu Zhen, Yilun Du, Chuang Gan
cs.AI

摘要

世界模型日益被用作策略在環的想像環境;在此類環境中,可靠的 rollout 需要對低階機器人動作具備細粒度可控性。在機器人領域擴展此類模型的一項關鍵障礙在於,動作並非像素空間中的通用語言:視覺環境、相機視角、機器人位置或具身形態的變化,會改變相同數值動作在視覺上的表現方式,導致混合訓練下的監督訊號衝突,以及部署時的脆弱泛化。我們提出 SyncWorld,一個動作條件世界模型,能在無需任何額外訓練的情況下,作為跨未見環境的零樣本模擬器。SyncWorld 利用視覺校準片段——由成對的影格與動作組成,展示所有可控自由度——以在上下文中指定特定設定的動作—視覺映射。以視覺校準上下文進行訓練,教導模型透過視覺證據解讀動作,並在明確校準不可用時利用互動歷史。實驗顯示,SyncWorld 能在先前未見的設定中準確模擬動作結果,且其模擬 rollout 的能力可在無需訓練的情況下實現測試時策略改進。
English
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual environment, camera view, robot placement, or embodiment alter how the same numerical action manifests visually, leading to conflicting supervision under mixed training and brittle generalization at deployment. We introduce SyncWorld, an action-conditioned world model that serves as a zero-shot simulator across unseen environments without any additional training. SyncWorld leverages a visual calibration episode---paired frames and actions that showcase all the controllable degrees of freedom---to specify the setup-specific Action--Visual Mapping in context. Training with visual calibration contexts teaches the model to interpret actions through visual evidence and to leverage interaction history when explicit calibration is unavailable. Experiments show that SyncWorld can accurately simulate action outcomes in previously unseen settings, and that its capability of simulating rollouts enables test-time policy improvement without training.