INTACT:無搜尋世界模型的同構意圖到行動學習
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models
July 28, 2026
作者: Junhan Sun, Hao Zhao, Guofeng Zhang
cs.AI
摘要
前瞻性潛在世界模型預測動作如何改變場景,但僅能透過昂貴的測試時搜尋來為期望的變化恢復動作。我們提出INTACT(意向到動作,INtent-To-ACTion),一種端到端聯合嵌入預測架構(JEPA),將帶動作標籤、無獎勵的軌跡轉化為可部署的意向到動作介面。每個轉移提供物理意向z_{t+1}-z_t,而未來目標則提供部署意向sg(z_g)-z_t。該架構在局部與目標動作意向骨幹輸入圖之間透過相同的四槽語法與共享參數實現同構,並在受支持的局部與目標動作意向族之間透過由同一預測器所誘導的動作規律語義實現同構,而非逐點潛在相等性。INTACT亦提供從RGB證據到動作有效潛在意向座標,以及從意向族到其對應動作規律族的完整遷移。非對稱端點梯度將物理後繼者落地,並將未來目標固定為錨點,在無需逐點潛在匹配或全局線性動力學的情況下結合表徵學習與控制。所產生的座標支持穩健的分布性動作規律:其條件均值可直接作為免搜尋策略,而取樣仍可用於多樣性或可選驗證。在四個官方LeWM任務上,單輪訓練、零搜尋模型分別達到85.78%、100.00%、97.67%與97.89%的成功率。以直接計畫為中心的可選局部交叉熵法(CEM)以384條而非9,000條候選序列達到96.86%的宏觀成功率,將取樣量減少23.44倍,同時比純CEM提升16.00個百分點。一個共享的四任務編碼器達到89.39%的E5直接宏觀成功率,並在每項任務上優於聯合訓練的LeWM,而預測-專家動作族kNN以r=0.954追蹤直接成功率。直接推論僅需2.9–5.5毫秒。
English
Forward latent world models predict how actions change a scene, but recover actions for a desired change only through expensive test-time search. We introduce INTACT (INtent-To-ACTion), an end-to-end JEPA that turns action-labeled, reward-free trajectories into a deployable intent-to-action interface. Each transition supplies physical intent z_{t+1}-z_t, while a future goal supplies deployment intent sg(z_g)-z_t. The architecture is isomorphic between the local and goal motion-intent backbone-input graphs through an identical four-slot grammar and shared parameters, and between supported local and goal motion-intent families through action-law semantics induced by the same predictor rather than pointwise latent equality. INTACT also provides intact transfer from RGB evidence to action-effective latent intent coordinates and from intent families to their corresponding action-law families. Asymmetric endpoint gradients ground physical successors and fix future goals as anchors, joining representation learning and control without pointwise latent matching or globally linear dynamics. The resulting coordinates support a robust distributional action law: its conditional mean serves directly as a search-free policy, while sampling remains available for diversity or optional verification. On the four official LeWM tasks, one-epoch, zero-search models reach 85.78\%, 100.00\%, 97.67\%, and 97.89\% success. Optional local CEM centered on the Direct plan reaches 96.86\% macro success using 384 instead of 9,000 candidate sequences, reducing sampling by 23.44times while improving pure CEM by 16.00 points. One shared four-task encoder reaches 89.39\% E5 Direct macro and improves every task over jointly trained LeWM, while predicted--expert action-family kNN tracks Direct success at r=0.954. Direct inference takes 2.9--5.5 ms.