잠재 행동을 의도로 활용한 세계 행동 모델의 효율적 미래 상상
Latent Action as Intention Enables Efficient Future Imagination for World Action Models
August 25, 2026
저자: Xiang Li, Yupeng Zheng, Songen Gu, Huailiang Ma, Feng Yu, Xian Nie, Shanshuai Yuan, Yujie Zang, Weize Li, Shuai Tian, Moyang Liu, Ya-Qin Zhang, Wenchao Ding
cs.AI
초록
세계 행동 모델(WAM)은 관찰이 어떻게 변화하는지 모델링하여 로봇 제어를 개선하지만, 테스트 시점에 미래 관찰을 생성하는 과정은 상당한 지연 시간을 유발한다. Fast-WAM은 효율성을 위해 이 과정을 제거한다. 그러나 동일 조건으로 구현했을 때, Fast-WAM은 미래 인지(future-aware) 대안들보다 낮은 일반화 성능을 보이며, 특히 로봇 시연 데이터가 부족한 상황과 분포 외(out-of-distribution) 시나리오에서 두드러진다. 이러한 격차를 해소하기 위해, 우리는 **LAWA**를 제안한다. LAWA는 미래 의도의 작동적 표현으로 컴팩트한 잠재 행동을 사용하는 WAM 아키텍처로, 미래 관찰을 생성하지 않고도 효율적인 테스트 시점 미래 상상을 가능하게 한다. 구체적으로, 행동 없는 사전학습(action-free pre-training)으로 강화된 이산 토크나이저가 조작 중심의 코드북 목표값을 생성한다. LAWA는 이러한 목표값에 고정된 연속 잠재 상태를 실행 가능한 행동 청크와 함께 공동으로 노이즈 제거하며, 추론 시 미래 비디오 분기는 생략한다. RoboCasa에서 LAWA는 퓨샷 및 전체 데이터 설정에서 각각 65.6%와 80.8%의 최고 수준(SOTA) 평균 성공률을 달성하여, 동일 조건의 Fast-WAM 기준선 대비 각각 9.6포인트와 4.5포인트 향상된 성능을 보인다. 또한 동일 조건의 Joint-WAM 변형의 성능 수준을 유지하면서도 추론 지연 시간을 42.9% 단축한다. LAWA는 LIBERO-Plus에서 경쟁력 있는 제로샷 강건성을 입증하고 실제 로봇 작업에서도 우수한 성능을 보여준다. 이러한 결과는 미래 상상이 폐기될 필요가 없음을 시사한다. 컴팩트한 잠재 행동을 통해 미래 상상을 유지하는 것은 성능, 일반화, 지연 시간 간의 효과적인 절충을 제공한다. 코드와 모델은 공개될 예정이다.
English
World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-WAM removes this process for efficiency; however, our matched implementations show lower generalization for Fast-WAM than for future-aware alternatives, especially with scarce robot demonstrations and in out-of-distribution scenarios. To bridge this gap, we introduce **LAWA**, a WAM architecture that uses compact latent actions as an operational representation of future intentions, enabling efficient test-time future imagination without generating future observations. Specifically, a discrete tokenizer enhanced by action-free pre-training produces manipulation-centric codebook targets. LAWA jointly denoises a continuous latent state anchored to these targets with executable action chunks while omitting the future-video branch at inference. On RoboCasa, LAWA achieves state-of-the-art average success rates of 65.6% and 80.8% in the few-shot and full data settings, improving over the matched Fast-WAM baseline by 9.6 and 4.5 points, respectively. It also preserves the performance level of the matched Joint-WAM variant while requiring 42.9% lower inference latency. LAWA also demonstrates competitive zero-shot robustness on LIBERO-Plus and superior performance on real-world tasks. These results show that future imagination need not be discarded: retaining it with compact latent actions yields an effective trade-off among performance, generalization, and latency. Code and models will be released.