프로그래밍 가능한 월드 모델
Programmable World Model
September 9, 2026
저자: Zheng-Hui Huang, Guixu Lin, Jiacheng Lin, Yi-Chuan Huang, Ruihan Yu, Muyao Niu, Siqi Yang, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
cs.AI
초록
최근의 비디오 월드 모델들은 점점 더 사실적이고 상호작용적인 시각 경험을 생성하지만, 장기 상호작용에 걸쳐 지속적인 월드 상태를 유지하고 프로그래밍 가능한 규칙을 강제할 신뢰할 수 있는 메커니즘이 부족하다. 우리는 월드 상태 진화를 시각적 관측 생성으로부터 분리하는 프레임워크인 프로그래머블 월드 모델(Programmable World Model)을 제안한다. 에이전트는 자연어 지시를 개체 상태와 상태 전이 규칙을 지정하는 실행 가능한 프로그램으로 변환하여, 개별 개체와 이들의 상호작용을 직접 제어할 수 있게 한다. 경량 엔진은 이러한 프로그램을 실행하여 화면 밖 개체와 비시각적 속성을 포함하는 명시적이고 지속적인 전역 월드 상태를 갱신하고 유지한다. 월드 상태를 시각 생성과 연결하기 위해, 우리는 상태 증강 3D 방향성 경계 상자(OBB)를 중간 표현으로 도입한다. 이 표현은 목표 카메라 궤적과 함께 생성형 렌더러 역할을 하는 사전 학습된 비디오 모델을 위한 픽셀 정렬 시공간 조건화 신호로 결정론적으로 컴파일된다. 이 설계는 사용자가 미리 정의된 메커닉을 가진 플레이 가능한 게임을 만들고, 개별 개체를 직접 제어하며, 게임플레이 전반에 걸쳐 지속적인 월드 상태를 유지할 수 있게 한다. 우리는 또한 프로그래머블 월드 모델을 평가하기 위한 벤치마크인 CombatStateBench를 소개한다. CombatStateBench에서 우리 방법은 94%의 개수 정확도(Count Accuracy)와 98%의 상태 정확도(State Accuracy)를 달성하여, 일관된 장기(long-horizon) 생성을 지원하면서 기존 상호작용형 비디오 월드 모델을 크게 능가한다. 이러한 결과는 지속적이고 프로그래밍 가능한 월드 구축을 위해 명시적 상태 진화를 생성형 렌더링으로부터 분리하는 것의 효과를 입증한다.
English
Recent video world models generate increasingly realistic and interactive visual experiences, yet lack reliable mechanisms for maintaining persistent world state and enforcing programmable rules over extended interactions. We introduce Programmable World Model, a framework that decouples world-state evolution from visual observation generation. An agent translates natural-language instructions into executable programs that specify entity states and state-transition rules, enabling direct control over individual entities and their interactions. A lightweight engine executes these programs to update and maintain an explicit, persistent global world state, including off-screen entities and non-visual attributes. To connect world state with visual generation, we introduce state-augmented 3D oriented bounding boxes (OBBs) as an intermediate representation. This representation, together with the target camera trajectory, is deterministically compiled into pixel-aligned spatiotemporal conditioning signals for a pretrained video model serving as the generative renderer. This design allows users to create playable games with predefined mechanics, direct control over individual entities, and persistent world state throughout gameplay. We further introduce CombatStateBench, a benchmark for evaluating programmable world models. On CombatStateBench, our method achieves 94% Count Accuracy and 98% State Accuracy, substantially outperforming existing interactive video world models while supporting coherent long-horizon generation. These results demonstrate the effectiveness of separating explicit state evolution from generative rendering for building persistent, programmable worlds.