Marionette: 세계 상태 예측, 지오메트리 렌더링, 외관 페인팅
Marionette: Predicting World States, Rendering Geometry, Painting Appearance
August 14, 2026
저자: Zian Meng, Zhen Li, Chuanhao Li, Qiang Li, Kaipeng Zhang
cs.AI
초록
대화형 게임 세계 모델은 일반적으로 시각적 관측을 픽셀 또는 잠재 공간에서 직접 자기회귀하므로, 자세, 기하학, 가림과 같은 구조적 속성이 동일한 생성 시퀀스에 의해 암묵적으로 유지되어야 한다. 긴 지평에 걸쳐 이러한 잠재 세계 속성의 오류가 누적되어 일관성과 제어 가능성이 취약해진다. 우리는 진화하는 세계 상태를 명시적으로 모델링하고, 정확한 기하학 계산을 매개변수가 없는 고정 렌더러에 위임하며, 신경망 모델은 외관 합성만 담당하게 한다. 이 아이디어를 관절형 캐릭터가 등장하는 대화형 게임용 세계 모델인 Marionette로 구현한다. 첫째, 2단계 자기회귀 동역학 모델이 다중 개체 관절형 골격, 미터 단위 루트 궤적, 회전을 포함하는 명시적이고 해석 가능한 276차원 3D 세계 상태를 예측한다. 둘째, 매개변수가 없는 그래픽 브리지가 예측된 상태를 자세 제어 비디오로 변환하고, 세계 공간 기하학과 가림을 닫힌 형태로 계산한다. 셋째, 제어 조건부 비디오 확산 관측 모델이 결과적으로 얻어진 구조적 제어로부터 사진실사 RGB 관측을 합성한다. 실험을 통해 Marionette의 두 가지 속성을 확인한다. 첫째, 예측된 세계 상태는 직접 제어 가능하다. 일치하지 않는 행동 스트림을 강제하면 48개의 보류 세그먼트에서 루트 정렬 관절 오차가 31% 변한다. 둘째, 장기적 행동은 상태에서 결정되며, 그 상태에서 수정할 수 있다. 방치할 경우 생성된 두 캐릭터는 21.2m까지 멀어지고(기록된 세션은 약 5m를 유지), 프레임의 3분의 1에서 지면 관통이 발생한다. 명시적 상태에 부과된 두 가지 규칙, 즉 지형 충돌체와 분리 상한은 관측 모델을 변경하지 않고도 관통을 66% 줄이고 두 캐릭터가 상호작용하도록 유지한다. 예측된 상태를 통해 외관을 라우팅하는 것은 FVD가 기록된 자세의 799와 비교하여 831로, 우리가 감지할 수 있는 충실도 손실이 없다.
English
Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two-stage autoregressive dynamics model predicts an explicit and interpretable 276-dimensional 3D world state comprising multi-entity articulated skeletons, metric root trajectories, and rotations. Second, a zero-parameter graphics bridge converts the predicted state into pose-control videos, computing world-space geometry and occlusion in closed form. Third, a control-conditioned video-diffusion observation model synthesizes photorealistic RGB observations from the resulting structured controls. Our experiments establish two properties of Marionette. First, the predicted world state is directly controllable. Forcing a mismatched action stream changes root-aligned joint error by 31% across 48 held-out segments. Second, long-horizon behaviour is determined in the state, and can be repaired there. Left free, the two generated characters drift to 21.2 m apart (recorded sessions stay near 5 m) and a third of frames show ground penetration. Two rules imposed on the explicit state, a terrain collider and a separation cap, cut penetration by 66% and keep the pair engaged, with no change to the observation model. Routing appearance through the predicted state costs no fidelity we can detect, at an FVD of 831 against 799 for recorded pose.