ChatPaper.aiChatPaper

StatePlay: 메커니즘 일관 생성을 위한 상태 인식 게임 세계 모델

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

July 29, 2026
저자: Zijun Lin, Zeqing Wang, Cheston Tan, Bihan Wen, Yeying Jin
cs.AI

초록

최근 게임 월드 모델은 플레이어 행동에 따라 시각적으로 사실적이고 상호작용 가능한 환경을 생성할 수 있다. 그러나 게임은 단순히 픽셀만으로 정의되지 않는다. 게임은 명시적 메커니즘, 즉 체력 감소, 스킬 활성화, 게임 종료를 제어하는 상태 의존적 규칙에 의해 운영된다. 이러한 메커니즘은 체력 포인트, 스킬 미터, 타이머와 같은 정밀한 내부 상태에 의존하며, 이는 시각적 관찰과 밀접하게 연결되어 게임플레이가 어떻게 전개될지를 결정한다. 이러한 상태 역학을 모델링하지 않으면 기존 게임 월드 모델은 시각적으로 그럴듯한 전개를 생성할 수 있지만, 근본적인 게임 규칙을 위반할 수 있다. 본 논문에서는 메커니즘 일관된 생성을 촉진하기 위해 시각적 콘텐츠와 게임 상태를 공동으로 예측하는 새로운 상태 인식 게임 월드 모델인 StatePlay를 제안한다. StatePlay는 혼합 트랜스포머(MoT) 스타일 아키텍처를 채택하여 특화된 시각 및 상태 표현을 유지하면서 교차 모달 상호작용을 가능하게 하여 예측된 상태가 프레임 생성을 안내하도록 한다. 각 분기는 해당 모달리티에 적합한 고유한 목표로 추가 최적화된다. 실험 결과 StatePlay는 상태 예측에 대해 평균 정규화 L1 거리가 0.06 미만을 달성한다. 또한, 명시적 상태 모델링이 없는 모델과 비교하여 우리의 방법은 생성된 게임 전개에서 메커니즘 충실도를 18.6% 향상시킨다. 전반적으로, 우리의 연구는 상태 인식 게임 월드 모델링의 중요성을 강조하며, 픽셀 수준의 사실성을 넘어 완전하고 메커니즘적으로 충실한 게임 생성으로 나아간다.
English
Recent game world models can generate visually realistic and interactive environments conditioned on player actions. However, games are not defined by pixels alone; they are governed by explicit mechanics, namely state-dependent rules that control health reduction, skill activation, and game termination. These mechanics depend on precise internal states, such as health points, skill meters, and timers, which are tightly coupled with visual observations and determine how gameplay evolves. Without modeling these state dynamics, existing game world models may generate visually plausible rollouts but violate the underlying game rules. In this paper, we propose StatePlay, a novel state-aware game world model that jointly predicts visual content and game states to promote mechanics-consistent generation. StatePlay adopts a mixture-of-transformers (MoT)-style architecture that preserves specialized visual and state representations while enabling cross-modal interaction, allowing predicted states to guide frame generation. Each branch is further optimized with a distinct objective suited to its modality. Experiments show that StatePlay achieves an average normalized L1 distance below 0.06 for state prediction. Furthermore, compared with models without explicit state modeling, our method improves mechanics fidelity in generated game rollouts by 18.6%. Overall, our work highlights the importance of state-aware game world modeling and advances beyond pixel-level realism toward complete and mechanically faithful game generation.