ChatPaper.aiChatPaper

Game2World 엔진: 실제 게임플레이 영상을 활용한 월드 모델 학습

Game2World Engine: Unlocking In-the-Wild Gameplay Videos for World Model Training

August 25, 2026
저자: Wenxuan Shen, Dongna Jin, Dongping Chen
cs.AI

초록

비디오 게임은 다양한 환경, 복잡한 상호작용, 풍부한 실제(in-the-wild) 게임플레이 비디오를 제공하여 비디오 세계 모델의 훈련 데이터를 확장 가능한 방식으로 공급한다. 그러나 원시 게임플레이 푸티지는 게임 세계와 화면 공간 인터페이스를 얽히게 하여 게임 특유의 편향과 무관한 동역학을 도입함으로써 세계 모델 훈련을 방해한다. 이 문제를 해결하기 위해 우리는 게임플레이 UI 그라운딩 및 제거를 정형화하는 풀스택 프레임워크인 GameUI-Taxonomy와 G2WEngine을 소개한다. G2WEngine은 실제 게임플레이 비디오에서 재사용 가능한 UI 자산을 자동으로 추출하고, 깨끗한 푸티지 위에 시간적으로 일관된 UI 오버레이를 합성한다. 이 엔진을 활용하여 우리는 정밀한 재구성 대상이 포함된 96K개의 합성 쌍(paired) 비디오와 실제 평가를 위한 303개 게임에서 수집한 1,079개의 실제(in-the-wild) 클립으로 구성된 Game2World를 구축한다. 이 데이터셋의 자산 라이브러리에는 1,010개의 대표 게임플레이 프레임에서 수집된 21개 분류 범주의 검증된 5,132개 UI 요소가 포함되어 있다. Game2World에 기반하여 우리는 다중 모달 의미 이해와 비디오 편집 기능을 결합하는 마스크 없는 게임플레이 UI 제거 모델인 GameCleaner를 제안한다. GameCleaner는 마스크 기반 방법과 달리 기본 장면 콘텐츠와 시간적 동역학을 보존하면서 다양한 HUD 요소를 직접 식별하고 제거한다. 통제된 예비 실험에서 UI 없는 게임플레이로 훈련된 세계 모델은 UI가 덮인 데이터로 훈련된 모델보다 전체 VideoReward를 6.83% 향상시켰다. UI 제거 평가에서 GameCleaner는 합성 비디오에서 평균 AAR 95.36을 달성하여 가장 강력한 시간적 마스크 기준선보다 57.3% 우수했으며, 99.8의 배경 보존과 함께 실제(in-the-wild) 환경 최고 AAR 80.05를 얻었다. 이러한 결과는 인터넷 게임플레이 비디오를 고품질 세계 모델 훈련 데이터로 변환하는 확장 가능한 잠재력을 입증한다. 코드, 데이터셋, 모델은 https://github.com/Dongping-Chen/Game2World에서 제공될 예정이다.
English
Video games provide a scalable source of training data for video world models, offering diverse environments, complex interactions, and abundant in-the-wild gameplay videos. However, raw gameplay footage entangles the game world with screen-space interfaces, introducing game-specific biases and irrelevant dynamics that hinder world-model training. To address this problem, we introduce GameUI-Taxonomy and G2WEngine, a full-stack framework that formalizes gameplay UI grounding and removal. G2WEngine automatically extracts reusable UI assets from real gameplay videos and synthesizes temporally coherent UI overlays on clean footage. Using this engine, we construct Game2World, comprising 96K synthetic paired videos with precise reconstruction targets and 1,079 in-the-wild clips from 303 games for realistic evaluation. Its asset library contains 5,132 verified UI elements across 21 taxonomy categories, collected from 1,010 representative gameplay frames. Based on Game2World, we propose GameCleaner, a mask-free gameplay UI removal model that combines multimodal semantic understanding with video editing capabilities. Unlike mask-based methods, GameCleaner directly identifies and removes diverse HUD elements while preserving the underlying scene content and temporal dynamics. In a controlled pilot, world models trained on UI-free gameplay improve overall VideoReward by 6.83% over those trained on UI-overlaid data. On UI-removal evaluation, GameCleaner achieves an average AAR of 95.36 on synthetic videos, outperforming the strongest temporal mask baseline by 57.3%, and obtains the best in-the-wild AAR of 80.05 with 99.8 background preservation. These results demonstrate the scalable potential of transforming Internet gameplay videos into high-quality world-model training data. Code, dataset, and model will be available at https://github.com/Dongping-Chen/Game2World.