H3-World: 언어 이해를 세계 제어로 전환하기
H3-World: Turning Language Understanding into World Control
September 1, 2026
저자: Danze Chen, Zeqing Wang, Ziyue Lin, Xingyi Yang, Yeying Jin
cs.AI
초록
본 논문에서는 33B MiniMax-H3 비디오 생성기를 대화형 세계 모델로 전환하는 효율적인 프레임워크인 H3-World를 제안한다. 핵심 발견은 대규모 비디오 생성기의 성능이 향상됨에 따라 언어가 제어를 위한 자연스러운 인터페이스로 부상하고 있다는 점이다. 예를 들어, MiniMax-H3는 자연어 명령을 통한 캐릭터 행동 및 카메라 움직임의 제로샷 제어를 이미 지원한다. 이를 기반으로 H3-World는 전용 행동 모듈을 도입하지 않고도 이러한 조악한 언어 인터페이스를 정밀하고 시간적으로 근거한 세계 제어로 전환한다. 구체적으로, 각 행동을 캐릭터 및 카메라 명령의 구조화된 조합으로 표현하고 이를 해당 시간적 비디오 잠재 표현과 정렬한다. 제어의 시간적 정밀성을 확보하기 위해 시간적 어텐션 라우팅을 추가로 도입하여 각 명령이 의도된 시간 구간에만 적용되도록 제한하고 행동 간 제어 누출을 줄인다. 중요하게도 H3-World는 대규모 비디오 사전학습 중 학습된 의미론적 표현을 직접 재사용하며 경량 적응만을 요구한다. 단 8,000개의 게임플레이 샘플, 10,000회의 LoRA 최적화 단계, 0.199%의 학습 가능한 파라미터만으로 H3-World는 강력한 생성 품질을 유지하면서 효과적인 캐릭터 및 카메라 제어를 달성한다. 또한 보지 못한 시나리오에도 일반화된다. 이러한 결과는 대규모 비디오 생성기에서 나타나는 제어 능력이 대화형 세계 제어로 효율적으로 전환될 수 있음을 보여준다.
English
We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.