動画 = 世界 + イベントストリーム
Video = World + Event Stream
July 16, 2026
著者: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Cheng Yu, Chen Liang, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zhiwei Lin, Zoubin Bi
cs.AI
要旨
我々はWan-Streamer v0.3を提案する。これは、単一の整理された見解の下でネイティブストリーミング対話モデルを再構築するものである:ビデオとは「世界」と「イベントストリーム」から成る。世界とは、ビデオが展開される持続的なコンテキストであり、環境、シーン、被写体、周囲の音響条件、音声特性、その他の比較的安定した状況を含む。イベントストリームとは、その世界の中で時間とともに変化するすべてのものであり、シーンや環境の変化、被写体の行動、発話、その他の音を含む。これにより、大量の実ビデオに対する汎用的な事前学習タスクが導かれる:与えられた世界と入力に対して、世界がどのように動き、変化し、リアルタイムで応答するかを予測する。得られた能力は、幅広いリアルタイム下流タスクに特化させることができる。我々はこれを、エージェントの発話と自由形式の行動からなるイベントストリームを用いた、リアルタイム全二重音声-視覚対話に実装する。機能的には、モデルのマルチモーダル理解プロセスは視覚-言語-行動に類似している:マルチモーダルなユーザー入力を、言語形式の発話と行動のアクションにマッピングする。Wan-Streamer v0.3はv0.2の動作点を維持している:640x368ビデオを25 FPS、160 msのストリーミング単位、約200 msのモデル側応答レイテンシ、および350 msの双方向ネットワーク予算の下で約550 msの総対話レイテンシ。
English
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.