影片 = 世界 + 事件流
Video = World + Event Stream
July 16, 2026
作者: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Cheng Yu, Chen Liang, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zhiwei Lin, Zoubin Bi
cs.AI
摘要
我們提出 Wan-Streamer v0.3,該版本將我們的原生串流互動模型重新框架化為一個統一的觀點:影片是一個世界加上一個事件流。世界是影片展開的持續背景,包含環境、場景、主體、環境聲學條件、語音特徵及其他相對穩定的條件。事件流是世界內隨時間變化的一切,包括場景或環境變化、主體行為、語音及其他聲音。這產生了一個基於大量真實影片的通用預訓練任務:給定一個世界和輸入,預測世界如何即時移動、變化與回應。所獲得的能力可專門應用於廣泛的即時下游任務。我們將其實例化於即時全雙工視聽互動場景,其中事件流是智能體的語音及自由形式的行為。功能上,該模型的多模態理解過程類似於視覺-語言-動作:它將多模態使用者輸入映射為語言形式的語音與行為動作。Wan-Streamer v0.3 保留了 v0.2 的操作設定:640×368 解析度影片、25 FPS、160 毫秒串流單元、約 200 毫秒模型端回應延遲,以及在 350 毫秒雙向網路預算下的約 550 毫秒總互動延遲。
English
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.