ChatPaper.aiChatPaper

视频 = 世界 + 事件流

Video = World + Event Stream

July 16, 2026
作者: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Cheng Yu, Chen Liang, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zhiwei Lin, Zoubin Bi
cs.AI

摘要

我们提出 Wan-Streamer v0.3,该版本将我们的原生流式交互模型重新定义为一个统一视角:视频即世界加事件流。世界是视频展开过程中持续存在的上下文,包括环境、场景、主体、环境声学条件、语音特征及其他相对稳定的状态。事件流则是该世界中随时间变化的一切,包括场景或环境变化、主体行为、语音及其他声音。由此,在大量真实视频上,我们构建了一个通用预训练任务:给定世界和输入信号,预测世界如何实时移动、变化并作出响应。由此获得的能力可被迁移至广泛的实时下游任务中。我们将其实例化于实时全双工音视频交互,其中事件流由智能体的语音及自由形式的行为构成。功能上,模型的多模态理解过程类似于视觉-语言-动作:它将多模态用户输入映射为语言形式的语音和行为动作。Wan-Streamer v0.3 保留了 v0.2 的运行指标:640×368 视频以 25 FPS 处理,160 毫秒的流式单元,约 200 毫秒的模型端响应延迟,以及在 350 毫秒双向网络预算下约 550 毫秒的总交互延迟。
English
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.