ChatPaper.aiChatPaper

영상 = 세계 + 이벤트 스트림

Video = World + Event Stream

July 16, 2026
저자: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Cheng Yu, Chen Liang, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zhiwei Lin, Zoubin Bi
cs.AI

초록

우리는 Wan-Streamer v0.3을 소개한다. 이는 기존의 네이티브 스트리밍 상호작용 모델을 단일 조직 관점, 즉 비디오는 세계(world)와 이벤트 스트림(event stream)의 결합이라는 관점으로 재구성한다. 세계(world)는 비디오가 전개되는 지속적인 맥락으로, 환경, 장면, 주체, 주변 음향 조건, 음성 특성 및 기타 상대적으로 안정적인 조건들을 포함한다. 이벤트 스트림(event stream)은 해당 세계 내에서 시간에 따라 변화하는 모든 것으로, 장면 또는 환경 변화, 주체 행동, 음성 및 기타 소리를 포함한다. 이는 대량의 실제 비디오에 대한 범용 사전학습 작업을 제공한다. 즉, 세계와 들어오는 입력이 주어졌을 때, 세계가 실시간으로 어떻게 움직이고, 변화하며, 반응하는지 예측하는 것이다. 그 결과 얻어진 역량은 광범위한 실시간 다운스트림 작업군에 특화될 수 있다. 우리는 이를 실시간 전이중(full-duplex) 시청각 상호작용에 구현하였으며, 여기서 이벤트 스트림은 에이전트의 음성과 자유형식 행동으로 구성된다. 기능적으로, 모델의 다중 모드 이해 과정은 시각-언어-행동(vision-language-action)과 유사하다. 즉, 다중 모드 사용자 입력을 언어 형태의 음성 및 행동 액션으로 매핑한다. Wan-Streamer v0.3은 v0.2의 운영 사양을 유지한다: 25 FPS의 640x368 비디오, 160ms 스트리밍 단위, 약 200ms의 모델 측 응답 지연 시간, 그리고 350ms 양방향 네트워크 예산 하에서 약 550ms의 총 상호작용 지연 시간.
English
We present Wan-Streamer v0.3, which reframes our native-streaming interaction model under a single organizing view: a video is a world plus an event stream. The world is the persistent context in which a video unfolds, including the environment, scene, subjects, ambient acoustic conditions, voice characteristics, and other relatively stable conditions. The event stream is everything that changes over time within that world, including scene or environmental changes, subject behavior, speech, and other sounds. This yields a general-purpose pretraining task over large amounts of real video: given a world and incoming input, predict how the world moves, changes, and responds in real time. The resulting competence can be specialized to a broad family of real-time downstream tasks. We instantiate it on real-time full-duplex audio-visual interaction, where the event stream is the agent's speech together with free-form behavior. Functionally, the model's multimodal understanding process is vision-language-action-like: it maps multimodal user input to language-form speech and behavior actions. Wan-Streamer v0.3 preserves the v0.2 operating point: 640x368 video at 25 FPS, a 160 ms streaming unit, approximately 200 ms model-side response latency, and approximately 550 ms total interaction latency under a 350 ms bidirectional network budget.