Wan-Streamer v0.2: より高い解像度、同じレイテンシ
Wan-Streamer v0.2: Higher Resolution, Same Latency
July 5, 2026
著者: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zoubin Bi
cs.AI
要旨
Wan-Streamer v0.2を紹介します。これは、ネイティブストリーミング型のエンドツーエンド音声・映像対話モデルにおいて、遅延を維持しながらアップグレードしたバージョンです。v0.2はv0.1のモデル化手法を踏襲しつつ、対話出力ストリームを192×336から640×368に引き上げ、25FPSにおいて約200ミリ秒のモデル側信号間遅延を維持しています。この高解像度ストリームにより、実時間対話中に姿勢、視線、手、近傍の物体、局所的なシーンレイアウトが明瞭に判別可能な、シーンに基づくミッドショットエージェントをサポートします。v0.2は、ユーザーに認識可能な遅延を増やすことなく、より大きな視覚ストリームに対応するため、シンカーを単一GPUの低遅延パスとして維持し、ストリーミング認識、生成キャッシュを構築する短い言語・状態Transformerパス、および最終デコードを担当します。パフォーマーは、高コストな次ユニット潜在生成のために、マルチGPUによるUlyssesスタイルのコンテキスト並列グループとなります。各パフォーマーランクは、入力されたK/Vを事前にシャーディングされたローカルキャッシュに書き込みます。長く高解像度の潜在ビデオシーケンスは、ランク間で分割されてノイズ除去され、Ulysses通信によって収集されます。一方、はるかに短いオーディオ潜在シーケンスは、シーケンスシャーディングなしで生成されます。この分割において、シンカーの言語・状態計算は、K/V条件付けとしてのみパフォーマーに到達するため、パフォーマーグループ内で別個の言語シーケンスを通信する必要はありません。これにより、コンパクトなシンカー・パフォーマー境界を維持しつつ、追加のハードウェアを視覚生成に集中させ、350ミリ秒の双方向ネットワーク予算を含めた場合、総リモートインタラクション遅延は約550ミリ秒となります。
English
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-side signal-to-signal latency at 25 FPS. The higher-resolution stream supports scene-grounded mid-shot agents whose posture, gaze, hands, nearby objects, and local scene layout remain legible during real-time conversation. To support the larger visual stream without adding user-visible delay, v0.2 keeps the thinker as a single-GPU low-latency path for streaming perception, the short language/state Transformer pass that builds the generation cache, and final decoding. The performer becomes a multi-GPU Ulysses-style context-parallel group for the expensive next-unit latent generation. Each performer rank writes incoming K/V into a pre-sharded local cache. The long high-resolution latent video sequence is split across ranks for denoising and gathered through Ulysses communication, while the much shorter audio latent sequence is generated without sequence sharding. In this split, the thinker's language/state computation reaches the performer only as K/V conditioning, so no separate language sequence has to be communicated inside the performer group. This concentrates additional hardware on visual generation while preserving the compact thinker-performer boundary, keeping total remote interaction latency at approximately 550 ms when a 350 ms bidirectional network budget is included.