ChatPaper.aiChatPaper

Wan-Streamer v0.2: 더 높은 해상도, 동일한 레이턴시

Wan-Streamer v0.2: Higher Resolution, Same Latency

July 5, 2026
저자: Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zoubin Bi
cs.AI

초록

본 논문에서는 Wan-Streamer v0.2를 소개한다. 이는 기본 스트리밍 방식의 종단 간 오디오-비주얼 상호작용 모델의 지연 시간을 유지하면서 성능을 개선한 버전이다. v0.2는 v0.1의 모델링 구조를 유지하지만, 상호작용 출력 스트림의 해상도를 192×336에서 640×368로 높이면서도 25 FPS에서 약 200ms의 모델 측 신호 간 지연 시간을 유지한다. 향상된 해상도의 스트림은 실시간 대화 중에도 자세, 시선, 손, 주변 사물 및 로컬 장면 구성이 식별 가능한 장면 기반 미드샷 에이전트를 지원한다. 더 큰 비주얼 스트림을 사용자 체감 지연 시간 없이 지원하기 위해, v0.2는 싱글 GPU 저지연 경로를 통해 스트리밍 인지를 담당하는 thinker, 생성 캐시를 구축하는 짧은 언어/상태 Transformer 패스, 그리고 최종 디코딩을 유지한다. performer는 값비싼 다음 유닛 잠재 생성(unit latent generation)을 위한 멀티 GPU Ulysses 스타일의 컨텍스트 병렬 그룹이 된다. 각 performer rank는 수신된 K/V를 사전 샤딩된 로컬 캐시에 기록한다. 긴 고해상도 잠재 비디오 시퀀스는 노이즈 제거를 위해 rank 간에 분할되고 Ulysses 통신을 통해 집계되는 반면, 훨씬 짧은 오디오 잠재 시퀀스는 시퀀스 샤딩 없이 생성된다. 이러한 분할에서 thinker의 언어/상태 연산은 K/V 조건으로만 performer에 도달하므로, performer 그룹 내에서 별도의 언어 시퀀스를 통신할 필요가 없다. 이는 thinker-performer 간의 컴팩트한 경계를 유지하면서 추가 하드웨어를 비주얼 생성에 집중시키며, 350ms 양방향 네트워크 예산을 포함할 때 총 원격 상호작용 지연 시간을 약 550ms로 유지한다.
English
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises the interactive output stream from 192x336 to 640x368 while preserving approximately 200 ms model-side signal-to-signal latency at 25 FPS. The higher-resolution stream supports scene-grounded mid-shot agents whose posture, gaze, hands, nearby objects, and local scene layout remain legible during real-time conversation. To support the larger visual stream without adding user-visible delay, v0.2 keeps the thinker as a single-GPU low-latency path for streaming perception, the short language/state Transformer pass that builds the generation cache, and final decoding. The performer becomes a multi-GPU Ulysses-style context-parallel group for the expensive next-unit latent generation. Each performer rank writes incoming K/V into a pre-sharded local cache. The long high-resolution latent video sequence is split across ranks for denoising and gathered through Ulysses communication, while the much shorter audio latent sequence is generated without sequence sharding. In this split, the thinker's language/state computation reaches the performer only as K/V conditioning, so no separate language sequence has to be communicated inside the performer group. This concentrates additional hardware on visual generation while preserving the compact thinker-performer boundary, keeping total remote interaction latency at approximately 550 ms when a 350 ms bidirectional network budget is included.