Vidu S1:一個即時互動式影片生成模型
Vidu S1: A Real-Time Interactive Video Generation Model
July 3, 2026
作者: Jintao Zhang, Kai Jiang, Jintao Chen, Xu Wang, Yang Luo, Yuji Wang, Dechuang Chen, Jungang Li, Chengyang Ye, Marco Chen, Hongzhou Zhu, Min Zhao, Yuxuan Jiang, Zhengkun Huang, Chendong Xiang, Kaiwen Zheng, Haoxu Wang, Xiaohang Wang, Qi Jia, Xin Chen, Yimin Chen, Youhe Jiang, Fangcheng Fu, Zhijie Deng, Fan Bao, Jianfei Chen, Jun Zhu
cs.AI
摘要
我們推出Vidu S1,這是一個支援語音控制數位角色的即時互動式影片生成模型。使用者可透過語音指令隨時控制影片生成內容。Vidu S1支援無限制長度的即時影片生成,無模糊、漂移或視覺失真。基於TurboDiffusion與TurboServe技術,Vidu S1能在一般消費級GPU上以高達42 FPS的幀率輸出540p即時影片。使用者可上傳真人、動漫及寵物的自訂圖像,並選擇不同的語調以獲得個人化體驗。實驗結果顯示,Vidu S1在所有測試指標中均達到最佳表現,同時完全滿足即時推論需求。線上示範版本可於 https://vidu.com/vidu-stream 體驗。
English
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual distortion. Built with TurboDiffusion and TurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regular consumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.