ChatPaper.aiChatPaper

GraphVid:交互式圖形可控視頻生成

GraphVid: Interactive Graph-Controllable Video Generation

July 23, 2026
作者: Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjio Yu, Adheesh Juvekar, Muntasir Waheed, Ismini Lourentzou
cs.AI

摘要

可控視頻生成仍具挑戰性,原因在於文本提示或主要約束像素運動的運動控制輸入難以精確指定多物件互動。實務上,基於軌跡的控制常要求使用者為多個物件繪製精確路徑,此做法在場景複雜度增加時擴展性不佳,且遇遮蔽或重疊時會產生歧義。為實現靈活且精確的多主體控制,我們提出GraphVid——一種以圖形為條件的影像轉視頻生成模型,能透過結構化交互圖實現互動式控制。我們進一步整理GraphVid-Bench,一個大規模以互動為核心的視頻資料集,內含結構化關係註釋,以利訓練具互動感知能力的視頻生成模型。儘管使用的訓練資料與可訓練參數遠少於先前的運動控制方法,GraphVid仍展現出強大的可控性與視頻品質。相較於Motion-I2V,GraphVid的FID降低了39.9%,FVD降低了37.6%,同時PSNR(9.87提升至15.98)與SSIM(0.38提升至0.61)皆有進步。我們的研究成果凸顯了結構化語義介面作為可控視頻生成強大典範的潛力。
English
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.