GraphVid:交互式图可控视频生成
GraphVid: Interactive Graph-Controllable Video Generation
July 23, 2026
作者: Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjio Yu, Adheesh Juvekar, Muntasir Waheed, Ismini Lourentzou
cs.AI
摘要
可控视频生成面临挑战,主要源于通过文本提示或主要约束像素运动的运动控制输入来精确指定多物体交互存在困难。在实际应用中,基于轨迹的控制通常要求用户为多个物体绘制精确轨迹,这种方式难以适应场景复杂度,且在物体遮挡或重叠时会产生歧义。为实现灵活且精确的多主体控制,我们提出了GraphVid——一种基于图条件约束的图像到视频生成模型,该模型通过结构化交互图实现交互式控制。我们进一步构建了GraphVid-Bench,这是一个以交互为中心的大规模视频数据集,包含结构化关系标注,用于训练具备交互感知能力的视频生成模型。尽管相比先前的运动控制方法,GraphVid使用的训练数据和可训练参数大幅减少,但其在可控性和视频质量上表现优异。与Motion-I2V相比,GraphVid将FID降低了39.9%,FVD降低了37.6%,同时提升了PSNR(从9.87提升至15.98)和SSIM(从0.38提升至0.61)。我们的研究结果凸显了结构化语义界面作为可控视频生成强大范式的潜力。
English
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.