ChatPaper.aiChatPaper

GraphVid: 인터랙티브 그래프 제어 가능 비디오 생성

GraphVid: Interactive Graph-Controllable Video Generation

July 23, 2026
저자: Vedant Shah, Onkar Susladkar, Tushar Prakash, Kiet Nguyen, Tianjio Yu, Adheesh Juvekar, Muntasir Waheed, Ismini Lourentzou
cs.AI

초록

제어 가능한 비디오 생성은 텍스트 프롬프트나 주로 픽셀 이동을 제약하는 모션 제어 입력만으로는 정밀한 다중 객체 상호작용을 명시하기 어렵기 때문에 여전히 도전적인 과제이다. 실제로 궤적 기반 제어는 사용자가 여러 객체에 대해 정확한 트랙을 그리도록 요구하며, 이는 장면 복잡도에 따라 확장성이 떨어지고 폐색이나 중첩 상황에서 모호해진다. 유연하면서도 정밀한 다중 객체 제어를 가능하게 하기 위해, 본 논문에서는 구조화된 상호작용 그래프를 통해 대화형 제어를 구현하는 그래프 조건부 이미지-비디오 생성 모델인 GraphVid를 제안한다. 또한, 상호작용 인식 비디오 생성 모델의 훈련을 가능하게 하기 위해 구조화된 관계형 주석이 포함된 대규모 상호작용 중심 비디오 데이터셋인 GraphVid-Bench를 구축한다. GraphVid는 기존 모션 제어 방법보다 훨씬 적은 학습 데이터와 학습 가능한 파라미터를 사용함에도 불구하고 강력한 제어성과 비디오 품질을 제공한다. Motion-I2V와 비교하여 GraphVid는 FID를 최대 39.9%, FVD를 37.6% 감소시키는 동시에 PSNR(9.87→15.98)과 SSIM(0.38→0.61)을 향상시킨다. 본 연구 결과는 구조화된 의미 인터페이스가 제어 가능한 비디오 생성을 위한 강력한 패러다임이 될 가능성을 강조한다.
English
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs. We further curate GraphVid-Bench, a large-scale interaction-centric video dataset with structured relational annotations to enable training of interaction-aware video generation models. Despite using substantially less training data and fewer trainable parameters than prior motion-control methods, GraphVid delivers strong controllability and video quality. Compared with Motion-I2V, GraphVid reduces FID by up to 39.9% and FVD by 37.6%, while improving PSNR (9.87=>15.98) and SSIM (0.38=>0.61). Our results highlight the potential of structured semantic interfaces as a powerful paradigm for controllable video generation.