ChatPaper.aiChatPaper

EffectLearner: 실세계 비디오 객체 제거를 위한 세계 인식 객체-효과 추론

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

August 6, 2026
저자: Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu
cs.AI

초록

비디오 객체 제거는 대상 객체뿐만 아니라 그로 인해 유발된 효과까지 제거하면서, 고충실도와 시공간적으로 일관된 복원을 유지해야 한다. 기존 방법들은 주로 사전 정의된 효과 범주와 고정된 데이터 분포로부터 객체-효과 대응 관계를 암시적으로 학습하므로, 복합적 효과, 공간적으로 분리되거나 약하게 상관된 효과, 롱테일 물리 현상, 동적으로 진화하는 상호작용을 포함하는 복잡한 실제 장면에 대한 일반화에 한계가 있다. 본 논문에서는 VLM 기반 객체-효과 추론기(Reasoner)와 DiT 기반 비디오 지우개(Video Eraser)를 결합한 의미론적 추론 강화 프레임워크인 EffectLearner를 제안한다. 구조화된 효과 분석 프롬프트의 안내에 따라 추론기는 대상 객체가 강조된 비디오에 대해 교차 모달 추론을 수행하고 간결한 효과 인식 컨텍스트를 추출하며, 이는 비디오 지우개가 포괄적인 객체-효과 제거를 수행하도록 안내한다. 모션 인식 마스크 가이드와 모션 일관성 감독은 객체의 움직임과 장면 역학의 진화 하에서 제거 범위와 시공간적 안정성을 추가로 향상시킨다. 복잡한 실제 시나리오에서 프레임워크를 완전히 활용하기 위해, 복잡한 객체 유발 효과를 위해 특별히 설계된 짝지어진 비디오 데이터셋인 EffectWorld를 구축하고, 공통 감독과 복합 효과 데이터를 결합한 점진적 훈련 커리큘럼을 도입한다. 표준 ROSE-Bench에서 EffectLearner는 대부분의 지표에서 기존 베이스라인을 능가하며, EffectWorld-Eval과 도전적인 EffectWorld-Wild에서도 명확한 우위를 보여 복잡한 실제 장면에서 고품질 비디오 객체 제거를 달성할 수 있음을 입증한다.
English
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.