ChatPaper.aiChatPaper

EffectLearner:面向真實世界影片物體移除的世界感知物體效應推理

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

August 6, 2026
作者: Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu
cs.AI

摘要

影片物件移除不僅需去除目標物件,還須消除其引發的效果,同時維持高保真度與時空連貫的還原。現有方法主要從預定義的效果類別與固定資料分布中隱式學習物件與效果之間的對應關係,因而限制其對複雜真實場景的泛化能力,包括組合效果、空間分離或弱相關效果、長尾物理現象,以及動態演變的互動。我們提出 EffectLearner,一個具語意推理增強的框架,結合了基於視覺語言模型(VLM)的物件-效果推理器與基於擴散變換器(DiT)的影片擦除器。在結構化效果分析提示的引導下,推理器對標記目標的影片進行跨模態推理,並提取精簡的效果感知上下文,進而引導影片擦除器達成全面的物件與效果移除。運動感知遮罩引導與運動一致性監督進一步提升在物件運動與場景動態演變下的移除覆蓋率與時空穩定性。為了在具挑戰性的真實場景中充分發揮此框架的效能,我們進一步建構 EffectWorld——一個專為複雜物件引發效果設計的成對影片資料集——並引入結合一般監督與複雜效果資料的漸進式訓練課程。在標準的 ROSE-Bench 上,EffectLearner 在多數指標上優於現有基線方法,並在 EffectWorld-Eval 及具挑戰性的 EffectWorld-Wild 上皆取得明顯優勢,展現其在複雜真實場景中提供高品質影片物件移除的能力。
English
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.