ChatPaper.aiChatPaper

EffectLearner:面向真实世界视频物体移除的世界感知物体效应推理

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

August 6, 2026
作者: Feier Wu, Wanke Xia, Xu He, Zilang Zhou, Si Chen, Dongxia Liu, Liyang Chen, Qimeng Wu, Zhengbo Zhang, Wenming Yang, Zhiyong Wu
cs.AI

摘要

视频物体移除不仅需要消除目标物体,还需移除其引发的效应,同时保持高保真且时空一致的恢复效果。现有方法主要从预定义的效应类别和固定数据分布中隐式学习物体-效应对应关系,限制了其对复杂真实场景的泛化能力,这些场景涉及组合效应、空间分离或弱相关的效应、长尾物理现象以及动态演化的交互。我们提出EffectLearner,一种语义推理增强框架,将基于VLM的物体-效应推理器与基于DiT的视频擦除器相结合。在结构化效应分析提示的引导下,推理器对目标高亮视频进行跨模态推理,提取紧凑的效应感知上下文,从而引导视频擦除器实现全面的物体-效应移除。运动感知掩码引导与运动一致性监督进一步提升了物体运动和场景动态演化条件下的移除覆盖率与时空稳定性。为在挑战性真实场景中充分利用该框架,我们进一步构建了EffectWorld,一个专门针对复杂物体引发效应设计的成对视频数据集,并引入了一种结合通用监督与复杂效应数据的渐进式训练课程。在标准ROSE-Bench基准上,EffectLearner在大多数指标上优于现有基线,并在EffectWorld-Eval及更具挑战性的EffectWorld-Wild上均展现出明显优势,证明了其在复杂真实场景中提供高质量视频物体移除的能力。
English
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.