ChatPaper.aiChatPaper

O-VAD:通过以对象为中心的跟踪与推理实现工业视频异常检测

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

July 20, 2026
作者: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu
cs.AI

摘要

工业视频异常检测(Industrial Video Anomaly Detection, IVAD)旨在识别工业过程中的异常物体与事件,这对现代制造业与质量控制体系至关重要。现有基于视觉语言模型(VLM)的异常推理方法虽能检测通用领域的开放式异常,但在工业场景中性能显著下降——此类场景常涉及复杂的物体变换、严格的物理规律及流程约束。为应对这种交互密集型检测的复杂性,我们提出了一种无需领域特定知识且无需训练的智能体框架,该框架强调像人类检验员那样追踪物体状态演化。其设计旨在捕捉被检测物体随时间变化的时空动态与潜在变换,进而基于逐物体的时间状态轨迹进行推理,以在定位帧中识别异常物体。本方法克服了先前方法的局限——这些方法需依赖正常片段重训练或在测试时注入领域知识作为上下文。在三个IVAD数据集上的大量实验表明,本方法不仅超越了前沿VLM、智能体框架以及在各自数据集上微调的传统视频异常检测(VAD)方法,还能针对异常过程与类型提供可解释的报告。
English
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.