O-VAD:基於物體中心追蹤與推理的工業影片異常檢測
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
July 20, 2026
作者: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu
cs.AI
摘要
工業視頻異常檢測(IVAD)旨在識別工業流程中的異常物體與事件,這對現代製造與品質控制系統至關重要。現有的基於VLM的異常推理方法雖能檢測通用領域的開放式異常,但在具有複雜物體變形、嚴格物理約束與程序限制的工業環境中表現下降。為應對此類高度互動檢測的複雜性,我們提出了一種無需訓練、無需領域特定知識的智能體框架,其運作方式類似於人類檢查員,著重於物體狀態的演變。該框架旨在追蹤檢測物體隨時間變化的時空動態與潛在轉變,並基於物體層面的時間狀態軌跡進行推理,以在定位框架中識別異常物體。我們的方法克服了先前方法需基於正常片段重新訓練或在測試推論時注入領域知識作為上下文的限制。在三個IVAD數據集上的廣泛實驗表明,我們的方法優於前沿的VLM、智能體框架及在各自數據集上微調的傳統VAD方法,同時能提供關於異常過程與類型的可解釋報告。
English
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.