O-VAD: オブジェクト中心の追跡と推論による産業用ビデオ異常検出
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
July 20, 2026
著者: Mei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu
cs.AI
要旨
産業ビデオ異常検出(IVAD)は、産業プロセスにおける異常な物体や事象を識別することを目的としており、現代の製造業や品質管理システムにおいて極めて重要である。既存のVLM(視覚言語モデル)ベースの異常推論手法は、一般的な領域における未定義の異常を検出できる。しかし、複雑な物体変形、厳格な物理法則、工程上の制約が特徴的な産業環境では、その性能は低下する。このような相互作用が頻繁な検出の複雑さに対処するため、我々は訓練不要でドメイン固有の知識を必要としないエージェンティックフレームワークを導入する。これは人間の検査員のように、物体の状態進化に重点を置く。本手法は、検出された物体の時間的・空間的ダイナミクスとその背後にある変形を経時的に追跡し、物体ごとの時間的状態軌跡に基づいて推論を行うことで、グラウンディングされたフレーム内の異常物体を識別する。本手法は、正常クリップでの再学習やテスト時推論のためのコンテキストとしてのドメイン知識の注入に依存する従来手法の限界を克服する。3つのIVADデータセットを用いた広範な実験により、本手法が最先端のVLM、エージェンティックフレームワーク、および各データセットで微調整された従来のVAD手法を上回る性能を示すとともに、異常プロセスとその種類に関する解釈可能なレポートを提供することを実証した。
English
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.