ChatPaper.aiChatPaper

TimeLens2: マルチモーダルLLMを用いた汎用的な動画時間的グラウンディング

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

July 19, 2026
著者: Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang
cs.AI

要旨

ビデオマルチモーダル大規模言語モデル(MLLM)は、ビデオ内で何が起こっているかを説明できるが、その根拠となる証拠がいつ発生するかを特定することはほとんどない。本研究では、汎用的なビデオ時間的グラウンディングを研究する。これは、単一のモデルがビデオの長さ、ドメイン、クエリ形式、視点を横断して、可変基数の証拠区間集合を予測するタスクである。既存の学習戦略は、この集合値タスクと整合していない。長尺ビデオのラベルは、脆弱な一回限りのアノテーションに依存することが多く、強化学習の報酬は、重複しない予測を区別できないか、または脆弱なセグメントマッチングを必要とするかのいずれかである。TimeLens2は、監視と最適化の全体を通じて、時間的証拠を区間集合として扱う。TimeLens2-93Kは、キャプション由来の提案、独立した位置特定、エージェント間のコンセンサス、意味検証、境界精緻化を通じて、信頼性の高いマルチスパン監視を構築する。我々の時間的ワッサーシュタイン報酬は、マージされた区間サポート上の一様分布間の正確な1次元W1を計算し、不等な基数や等価な断片化の下で、密でマッチング不要なフィードバックを提供する。時間的IoUは、正確なオーバーラップフィードバックでこれを補完する。7つのベンチマークにおいて、TimeLens2-2Bはすべてのサイズが一致するベースラインをすべてのベンチマークで上回り、4Bおよび8Bのバリアントは、最大397Bパラメータのオープンソースモデルを凌駕する最先端の性能を達成する。2B、4B、8Bのバリアントは、それぞれQwen3-VLバックボーンに対して14.2、13.0、18.1 mIoUポイントの改善を示している。
English
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.