ChatPaper.aiChatPaper

TimeLens2: 멀티모달 LLM을 활용한 범용 비디오 시간적 근거 찾기

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

July 19, 2026
저자: Yuhan Zhu, Changlian Ma, Xiangyu Zeng, Xinhao Li, Zhiqiu Zhang, Songze Li, Jun Zhang, Tianxiang Jiang, Yuandong Yang, Ziang Yan, Zikang Wang, Xinyu Chen, Haoran Chen, Shaowei Zhang, Limin Wang
cs.AI

초록

비디오 멀티모달 대규모 언어 모델(MLLM)은 비디오에서 발생하는 일을 설명할 수 있지만, 해당 증거가 발생하는 시점을 식별하는 경우는 드물다. 우리는 하나의 모델이 비디오 길이, 도메인, 쿼리 형식, 시점에 걸쳐 가변 개수의 증거 구간 집합을 예측하는 범용 비디오 시간적 근거 찾기(generalist video temporal grounding)를 연구한다. 기존 학습 전략은 이러한 집합값(set-valued) 작업과 정렬되지 않는다. 긴 비디오 레이블은 종종 취약한 일회성 주석에 의존하는 반면, 강화 학습 보상은 겹치지 않는 예측을 구분하지 못하거나 깨지기 쉬운 세그먼트 매칭을 요구한다. TimeLens2는 감독과 최적화 전반에 걸쳐 시간적 증거를 구간 집합으로 취급한다. TimeLens2-93K는 캡션 기반 제안, 독립적 위치 추정, 교차 에이전트 합의, 의미 검증 및 경계 개선을 통해 신뢰할 수 있는 다중 구간 감독을 구축한다. 시간적 Wasserstein 보상은 병합된 구간 지지 집합 위의 균등 분포 간의 정확한 1차원 \(W_1\)을 계산하여, 불균등한 개수와 동등한 분할 하에서 밀집되고 매칭이 필요 없는 피드백을 제공하며, 시간적 IoU는 정밀한 중첩 피드백으로 이를 보완한다. 일곱 개의 벤치마크에 걸쳐 TimeLens2-2B는 모든 크기 일치 기준선을 각 벤치마크에서 능가하며, 4B 및 8B 변종은 최첨단 성능을 달성하여 최대 397B 파라미터의 오픈소스 모델을 능가한다. 2B, 4B, 8B 변종은 각각 Qwen3-VL 백본 대비 14.2, 13.0, 18.1 mIoU 포인트 개선을 보인다.
English
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.