ChatPaper.aiChatPaper

비디오에서 무엇이든 찾기: 효율적 생성적 시공간 비디오 그라운딩의 재고찰

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

August 28, 2026
저자: Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
cs.AI

초록

시공간 비디오 그라운딩(STVG)은 지시된 사건이 발생하는 시점을 식별하고 해당 시간 구간 전체에 걸쳐 대상 개체의 위치를 특정해야 한다. 기존의 멀티모달 대규모 언어 모델은 일반적으로 밀집된 위치 특정 궤적을 자기회귀적으로 직렬화하므로, 디코딩 지연 시간이 튜브 길이에 따라 증가하고 위치 특정 오류가 시간에 걸쳐 전파될 수 있다. 우리는 병렬 튜브 디코딩(PTD)을 소개한다. PTD는 그라운딩을 시간 블록과 뒤이은 시간 조건부 공간 블록들로 분해하고, 이 공간 블록들을 동시에 디코딩하는 생성적 정식화이다. 이는 토큰 수준 및 궤적 수준 의존성을 모두 제거하여 순차 디코딩 깊이를 튜브 길이와 무관하게 고정된 1+1 라운드로 줄인다. 병렬 공간 생성을 위해, 우리는 공유된 비디오-쿼리 맥락에 대한 접근을 유지하면서 박스 간 의존성을 제거하는 분리 블록 어텐션(Decoupled Block Attention)과, 시간 경계 및 공간 기하에 대한 위치 특정 인식 정책 최적화(localization-aware policy optimization)를 도입한다. VidSTG에서 PTD는 표준 자기회귀 디코딩 대비 튜브 완성 지연 시간(Tube Completion Latency)을 79배 줄이고 공간 디코딩 처리량을 92배 증가시키며, 그라운딩 정확도도 향상시킨다. 컴팩트한 4B 백본으로 우리 모델은 VidSTG와 HC-STVG에서 우수한 성능을 보이고, 시간적 그라운딩, 그라운딩 비디오 QA, 지시적 비디오 객체 추적에 제로샷으로 일반화한다. 우리의 결과는 병렬 튜브 생성이 비디오에서 자기회귀적 위치 특정의 효율적이고 효과적인 대안임을 보여준다.
English
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed 1 + 1 rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.