ChatPaper.aiChatPaper

AnyGroundBench: Een domeinspecifieke benchmark voor videogronding in visie-taalmodelen

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

July 2, 2026
Auteurs: Rintaro Otsubo, Ryo Fujii, Reina Ishikawa, Taiki Kanaya, Kanta Sawafuji, Hiroki Kajita, Shigeki Sakai, Hideo Saito, Ryo Hachiuma
cs.AI

Samenvatting

Visie-taalmodellen (VLMs) hebben aanzienlijke belofte getoond in ruimtelijk-temporele video-grounding (STVG). Echter, huidige evaluatieprotocollen zijn grotendeels beperkt tot zero-shot-beoordelingen op algemene, dagelijkse benchmarks. Dit creëert een kritieke kloof met praktijktoepassingen in gespecialiseerde domeinen, waar modellen onvermijdelijk te maken krijgen met zeldzame visuele concepten en complexe ruimtelijk-temporele dynamieken. Aangezien uitputtende pre-training over oneindige datadistributies onhaalbaar is, is het vermogen om zich aan te passen aan nieuwe domeinen essentieel. Om deze kloof te overbruggen, introduceren we AnyGroundBench, een domeinaanpassingsbenchmark die is ontworpen om het STVG-evaluatieparadigma te verschuiven van statische zero-shot-testen naar rigoureuze domeinaanpassing. AnyGroundBench richt zich op vijf gespecialiseerde domeinen (dier, industrie, sport, chirurgie en openbare veiligheid) en koppelt nieuw opgenomen video's, zoals door experts geannoteerd muizengedrag, aan bestaande datasets, en verenigt ze door middel van dichte, hoogwaardige ruimtelijk-temporele annotaties. Essentieel is dat de benchmark speciale trainingssubsets biedt om de domeinaanpasbaarheid systematisch te meten. We evalueren uitgebreid 15 state-of-the-art VLMs, waarbij we hun zero-shot-generalisatie en in-context-leren (ICL)-capaciteiten beoordelen onder praktische rekenkundige beperkingen. Uiteindelijk tonen onze bevindingen aan dat huidige modellen falen in zowel zero-shot- als ICL-gebaseerde aanpassing wanneer ze geconfronteerd worden met gespecialiseerde domeinen, wat kritieke tekortkomingen in het ruimtelijk-temporele redeneren blootlegt die toekomstig onderzoek moet aanpakken.
English
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation protocols are largely confined to zero-shot assessments on general, daily-life benchmarks. This creates a critical disconnect from real-world applications in specialized fields, where models inevitably encounter rare visual concepts and complex spatio-temporal dynamics. Since exhaustive pre-training across infinite data distributions is infeasible, the ability to adapt to novel domains is essential. To bridge this gap, we introduce AnyGroundBench, a domain-adaptation benchmark designed to shift the STVG evaluation paradigm from static zero-shot testing to rigorous domain adaptation. Targeting five specialized domains (animal, industry, sports, surgery, and public security), AnyGroundBench pairs newly captured videos such as expert-annotated mouse behaviors with established datasets, unifying them through dense, high-fidelity spatio-temporal annotations. Crucially, the benchmark provides dedicated training subsets to systematically measure domain adaptability. We extensively evaluate 15 state-of-the-art VLMs, assessing their zero-shot generalization and In-Context Learning (ICL) capabilities under practical computational constraints. Ultimately, our findings reveal that current models fail in both zero-shot and ICL-based adaptation when confronted with specialized domains, exposing critical flaws in spatio-temporal reasoning that future research must address.