VIABench: 視覚障害者から収集された視覚障害支援のための包括的なビデオベンチマーク
VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
July 16, 2026
著者: Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
cs.AI
要旨
視覚障碍者(VIIs)は、視覚情報へのアクセスが限られていることから、日常生活において多くの困難に直面している。マルチモーダル大規模言語モデル(MLLMs)は、一般的な視覚・言語タスクにおいて目覚ましい成果を挙げているものの、実際の視覚障碍者支援における実用的な有用性はまだ十分に探求されていない。このギャップを埋めるため、我々はVIABenchを提案する。これは、視覚障碍者自身が記録または共有した一人称視点の映像を用いて、視覚障碍者支援シナリオにおけるMLLMsを評価するために特別に設計された包括的なビデオベンチマークである。VIABenchは、視覚支援におけるそれぞれ異なる要件を対象とした3つの中核タスクを定義している。プロアクティブリマインダー:モデルが進行中のビデオコンテンツを解釈し、同時に将来のナビゲーション上重要なイベントを予測して口頭で説明する能力を評価する。視覚質問応答(VQA):モデルがビデオ内の環境や物体に関するユーザーの質問に答える能力を評価する。視覚誘導インタラクション:ユーザーと環境の間の意図的なインタラクションを達成するための文脈認識推論をテストする。堅牢で公平な評価を確保するため、オンライン(リアルタイム)およびオフラインの両方の設定をサポートする厳格なベンチマークパイプラインを提案する。実験の結果、現在のMLLMsは、特に正確な予測とリアルタイム応答性が求められるプロアクティブリマインダータスクにおいて、視覚障碍者への包括的な支援を提供するには依然として困難があることが示された。VIABenchが、実世界の支援向けにカスタマイズされたMLLMsの開発に向けた今後の研究を促進し、最終的に視覚障碍者のナビゲーションとインタラクション体験を向上させることを期待する。コードとデータは https://github.com/MCG-NJU/VIABench で公開予定である。
English
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.