ChatPaper.aiChatPaper

VIABench:一個由盲人收集的綜合性影片基準,用於視覺障礙輔助

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

July 16, 2026
作者: Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
cs.AI

摘要

視覺障礙者因獲取視覺資訊受限,在日常生活中面臨諸多重大挑戰。儘管多模態大型語言模型(MLLMs)在通用視覺與語言任務上已取得顯著成果,但其在真實盲人輔助場景中的實用性仍鮮少被深入探討。為填補此一缺口,我們提出VIABench,一個專門設計用以評估MLLMs於視覺障礙輔助情境表現的綜合性影片基準測試,該測試採用視覺障礙者自身錄製或分享的第一人稱影片資料。VIABench定義了三項核心任務,分別對應視覺輔助中的不同需求。**主動提醒**:評估模型在解讀即時影片內容的同時,能否主動預測並以口語描述即將發生的導航關鍵事件;**視覺問答**:評估模型回答使用者針對影片中環境或物體所提出問題的能力;**視覺引導互動**:測試模型具備情境感知推理能力,以完成使用者與環境間的有意識互動。為確保評估的穩健性與公平性,我們提出一套嚴謹的基準測試流程,同時支援線上(即時)與離線兩種設定。實驗結果顯示,現有MLLMs仍難以為視覺障礙者提供全面支援,尤其在需要準確預測與即時回應的主動提醒任務中表現不足。我們期望VIABench能推動未來研究朝向開發專為真實輔助情境客製化的MLLMs邁進,最終提升視覺障礙者的導航與互動體驗。程式碼與資料將於 https://github.com/MCG-NJU/VIABench 公開釋出。
English
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.