ChatPaper.aiChatPaper

VIABench:一项从盲人个体收集的面向视觉障碍辅助的综合性视频基准测试集

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

July 16, 2026
作者: Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
cs.AI

摘要

视觉障碍人士(VIIs)因视觉信息获取受限,在日常活动中面临重大挑战。尽管多模态大语言模型(MLLMs)在通用视觉与语言任务上取得了显著成果,但其在真实场景盲人辅助中的实际应用价值仍鲜少被探索。为填补这一空白,我们提出VIABench——一个专门用于评估MLLMs在视觉障碍辅助场景中表现的综合视频基准测试集,其数据源自视觉障碍人士自身拍摄或共享的第一人称视频。VIABench定义了三项核心任务,分别针对视觉辅助中的不同需求:**主动提醒**(评估模型在实时解读视频内容的同时,主动预判并口头描述即将发生的导航关键事件的能力)、**视觉问答**(评价模型回答用户关于视频环境或物体提问的能力)、**视觉引导交互**(测试模型在用户与环境间实现意图交互时的上下文感知推理能力)。为确保评估的严谨性和公平性,我们设计了支持在线(实时)与离线两种模式的标准化基准测试流程。实验表明,当前MLLMs仍难以全面支持视觉障碍人士,尤其在需要精准预判和实时响应的主动提醒任务中表现不足。我们期望VIABench能够推动未来研究聚焦于定制化MLLMs的开发,从而切实改善视觉障碍人士的导航与交互体验。代码与数据将在https://github.com/MCG-NJU/VIABench公开。
English
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.