ChatPaper.aiChatPaper

VIABench: 시각 장애인으로부터 수집된 시각 장애 보조를 위한 종합 비디오 벤치마크

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

July 16, 2026
저자: Yunfeng Liu, Yuandong Yang, Jiarui Han, Zhenpeng Huang, Yuqing Tang, Xiangyu Zeng, Gangshan Wu, Limin Wang
cs.AI

초록

시각 장애인(Visually Impaired Individuals, VIIs)은 시각 정보에 대한 접근이 제한되어 일상생활에서 심각한 어려움을 겪는다. 다중 모달 대규모 언어 모델(Multimodal Large Language Models, MLLMs)이 일반적인 시각 및 언어 과제에서 인상적인 성과를 거두었지만, 실제 시각 장애인 지원 환경에서의 실용적 유용성은 여전히 크게 탐구되지 않은 상태이다. 이러한 격차를 해소하기 위해, 본 연구는 시각 장애인이 직접 촬영하거나 공유한 1인칭 시점 영상을 활용하여 시각 장애인 지원(Visually Impaired Assistance) 시나리오에서 MLLM을 평가하도록 특별히 설계된 포괄적 비디오 벤치마크인 VIABench를 제안한다. VIABench는 시각 지원에서 각기 다른 요구 사항을 목표로 하는 세 가지 핵심 과제를 정의한다. **사전 알림(Proactive Reminder)**: 모델이 진행 중인 비디오 콘텐츠를 해석하고, 다가오는 내비게이션에 중요한 이벤트를 사전에 예측하여 구두로 설명하는 능력을 평가한다. **시각 질의 응답(Visual Question Answering, VQA)**: 비디오 내 환경이나 객체에 대한 사용자의 질문에 답변하는 모델의 능력을 평가한다. **시각 유도 상호작용(Vision-Guided Interaction)**: 사용자와 환경 간의 의도적 상호작용을 수행하기 위해 맥락 인식 추론을 테스트한다. 강력하고 공정한 평가를 보장하기 위해, 온라인(실시간) 및 오프라인 설정을 모두 지원하는 엄격한 벤치마킹 파이프라인을 제안한다. 실험 결과, 현재의 MLLM은 특히 정확한 예측과 실시간 응답성을 요구하는 사전 알림(Proactive Reminder) 과제에서 시각 장애인에게 포괄적 지원을 제공하는 데 여전히 어려움을 겪고 있음을 보여준다. VIABench가 실제 지원을 위한 맞춤형 MLLM 개발을 위한 향후 연구를 촉진하여, 궁극적으로 시각 장애인의 내비게이션 및 상호작용 경험을 개선할 수 있기를 기대한다. 코드와 데이터는 https://github.com/MCG-NJU/VIABench에서 공개될 예정이다.
English
Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.