ChatPaper.aiChatPaper

多模态大语言模型能否理解OCT?

Can Multimodal Large Language Models Understand OCT?

July 18, 2026
作者: Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song
cs.AI

摘要

光学相干断层扫描(OCT)成像对于视网膜疾病的诊断和治疗至关重要。尽管多模态大语言模型(MLLMs)在医学图像分析中展现出巨大潜力,但现有基准测试大多将OCT理解简化为粗粒度疾病分类或孤立的视觉问答,未能充分评估从视觉感知到临床推理的完整认知过程。为弥补这一不足,我们提出了OCT-Bench——一个专注于OCT图像理解的综合性基准测试。OCT-Bench包含来自七个公开数据集、基于4,137张OCT图像构建的10,076道高质量选择题。遵循真实临床判读流程,我们建立了一个层次化能力分类体系,包含三个维度(感知、认知、推理)下的20项细粒度任务。这些任务涵盖了广泛的能力,包括成像属性、视网膜解剖结构、病灶特征、空间关系、疾病评估、治疗决策和预后管理。我们系统评估了20个代表性MLLMs,包括商业模型、开源通用模型和医学领域模型。实验结果表明,当前模型在可靠的OCT理解方面仍有较大差距。此外,无论是医学领域适配还是模型规模扩大,都未能持续提升各能力层次的表现。OCT-Bench能够对MLLMs进行全面细粒度的评估,为识别能力瓶颈、推动基于临床实践的OCT理解奠定基础。
English
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.