多模态大型语言模型能否理解OCT(光学相干断层扫描)?
Can Multimodal Large Language Models Understand OCT?
July 18, 2026
作者: Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song
cs.AI
摘要
光學相干斷層掃描(OCT)成像對於視網膜疾病的診斷與治療至關重要。儘管多模態大語言模型(MLLMs)在醫學影像分析中展現出相當大的潛力,現有基準大多將OCT理解簡化為粗粒度的疾病分類或孤立的視覺問答,使得從視覺感知到臨床推理的完整認知過程未能得到充分評估。為了解決這一局限性,我們提出OCT-Bench,這是一個專注於OCT影像理解的綜合性基準。OCT-Bench包含10,076道高品質選擇題,這些題目來自七個公開數據集中的4,137張OCT影像。遵循真實世界的臨床判讀流程,我們建立了一個分層能力分類體系,涵蓋感知、認知與推理三個維度,共20項細粒度任務。這些任務涵蓋廣泛的能力,包括成像屬性、視網膜解剖、病灶特徵、空間關係、疾病評估、治療決策及預後管理。我們系統性地評估了20個具代表性的MLLMs,涵蓋商業模型、開源通用模型以及醫學領域模型。實驗結果表明,當前模型在可靠的OCT理解方面仍遠遠不足。此外,無論是醫學領域適應還是增加模型規模,都未能持續提升各能力層次的表現。OCT-Bench能夠對MLLMs進行全面且細粒度的評估,為識別能力瓶頸及推動基於臨床的OCT理解奠定了基礎。
English
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.