다중모드 대규모 언어 모델은 OCT를 이해할 수 있는가?
Can Multimodal Large Language Models Understand OCT?
July 18, 2026
저자: Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song
cs.AI
초록
광간섭단층촬영(OCT) 영상은 망막 질환의 진단과 치료에 필수적이다. 다중 모달 대규모 언어 모델(MLLM)이 의료 영상 분석에서 상당한 잠재력을 입증했지만, 기존 벤치마크는 대부분 OCT 이해를 거친 수준의 질병 분류나 고립된 시각 질의응답으로 축소하여, 시각적 지각에서 임상적 추론에 이르는 완전한 인지 과정을 충분히 평가하지 못한다. 이러한 한계를 해결하기 위해, 우리는 OCT 영상 이해에 특화된 포괄적인 벤치마크인 OCT-Bench를 소개한다. OCT-Bench는 7개의 공개 데이터셋에서 수집한 4,137개의 OCT 영상으로부터 구성된 10,076개의 고품질 객관식 질문으로 이루어져 있다. 실제 임상 판독 워크플로우를 따라, 우리는 지각, 인지, 추론의 세 가지 차원에 걸쳐 20개의 세분화된 작업으로 구성된 계층적 능력 분류 체계를 구축했다. 이러한 작업은 영상 속성, 망막 해부학, 병변 특성, 공간 관계, 질병 평가, 치료 의사 결정, 예후 관리 등 광범위한 능력을 포괄한다. 우리는 독점 모델, 오픈소스 범용 모델, 의료 분야 모델을 포함한 20개의 대표적인 MLLM을 체계적으로 평가했다. 실험 결과, 현재 모델은 신뢰할 수 있는 OCT 이해에 여전히 크게 미치지 못함을 보여준다. 또한, 의료 분야 적응이나 모델 규모 증가 모두 능력 수준 전반에 걸쳐 일관된 성능 향상을 가져오지 않았다. OCT-Bench는 MLLM에 대한 포괄적이고 세분화된 평가를 가능하게 하여, 능력 병목 지점을 식별하고 임상 기반의 OCT 이해를 발전시키는 기반을 제공한다.
English
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.