ChatPaper.aiChatPaper

マルチモーダル大規模言語モデルはOCTを理解できるのか?

Can Multimodal Large Language Models Understand OCT?

July 18, 2026
著者: Baochen Fu, Wenzhi Deng, Baihao Jin, Yang Li, Zihan Nie, Kailin Jiang, Yuntao Du, Weiye Song
cs.AI

要旨

光干渉断層撮影(OCT)イメージングは、網膜疾患の診断と治療に不可欠である。マルチモーダル大規模言語モデル(MLLM)は医用画像解析において大きな可能性を示しているものの、既存のベンチマークはOCT理解を粗い粒度の疾患分類や孤立した視覚的質問応答に還元しており、視覚的知覚から臨床推論に至る完全な認知プロセスを十分に評価できていない。この制約に対処するため、我々はOCT画像理解に特化した包括的ベンチマークであるOCT-Benchを導入する。OCT-Benchは、7つの公開データセットから収集した4,137枚のOCT画像に基づいて構築された10,076問の高品質な多肢選択問題で構成される。実際の臨床解釈ワークフローに従い、知覚、認知、推論の3次元にわたる20の細粒度タスクからなる階層的能力分類体系を構築する。これらのタスクは、画像属性、網膜解剖、病変特性、空間関係、疾患評価、治療判断、予後管理など、幅広い能力をカバーする。我々は、プロプライエタリモデル、オープンソース汎用モデル、医療ドメインモデルを含む20の代表的なMLLMを体系的に評価する。実験結果は、現在のモデルが信頼できるOCT理解にはほど遠いことを示している。さらに、医療ドメインへの適応もモデルスケールの増大も、能力レベル全体で一貫した性能向上をもたらさない。OCT-BenchはMLLMの包括的かつ細粒度な評価を可能にし、能力のボトルネックを特定し、臨床に根ざしたOCT理解を促進する基盤を提供する。
English
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.