ClinFusion: 以视觉为中心的多模态大语言模型系统,用于整体医学理解
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
July 27, 2026
作者: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
cs.AI
摘要
多模态大语言模型(MLLMs)在革新临床实践中蕴含巨大潜力,但在医学领域的部署本质上是以视觉为核心的挑战:模型必须从异质的2D和3D医学影像中吸收知识,评估方案需与放射科医生的临床实践对齐,并提供准确、细粒度且基于事实的评估。本文提出ClinFusion——一种以视觉为中心、面向整体医学理解的多模态大语言模型,系统性地应对了上述限制。我们提出一种组合式级联视觉编码器架构,其中包含级联空间感知局部融合算子,可在融合编码器中统一处理多样化的2D与原生3D医学图像理解。我们还提出一种基于视觉的评估框架,包括用于指令遵循评估的MedIF-Bench基准,以及基于感兴趣区域的方法,用于临床对齐且基于事实的报告生成评估。研究表明,ClinFusion在涵盖视觉问答、报告生成、指令遵循以及医学文本任务的全面2D与3D多模态医学基准测试中均达到最优性能——在24项基准中的20项上超越领先的开源医学多模态大模型(如Hulu-Med、Lingshu),并在16项基准中的13项上展现出优于GPT-5.2和Gemini-3-Flash等强大商用模型的多模态能力。该模型还可通过智能体工具调用进一步增强,实现检索增强和工具辅助的临床工作流。由持有执业认证的放射科医生开展的双盲评估证实,ClinFusion可生成评级最高的报告,同时验证了基于感兴趣区域的评估指标在所有自动评估指标中实现了与专家判断的最强相关性。
English
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.