ChatPaper.aiChatPaper

ClinFusion:一個以視覺為核心的多模態大型語言模型系統,用於全面醫療理解

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

July 27, 2026
作者: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
cs.AI

摘要

多模態大型語言模型(MLLMs)在革新臨床實務方面擁有巨大潛力,然而將其部署至醫療領域本質上是一項以視覺為核心的挑戰:模型必須從異質的二維與三維醫學影像中吸收知識,且評估協議須與放射科醫師的臨床實務一致,並提供準確、細粒度且以事實性為導向的評量。本文提出 ClinFusion,一個以視覺為核心、專為全面醫學理解所設計的多模態大型語言模型,系統性地解決上述限制。我們提出一種組合式與級聯視覺編碼器架構,其中包含級聯空間感知局部融合運算子,能在單一融合編碼器中統一處理異質的二維與原生三維醫學影像理解。我們進一步引入以視覺為基礎的評估框架,包括用於指令遵循評估的 MedIF-Bench,以及一種基於興趣區域的方法,用於臨床一致且事實性驅動的報告生成評估。我們證明,ClinFusion 在涵蓋視覺問答、報告生成與指令遵循的全面二維與三維多模態醫學基準測試套件,以及文字類醫療任務中,樹立了新的業界最佳標準——在 24 項基準中於 20 項超越領先的開源醫療多模態大型語言模型(如 Hulu-Med、Lingshu),並在 16 項基準中有 13 項展現出優於 GPT-5.2 與 Gemini-3-Flash 等強大專有模型的多模態能力;此外,它還能進一步透過代理工具使用進行增強,以實現檢索增強與工具輔助的臨床工作流程。由具專科證書的放射科醫師所進行的盲審評估證實,ClinFusion 產生的報告排名最高,同時也驗證了我們所提出的以興趣區域為基礎的指標,在所有受檢驗的自動評估指標中,與專家判斷之間具有最強的相關性。
English
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.