ChatPaper.aiChatPaper

ClinFusion: 包括的医療理解のためのビジョン中心マルチモーダルLLMシステム

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

July 27, 2026
著者: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
cs.AI

要旨

マルチモーダル大規模言語モデル(MLLM)は臨床実践に革命をもたらす計り知れない可能性を秘めているが、医療領域への展開は根本的に視覚中心の課題である。すなわち、モデルは異種の2Dおよび3D医用画像から知識を吸収しなければならず、評価プロトコルは放射線科医の臨床実践と整合し、正確かつきめ細かく事実性に基づく評価を提供しなければならない。本稿では、これらの制限を体系的に解決する、総合的な医用理解のために設計された視覚中心のMLLMであるClinFusionを紹介する。我々は、カスケード空間認識局所性融合演算子を特徴とする構成可能でカスケード型の視覚エンコーダアーキテクチャを提案する。この演算子は、多様な2Dおよびネイティブ3D医用画像理解を融合エンコーダ内で統合する。さらに、指示追従評価のためのMedIF-Benchと、臨床的に整合し事実性に基づくレポート生成評価のための関心領域に基づく手法を含む、視覚に基づく評価フレームワークを導入する。ClinFusionは、視覚的質問応答、レポート生成、指示追従にわたる2Dおよび3Dマルチモーダル医療ベンチマークの包括的なスイート、ならびにテキスト医療タスクにおいて新たな最先端を達成し、24ベンチマーク中20で主要なオープンソース医療MLLM(例:Hulu-Med、Lingshu)を上回り、16ベンチマーク中13でGPT-5.2やGemini-3-Flashなどの強力なプロプライエタリモデルよりも優れたマルチモーダル能力を示すことを実証する。さらに、検索拡張型およびツール支援型の臨床ワークフロー向けに、エージェント的ツール使用によって拡張可能である。認定放射線科医による盲検評価により、ClinFusionが最も高評価のレポートを生成することが確認され、我々のRoIに基づく指標が、検討されたすべての自動評価指標の中で専門家の判断との最も強い相関を達成することが検証された。
English
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.