ChatPaper.aiChatPaper

ClinFusion: 총체적 의학 이해를 위한 시각 중심 멀티모달 LLM 시스템

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

July 27, 2026
저자: Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu, Lirong Wu, Sicheng Yang, Jinwang Wang, Pengju Wang, Zhitao Zeng, Yizeng Han, Yan Xing, Shengxuan Luo, Tao Feng, Qing Xie, Weigen Yao, Yi Yang, Zuozhu Liu, Jiasheng Tang, Shaocheng Wang, Jitao Wang, Jiahong Dong, Weihua Chen, Feng Xu, Fan Wang
cs.AI

초록

다중 모달 거대 언어 모델(MLLM)은 임상 진료에 혁명을 일으킬 엄청난 잠재력을 지니고 있지만, 이를 의료 영역에 배포하는 것은 근본적으로 시각 중심의 도전 과제이다. 즉, 모델은 이질적인 2D 및 3D 의료 영상으로부터 지식을 흡수해야 하며, 평가 프로토콜은 방사선과 의사의 임상 진료와 일치하고 정확하고 세밀하며 사실성 기반의 평가를 제공해야 한다. 본 논문에서는 이러한 한계점을 체계적으로 해결하기 위해 설계된 시각 중심의 MLLM인 ClinFusion을 소개한다. 우리는 구성적이고 계층적인 비전 인코더 아키텍처를 제안하며, 이는 다양한 2D 및 네이티브 3D 의료 영상 이해를 통합 인코더 내에서 하나로 결합하는 Cascade Spatial-Aware Locality Fusion 연산자를 특징으로 한다. 또한, 지시 수행 평가를 위한 MedIF-Bench와 임상적으로 정렬되고 사실성 기반의 보고서 생성 평가를 위한 관심 영역 기반 방법을 포함하는, 시각에 기반한 평가 프레임워크를 도입한다. ClinFusion은 2D 및 3D 다중 모달 의료 벤치마크(시각 질의응답, 보고서 생성, 지시 수행)와 텍스트 기반 의료 작업을 아우르는 포괄적인 평가 세트에서 최신 기술 수준을 확립하여, 24개 벤치마크 중 20개에서 선도적인 오픈소스 의료 MLLM(예: Hulu-Med, Lingshu)을 능가하고, 16개 벤치마크 중 13개에서 GPT-5.2, Gemini-3-Flash와 같은 강력한 독점 모델보다 뛰어난 다중 모달 성능을 보여준다. 또한, 에이전트 기반 도구 사용을 통해 검색 증강 및 도구 지원 임상 워크플로우로 확장될 수 있다. 전문의 자격을 갖춘 방사선과 의사에 의한 맹검 평가는 ClinFusion이 최고 순위의 보고서를 생성함을 확인하고, 제안된 RoI 기반 평가 지표가 검토된 모든 자동 평가 지표 중에서 전문가 판단과 가장 강한 상관관계를 달성함을 입증한다.
English
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understanding within a fused encoder. We further introduce a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest-grounded method for clinically aligned and factualness-driven report generation evaluation. We show that ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks---spanning visual question answering, report generation, and instruction following---as well as textual medical tasks, outperforming leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrating multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks, and can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirms that ClinFusion produces the highest-ranked reports, and validates our RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.