ChatPaper.aiChatPaper

MedPMC:用于扩展基础模型的高保真医学多模态数据的系统框架

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

July 8, 2026
作者: Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
cs.AI

摘要

医学本质上是多模态的,要求临床医生整合来自不同数据流的信息。然而,多模态基础模型的开发因大规模高质量临床数据获取受限而受到制约。尽管PubMed Central(PMC)提供了专家撰写的图像-文本数据作为补充来源,但现有基于PMC的数据集在保真度、可重复性和临床验证方面仍存在局限。我们提出MedPMC——一个自动化、可持续更新的框架,可将许可开放的文献转化为医学多模态模型的高保真基础设施。应用于610万篇PMC文章后,MedPMC整理出1100万对医学图像-文本数据。组件评估显示,在初始筛选(F1=93.2)、多面板图形检测(F1=96.5)、图形分离(mAP=89.8)、标题分离与对齐(F1=81.4;ROUGE-L=85.3)以及医学图形分类(F1=96.5)方面均表现优异。由五位标注员(其中三位具有医学背景)进行的人工审核发现,MedPMC图像中95.3%具有医学相关性,而此前基于PMC的数据集该比例仅为19.7%。在涵盖11个专科的26项基准测试中,使用MedPMC训练的CLIP风格模型在零样本AUC上平均比最强的架构匹配生物医学CLIP基线高出7.1个百分点,尽管其使用的图像-文本对数量不足后者一半。作为多模态大语言模型的视觉编码器,该模型在两个基准测试中将医学视觉问答性能分别提升了1.9和16.9个百分点。在耶鲁纽黑文健康系统的10524张皮肤病学照片中,该模型将形态学到图像检索的Recall@5提升了11.7个百分点。这些发现表明,高保真文献整理能够强化医学多模态基础模型在基准测试和临床场景中的表现。我们公开发布该框架、语料库、基准测试及预训练模型。
English
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.