MedPMC:一種擴展基礎模型高保真醫學多模態資料的系統化框架
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
July 8, 2026
作者: Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
cs.AI
摘要
醫學本質上即為多模態,臨床醫師需綜合來自不同數據流的信息。然而,多模態基礎模型的開發受到大規模、高質量臨床數據取得限制。雖然PubMed Central(PMC)提供了專家撰寫的圖文數據作為補充來源,但現有PMC衍生資源在保真度、可重複性及臨床驗證方面仍有限。我們提出MedPMC,這是一個自動化、可持續更新的框架,可將許可寬鬆的文獻轉化為醫學多模態模型的高保真基礎設施。應用於610萬篇PMC文章後,MedPMC整理出1100萬對醫學圖文數據。元件評估顯示,初步篩選(F1=93.2)、多面板圖檢測(F1=96.5)、圖像分離(mAP=89.8)、標題分離與對齊(F1=81.4;ROUGE-L=85.3)及醫學圖像分類(F1=96.5)均表現優異。經五位註釋員(其中三位具醫學背景)人工審查,MedPMC中95.3%的圖像具醫學相關性,而先前PMC衍生數據集僅為19.7%。在涵蓋11個專科的26項基準測試中,以MedPMC訓練的CLIP風格模型平均零樣本AUC較最強的架構匹配生物醫學CLIP基線提升7.1個百分點,即便其使用的圖文對數量不到後者一半。作為多模態大型語言模型的視覺編碼器,該模型在兩項醫學視覺問答基準測試中分別提升1.9及16.9個百分點。在耶魯紐黑文健康系統(Yale New Haven Health System)的10,524張皮膚科照片中,其將型態至圖像檢索的Recall@5提升11.7個百分點。這些結果顯示,高保真文獻整理能強化醫學多模態基礎模型在基準測試與臨床場景中的表現。我們公開釋出此框架、語料庫、基準測試及預訓練模型。
English
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.