MedPMC: 기초 모델을 위한 고충실도 의료 멀티모달 데이터 확장 체계적 프레임워크
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
July 8, 2026
저자: Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
cs.AI
초록
의료는 본질적으로 다중 양식(multimodal)적이며, 임상의는 다양한 데이터 스트림에 걸쳐 정보를 종합해야 한다. 그러나 다중 양식 기반 모델(multimodal foundation model)의 개발은 대규모 고품질 임상 데이터에 대한 제한된 접근성으로 인해 제약을 받는다. PubMed Central(PMC)은 전문가가 작성한 이미지-텍스트 데이터의 보완적 출처를 제공하지만, 기존의 PMC 기반 리소스는 충실도(fidelity), 재현성(reproducibility), 임상 검증(clinical validation) 측면에서 여전히 제한적이다. 본 연구에서는 MedPMC를 소개한다. 이는 자동화되고 지속적으로 업데이트 가능한 프레임워크로, 허용 라이선스(permissively licensed) 문헌을 의료 다중 양식 모델을 위한 고충실도 인프라로 변환한다. 610만 개의 PMC 논문에 적용하여 MedPMC는 1100만 개의 의료 이미지-텍스트 쌍을 선별하였다. 구성 요소 평가에서 초기 스크리닝(F1 = 93.2), 다중 패널 그림 탐지(F1 = 96.5), 그림 분리(mAP = 89.8), 캡션 분리 및 정렬(F1 = 81.4; ROUGE-L = 85.3), 의료 그림 분류(F1 = 96.5)에서 우수한 성능을 보였다. 5명의 주석자(이 중 3명은 의학 훈련을 받음)의 수동 검토 결과, MedPMC 이미지의 95.3%가 의학적으로 관련 있는 것으로 나타난 반면, 이전 PMC 기반 데이터셋에서는 19.7%에 불과했다. 11개 전문 분야에 걸친 26개 벤치마크 실험에서, MedPMC로 훈련된 CLIP 스타일 모델은 절반 미만의 이미지-텍스트 쌍을 사용했음에도 불구하고, 가장 강력한 아키텍처 일치 생의학 CLIP 기준선(baseline) 대비 평균 제로샷 AUC를 7.1% 포인트 향상시켰다. 다중 양식 대규모 언어 모델(multimodal large language model)의 비전 인코더로 사용되었을 때, 두 개의 벤치마크에서 의료 시각 질의응답(medical visual question-answering) 성능을 각각 1.9% 포인트 및 16.9% 포인트 향상시켰다. Yale New Haven Health System의 10,524장 피부과 사진에서, 형태-이미지 검색 Recall@5를 11.7% 포인트 개선하였다. 이러한 결과는 고충실도 문헌 선별이 벤치마크 및 임상 환경 전반에서 의료 다중 양식 기반 모델을 강화함을 보여준다. 본 연구에서는 프레임워크, 코퍼스, 벤치마크 및 사전 훈련된 모델을 공개한다.
English
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.