MedPMC: 基盤モデル向けに高忠実度医療マルチモーダルデータをスケーリングするための体系的枠組み
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
July 8, 2026
著者: Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
cs.AI
要旨
医学は本質的にマルチモーダルであり、臨床医は多様なデータストリームを統合して情報を合成する必要がある。しかし、マルチモーダル基盤モデルの開発は、大規模で高品質な臨床データへのアクセス制限によって制約されている。PubMed Central(PMC)は専門家が執筆した画像テキストデータの補完的な情報源を提供するが、既存のPMC由来リソースは、忠実性、再現性、臨床検証の点で依然として限定的である。本稿では、許諾の得られた文献を医療用マルチモーダルモデルのための高忠実性インフラストラクチャに変換する、自動化された継続的更新可能なフレームワークMedPMCを紹介する。610万のPMC論文に適用し、MedPMCは1100万の医療用画像テキストペアをキュレーションした。構成要素の評価では、初期スクリーニング(F1 = 93.2)、マルチパネル図検出(F1 = 96.5)、図分割(mAP = 89.8)、キャプション分割と整列(F1 = 81.4;ROUGE-L = 85.3)、医療用図分類(F1 = 96.5)において高い性能を示した。5名のアノテーター(うち3名は医学トレーニングを受けた)による手動レビューでは、MedPMC画像の95.3%が医学的に関連性があると判断されたのに対し、以前のPMC由来データセットでは19.7%であった。11の専門分野にわたる26のベンチマークにおいて、MedPMCで訓練されたCLIPスタイルのモデルは、半数未満の画像テキストペアしか使用していないにもかかわらず、最も強力なアーキテクチャ一致の生物医学CLIPベースラインと比較して、平均ゼロショットAUCを7.1パーセントポイント改善した。マルチモーダル大規模言語モデルの視覚エンコーダとして使用した場合、2つのベンチマークにおいて医療用視覚質問応答をそれぞれ1.9および16.9パーセントポイント改善した。イェール・ニューヘブン・ヘルスシステムの10,524件の皮膚科写真において、形態から画像への検索のRecall@5を11.7パーセントポイント改善した。これらの知見は、高忠実性の文献キュレーションが、ベンチマークおよび臨床環境の両方において医療用マルチモーダル基盤モデルを強化することを示している。本フレームワーク、コーパス、ベンチマーク、事前訓練済みモデルを公開する。
English
Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundation models is constrained by limited access to large-scale, high-quality clinical data. Although PubMed Central (PMC) offers a complementary source of expert-authored image-text data, existing PMC-derived resources remain limited in fidelity, reproducibility, and clinical validation. We introduce MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models. Applied to 6.1 million PMC articles, MedPMC curated 11 million medical image-text pairs. Component evaluations showed strong performance for initial screening (F1 = 93.2), multi-panel figure detection (F1 = 96.5), figure separation (mAP = 89.8), caption separation and alignment (F1 = 81.4; ROUGE-L = 85.3), and medical figure classification (F1 = 96.5). Manual review by five annotators, three with medical training, found 95.3% of MedPMC images medically relevant, versus 19.7% in a prior PMC-derived dataset. Across 26 benchmarks spanning 11 specialties, a MedPMC-trained CLIP-style model improved average zero-shot AUC by 7.1 percentage points over the strongest architecture-matched biomedical CLIP baseline despite using fewer than half as many image-text pairs. As the vision encoder in a multimodal large language model, it improved medical visual question-answering by 1.9 and 16.9 percentage points across two benchmarks. In 10,524 Yale New Haven Health System dermatology photographs, it improved morphology-to-image retrieval Recall@5 by 11.7 percentage points. These findings show that high-fidelity literature curation strengthens medical multimodal foundation models across benchmark and clinical settings. We publicly release the framework, corpus, benchmarks, and pretrained models.