CardioState-JEPA:具延遲感知之跨模態共享心臟表徵學習
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
August 13, 2026
作者: Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed
cs.AI
摘要
心電圖(ECG)、光體積變化描記圖(PPG)與心音圖(PCG)提供同一心臟週期的互補視角,然而現有的心臟基礎模型僅針對單一感測模態進行訓練,未能利用感測器之間共有的生理資訊。我們提出 CardioState-JEPA,這是一個心臟基礎模型,旨在跨 ECG、PPG 與 PCG 聯合學習單一共享表徵,其架構建立在具有生理感知能力的聯合嵌入預測架構之上。該模型將異質波形映射至共同的詞元空間,以單一共享 Transformer 編碼器進行處理,並透過預測遮罩的潛在心臟狀態來學習,將預訓練目標置於共享生理而非感測器特定的波形外觀上。為處理電學、力學與血流動力學事件之間的時間偏移,跨模態預測採用學習式延遲對齊器,在對應的心臟時間點匹配訊號。由於同步多感測器記錄稀少,CardioState-JEPA 首先從大量單模態資料中學習模態內部結構,再利用配對資料在潛在心臟時間中對齊各模態。在涵蓋 ECG、PPG 與 PCG 的 25 項下游任務中,以凍結編碼器進行評估,我們的編碼器在平均 PPG 分類上較最佳自監督訊號基線提升 8.2 個 AUROC 百分點,PCG 心雜音偵測提升 18.8 個 AUROC 百分點,ECG 分類提升 15.5 個 AUROC 百分點,並在多項 ECG 基準上達到或超越以特權臨床文本或監督標籤訓練的心臟模型。這些結果確立了異質心臟訊號能夠相互監督單一心臟生理基礎模型的學習。
English
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.