ChatPaper.aiChatPaper

CardioState-JEPA:面向延迟感知的跨模态共享心脏表征学习

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

August 13, 2026
作者: Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed
cs.AI

摘要

心电图(ECG)、光电容积脉搏波描记(PPG)和心音描记(PCG)为同一心动周期提供了互补的视角,然而现有心脏基础模型均针对单一传感模态训练,未能利用不同传感器之间共享的生理信息。我们提出CardioState-JEPA,一种心脏基础模型,基于生理感知的联合嵌入预测架构,在ECG、PPG和PCG三种信号上联合学习单一共享表示。该模型将异构波形映射到公共词元空间,使用单一共享Transformer编码器进行处理,并通过预测掩码的心脏潜状态进行学习,将预训练目标置于共享生理信息而非传感器特定的波形外观之上。为处理电学、力学和血流动力学事件之间的时间偏移,跨模态预测采用学习得到的延迟对齐器,在相应的心脏时相上对齐信号。由于同步多传感器记录数据稀缺,CardioState-JEPA首先从丰富的单模态数据中学习模态内结构,然后利用配对数据在潜在心脏时相上对齐各模态。作为冻结编码器,我们在涵盖ECG、PPG和PCG的25项下游任务上进行了评估,与最优自监督信号基线相比,我们的编码器在PPG分类平均AUROC上提升8.2个百分点,在PCG心脏杂音检测上提升18.8个百分点,在ECG分类上提升15.5个百分点,并在若干ECG基准任务上达到或超越了使用特权临床文本或监督标签训练的心脏模型。这些结果证明,异质心脏信号可以相互监督训练出一个统一的心脏生理基础模型。
English
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.