ChatPaper.aiChatPaper

CardioState-JEPA: 공유 심장 표현을 위한 지연 인지 교차 모달 학습

CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

August 13, 2026
저자: Hamza Shafiq, Hung Manh Pham, Bin Zhu, Pan Zhou, Jun Hu, Aaqib Saeed
cs.AI

초록

심전도(ECG), 광용적맥파(PPG), 심음도(PCG)는 동일한 심장 주기에 대한 상보적 관점을 제공하지만, 기존 심장 기반 모델들은 단일 센서 방식으로만 학습되어 센서 간 공유되는 생리학적 특성을 활용하지 못한다. 우리는 ECG, PPG, PCG에 걸쳐 단일 공유 표현을 공동으로 학습하는 심장 기반 모델인 CardioState-JEPA를 소개한다. 이 모델은 생리학적 특성을 반영한 공동 임베딩 예측 구조를 기반으로 한다. 이 모델은 이질적인 파형을 공통 토큰 공간으로 매핑하고, 단일 공유 트랜스포머 인코더로 처리하며, 마스킹된 잠재 심장 상태를 예측하는 방식으로 학습한다. 이를 통해 사전학습 목표를 센서 특유의 파형 형태가 아닌 공유된 생리학에 둔다. 전기적, 기계적, 혈역학적 사건 사이의 시간적 오프셋을 처리하기 위해, 교차 모달 예측은 신호를 해당 심장 시점에 맞추는 학습된 지연 정렬기를 사용한다. 동기화된 다중 센서 기록은 드물기 때문에, CardioState-JEPA는 먼저 풍부한 단일 모달 데이터에서 모달 내부 구조를 학습한 다음, 쌍을 이룬 데이터를 사용하여 잠재 심장 시간에서 모달 간을 정렬한다. 우리의 인코더는 ECG, PPG, PCG를 포괄하는 25개 다운스트림 작업에서 고정 인코더로 평가되었으며, 최상의 자기 지도 신호 기준 모델 대비 평균 PPG 분류에서 8.2 AUROC 포인트, PCG 심잡음 검출에서 18.8 AUROC 포인트, ECG 분류에서 15.5 AUROC 포인트의 성능 향상을 보였다. 또한 여러 ECG 벤치마크에서 특권 임상 텍스트나 지도 레이블로 학습된 심장 모델과 동등하거나 더 나은 성능을 달성했다. 이러한 결과는 이질적인 심장 신호가 심장 생리에 대한 단일 기반 모델을 상호 지도할 수 있음을 입증한다.
English
Electrocardiography (ECG), photoplethysmography (PPG), and phonocardiography (PCG) provide complementary views of the same cardiac cycle, yet existing cardiac foundation models are trained for a single sensing modality, leaving the shared physiology across sensors unexploited. We introduce CardioState-JEPA, a cardiac foundation model to learn a single shared representation jointly across ECG, PPG, and PCG, built on a physiology-aware joint-embedding predictive architecture. The model maps heterogeneous waveforms into a common token space, processes them with a single shared Transformer encoder, and learns by predicting masked latent cardiac states, placing the pretraining target on shared physiology rather than sensor-specific waveform appearance. To handle the temporal offsets between electrical, mechanical, and hemodynamic events, cross-modal prediction uses a learned delay aligner that matches signals at the corresponding cardiac time. Because synchronized multi-sensor recordings are scarce, CardioState-JEPA first learns within-modality structure from abundant unimodal data and then uses paired data to align modalities in latent cardiac time. Evaluated as a frozen encoder across 25 downstream tasks spanning ECG, PPG, and PCG, our encoder improves average PPG classification by 8.2 AUROC points, PCG murmur detection by 18.8 AUROC points, and ECG classification by 15.5 AUROC points over the best self-supervised signal baseline and matches or exceeds cardiac models trained with privileged clinical text or supervised labels on several ECG benchmarks. These results establish that heterogeneous cardiac signals can mutually supervise a single foundation model of cardiac physiology.