Self-Supervised ECG Learning for Multimodal Clinical Tasks
Name
chen-peili-meng-eecs-2025-thesis.pdf
Description
Thesis PDF
Size
2.14 MB
Format
Adobe PDF
Checksum (MD5)
4e7a451ef8540d97efaf71d2256ab73c
Author(s)
Chen, Peilin
Advisor(s)
Liang, Paul
Date Issued
May 2025
Publisher
Massachusetts Institute of Technology
Abstract
We present a multimodal clinical AI framework that integrates time series, images, and text to support robust diagnostic reasoning across diverse input combinations. We first introduce ECG-JEPA, a self-supervised encoder pretrained on multiple ECG datasets to learn generalizable time series representations. This unimodal pretraining improves ECG classification, achieving a 23-point AUC gain on the underrepresented Ga dataset. We then align and fuse these ECG embeddings with chest X-rays and EHR text using a vision–language model backbone, enabling end-to-end multimodal inference. Our results show that incorporating ECG signals meaningfully improves diagnostic performance, highlighting the value of multitask time series pretraining and modular fusion for clinical AI.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link