Inferring multimodal latent topics from electronic health records
Name
s41467-020-16378-3.pdf
Description
Published version
Size
3.9 MB
Format
Adobe PDF
Checksum (MD5)
8c9d522c9af205cb13e8d96de733b29d
Author(s) • • • • • • • • •
Li, Yue
Nair, Pratheeksha
Lu, Xing Han
Wen, Zhi
Wang, Yuening
Dehaghi, Amir Ardalan Kalantari
Miao, Yan
Liu, Weiqi
Ordog, Tamas
Biernacka, Joanna M
Date Issued
2020
Journal
Nature Communications
Publisher
Springer Science and Business Media LLC
Version
Final published version
Abstract
© 2020, The Author(s). Electronic health records (EHR) are rich heterogeneous collections of patient health information, whose broad adoption provides clinicians and researchers unprecedented opportunities for health informatics, disease-risk prediction, actionable clinical recommendations, and precision medicine. However, EHRs present several modeling challenges, including highly sparse data matrices, noisy irregular clinical notes, arbitrary biases in billing code assignment, diagnosis-driven lab tests, and heterogeneous data types. To address these challenges, we present MixEHR, a multi-view Bayesian topic model. We demonstrate MixEHR on MIMIC-III, Mayo Clinic Bipolar Disorder, and Quebec Congenital Heart Disease EHR datasets. Qualitatively, MixEHR disease topics reveal meaningful combinations of clinical features across heterogeneous data types. Quantitatively, we observe superior prediction accuracy of diagnostic codes and lab test imputations compared to the state-of-art methods. We leverage the inferred patient topic mixtures to classify target diseases and predict mortality of patients in critical conditions. In all comparison, MixEHR confers competitive performance and reveals meaningful disease-related topics.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1038/S41467-020-16378-3