MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models
Name
3627673.3679962.pdf
Size
1.12 MB
Format
Adobe PDF
Checksum (MD5)
be04af83c75a460ab3b321a62fe8d0b4
Author(s) • • • • • • • • •
Thao, Phan Nguyen Minh
Dao, Cong-Tinh
Wu, Chenwei
Wang, Jian-Zhe
Liu, Shun
Ding, Jun-En
Restrepo, David
Liu, Feng
Hung, Fang-Ming
Peng, Wen-Chih
Date Issued
October 21, 2024
Publisher
ACM|Proceedings of the 33rd ACM International Conference on Information and Knowledge Management
Citation
Thao, Phan Nguyen Minh, Dao, Cong-Tinh, Wu, Chenwei, Wang, Jian-Zhe, Liu, Shun et al. 2024. "MEDFuse: Multimodal EHR Data Fusion with Masked Lab-Test Modeling and Large Language Models."
Version
Final published version
Abstract
Electronic health records (EHRs) are multimodal by nature, consisting of structured tabular features like lab tests and unstructured clinical notes. In real-life clinical practice, doctors use complementary multimodal EHR data sources to get a clearer picture of patients' health and support clinical decision-making. However, most EHR predictive models do not reflect these procedures, as they either focus on a single modality or overlook the inter-modality interactions/redundancy. In this work, we propose MEDFuse, a Multimodal EHR Data Fusion framework that incorporates masked lab-test modeling and large language models (LLMs) to effectively integrate structured and unstructured medical data. MEDFuse leverages multimodal embeddings extracted from two sources: LLMs fine-tuned on free clinical text and masked tabular transformers trained on structured lab test results. We design a disentangled transformer module, optimized by a mutual information loss to 1) decouple modality-specific and modality-shared information and 2) extract useful joint representation from the noise and redundancy present in clinical notes. Through comprehensive validation on the public MIMIC-III dataset and the in-house FEMH dataset, MEDFuse demonstrates great potential in advancing clinical predictions, achieving over 90% F1 score in the 10-disease multi-label classification task.
Description
CIKM ’24, October 21–25, 2024, Boise, ID, USA
MIT Department
Massachusetts Institute of Technology. Institute for Medical Engineering & Science
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3627673.3679962