A Pilot Study in Surveying Clinical Judgments to Evaluate Radiology Report Generation
Name
3442188.3445909.pdf
Size
1.54 MB
Format
Adobe PDF
Checksum (MD5)
42cc0af5ff32d00278d2c9dae0c7b3f9
Author(s) • • • •
Boag, William
Kane, Hassan
Rawat, Saumya
Wei, Jesse
Goehler, Alexander
Date Issued
March 1, 2021
Publisher
Association for Computing Machinery
Citation
William Boag, Hassan Kané, Saumya Rawat, Jesse Wei, and Alexander Goehler. 2021. A Pilot Study in Surveying Clinical Judgments to Evaluate Radiology Report Generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT '21). Association for Computing Machinery, New York, NY, USA, 458–465.
Version
Final published version
Abstract
The recent release of many Chest X-Ray datasets has prompted a lot of interest in radiology report generation. To date, this has been framed as an image captioning task, where the machine takes an RGB image as input and generates a 2-3 sentence summary of findings as output. The quality of these reports has been canonically measured using metrics from the NLP community for language generation such as Machine Translation and Summarization. However, the evaluation metrics (e.g. BLEU, CIDEr) are inappropriate for the medical domain, where clinical correctness is critical. To address this, our team brought together machine learning experts with radiologists for a pilot study in co-designing a better metric for evaluating the quality of an algorithmically-generated radiology report. The interdisciplinary collaborative process involved multiple interviews, outreach, and preliminary annotation to design a larger scale study - which is now underway - to build a more meaningful evaluation tool.
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3442188.3445909