Student Research Abstract: Evaluating Dialogue Summarization Using LLMs
Name
3672608.3707999.pdf
Size
1.35 MB
Format
Adobe PDF
Checksum (MD5)
ca9f45ad61604779303c2a5500f903b3
Author(s)
Wang, Alison
Date Issued
May 14, 2025
Publisher
ACM|The 40th ACM/SIGAPP Symposium on Applied Computing
Citation
Wang, Alison. 2025. "Student Research Abstract: Evaluating Dialogue Summarization Using LLMs."
Version
Final published version
Abstract
With the surge in audio data available today, there is a growing need for effective dialogue summarization. This study conducts two experiments using two LLMs, BART and Mistral, to assess dialogue summarization. The first experiment evaluates model performance, while the second examines the impact of upstream errors from Automatic Speech Recognition (ASR) and Machine Translation (MT) on summarization performance. Results indicate that SummaC, a commonly used evaluation metric, is unreliable for dialogue summarization. Additionally, Mistral's summarization performance is more sensitive to upstream errors than BART's.
Description
SAC ’25, March 31-April 4, 2025, Catania, Italy
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Article is made available in accordance with the publisher's policy and may be subject to US copyright law. Please refer to the publisher's site for terms of use.
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3672608.3707999