Evaluating Differences in GPT-4 Treatment by Gender in Healthcare Applications
Name
pan-eileenp-sm-eecs-2025-thesis.pdf
Description
Thesis PDF
Size
1.93 MB
Format
Adobe PDF
Checksum (MD5)
eb321259b1b2d755b3e2c7b7f7d0a0a1
Author(s)
Pan, Eileen
Advisor(s)
Ghassemi, Marzyeh
Date Issued
May 2025
Publisher
Massachusetts Institute of Technology
Abstract
LLMs already permeate medical settings, supporting patient messaging, medical scribing, and chatbots. While prior work has examined bias in medical LLMs, few studies focus on realistic use cases or analyze the source of the bias. To assess whether medical LLMs exhibit differential performance by gender, we audit their responses and investigate whether the disparities stem from implicit or explicit gender cues. We conduct a large-scale human evaluation of GPT-4 responses to medical questions, including counterfactual gender pairs for each question. Our findings reveal differential treatment based on the original patient gender. Specifically, responses for women more often recommend supportive resources, while those for men advise emergency care. Additionally, LLMs tend to downplay medical urgency for female patients and escalate it for male patients. Given rising interest in “LLM-as-a-judge” approaches, we also evaluate whether LLMs can serve as a proxy for human annotators in identifying disparities. We find that LLM-generated annotations diverge from human assessments in heterogeneous ways, particularly regarding error detection and relative urgency.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link