Truthfulness in Large Language Models
Name
liu-kevliu-meng-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
1.54 MB
Format
Adobe PDF
Checksum (MD5)
70962763d1a7e02aee559f20cecd101d
Author(s)
Liu, Kevin
Advisor(s)
Andreas, Jacob
Hadfield-Menell, Dylan
Date Issued
June 2023
Publisher
Massachusetts Institute of Technology
Abstract
Large language models (LLMs) have been experiencing a rapid rise in utility, accessibility, and popularity, but there are still many areas in which they can improve. One such area for improvement is their truthfulness. We seek to improve the truthfulness of LLMs by probing their internal representations. We find that a linear probe on the last hidden layer representation is able to improve a model’s accuracy by reducing its confidence in incorrect answers. However, this probe is less effective at perturbing the model to change its behavior and driving the model towards correct answers.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link