Code Summarization and Program Synthesis with Large Language Models
Name
lam-lamk-meng-eecs-2024-thesis.pdf
Description
Thesis PDF
Size
559.24 KB
Format
Adobe PDF
Checksum (MD5)
5fb97753cfcc28074cf0b54f42525a5a
Author(s)
Lam, Kelly
Advisor(s)
Cafarella, Michael
Date Issued
May 2024
Publisher
Massachusetts Institute of Technology
Abstract
Automatic source code summarization and generation are naturally complimentary operations because they bridge the gap between natural-language text and executable programs, allowing users to flow between the two modes. Even though large language models, have become increasingly popular, it is unclear how effective they are with code summarization and generation, especially as we examine longer source code segments or more complicated prompts for generation. In this thesis, we will formalize the automatic code summarization and generation problems, identify some cases where large-language models can perform poorly, propose some techniques to correct the initial bad results, and evaluate our results against appropriate baselines using suitable evaluation metrics.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link