Prediction and analysis of degree of suicidal ideation in online content
Name
1193027006-MIT.pdf
Size
11.38 MB
Format
Adobe PDF
Checksum (MD5)
96b8e33cd2e2fbb1fe6148ba9d8aed35
Author(s)
Jones, Noah C.(Noah Corinthian)
Advisor(s)
Rosalind Picard.
Date Issued
2020
Publisher
Massachusetts Institute of Technology
Abstract
Machine learning (ML) has increasingly been used to address the growing burden of mental illness and lack of access to quality mental health care. Recently such models have been applied to online data, such as social media postings to augment mental health screening. Despite the potential of these methods, online ML classifiers still perform poorly in multi-class settings. In this thesis, we propose the usage of novel document embeddings and mental health based user embeddings for triaged suicide risk screening. Machine learning to infer suicide risk and urgency is applied to a dataset of Reddit users in which the risk and urgency labels were derived from crowdsource consensus. We show that the document embedding approach outperforms count-based baselines and a method based on word importance, where important words were identified by domain experts. We examine interpretable features and methods that help to discern and explain risk labels. Finally, we find, using a Latent Dirichlet Allocation (LDA) topic model, that users labeled at-risk for suicide post about different topics to the rest of Reddit than non-suicidal users.
Description
Thesis: S.M., Massachusetts Institute of Technology, School of Architecture and Planning, Program in Media Arts and Sciences, May, 2020
Cataloged from the official PDF of thesis.
Includes bibliographical references (pages 51-57).
Subjects
Program in Media Arts and Sciences
MIT Department
Program in Media Arts and Sciences (Massachusetts Institute of Technology)
Terms of Use
MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided.
Persistent DSpace Link