Perturbation-invariant Speech Representation Learning by Online Clustering
Name
chang-hengjui-sm-eecs-2024-thesis.pdf
Description
Thesis PDF
Size
4.1 MB
Format
Adobe PDF
Checksum (MD5)
ab5a2f2331d153b74a5537505f4b599a
Author(s)
Chang, Heng-Jui
Advisor(s)
Glass, James R.
Date Issued
February 2024
Publisher
Massachusetts Institute of Technology
Abstract
Despite success across various tasks, self-supervised speech models face significant challenges in enhancing content-related performance with unlabeled data, requiring substantial computational resources. Meanwhile, learning from clustered discrete units has been shown to facilitate accurate phonetic representations. Thus, this thesis investigates speaker and noise-invariant speech representations. First, Speaker-invariant Clustering (Spin) is proposed to extract content representations through online clustering and speaker-invariant cross-view prediction. Second, Robust Spin (R-Spin) is devised to extend Spin to handle more distorted speech signals by leveraging acoustic pieces. Furthermore, this thesis includes a diverse set of evaluation and visualization techniques to quantitatively and qualitatively analyze the perturbation invariability of the proposed methods. This thesis offers approaches to producing perturbation-invariant speech representations and deeply investigates the characteristics of the learned representations, providing insights into these models and cultivating future extension possibilities.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link