Learning An Invariant Speech Representation
Name
CBMM-Memo-022.pdf
Size
1.81 MB
Format
Adobe PDF
Checksum (MD5)
20ce33f5b82fab337326ef509c6aa73e
Author(s) • • • •
Evangelopoulos, Georgios
Voinea, Stephen
Zhang, Chiyuan
Rosasco, Lorenzo
Poggio, Tomaso
Date Issued
June 15, 2014
Publisher
Center for Brains, Minds and Machines (CBMM), arXiv
Citation
arXiv:1406.3884v1
Series/Report no.
CBMM Memo Series;022
Abstract
Recognition of speech, and in particular the ability to generalize and learn from small sets of labelled examples like humans do, depends on an appropriate representation of the acoustic input. We formulate the problem of finding robust speech features for supervised learning with small sample complexity as a problem of learning representations of the signal that are maximally invariant to intraclass transformations and deformations. We propose an extension of a theory for unsupervised learning of invariant visual representations to the auditory domain and empirically evaluate its validity for voiced speech sound classification. Our version of the theory requires the memory-based, unsupervised storage of acoustic templates — such as specific phones or words — together with all the transformations of each that normally occur. A quasi-invariant representation for a speech segment can be obtained by projecting it to each template orbit, i.e., the set of transformed signals, and computing the associated one-dimensional empirical probability distributions. The computations can be performed by modules of filtering and pooling, and extended to hierarchical architectures. In this paper, we apply a single-layer, multicomponent representation for phonemes and demonstrate improved accuracy and decreased sample complexity for vowel classification compared to standard spectral, cepstral and perceptual features.
Subjects
Speech Recognition
Invariance
Machine Learning
Language
Terms of Use
Attribution-NonCommercial 3.0 United States
Persistent DSpace Link