Permutation Tests for Classification

Mukherjee, Sayan; Golland, Polina; Panchenko, Dmitry

Author(s)

Mukherjee, Sayan; Golland, Polina; Panchenko, Dmitry

DownloadMIT-CSAIL-TR-2003-016.ps (22340Kb)

Additional downloads

Metadata

Show full item record

Abstract

We introduce and explore an approach to estimating statisticalsignificance of classification accuracy, which is particularly usefulin scientific applications of machine learning where highdimensionality of the data and the small number of training examplesrender most standard convergence bounds too loose to yield ameaningful guarantee of the generalization ability of theclassifier. Instead, we estimate statistical significance of theobserved classification accuracy, or the likelihood of observing suchaccuracy by chance due to spurious correlations of thehigh-dimensional data patterns with the class labels in the giventraining set. We adopt permutation testing, a non-parametric techniquepreviously developed in classical statistics for hypothesis testing inthe generative setting (i.e., comparing two probabilitydistributions). We demonstrate the method on real examples fromneuroimaging studies and DNA microarray analysis and suggest atheoretical analysis of the procedure that relates the asymptoticbehavior of the test to the existing convergence bounds.

Date issued

2003-08-28

URI

http://hdl.handle.net/1721.1/30408

Other identifiers

MIT-CSAIL-TR-2003-016

AIM-2003-019

Series/Report no.

Massachusetts Institute of Technology Computer Science and Artificial Intelligence Laboratory

Keywords

AI, Classification, Permutation testing, Statistical significance.

Collections

CSAIL Technical Reports (July 1, 2003 - present)