Detecting Novel Associations in Large Data Sets
Name
Lander_Detecting novel.pdf
Size
1.62 MB
Format
Adobe PDF
Checksum (MD5)
71bee4b1ec77986cfd661cdb674025bd
Author(s) • • • • • • • •
Reshef, David N.
Reshef, Yakir
Grossman, Sharon Rachel
Finucane, Hilary Kiyo
McVean, Gilean
Turnbaugh, Peter J.
Mitzenmacher, Michael
Sabeti, Pardis C.
Lander, Eric Steven
Date Issued
December 2011
Journal
Science
Publisher
American Association for the Advancement of Science (AAAS)
Citation
Reshef, D. N., Y. A. Reshef, H. K. Finucane, S. R. Grossman, G. McVean, P. J. Turnbaugh, E. S. Lander, M. Mitzenmacher, and P. C. Sabeti. “Detecting Novel Associations in Large Data Sets.” Science 334, no. 6062 (December 15, 2011): 1518-1524.
Version
Author's final manuscript
Abstract
Identifying interesting relationships between pairs of variables in large data sets is increasingly important. Here, we present a measure of dependence for two-variable relationships: the maximal information coefficient (MIC). MIC captures a wide range of associations both functional and not, and for functional relationships provides a score that roughly equals the coefficient of determination (R[superscript 2]) of the data relative to the regression function. MIC belongs to a larger class of maximal information-based nonparametric exploration (MINE) statistics for identifying and classifying relationships. We apply MIC and MINE to data sets in global health, gene expression, major-league baseball, and the human gut microbiota and identify known and novel relationships.
MIT Department
Whitaker College of Health Sciences and Technology
Massachusetts Institute of Technology. Department of Biology
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike 3.0
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1126/science.1205438