Growing a list
Name
Rudin_Growing a list.pdf
Size
785.28 KB
Format
Adobe PDF
Checksum (MD5)
d7ef0cad68681c861d2fe1df5e50a358
Author(s) • •
Letham, Benjamin
Rudin, Cynthia
Heller, Katherine A.
Date Issued
July 2013
Journal
Data Mining and Knowledge Discovery
Publisher
Springer-Verlag
Citation
Letham, Benjamin, Cynthia Rudin, and Katherine A. Heller. “Growing a List.” Data Mining and Knowledge Discovery 27, no. 3 (November 2013): 372–95.
Version
Author's final manuscript
Abstract
It is easy to find expert knowledge on the Internet on almost any topic, but obtaining a complete overview of a given topic is not always easy: information can be scattered across many sources and must be aggregated to be useful. We introduce a method for intelligently growing a list of relevant items, starting from a small seed of examples. Our algorithm takes advantage of the wisdom of the crowd, in the sense that there are many experts who post lists of things on the Internet. We use a collection of simple machine learning components to find these experts and aggregate their lists to produce a single complete and meaningful list. We use experiments with gold standards and open-ended experiments without gold standards to show that our method significantly outperforms the state of the art. Our method uses the ranking algorithm Bayesian Sets even when its underlying independence assumption is violated, and we provide a theoretical generalization bound to motivate its use.
MIT Department
Massachusetts Institute of Technology. Operations Research Center
Sloan School of Management
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1007/s10618-013-0329-7