Approximating the crowd
Name
10618_2014_354_ReferencePDF.pdf
Size
396.89 KB
Format
Adobe PDF
Checksum (MD5)
7859956b930ce745bc224b0c4c2df846
Author(s) • • •
Ertekin, Şeyda
Rudin, Cynthia
Hirsh, Haym
Ertekin, Seyda
Date Issued
June 2014
Journal
Data Mining and Knowledge Discovery
Publisher
Springer US
Citation
Ertekin, Şeyda, Cynthia Rudin, and Haym Hirsh. “Approximating the Crowd.” Data Min Knowl Disc 28, no. 5–6 (June 14, 2014): 1189–1221.
Version
Author's final manuscript
Abstract
The problem of “approximating the crowd” is that of estimating the crowd’s majority opinion by querying only a subset of it. Algorithms that approximate the crowd can intelligently stretch a limited budget for a crowdsourcing task. We present an algorithm, “CrowdSense,” that works in an online fashion where items come one at a time. CrowdSense dynamically samples subsets of the crowd based on an exploration/exploitation criterion. The algorithm produces a weighted combination of the subset’s votes that approximates the crowd’s opinion. We then introduce two variations of CrowdSense that make various distributional approximations to handle distinct crowd characteristics. In particular, the first algorithm makes a statistical independence approximation of the labelers for large crowds, whereas the second algorithm finds a lower bound on how often the current subcrowd agrees with the crowd’s majority vote. Our experiments on CrowdSense and several baselines demonstrate that we can reliably approximate the entire crowd’s vote by collecting opinions from a representative subset of the crowd.
MIT Department
Massachusetts Institute of Technology. Center for Collective Intelligence
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Sloan School of Management
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1007/s10618-014-0354-1