Generating Component-based Supervised Learning Programs From Crowdsourced Examples
Name
MIT-CSAIL-TR-2017-015.pdf
Size
1.6 MB
Format
Adobe PDF
Checksum (MD5)
42e40a68542c585ce0d86ca5fcd9f110
Download all files submitted through automated deposit
Author(s) •
Cambronero, Jose
Rinard, Martin
Advisor(s)
Martin Rinard
Date Issued
December 21, 2017
Series/Report no.
MIT-CSAIL-TR-2017-015
Abstract
We present CrowdLearn, a new system that processes an existing corpus of crowdsourced machine learning programs to learn how to generate effective pipelines for solving supervised machine learning problems. CrowdLearn uses a probabilistic model of program likelihood, conditioned on the current sequence of pipeline components and on the characteristics of the input data to the next component in the pipeline, to predict candidate pipelines. Our results highlight the effectiveness of this technique in leveraging existing crowdsourced programs to generate pipelines that work well on a range of supervised learning problems.
Subjects
program synthesis
automated machine learning
code mining
Terms of Use
Creative Commons Attribution 4.0 International
Persistent DSpace Link