| dc.contributor.author | Zogaj, Fatjon | |
| dc.contributor.author | Cambronero, José Pablo | |
| dc.contributor.author | Rinard, Martin C | |
| dc.contributor.author | Cito, Jürgen | |
| dc.date.accessioned | 2022-07-19T14:34:55Z | |
| dc.date.available | 2022-07-19T14:34:55Z | |
| dc.date.issued | 2021 | |
| dc.identifier.uri | https://hdl.handle.net/1721.1/143854 | |
| dc.description.abstract | <jats:p>Automated machine learning (AutoML) promises to democratize machine learning by automatically generating machine learning pipelines with little to no user intervention. Typically, a search procedure is used to repeatedly generate and validate candidate pipelines, maximizing a predictive performance metric, subject to a limited execution time budget. While this approach to generating candidates works well for small tabular datasets, the same procedure does not directly scale to larger tabular datasets with 100,000s of observations, often producing fewer candidate pipelines and yielding lower performance, given the same execution time budget. We carry out an extensive empirical evaluation of the impact that downsampling - reducing the number of rows in the input tabular dataset - has on the pipelines produced by a genetic-programming-based AutoML search for classification tasks.</jats:p> | en_US |
| dc.language.iso | en | |
| dc.publisher | VLDB Endowment | en_US |
| dc.relation.isversionof | 10.14778/3476249.3476262 | en_US |
| dc.rights | Creative Commons Attribution-NonCommercial-NoDerivs License | en_US |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | en_US |
| dc.source | VLDB Endowment | en_US |
| dc.title | Doing more with less: characterizing dataset downsampling for AutoML | en_US |
| dc.type | Article | en_US |
| dc.identifier.citation | Zogaj, Fatjon, Cambronero, José Pablo, Rinard, Martin C and Cito, Jürgen. 2021. "Doing more with less: characterizing dataset downsampling for AutoML." Proceedings of the VLDB Endowment, 14 (11). | |
| dc.contributor.department | Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science | |
| dc.relation.journal | Proceedings of the VLDB Endowment | en_US |
| dc.eprint.version | Final published version | en_US |
| dc.type.uri | http://purl.org/eprint/type/ConferencePaper | en_US |
| eprint.status | http://purl.org/eprint/status/NonPeerReviewed | en_US |
| dc.date.updated | 2022-07-19T14:17:14Z | |
| dspace.orderedauthors | Zogaj, F; Cambronero, JP; Rinard, MC; Cito, J | en_US |
| dspace.date.submission | 2022-07-19T14:17:15Z | |
| mit.journal.volume | 14 | en_US |
| mit.journal.issue | 11 | en_US |
| mit.license | PUBLISHER_CC | |
| mit.metadata.status | Authority Work and Publication Information Needed | en_US |