Query Optimization for Dynamic Imputation
Name
p1310-feser.pdf
Description
Published version
Size
898.53 KB
Format
Adobe PDF
Checksum (MD5)
0fc8764b233efd5a85f36b050eadb432
Author(s) • • •
Cambronero, José
Feser, John K.
Smith, Micah J.
Madden, Samuel
Date Issued
August 2017
Publisher
VLDB Endowment
Citation
Cambronero, José, Feser, John K., Smith, Micah J. and Madden, Samuel. 2017. "Query Optimization for Dynamic Imputation." 10 (11).
Version
Final published version
Abstract
© 2017 VLDB. Missing values are common in data analysis and present a usability challenge. Users are forced to pick between removing tuples withmissing values or creating a cleaned version of their data by applying a relatively expensive imputation strategy. Our system, ImputeDB, incorporates imputation into a costbased query optimizer, performing necessary imputations onthefly for each query. This allows users to immediately explore their data, while the system picks the optimal placement of imputation operations. We evaluate this approach on three real-world survey-based datasets. Our experiments show that our query plans execute between 10 and 140 times faster than first imputing the base tables. Furthermore, we show that the query results from on-the-fly imputation differ from the traditional base-table imputation approach by 0-8%. Finally, we show that while dropping tuples with missing values that fail query constraints discards 6-78% of the data, on-the-fly imputation loses only 0-21%.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Laboratory for Information and Decision Systems
Terms of Use
Creative Commons Attribution-NonCommercial-NoDerivs License
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.14778/3137628.3137641