Tuplex: Robust, Efficient Analytics When Python Rules
Name
3352063.3352109.pdf
Description
Published version
Size
444.36 KB
Format
Unknown
Checksum (MD5)
47c4493c012cae2ee31c40993000f4ce
Author(s) •
Spiegelberg, Leonhard F
Kraska, Tim
Date Issued
2019
Journal
Proceedings of the VLDB Endowment
Publisher
VLDB Endowment
Version
Final published version
Abstract
© 2019 VLDB Endowment. Spark became the defacto industry standard as an execution engine for data preparation, cleaning, distributed machine learning, streaming and, warehousing over raw data. However, with the success of Python the landscape is shifting again; there is a strong demand for tools which better integrate with the Python landscape and do not have the impedance mismatch like Spark. In this paper, we demonstrate Tuplex (short for tuples and exceptions), a Pythonnative data preparation framework that allows users to develop and deploy pipelines faster and more robustly while providing bare-metal execution times through code compilation whenever possible.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution-NonCommercial-NoDerivs License
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.14778/3352063.3352109