Data formats in analytical DBMSs: performance trade-offs and future directions
Name
778_2025_Article_911.pdf
Size
3.11 MB
Format
Adobe PDF
Checksum (MD5)
d5db295f4622d14ee862052eb9902cc7
Author(s) • • •
Liu, Chunwei
Pavlenko, Anna
Interlandi, Matteo
Haynes, Brandon
Date Issued
March 19, 2025
Journal
The VLDB Journal
Publisher
Springer Berlin Heidelberg
Citation
Liu, C., Pavlenko, A., Interlandi, M. et al. Data formats in analytical DBMSs: performance trade-offs and future directions. The VLDB Journal 34, 30 (2025).
Version
Final published version
Abstract
This paper evaluates the suitability of Apache Arrow, Parquet, and ORC as formats for subsumption in an analytical DBMS. We systematically identify and explore the high-level features that are important to support efficient querying in modern OLAP DBMSs and evaluate the ability of each format to support these features. We find that each format has trade-offs that make it more or less suitable for use as a format in a DBMS and identify opportunities to more holistically co-design a unified in-memory and on-disk data representation. Notably, for certain popular machine learning tasks, none of these formats perform optimally, highlighting significant opportunities for advancing format design. Our hope is that this study can be used as a guide for system developers designing and using these formats, as well as provide the community with directions to pursue for improving these common open formats.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1007/s00778-025-00911-1