Learned Encodings in SageDB
Name
Cen-lujing-meng-eecs-2021-thesis.pdf
Description
Thesis PDF
Size
785.83 KB
Format
Adobe PDF
Checksum (MD5)
9af430c413ada4f5b2887192ae244abc
Author(s)
Cen, Lujing
Advisor(s)
Kraska, Tim
Date Issued
June 2021
Publisher
Massachusetts Institute of Technology
Abstract
As the demand for data outpaces diminishing improvements in the hardware used to store and query them, we must find intelligent ways to increase database performance on existing systems. This project is focused on integrating learned encodings into SageDB, a database capable of accelerating queries by analyzing and adapting to different workloads. Encodings improve query performance through lossless compression, thereby reducing I/O time during scans. Different encoding types exhibit different characteristics depending on properties of the underlying data and the hardware on which queries are executed. We implement a variety of common encodings in SageDB and propose a learning-based approach to select the optimal encoding for a given data block by combining block-level statistics with sampling. In addition, we demonstrate how to leverage properties of encoded data along with vectorized processing units in modern CPUs to more efficiently execute queries without the need to decode every value.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright MIT
Persistent DSpace Link