Stage: Query Execution Time Prediction in Amazon Redshift
Name
3626246.3653391.pdf
Size
1.97 MB
Format
Adobe PDF
Checksum (MD5)
ec74e0a77726e9b28f0e094a8c19b2b9
Author(s) • • • • • • • • •
Wu, Ziniu
Marcus, Ryan
Liu, Zhengchun
Negi, Parimarjan
Nathan, Vikram
Pfeil, Pascal
Saxena, Gaurav
Rahman, Mohammad
Narayanaswamy, Balakrishnan
Kraska, Tim
Date Issued
June 9, 2024
Publisher
ACM|Companion of the 2024 International Conference on Management of Data
Citation
Wu, Ziniu, Marcus, Ryan, Liu, Zhengchun, Negi, Parimarjan, Nathan, Vikram et al. 2024. "Stage: Query Execution Time Prediction in Amazon Redshift."
Version
Final published version
Abstract
Query performance (e.g., execution time) prediction is a critical component of modern DBMSes. As a pioneering cloud data warehouse, Amazon Redshift relies on an accurate execution time prediction for many downstream tasks, ranging from high-level optimizations, such as automatically creating materialized views, to low-level tasks on the critical path of query execution, such as admission, scheduling, and execution resource control. Unfortunately, many existing execution time prediction techniques, including those used in Redshift, suffer from cold start issues, inaccurate estimation, and are not robust against workload/data changes.
In this paper, we propose a novel hierarchical execution time predictor: the Stage predictor. The Stage predictor is designed to leverage the unique characteristics and challenges faced by Redshift. The Stage predictor consists of three model states: an execution time cache, a lightweight local model optimized for a specific DB instance with uncertainty measurement, and a complex global model that is transferable across all instances in Redshift. We design a systematic approach to use these models that best leverages optimality (cache), instance-optimization (local model), and transferable knowledge about Redshift (global model). Experimentally, we show that the Stage predictor makes more accurate and robust predictions while maintaining a practical inference latency and memory overhead. Overall, the Stage predictor can improve the average query execution latency by 20% on these instances compared to the prior query performance predictor in Redshift.
MIT Department
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3626246.3653391