BlinkDB: queries with bounded errors and bounded response times on very large data

Agarwal, Sameer; Mozafari, Barzan; Panda, Aurojit; Milner, Henry; Madden, Samuel; Stoica, Ion

Author(s)

Agarwal, Sameer; Mozafari, Barzan; Panda, Aurojit; Milner, Henry; Stoica, Ion; ... Show more

DownloadMadden_BlinkDB.pdf (876.5Kb)

OPEN_ACCESS_POLICY

Terms of use

Creative Commons Attribution-Noncommercial-Share Alike http://creativecommons.org/licenses/by-nc-sa/4.0/

Metadata

Show full item record

Abstract

In this paper, we present BlinkDB, a massively parallel, approximate query engine for running interactive SQL queries on large volumes of data. BlinkDB allows users to trade-off query accuracy for response time, enabling interactive queries over massive data by running queries on data samples and presenting results annotated with meaningful error bars. To achieve this, BlinkDB uses two key ideas: (1) an adaptive optimization framework that builds and maintains a set of multi-dimensional stratified samples from original data over time, and (2) a dynamic sample selection strategy that selects an appropriately sized sample based on a query's accuracy or response time requirements. We evaluate BlinkDB against the well-known TPC-H benchmarks and a real-world analytic workload derived from Conviva Inc., a company that manages video distribution over the Internet. Our experiments on a 100 node cluster show that BlinkDB can answer queries on up to 17 TBs of data in less than 2 seconds (over 200 x faster than Hive), within an error of 2-10%.

Date issued

2013-04

URI

http://hdl.handle.net/1721.1/100911

Department

Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory; Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science

Journal

Proceedings of the 8th ACM European Conference on Computer Systems (EuroSys '13)

Publisher

Association for Computing Machinery (ACM)

Citation

Sameer Agarwal, Barzan Mozafari, Aurojit Panda, Henry Milner, Samuel Madden, and Ion Stoica. 2013. BlinkDB: queries with bounded errors and bounded response times on very large data. In Proceedings of the 8th ACM European Conference on Computer Systems (EuroSys '13). ACM, New York, NY, USA, 29-42.

Version: Author's final manuscript

ISBN

9781450319942

Collections

MIT Open Access Articles

DSpace@MIT