EDBT 2026 Demo / reviewers in the wild / expert
Andrey Gubarev
dblp:23/8510
· DBLP profile ↗
6ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Database system architecture and tuning · 34% Distributed and cloud data management · 28% Query processing and optimization · 20% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 83% Cloud and datacenter computing · 17% |
Topics — the 11 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Database system architecture and tuning
disaggregated storage and compute |
0.4 | 1 | 2020 | Dremel: A Decade of Interactive SQL Analysis at Web Scale · Proc. VLDB Endow. 2020 |
Indexing and storage engines
columnar storage |
0.4 | 2 | 2017 | Spanner: Becoming a SQL System · SIGMOD Conference 2017 Dremel: Interactive Analysis of Web-Scale Datasets · Proc. VLDB Endow. 2010 |
Distributed systems
distributed database |
0.3 | 2 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 Spanner: Google's Globally-Distributed Database · OSDI 2012 |
Distributed and cloud data management
distributed query processing |
0.3 | 1 | 2017 | Spanner: Becoming a SQL System · SIGMOD Conference 2017 |
Distributed systems
consensus |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Distributed systems
replication |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Distributed systems › replication › update propagation
synchronous replication |
0.2 | 1 | 2013 | Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013 |
Distributed and cloud data management › data replication
replica consistency |
0.1 | 1 | 2017 | Spanner: Becoming a SQL System · SIGMOD Conference 2017 |
Distributed systems
fault tolerance |
0.1 | 1 | 2014 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing · Proc. VLDB Endow. 2014 |
Distributed systems › distributed coordination and fault tolerance
consensus and replication |
0.0 | 1 | 2012 | Spanner: Google's Globally-Distributed Database · OSDI 2012 |
Distributed and cloud data management
mapreduce |
0.0 | 1 | 2010 | Dremel: Interactive Analysis of Web-Scale Datasets · Proc. VLDB Endow. 2010 |
Methods — techniques the papers use, named apart from their topics
columnar storage · 0.9in-situ analysis · 0.4in situ analysis · 0.4near real-time ingestion · 0.4geo-replication · 0.4range extraction · 0.3query restart · 0.3truetime · 0.2multi-level execution trees · 0.1columnar data layout · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Dremel: A Decade of Interactive SQL Analysis at Web ScaleabstractGoogle's Dremel was one of the first systems that combined a set of architectural principles that have become a common practice in today's cloud-native analytics tools, including disaggregated storage and compute, in situ analysis, and columnar storage for semistructured data. In this paper, we discuss how these ideas evolved in the past decade and became the foundation for Google BigQuery. Sergey Melnik 0001, Andrey Gubarev, Jing Jing Long, Geoffrey Romer, Shiva Shivakumar, Matt Tolton, Theo Vassilakis, Hossein Ahmadi 0001, Dan Delorey, Slava Min, Mosha Pasumansky, Jeff Shute |
Proc. VLDB Endow. | 2 |
| 2017 | Spanner: Becoming a SQL SystemabstractSpanner is a globally-distributed data management system that backs hundreds of mission-critical services at Google. Spanner is built on ideas from both the systems and database communities. The first Spanner paper published at OSDI'12 focused on the systems aspects such as scalability, automatic sharding, fault tolerance, consistent replication, external consistency, and wide-area distribution. This paper highlights the database DNA of Spanner. We describe distributed query execution in the presence of resharding, query restarts upon transient failures, range extraction that drives query routing and index seeks, and the improved blockwise-columnar storage format. We touch upon migrating Spanner to the common SQL dialect shared with other systems at Google. David F. Bacon, Nathan Bales, Nicolas Bruno, Brian F. Cooper, Adam Dickinson, Andrew Fikes, Campbell Fraser, Andrey Gubarev, Milind Joshi, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, Rajesh Rao, David Shue, Marcel van der Holst, Dale Woodford |
SIGMOD Conference | 8 |
| 2014 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data WarehousingabstractMesa is a highly scalable analytic data warehousing system that stores critical measurement data related to Google's Internet advertising business. Mesa is designed to satisfy a complex and challenging set of user and systems requirements, including near real-time data ingestion and queryability, as well as high availability, reliability, fault tolerance, and scalability for large data and query volumes. Specifically, Mesa handles petabytes of data, processes millions of row updates per second, and serves billions of queries that fetch trillions of rows per day. Mesa is geo-replicated across multiple datacenters and provides consistent and repeatable query answers at low latency, even when an entire datacenter fails. This paper presents the Mesa system and reports the performance and scale that it achieves. Jason Govig, Adam Kirsch, Kelvin Chan, Sandeep Govind Dhoot, Abhilash Rajesh Kumar, Ankur Agiwal, Sanjay Bhansali, Mingsheng Hong, Jamie Cameron, Masood Siddiqi, Jeff Shute, Andrey Gubarev, Shivakumar Venkataraman, Divyakant Agrawal |
Proc. VLDB Endow. | 17 |
| 2013 | Spanner: Google's Globally Distributed DatabaseabstractSpanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner. James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
ACM Trans. Comput. Syst. | 8 |
| 2012 | Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford |
OSDI | 8 |
| 2010 | Dremel: Interactive Analysis of Web-Scale DatasetsabstractDremel is a scalable, interactive ad-hoc query system for analysis of read-only nested data. By combining multi-level execution trees and columnar data layout, it is capable of running aggregation queries over trillion-row tables in seconds. The system scales to thousands of CPUs and petabytes of data, and has thousands of users at Google. In this paper, we describe the architecture and implementation of Dremel, and explain how it complements MapReduce-based computing. We present a novel columnar storage representation for nested records and discuss experiments on few-thousand node instances of the system. Sergey Melnik 0001, Andrey Gubarev, Jing Jing Long, Geoffrey Romer, Shiva Shivakumar, Matt Tolton, Theo Vassilakis |
Proc. VLDB Endow. | 2 |