Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Andrey Gubarev

dblp:23/8510 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Database system architecture and tuning · 34% Distributed and cloud data management · 28% Query processing and optimization · 20%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 83% Cloud and datacenter computing · 17%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
disaggregated storage and compute
0.412020
Dremel: A Decade of Interactive SQL Analysis at Web Scale · Proc. VLDB Endow. 2020
Indexing and storage engines
columnar storage
0.422017
Spanner: Becoming a SQL System · SIGMOD Conference 2017
Dremel: Interactive Analysis of Web-Scale Datasets · Proc. VLDB Endow. 2010
Distributed systems
distributed database
0.322013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Spanner: Google's Globally-Distributed Database · OSDI 2012
Distributed and cloud data management
distributed query processing
0.312017
Spanner: Becoming a SQL System · SIGMOD Conference 2017
Distributed systems
consensus
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Distributed systems
replication
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Distributed systems › replication › update propagation
synchronous replication
0.212013
Spanner: Google's Globally Distributed Database · ACM Trans. Comput. Syst. 2013
Distributed and cloud data management › data replication
replica consistency
0.112017
Spanner: Becoming a SQL System · SIGMOD Conference 2017
Distributed systems
fault tolerance
0.112014
Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing · Proc. VLDB Endow. 2014
Distributed systems › distributed coordination and fault tolerance
consensus and replication
0.012012
Spanner: Google's Globally-Distributed Database · OSDI 2012
Distributed and cloud data management
mapreduce
0.012010
Dremel: Interactive Analysis of Web-Scale Datasets · Proc. VLDB Endow. 2010

Methods — techniques the papers use, named apart from their topics

columnar storage · 0.9in-situ analysis · 0.4in situ analysis · 0.4near real-time ingestion · 0.4geo-replication · 0.4range extraction · 0.3query restart · 0.3truetime · 0.2multi-level execution trees · 0.1columnar data layout · 0.1
YearPublicationVenuePosition
2020 Dremel: A Decade of Interactive SQL Analysis at Web Scale
abstract
Google's Dremel was one of the first systems that combined a set of architectural principles that have become a common practice in today's cloud-native analytics tools, including disaggregated storage and compute, in situ analysis, and columnar storage for semistructured data. In this paper, we discuss how these ideas evolved in the past decade and became the foundation for Google BigQuery.
Sergey Melnik 0001, Andrey Gubarev, Jing Jing Long, Geoffrey Romer, Shiva Shivakumar, Matt Tolton, Theo Vassilakis, Hossein Ahmadi 0001, Dan Delorey, Slava Min, Mosha Pasumansky, Jeff Shute
Proc. VLDB Endow.2
2017 Spanner: Becoming a SQL System
abstract
Spanner is a globally-distributed data management system that backs hundreds of mission-critical services at Google. Spanner is built on ideas from both the systems and database communities. The first Spanner paper published at OSDI'12 focused on the systems aspects such as scalability, automatic sharding, fault tolerance, consistent replication, external consistency, and wide-area distribution. This paper highlights the database DNA of Spanner. We describe distributed query execution in the presence of resharding, query restarts upon transient failures, range extraction that drives query routing and index seeks, and the improved blockwise-columnar storage format. We touch upon migrating Spanner to the common SQL dialect shared with other systems at Google.
David F. Bacon, Nathan Bales, Nicolas Bruno, Brian F. Cooper, Adam Dickinson, Andrew Fikes, Campbell Fraser, Andrey Gubarev, Milind Joshi, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, Rajesh Rao, David Shue, Marcel van der Holst, Dale Woodford
SIGMOD Conference8
2014 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing
abstract
Mesa is a highly scalable analytic data warehousing system that stores critical measurement data related to Google's Internet advertising business. Mesa is designed to satisfy a complex and challenging set of user and systems requirements, including near real-time data ingestion and queryability, as well as high availability, reliability, fault tolerance, and scalability for large data and query volumes. Specifically, Mesa handles petabytes of data, processes millions of row updates per second, and serves billions of queries that fetch trillions of rows per day. Mesa is geo-replicated across multiple datacenters and provides consistent and repeatable query answers at low latency, even when an entire datacenter fails. This paper presents the Mesa system and reports the performance and scale that it achieves.
Jason Govig, Adam Kirsch, Kelvin Chan, Sandeep Govind Dhoot, Abhilash Rajesh Kumar, Ankur Agiwal, Sanjay Bhansali, Mingsheng Hong, Jamie Cameron, Masood Siddiqi, Jeff Shute, Andrey Gubarev, Shivakumar Venkataraman, Divyakant Agrawal
Proc. VLDB Endow.17
2013 Spanner: Google's Globally Distributed Database
abstract
Spanner is Google’s scalable, multiversion, globally distributed, and synchronously replicated database. It is the first system to distribute data at global scale and support externally-consistent distributed transactions. This article describes how Spanner is structured, its feature set, the rationale underlying various design decisions, and a novel time API that exposes clock uncertainty. This API and its implementation are critical to supporting external consistency and a variety of powerful features: nonblocking reads in the past, lock-free snapshot transactions, and atomic schema changes, across all of Spanner.
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
ACM Trans. Comput. Syst.8
2012 Spanner: Google's Globally-Distributed Database
James C. Corbett, Jeffrey Dean, Michael Epstein, Andrew Fikes, Christopher Frost 0001, J. J. Furman, Sanjay Ghemawat, Andrey Gubarev, Christopher Heiser, Peter Hochschild, Wilson C. Hsieh, Sebastian Kanthak, Eugene Kogan, Alexander Lloyd, Sergey Melnik 0001, David Mwaura, David Nagle, Sean Quinlan, Rajesh Rao, Lindsay Rolig, Yasushi Saito, Michal Szymaniak, Ruth Wang, Dale Woodford
OSDI8
2010 Dremel: Interactive Analysis of Web-Scale Datasets
abstract
Dremel is a scalable, interactive ad-hoc query system for analysis of read-only nested data. By combining multi-level execution trees and columnar data layout, it is capable of running aggregation queries over trillion-row tables in seconds. The system scales to thousands of CPUs and petabytes of data, and has thousands of users at Google. In this paper, we describe the architecture and implementation of Dremel, and explain how it complements MapReduce-based computing. We present a novel columnar storage representation for nested records and discuss experiments on few-thousand node instances of the system.
Sergey Melnik 0001, Andrey Gubarev, Jing Jing Long, Geoffrey Romer, Shiva Shivakumar, Matt Tolton, Theo Vassilakis
Proc. VLDB Endow.2