VLDB 2026 Research / reviewers in the wild / expert
Jeff Shute
dblp:31/11411
· DBLP profile ↗
9ranked-venue papers
4as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 4 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Data models and query languages · 46% Database system architecture and tuning · 29% Distributed and cloud data management · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Cloud and datacenter computing · 53% Distributed systems · 47% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data models and query languages
query language design |
0.8 | 1 | 2024 | SQL has problems. We can fix them: Pipe syntax in SQL · Proc. VLDB Endow. 2024 |
Data models and query languages
SQL |
0.8 | 1 | 2024 | SQL has problems. We can fix them: Pipe syntax in SQL · Proc. VLDB Endow. 2024 |
Data models and query languages › SQL
SQL extension |
0.8 | 1 | 2024 | SQL has problems. We can fix them: Pipe syntax in SQL · Proc. VLDB Endow. 2024 |
Database system architecture and tuning
disaggregated storage and compute |
0.4 | 1 | 2020 | Dremel: A Decade of Interactive SQL Analysis at Web Scale · Proc. VLDB Endow. 2020 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.4 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Cloud and datacenter computing
database-as-a-service |
0.4 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Distributed and cloud data management › federated database
federated query processing |
0.3 | 1 | 2018 | F1 Query: Declarative Querying at Scale · Proc. VLDB Endow. 2018 |
Distributed systems
distributed database |
0.3 | 2 | 2013 | Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 F1: the fault-tolerant distributed RDBMS supporting google's ad business · SIGMOD Conference 2012 |
Distributed and cloud data management › distributed database architecture
distributed relational database |
0.2 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Data models and query languages › schema management
schema evolution |
0.2 | 1 | 2013 | Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 |
Distributed systems
replication and consistency |
0.2 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Database system architecture and tuning
relational database system |
0.1 | 1 | 2012 | F1: the fault-tolerant distributed RDBMS supporting google's ad business · SIGMOD Conference 2012 |
Distributed and cloud data management
federated database |
0.1 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Distributed systems
fault tolerance |
0.1 | 2 | 2014 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing · Proc. VLDB Endow. 2014 Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 |
Distributed and cloud data management › distributed query processing
distributed query engine |
0.0 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Distributed and cloud data management
distributed data store |
0.0 | 1 | 2012 | F1: the fault-tolerant distributed RDBMS supporting google's ad business · SIGMOD Conference 2012 |
Methods — techniques the papers use, named apart from their topics
columnar storage · 0.9declarative querying · 0.7SQL · 0.7in-situ analysis · 0.4in situ analysis · 0.4near real-time ingestion · 0.4geo-replication · 0.4formal model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic Data Modeling, Graph Query, and SQL, Together at Last?
Jeff Shute, Colin Zheng, Romit Kudtarkar |
CIDR | 1 |
| 2024 | SQL has problems. We can fix them: Pipe syntax in SQLabstractSQL has been extremely successful as the de facto standard language for working with data. Virtually all mainstream database-like systems use SQL as their primary query language. But SQL is an old language with significant design problems, making it difficult to learn, difficult to use, and difficult to extend. Many have observed these challenges with SQL, and proposed solutions involving new languages. New language adoption is a significant obstacle for users, and none of the potential replacements have been successful enough to displace SQL. In GoogleSQL, we've taken a different approach - solving SQL's problems by extending SQL. Inspired by a pattern that works well in other modern data languages, we added piped data flow syntax to SQL. The results are transformative - SQL becomes a flexible language that's easier to learn, use and extend, while still leveraging the existing SQL ecosystem and existing userbase. Improving SQL from within allows incrementally adopting new features, without migrations and without learning a new language, making this a more productive approach to improve on standard SQL. Jeff Shute, Shannon Bales, Matthew Brown 0006, Jean-Daniel Browne, Brandon Dolphin, Romit Kudtarkar, Andrey Litvinov, Jingchi Ma, John D. Morcos, Michael Shen, David Wilhite, Lulan Yu |
Proc. VLDB Endow. | 1 |
| 2020 | Dremel: A Decade of Interactive SQL Analysis at Web ScaleabstractGoogle's Dremel was one of the first systems that combined a set of architectural principles that have become a common practice in today's cloud-native analytics tools, including disaggregated storage and compute, in situ analysis, and columnar storage for semistructured data. In this paper, we discuss how these ideas evolved in the past decade and became the foundation for Google BigQuery. Sergey Melnik 0001, Andrey Gubarev, Jing Jing Long, Geoffrey Romer, Shiva Shivakumar, Matt Tolton, Theo Vassilakis, Hossein Ahmadi 0001, Dan Delorey, Slava Min, Mosha Pasumansky, Jeff Shute |
Proc. VLDB Endow. | 12 |
| 2020 | F1 Lightning: HTAP as a ServiceabstractThe ongoing and increasing interest in HTAP (Hybrid Transactional and Analytical Processing) systems documents the intense interest from data owners in simultaneously running transactional and analytical workloads over the same data set. Much of the reported work on HTAP has arisen in the context of "greenfield" systems, answering the question "if we could design a system for HTAP from scratch, what would it look like?" While there is great merit in such an approach, and a lot of valuable technology has been developed with it, we found ourselves facing a different challenge: one in which there is a great deal of transactional data already existing in several transactional systems, heavily queried by an existing federated engine that does not "own" the transactional systems, supporting both new and legacy applications that demand transparent fast queries and transactions from this combination. This paper reports on our design and experiences with F1 Lightning, a system we built and deployed to meet this challenge. We describe our design decisions, some details of our implementation, and our experience with the system in production for some of Google's most demanding applications. Ian Rae, Jeff Shute, Zhan Yuan, Kelvin Lau, Qilin Dong, Junxiong Zhou 0002, Jeremy Wood, Goetz Graefe, Jeffrey F. Naughton, John Cieslewicz |
Proc. VLDB Endow. | 4 |
| 2018 | F1 Query: Declarative Querying at ScaleabstractF1 Query is a stand-alone, federated query processing platform that executes SQL queries against data stored in different file-based formats as well as different storage systems at Google (e.g., Bigtable, Spanner, Google Spreadsheets, etc.). F1 Query eliminates the need to maintain the traditional distinction between different types of data processing workloads by simultaneously supporting: (i) OLTP-style point queries that affect only a few records; (ii) low-latency OLAP querying of large amounts of data; and (iii) large ETL pipelines. F1 Query has also significantly reduced the need for developing hard-coded data processing pipelines by enabling declarative queries integrated with custom business logic. F1 Query satisfies key requirements that are highly desirable within Google: (i) it provides a unified view over data that is fragmented and distributed over multiple data sources; (ii) it leverages datacenter resources for performant query processing with high throughput and low latency; (iii) it provides high scalability for large data sizes by increasing computational parallelism; and (iv) it is extensible and uses innovative approaches to integrate complex business logic in declarative query processing. This paper presents the end-to-end design of F1 Query. Evolved out of F1, the distributed database originally built to manage Google's advertising data, F1 Query has been in production for multiple years at Google and serves the querying needs of a large number of users and systems. Bart Samwel, John Cieslewicz, Ben Handy, Jason Govig, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Himani Apte, Felix Weigel, David Wilhite, Jiexing Li, Zhan Yuan, Craig Chasseur, Ian Rae, Anurag Biyani, Andrew Harn, Andrey Gubichev, Amr El-Helw, Orri Erling, Zhepeng Yan, Mohan Yang, Yiqun Wei, Thanh Do, Colin Zheng, Goetz Graefe, Somayeh Sardashti, Ahmed M. Aly, Divyakant Agrawal, Shivakumar Venkataraman |
Proc. VLDB Endow. | 8 |
| 2014 | Mesa: Geo-Replicated, Near Real-Time, Scalable Data WarehousingabstractMesa is a highly scalable analytic data warehousing system that stores critical measurement data related to Google's Internet advertising business. Mesa is designed to satisfy a complex and challenging set of user and systems requirements, including near real-time data ingestion and queryability, as well as high availability, reliability, fault tolerance, and scalability for large data and query volumes. Specifically, Mesa handles petabytes of data, processes millions of row updates per second, and serves billions of queries that fetch trillions of rows per day. Mesa is geo-replicated across multiple datacenters and provides consistent and repeatable query answers at low latency, even when an entire datacenter fails. This paper presents the Mesa system and reports the performance and scale that it achieves. Jason Govig, Adam Kirsch, Kelvin Chan, Sandeep Govind Dhoot, Abhilash Rajesh Kumar, Ankur Agiwal, Sanjay Bhansali, Mingsheng Hong, Jamie Cameron, Masood Siddiqi, Jeff Shute, Andrey Gubarev, Shivakumar Venkataraman, Divyakant Agrawal |
Proc. VLDB Endow. | 16 |
| 2013 | Online, Asynchronous Schema Change in F1abstractWe introduce a protocol for schema evolution in a globally distributed database management system with shared data, stateless servers, and no global membership. Our protocol is asynchronous--it allows different servers in the database system to transition to a new schema at different times--and online--all servers can access and update all data during a schema change. We provide a formal model for determining the correctness of schema changes under these conditions, and we demonstrate that many common schema changes can cause anomalies and database corruption. We avoid these problems by replacing corruption-causing schema changes with a sequence of schema changes that is guaranteed to avoid corrupting the database so long as all servers are no more than one schema version behind at any time. Finally, we discuss a practical implementation of our protocol in F1, the database management system that stores data for Google AdWords. Ian Rae, Eric Rollins, Jeff Shute, Sukhdeep S. Sodhi, Radek Vingralek |
Proc. VLDB Endow. | 3 |
| 2013 | F1: A Distributed SQL Database That ScalesabstractF1 is a distributed relational database system built at Google to support the AdWords business. F1 is a hybrid database that combines high availability, the scalability of NoSQL systems like Bigtable, and the consistency and usability of traditional SQL databases. F1 is built on Spanner, which provides synchronous cross-datacenter replication and strong consistency. Synchronous replication implies higher commit latency, but we mitigate that latency by using a hierarchical schema model with structured data types and through smart application design. F1 also includes a fully functional distributed SQL query engine and automatic change tracking and publishing. Jeff Shute, Radek Vingralek, Bart Samwel, Ben Handy, Chad Whipkey, Eric Rollins, Mircea Oancea, Kyle Littlefield, David Menestrina, Stephan Ellner, John Cieslewicz, Ian Rae, Traian Stancescu, Himani Apte |
Proc. VLDB Endow. | 1 |
| 2012 | F1: the fault-tolerant distributed RDBMS supporting google's ad businessabstractMany of the services that are critical to Google's ad business have historically been backed by MySQL. We have recently migrated several of these services to F1, a new RDBMS developed at Google. F1 implements rich relational database features, including a strictly enforced schema, a powerful parallel SQL query engine, general transactions, change tracking and notification, and indexing, and is built on top of a highly distributed storage system that scales on standard hardware in Google data centers. The store is dynamically sharded, supports transactionally-consistent replication across data centers, and is able to handle data center outages without data loss. Jeff Shute, Mircea Oancea, Stephan Ellner, Ben Handy, Eric Rollins, Bart Samwel, Radek Vingralek, Chad Whipkey, Beat Jegerlehner, Kyle Littlefield, Phoenix Tong |
SIGMOD Conference | 1 |