VLDB 2026 Research / reviewers in the wild / expert
Ian Rae
dblp:87/7992
· DBLP profile ↗
9ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0001-8790-1925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-authorSystems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
7 papers |
Database system architecture and tuning · 30% Distributed and cloud data management · 26% Information retrieval · 21% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 58% Distributed systems · 42% |
Topics — the 17 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Database system architecture and tuning
hybrid transactional and analytical processing |
0.4 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Cloud and datacenter computing
database-as-a-service |
0.4 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Distributed and cloud data management › federated database
federated query processing |
0.3 | 1 | 2018 | F1 Query: Declarative Querying at Scale · Proc. VLDB Endow. 2018 |
Information retrieval › keyword search
keyword search over relational databases |
0.2 | 2 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 Toward industrial-strength keyword search systems over relational data · ICDE 2010 |
Information retrieval › indexing
inverted index |
0.2 | 1 | 2014 | In-RDBMS inverted indexes revisited · ICDE 2014 |
Distributed and cloud data management › distributed database architecture
distributed relational database |
0.2 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Data models and query languages › schema management
schema evolution |
0.2 | 1 | 2013 | Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 |
Distributed systems
distributed database |
0.2 | 1 | 2013 | Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 |
Distributed systems
replication and consistency |
0.2 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Distributed and cloud data management
federated database |
0.1 | 1 | 2020 | F1 Lightning: HTAP as a Service · Proc. VLDB Endow. 2020 |
Information retrieval
keyword search |
0.1 | 1 | 2010 | Toward industrial-strength keyword search systems over relational data · ICDE 2010 |
Query processing and optimization › online query processing
time-constrained query processing |
0.1 | 1 | 2010 | Toward industrial-strength keyword search systems over relational data · ICDE 2010 |
Query processing and optimization › interactive data exploration
answer space exploration |
0.1 | 2 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 Toward industrial-strength keyword search systems over relational data · ICDE 2010 |
Query processing and optimization › join processing
join algorithms |
0.1 | 1 | 2014 | In-RDBMS inverted indexes revisited · ICDE 2014 |
Distributed and cloud data management › distributed query processing
distributed query engine |
0.0 | 1 | 2013 | F1: A Distributed SQL Database That Scales · Proc. VLDB Endow. 2013 |
Distributed systems
fault tolerance |
0.0 | 1 | 2013 | Online, Asynchronous Schema Change in F1 · Proc. VLDB Endow. 2013 |
Information retrieval
search engines |
0.0 | 1 | 2010 | Toward Scalable Keyword Search over Relational Data · Proc. VLDB Endow. 2010 |
Methods — techniques the papers use, named apart from their topics
declarative querying · 0.7SQL · 0.7formal model · 0.3query forms · 0.2phrase query processing · 0.2conjunctive query processing · 0.2time-limited answer generation · 0.1proof of concept · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | F1 Lightning: HTAP as a ServiceabstractThe ongoing and increasing interest in HTAP (Hybrid Transactional and Analytical Processing) systems documents the intense interest from data owners in simultaneously running transactional and analytical workloads over the same data set. Much of the reported work on HTAP has arisen in the context of "greenfield" systems, answering the question "if we could design a system for HTAP from scratch, what would it look like?" While there is great merit in such an approach, and a lot of valuable technology has been developed with it, we found ourselves facing a different challenge: one in which there is a great deal of transactional data already existing in several transactional systems, heavily queried by an existing federated engine that does not "own" the transactional systems, supporting both new and legacy applications that demand transparent fast queries and transactions from this combination. This paper reports on our design and experiences with F1 Lightning, a system we built and deployed to meet this challenge. We describe our design decisions, some details of our implementation, and our experience with the system in production for some of Google's most demanding applications. Ian Rae, Jeff Shute, Zhan Yuan, Kelvin Lau, Qilin Dong, Junxiong Zhou 0002, Jeremy Wood, Goetz Graefe, Jeffrey F. Naughton, John Cieslewicz |
Proc. VLDB Endow. | 2 |
| 2018 | F1 Query: Declarative Querying at ScaleabstractF1 Query is a stand-alone, federated query processing platform that executes SQL queries against data stored in different file-based formats as well as different storage systems at Google (e.g., Bigtable, Spanner, Google Spreadsheets, etc.). F1 Query eliminates the need to maintain the traditional distinction between different types of data processing workloads by simultaneously supporting: (i) OLTP-style point queries that affect only a few records; (ii) low-latency OLAP querying of large amounts of data; and (iii) large ETL pipelines. F1 Query has also significantly reduced the need for developing hard-coded data processing pipelines by enabling declarative queries integrated with custom business logic. F1 Query satisfies key requirements that are highly desirable within Google: (i) it provides a unified view over data that is fragmented and distributed over multiple data sources; (ii) it leverages datacenter resources for performant query processing with high throughput and low latency; (iii) it provides high scalability for large data sizes by increasing computational parallelism; and (iv) it is extensible and uses innovative approaches to integrate complex business logic in declarative query processing. This paper presents the end-to-end design of F1 Query. Evolved out of F1, the distributed database originally built to manage Google's advertising data, F1 Query has been in production for multiple years at Google and serves the querying needs of a large number of users and systems. Bart Samwel, John Cieslewicz, Ben Handy, Jason Govig, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Himani Apte, Felix Weigel, David Wilhite, Jiexing Li, Zhan Yuan, Craig Chasseur, Ian Rae, Anurag Biyani, Andrew Harn, Andrey Gubichev, Amr El-Helw, Orri Erling, Zhepeng Yan, Mohan Yang, Yiqun Wei, Thanh Do, Colin Zheng, Goetz Graefe, Somayeh Sardashti, Ahmed M. Aly, Divyakant Agrawal, Shivakumar Venkataraman |
Proc. VLDB Endow. | 19 |
| 2015 | Database Optimization in the Cloud: Where Costs, Partial Results, and Consumer Choice Meet
Willis Lang, Rimma V. Nehme, Ian Rae |
CIDR | 3 |
| 2015 | The IceProd framework: Distributed data processing for the IceCube neutrino observatory
Mark G. Aartsen, Rasha U. Abbasi, Markus Ackermann 0003, Jenni Adams, Juan Antonio Aguilar Sánchez, Markus Ahlers, David Altmann, Carlos A. Argüelles Delgado, Jan Auffenberg, Xinhua Bai, Michael F. Baker, Steven W. Barwick, Volker Baum, Ryan Bay, James J. Beatty, Julia K. Becker Tjus, Karl-Heinz Becker, Segev BenZvi, Patrick Berghaus, David Berley, Elisa Bernardini, Anna Bernhard, David Z. Besson, G. Binder, Daniel Bindig, Martin Bissok, Erik Blaufuss, Jan Blumenthal, David J. Boersma, Christian Bohm, Debanjan Bose, Sebastian Böser, Olga Botner, Lionel Brayeur, Hans-Peter Bretz, Anthony M. Brown, Ronald Bruijn, James Casey, Martin Casier, Dmitry Chirkin, Asen Christov, Brian John Christy, Ken Clark, Lew Classen, Fabian Clevermann, Stefan Coenders, Shirit Cohen, Doug F. Cowen, Angel H. Cruz Silva, Matthias Danninger, Jacob Daughhetee, James C. Davis 0002, Melanie Day, Catherine De Clercq, Sam De Ridder, Paolo Desiati, Krijn D. de Vries, Meike de With, Tyce DeYoung, Juan Carlos Díaz-Vélez, Matthew Dunkman, Ryan Eagan, Benjamin Eberhardt, Björn Eichmann, Jonathan Eisch, Sebastian Euler, Paul A. Evenson, Oladipo O. Fadiran, Ali R. Fazely, Anatoli Fedynitch, Jacob Feintzeig, Tom Feusels, Kirill Filimonov, Chad Finley, Tobias Fischer-Wasels, Samuel Flis, Anna Franckowiak, Katharina Frantzen, Tomasz Fuchs, Thomas K. Gaisser, Joseph S. Gallagher, Lisa Gerhardt, Laura E. Gladstone, Thorsten Glüsenkamp, Azriel Goldschmidt, Geraldina Golup, Javier G. González, Jordan A. Goodman, Dariusz Góra, Dylan T. Grandmont, Darren Grant, Pavel Gretskov, John C. Groh, Andreas Groß, Chang Hyon Ha, Abd Al Karim Haj Ismail, Patrick Hallen, Allan Hallgren, Francis Halzen, Kael D. Hanson, Dustin Hebecker, David Heereman, Dirk Heinen, Klaus Helbing, Robert Eugene Hellauer III, Stephanie Virginia Hickford, Gary C. Hill, Kara D. Hoffman, Ruth Hoffmann, Andreas Homeier, Kotoyo Hoshina, Feifei Huang, Warren Huelsnitz, Per Olof Hulth, Klas Hultqvist, Aya Ishihara, Emanuel Jacobi, John E. Jacobsen, Kai Jagielski, George S. Japaridze, Kyle Jero, Ola Jlelati, Basho Kaminsky, Alexander Kappes, Timo Karg, Albrecht Karle, Matthew Kauer, John Lawrence Kelley, Joanna Kiryluk, J. Kläs, Spencer R. Klein, Jan-Hendrik Köhne, Georges Kohnen, Hermann Kolanoski, Lutz Köpke, Claudio Kopper, Sandro Kopper, D. Jason Koskinen, Marek Kowalski, Mark Krasberg, Anna Kriesten, Kai Michael Krings, Gösta Kroll, Jan Kunnen, Naoko Kurahashi, Takao Kuwabara, Mathieu L. M. Labare, Hagar Landsman, Michael James Larson, Mariola Lesiak-Bzdak, Martin Leuermann, Julia Leute, Jan Lünemann, Oscar A. Macías-Ramírez, James Madsen, Giuliano Maggi, Reina Maruyama, Keiichi Mase, Howard S. Matis, Frank McNally, Kevin James Meagher, Martin Merck, Gonzalo Merino, Thomas Meures, Sandra Miarecki, Eike Middell, Natalie Milke, John Lester Miller, Lars Mohrmann, Teresa Montaruli, Robert M. Morse, Rolf Nahnhauer, Uwe Naumann, Hans Niederhausen, Sarah C. Nowicki, David R. Nygren, Anna Pollmann, Sirin Odrowski, Alex Olivas, Ahmad Omairat, Aongus Starbuck Ó Murchadha, Larissa Paul, Joshua A. Pepper, Carlos Pérez de los Heros, Carl Pfendner, Damian Pieloth, Elisa Pinat, Jonas Posselt, P. Buford Price, Gerald T. Przybylski, Melissa Quinnan, Leif Rädel, Ian Rae, Mohamed Rameez, Katherine Rawlins, Peter Christian Redl, René Reimann, Elisa Resconi, Wolfgang Rhode, Mathieu Ribordy, Michael Richman, Benedikt Riedel, J. P. Rodrigues, Carsten Rott, Tim Ruhe, Bakhtiyar Ruzybayev, Dirk Ryckbosch, Sabine M. Saba, Heinz-Georg Sander, Juan Marcos Santander, Subir Sarkar 0002, Kai Schatto, Florian Scheriau, Torsten Schmidt, Martin Schmitz 0004, Sebastian Schoenen, Sebastian Schöneberg, Arne Schönwald, Anne Schukraft, Lukas Schulte, David Schultz, Olaf Schulz, David Seckel, Yolanda Sestayo de la Cerra, Surujhdeo Seunarine, Rezo Shanidze, Chris Sheremata, Miles W. E. Smith, Dennis Soldin, Glenn M. Spiczak, Christian Spiering, Michael Stamatikos, Todor Stanev, Nick A. Stanisha, Alexander Stasik, Thorsten Stezelberger, Robert G. Stokstad, Achim Stößl, Erik A. Strahler, Rickard Ström, Nora Linn Strotjohann, Gregory W. Sullivan, Henric Taavola, Ignacio J. Taboada, Alessio Tamburro, Andreas Tepe, Samvel Ter-Antonyan, Gordana Tesic, Serap Tilav, Patrick A. Toale, Moriah Natasha Tobin, Simona Toscano, Maria Tselengidou, Elisabeth Unger, Marcel Usner, Sofia Vallecorsa, Nick van Eijndhoven, Arne Van Overloop, Jakob van Santen, Markus Vehring, Markus Voge, Matthias Vraeghe, Christian Walck, Tilo Waldenmaier, Marius Wallraff, Christopher Weaver 0001, Mark T. Wellons, Christopher H. Wendt, Stefan Westerhoff, Nathan Whitehorn, Klaus Wiebe, Christopher H. Wiebusch, Dawn R. Williams, Henrike Wissing, Martin Wolf 0007, Terri R. Wood, Kurt Woschnagg, Donglian Xu, Xianwu Xu, Juan Pablo Yáñez, Gaurang B. Yodh, Shigeru Yoshida, Pavel Zarzhitsky, Jan Ziemann, Simon Zierke, Marcel Zoll |
J. Parallel Distributed Comput. | 194 |
| 2014 | In-RDBMS inverted indexes revisitedabstractEvery major open-source and commercial RDBMS offers some form of support for full-text search using inverted indexes. When providing this support, some developers have implemented specialized indexes that adapt techniques from the Information Retrieval (IR) community to work in a database setting, while others have opted to rely on the standard relational query engine to process inverted index lookups. This choice is an important one, since the storage formats and algorithms used can vary greatly between a specialized index and a relational index, but these alternatives have not been thoroughly compared in the same system. Our work explores the differences in implementation and performance of three representative environments for an in-RDBMS inverted index: an in-RDBMS IR engine, a row-oriented relational query engine, and a column-oriented relational query engine. We found that a specialized IR engine integrated into the RDBMS can provide more than an order of magnitude speedup over both the row- and column-oriented relational query engines for conjunctive and phrase queries. For warm queries, this advantage is largely algorithmic, and we show that by using ZigZag merge join to accelerate conjunctive and phrase query processing, relational inverted indexes can provide performance comparable to a specialized in-RDBMS IR engine with no change to the underlying storage format. Compression and index format, in contrast, have more impact on cold queries, where the IR and column-oriented engines are able to outperform the row-oriented engine, even with ZigZag merge join. Ian Rae, Alan Halverson, Jeffrey F. Naughton |
ICDE | 1 |
| 2013 | Online, Asynchronous Schema Change in F1abstractWe introduce a protocol for schema evolution in a globally distributed database management system with shared data, stateless servers, and no global membership. Our protocol is asynchronous--it allows different servers in the database system to transition to a new schema at different times--and online--all servers can access and update all data during a schema change. We provide a formal model for determining the correctness of schema changes under these conditions, and we demonstrate that many common schema changes can cause anomalies and database corruption. We avoid these problems by replacing corruption-causing schema changes with a sequence of schema changes that is guaranteed to avoid corrupting the database so long as all servers are no more than one schema version behind at any time. Finally, we discuss a practical implementation of our protocol in F1, the database management system that stores data for Google AdWords. Ian Rae, Eric Rollins, Jeff Shute, Sukhdeep S. Sodhi, Radek Vingralek |
Proc. VLDB Endow. | 1 |
| 2013 | F1: A Distributed SQL Database That ScalesabstractF1 is a distributed relational database system built at Google to support the AdWords business. F1 is a hybrid database that combines high availability, the scalability of NoSQL systems like Bigtable, and the consistency and usability of traditional SQL databases. F1 is built on Spanner, which provides synchronous cross-datacenter replication and strong consistency. Synchronous replication implies higher commit latency, but we mitigate that latency by using a hierarchical schema model with structured data types and through smart application design. F1 also includes a fully functional distributed SQL query engine and automatic change tracking and publishing. Jeff Shute, Radek Vingralek, Bart Samwel, Ben Handy, Chad Whipkey, Eric Rollins, Mircea Oancea, Kyle Littlefield, David Menestrina, Stephan Ellner, John Cieslewicz, Ian Rae, Traian Stancescu, Himani Apte |
Proc. VLDB Endow. | 12 |
| 2010 | Toward industrial-strength keyword search systems over relational dataabstractKeyword search (KWS) over relational data, where the answers are multiple tuples connected via joins, has received significant attention in the past decade. Numerous solutions have been proposed and many prototypes have been developed. Building on this rapid progress and on growing user needs, recently several RDBMS and Web companies as well as academic research groups have started to examine how to build industrial-strength keywords search systems. This task clearly requires addressing many issues, including robustness, accuracy, reliability, and privacy, among others. A major emerging issue, however, appears to be performance related: current KWS systems have unpredictable run time. In particular, for certain queries it takes too long to produce answers, and for others the system may even fail to return (e.g., after exhausting memory). In this paper we begin by examining the above problem and arguing that it is a fundamental problem unlikely to be solved in the near future by software and hardware advances. Next, we argue that in an industrial-strength setting, to ensure real-time interaction and facilitate user adoption, KWS systems should produce answers under an absolute time limit and then provide users with a description of what could be done next, should he or she choose to continue. Next, we show how to realize these requirements for DISCOVER, an exemplar of a recent KWS solution approach. Our basic idea is to produce answers as in today's KWS systems up to the time limit, then show users these answers as well as query forms that characterize the unexplored portion of the answer space. Finally, we present some preliminary experiments over real-world data to demonstrate the feasibility of the proposed solution approach. Akanksha Baid, Ian Rae, AnHai Doan, Jeffrey F. Naughton |
ICDE | 2 |
| 2010 | Toward Scalable Keyword Search over Relational DataabstractKeyword search (KWS) over relational databases has recently received significant attention. Many solutions and many prototypes have been developed. This task requires addressing many issues, including robustness, accuracy, reliability, and privacy. An emerging issue, however, appears to be performance related: current KWS systems have unpredictable running times. In particular, for certain queries it takes too long to produce answers, and for others the system may even fail to return (e.g., after exhausting memory). In this paper we argue that as today's users have been "spoiled" by the performance of Internet search engines, KWS systems should return whatever answers they can produce quickly and then provide users with options for exploring any portion of the answer space not covered by these answers. Our basic idea is to produce answers that can be generated quickly as in today's KWS systems, then to show users query forms that characterize the unexplored portion of the answer space. Combining KWS systems with forms allows us to bypass the performance problems inherent to KWS without compromising query coverage. We provide a proof of concept for this proposed approach, and discuss the challenges encountered in building this hybrid system. Finally, we present experiments over real-world datasets to demonstrate the feasibility of the proposed solution. Akanksha Baid, Ian Rae, Jiexing Li, AnHai Doan, Jeffrey F. Naughton |
Proc. VLDB Endow. | 2 |