EDBT 2026 Demo / reviewers in the wild / expert
Ruby Y. Tahboub
dblp:18/8896
· DBLP profile ↗
10ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-authorArtificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Query processing and optimization · 61% Spatial and temporal data management · 28% Machine learning and data management · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 77% Cloud and datacenter computing · 23% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 6 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
query compilation |
1.1 | 3 | 2020 | Architecting a Query Compiler for Spatial Workloads · SIGMOD Conference 2020 Flare & Lantern: Efficiently Swapping Horses Midstream · Proc. VLDB Endow. 2019 How to Architect a Query Compiler, Revisited · SIGMOD Conference 2018 |
Spatial and temporal data management
spatial query processing |
0.7 | 2 | 2020 | Architecting a Query Compiler for Spatial Workloads · SIGMOD Conference 2020 Similarity Group-by Operators for Multi-Dimensional Relational Data · IEEE Trans. Knowl. Data Eng. 2016 |
Machine learning and data management
data management for machine learning |
0.4 | 1 | 2019 | Flare & Lantern: Efficiently Swapping Horses Midstream · Proc. VLDB Endow. 2019 |
Query processing and optimization › query compilation
code generation for query execution |
0.3 | 1 | 2018 | How to Architect a Query Compiler, Revisited · SIGMOD Conference 2018 |
High-performance computing
performance optimization |
0.3 | 1 | 2018 | Flare: Optimizing Apache Spark with Native Compilation for Scale-Up Architectures and Medium-Size Data · OSDI 2018 |
Spatial and temporal data management
spatial indexing |
0.1 | 1 | 2020 | Architecting a Query Compiler for Spatial Workloads · SIGMOD Conference 2020 |
Methods — techniques the papers use, named apart from their topics
runtime compilation · 0.8native code generation · 0.8partial evaluation · 0.4generative programming · 0.4partitioning · 0.2indexing · 0.2distance-to-any grouping · 0.2clique grouping · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | On-stack replacement for program generators and source-to-source compilersabstractOn-stack replacement (OSR) describes the ability to replace currently executing code with a different version, either a more optimized one (tiered execution) or a more general one (deoptimization to undo speculative optimization). While OSR is a key component in all modern VMs for languages like Java or JavaScript, OSR has only recently been studied as a more abstract program transformation, independent of language VMs. Still, previous work has only considered OSR in the context of low-level execution models based on stack frames, labels, and jumps. Grégory M. Essertel, Ruby Y. Tahboub, Tiark Rompf |
GPCE | 2 |
| 2020 | Architecting a Query Compiler for Spatial WorkloadsabstractModern location-based applications rely extensively on the efficient processing of spatial data and queries. Spatial query engines are commonly engineered as an extension to a relational database or a cluster-computing framework. Large parts of the spatial processing runtime is spent on evaluating spatial predicates and traversing spatial indexing structures. Typical high-level implementations of these spatial structures incur significant interpretive overhead, which increases latency and lowers throughput. A promising idea to improve the performance of spatial workloads is to leverage native code generation techniques that have become popular in relational query engines. However, architecting a spatial query compiler is challenging since spatial processing has fundamentally different execution characteristics from relational workloads in terms of data dimensionality, indexing structures, and predicate evaluation. In this paper, we discuss the underlying reasons why standard query compilation techniques are not fully effective when applied to spatial workloads, and we demonstrate how a particular style of query compilation based on techniques borrowed from partial evaluation and generative programming manages to avoid most of these difficulties by extending the scope of custom code generation into the data structures layer. We extend the LB2 main-memory query compiler, a relational engine developed in this style, with spatial data types, predicates, indexing structures, and operators. We show that the spatial extension matches the performance of specialized library code and outperforms relational and map-reduce extensions. Ruby Y. Tahboub, Tiark Rompf |
SIGMOD Conference | 1 |
| 2019 | Flare & Lantern: Efficiently Swapping Horses MidstreamabstractRunning machine learning (ML) workloads at scale is as much a data management problem as a model engineering problem. Big performance challenges exist when data management systems invoke ML classifiers as user-defined functions (UDFs) or when stand-alone ML frameworks interact with data stores for data loading and pre-processing (ETL). In particular, UDFs can be precompiled or simply a black box for the data management system and the data layout may be completely different from the native layout, thus adding overheads at the boundaries. In this demo, we will show how bottlenecks between existing systems can be eliminated when their engines are designed around runtime compilation and native code generation, which is the case for many state-of-the-art relational engines as well as ML frameworks. We demonstrate an integration of Flare (an accelerator for Spark SQL), and Lantern (an accelerator for TensorFlow and PyTorch) that results in a highly optimized end-to-end compiled data path, switching between SQL and ML processing with negligible overhead. Grégory M. Essertel, Ruby Y. Tahboub, Fei Wang 0046, James M. Decker, Tiark Rompf |
Proc. VLDB Endow. | 2 |
| 2018 | Flare: Optimizing Apache Spark with Native Compilation for Scale-Up Architectures and Medium-Size Data
Grégory M. Essertel, Ruby Y. Tahboub, James M. Decker, Kevin J. Brown, Kunle Olukotun, Tiark Rompf |
OSDI | 2 |
| 2018 | How to Architect a Query Compiler, RevisitedabstractTo leverage modern hardware platforms to their fullest, more and more database systems embrace compilation of query plans to native code. In the research community, there is an ongoing debate about the best way to architect such query compilers. This is perceived to be a difficult task, requiring techniques fundamentally different from traditional interpreted query execution. Ruby Y. Tahboub, Grégory M. Essertel, Tiark Rompf |
SIGMOD Conference | 1 |
| 2016 | On supporting compilation in spatial query engines: (vision paper)abstractToday's 'Big' spatial computing and analytics are largely processed in-memory. Still, evaluation in prominent spatial query engines is neither fully optimized for modern-class platforms nor taking full advantage of compilation (i.e., generating low-level query code). Query compilation has been rapidly rising inside in-memory relational database management systems (RDBMSs) achieving remarkable speedups; how can we bring similar benefits to spatial query engines? Ruby Y. Tahboub, Tiark Rompf |
SIGSPATIAL/GIS | 1 |
| 2016 | Similarity Group-By operators for multi-dimensional relational dataabstractThe SQL group-by operator plays an important role in summarizing and aggregating large datasets in a data analytics stack. The Similarity SQL-based Group-By operator (SGB, for short) extends the semantics of the standard SQL Group-by by grouping data with similar but not necessarily equal values. While existing similarity-based grouping operators efficiently realize these approximate semantics, they primarily focus on one-dimensional attributes and treat multi-dimensional attributes independently. However, correlated attributes, such as in spatial data, are processed independently, and hence, groups in the multi-dimensional space are not detected properly. To address this problem, we introduce two new SGB operators for multi-dimensional data. The first operator is the clique (or distance-to-all) SGB, where all the tuples in a group are within some distance from each other. The second operator is the distance-to-any SGB, where a tuple belongs to a group if the tuple is within some distance from any other tuple in the group. Since a tuple may satisfy the membership criterion of multiple groups, we introduce three different semantics to deal with such a case: (i) eliminate the tuple, (ii) put the tuple in any one group, and (iii) create a new group for this tuple. We implement and test the new SGB operators and their algorithms inside PostgreSQL. The overhead introduced by these operators proves to be minimal and the execution times are comparable to those of the standard Group-by. The experimental study, based on TPC-H and a social check-in data, demonstrates that the proposed algorithms can achieve up to three orders of magnitude enhancement in performance over baseline methods developed to solve the same problem. MingJie Tang, Ruby Y. Tahboub, Walid G. Aref, Mikhail J. Atallah, Qutaibah M. Malluhi, Mourad Ouzzani, Yasin N. Silva |
ICDE | 2 |
| 2016 | Similarity Group-by Operators for Multi-Dimensional Relational DataabstractThe SQL group-by operator plays an important role in summarizing and aggregating large datasets in a data analytics stack. While the standard group-by operator, which is based on equality, is useful in several applications, allowing similarity aware grouping provides a more realistic view on real-world data that could lead to better insights. The Similarity SQL-based Group-By operator (SGB, for short) extends the semantics of the standard SQL Group-by by grouping data with similar but not necessarily equal values. While existing similarity-based grouping operators efficiently realize these approximate semantics, they primarily focus on one-dimensional attributes and treat multi-dimensional attributes independently. However, correlated attributes, such as in spatial data, are processed independently, and hence, groups in the multi-dimensional space are not detected properly. To address this problem, we introduce two new SGB operators for multi-dimensional data. The first operator is the clique (or distance-to-all) SGB, where all the tuples in a group are within some distance from each other. The second operator is the distance-to-any SGB, where a tuple belongs to a group if the tuple is within some distance from any other tuple in the group. Since a tuple may satisfy the membership criterion of multiple groups, we introduce three different semantics to deal with such a case: (i) eliminate the tuple, (ii) put the tuple in any one group, and (iii) create a new group for this tuple. We implement and test the new SGB operators and their algorithms inside PostgreSQL. The overhead introduced by these operators proves to be minimal and the execution times are comparable to those of the standard Group-by. The experimental study, based on TPC-H and a social check-in data, demonstrates that the proposed algorithms can achieve up to three orders of magnitude enhancement in performance over baseline methods developed to solve the same problem. MingJie Tang, Ruby Y. Tahboub, Walid G. Aref, Mikhail J. Atallah, Qutaibah M. Malluhi, Mourad Ouzzani, Yasin N. Silva |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | On map-centric programming environments: vision paperabstract2D maps and 3D globes can be used as programming toys to help students learn programming in contrast to using robots, visual, or multimedia components in Computer Science I introductory programming courses. This paper studies research challenges related to supporting this concept, and presents one instance of a 2D and 3D map-based programming environment to motivate these challenges. Walid G. Aref, Sunil Prabhakar 0001, Jaewoo Shin 0001, Ruby Y. Tahboub, Aya Abdelsalam, Jalaleldeen W. Aref |
SIGSPATIAL/GIS | 4 |
| 2015 | LIMO: learning programming using interactive map activitiesabstractAdvances in geographic information, interactive two- and three-dimensional map visualization accompanied with the proliferation of mobile devices and location data have tremendously benefited the development of geo-educational applications. We demonstrate LIMO; a web-based programming environment that is centered around operations on interactive geographical maps, location-oriented data, and the operations of synthetic objects that move on the maps. LIMO materializes a low-cost open-ended environment that integrates interactive maps and spatial data (e.g., OpenStreetMap). The unique advantage of LIMO is that it relates programming concepts to interactive geographical maps and location data. LIMO offers an environment for students to learn how to program by providing: 1. An easy-to-program library of map and spatial operations, 2. High-quality interactive map graphics, and 3. Example programs that introduce users to writing programs in the LIMO environment. Ruby Y. Tahboub, Jaewoo Shin 0001, Aya Abdelsalam, Jalaleldeen W. Aref, Walid G. Aref, Sunil Prabhakar 0001 |
SIGSPATIAL/GIS | 1 |