EDBT 2026 Demo / reviewers in the wild / expert
Zekai J. Gao
dblp:147/1259
· DBLP profile ↗
7ranked-venue papers
3as first author
0since 2021 · last 2019
0000-0002-1886-1912ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Query processing and optimization · 38% Database system architecture and tuning · 31% Machine learning and data management · 28% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 54% Performance modeling and evaluation · 36% Cloud and datacenter computing · 11% |
Topics — the 11 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query optimization
cost-based optimization |
0.7 | 2 | 2019 | Scalable Linear Algebra on a Relational Database System · IEEE Trans. Knowl. Data Eng. 2019 Scalable Linear Algebra on a Relational Database System · ICDE 2017 |
Machine learning and data management
linear algebra over relational data |
0.4 | 1 | 2019 | Scalable Linear Algebra on a Relational Database System · IEEE Trans. Knowl. Data Eng. 2019 |
Query processing and optimization
query optimization |
0.4 | 1 | 2019 | Declarative Parameterizations of User-Defined Functions for Large-Scale Machine Learning and Optimization · IEEE Trans. Knowl. Data Eng. 2019 |
Database system architecture and tuning
relational database system |
0.4 | 1 | 2019 | Scalable Linear Algebra on a Relational Database System · IEEE Trans. Knowl. Data Eng. 2019 |
Distributed systems
distributed machine learning |
0.3 | 1 | 2017 | The BUDS Language for Distributed Bayesian Machine Learning · SIGMOD Conference 2017 |
Machine learning and data management › scalable machine learning
distributed learning |
0.2 | 1 | 2014 | A comparison of platforms for implementing and running very large scale machine learning algorithms · SIGMOD Conference 2014 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2014 | A comparison of platforms for implementing and running very large scale machine learning algorithms · SIGMOD Conference 2014 |
Machine learning and data management
data management for machine learning |
0.1 | 1 | 2019 | Declarative Recursive Computation on an RDBMS · Proc. VLDB Endow. 2019 |
Database system architecture and tuning
parallel database system |
0.1 | 1 | 2019 | Scalable Linear Algebra on a Relational Database System · IEEE Trans. Knowl. Data Eng. 2019 |
Compilers and program optimization
domain-specific compilation |
0.1 | 1 | 2017 | The BUDS Language for Distributed Bayesian Machine Learning · SIGMOD Conference 2017 |
Cloud and datacenter computing
cloud platform |
0.1 | 1 | 2014 | A comparison of platforms for implementing and running very large scale machine learning algorithms · SIGMOD Conference 2014 |
Methods — techniques the papers use, named apart from their topics
bayesian machine learning · 0.6recursive query · 0.4distributed query execution · 0.4dataflow operators · 0.4cost-based optimization · 0.4SimSQL · 0.4SQL · 0.4hierarchical model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Declarative Recursive Computation on an RDBMSabstractA number of popular systems, most notably Google's TensorFlow, have been implemented from the ground up to support machine learning tasks. We consider how to make a very small set of changes to a modern relational database management system (RDBMS) to make it suitable for distributed learning computations. Changes include adding better support for recursion, and optimization and execution of very large compute plans. We also show that there are key advantages to using an RDBMS as a machine learning platform. In particular, learning based on a database management system allows for trivial scaling to large data sets and especially large models, where different computational units operate on different parts of a model that may be too large to fit into RAM. Dimitrije Jankov, Shangyu Luo, Binhang Yuan, Zhuhua Cai, Jia Zou 0001, Chris Jermaine, Zekai J. Gao |
Proc. VLDB Endow. | 7 |
| 2019 | Declarative Parameterizations of User-Defined Functions for Large-Scale Machine Learning and OptimizationabstractLarge-scale optimization has become an important application for data management systems, particularly in the context of statistical machine learning. In this paper, we consider how one might implement the join-and-co-group pattern in the context of a fully declarative data processing system. The join-and-co-group pattern is ubiquitous in iterative, large-scale optimization. In the join-and-co-group pattern, a user-defined function g is parameterized with a data object x as well as the subset of the statistical model Θxthat applies to that object, so that g(x|Θx) can be used to compute a partial update of the model. This is repeated for every x in the full data set X. All partial updates are then aggregated and used to perform a complete update of the model. The join-and-co-group pattern has several implementation challenges, including the potential for a massive blow-up in the size of a fully parameterized model. Thus, unless the correct physical execution plan be chosen for implementing the join-and-co-group pattern, it is easily possible to have an execution that takes a very long time or even fails to complete. In this paper, we carefully consider the alternatives for implementing the join-and-co-group pattern on top of a declarative system, as well as how the best alternative can be selected automatically. Our focus is on the SimSQL database system, which is an SQL-based system with special facilities for large-scale, iterative optimization. Since it is an SQL-based system with a query optimizer, those choices can be made automatically. Zekai J. Gao, Niketan Pansare, Chris Jermaine |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Scalable Linear Algebra on a Relational Database SystemabstractAs data analytics has become an important application for modern data management systems, a new category of data management system has appeared recently: the scalable linear algebra system. In this paper, we argue that a parallel or distributed database system is actually an excellent platform upon which to build such functionality. Most relational systems already have support for cost-based optimization-which is vital to scaling linear algebra computations-and it is well-known how to make relational systems scale. We show that by making just a few changes to a parallel/distributed relational database system, such a system can be a competitive platform for scalable linear algebra. Taken together, our results should at least raise the possibility that brand new systems designed from the ground up to support scalable linear algebra are not absolutely necessary, and that such systems could instead be built on top of existing relational technology. Our results also suggest that if scalable linear algebra is to be added to a modern dataflow platform such as Spark, they should be added on top of the system's more structured (relational) data abstractions, rather than being constructed directly on top of the system's raw dataflow operators. Shangyu Luo, Zekai J. Gao, Michael N. Gubanov, Luis Leopoldo Perez, Chris Jermaine |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Scalable Linear Algebra on a Relational Database SystemabstractAs data analytics has become an important application for modern data management systems, a new category of data management system has appeared recently: the scalable linear algebra system. In this paper, we argue that a parallel or distributed database system is actually an excellent platform upon which to build such functionality. Most relational systems already have support for cost-based optimization-which is vital to scaling linear algebra computations-and it is well-known how to make relational systems scale. We show that by making just a few changes to a parallel/ distributed relational database system, such a system can be a competitive platform for scalable linear algebra. Taken together, our results should at least raise the possibility that brand new systems designed from the ground up to support scalable linear algebra are not absolutely necessary, and that such systems could instead be built on top of existing relational technology. Our results also suggest that if scalable linear algebra is to be added to a modern dataflow platform such as Spark, they should be added on top of the system's more structured (relational) data abstractions, rather than being constructed directly on top of the system's raw dataflow operators. Shangyu Luo, Zekai J. Gao, Michael N. Gubanov, Luis Leopoldo Perez, Chris Jermaine |
ICDE | 2 |
| 2017 | The BUDS Language for Distributed Bayesian Machine LearningabstractWe describe BUDS, a declarative language for succinctly and simply specifying the implementation of large-scale machine learning algorithms on a distributed computing platform. The types supported in BUDS--vectors, arrays, etc.--are simply logical abstractions useful for programming, and do not correspond to the actual implementation. In fact, BUDS automatically chooses the physical realization of these abstractions in a distributed system, by taking into account the characteristics of the data. Likewise, there are many available implementations of the abstract operations offered by BUDS (matrix multiplies, transposes, Hadamard products, etc.). These are tightly coupled with the physical representation. In BUDS, these implementations are co-optimized along with the representation. All of this allows for the BUDS compiler to automatically perform deep optimizations of the user's program, and automatically generate efficient implementations. Zekai J. Gao, Shangyu Luo, Luis Leopoldo Perez, Chris Jermaine |
SIGMOD Conference | 1 |
| 2016 | Distributed Algorithms for Computing Very Large Thresholded Covariance MatricesabstractComputation of covariance matrices from observed data is an important problem, as such matrices are used in applications such as principal component analysis (PCA), linear discriminant analysis (LDA), and increasingly in the learning and application of probabilistic graphical models. However, computing an empirical covariance matrix is not always an easy problem. There are two key difficulties associated with computing such a matrix from a very high-dimensional dataset. The first problem is over-fitting. For a p -dimensional covariance matrix, there are p ( p − 1)/2 unique, off-diagonal entries in the empirical covariance matrix Ŝ for large p (say, p > 10 5 ), the size n of the dataset is often much smaller than the number of covariances to compute. Over-fitting is a concern in any situation in which the number of parameters learned can greatly exceed the size of the dataset. Thus, there are strong theoretical reasons to expect that for high-dimensional data—even Gaussian data—the empirical covariance matrix is not a good estimate for the true covariance matrix underlying the generative process. The second problem is computational. Computing a covariance matrix takes O ( np 2 ) time. For large p (greater than 10,000) and n much greater than p , this is debilitating. In this article, we consider how both of these difficulties can be handled simultaneously. Specifically, a key regularization technique for high-dimensional covariance estimation is thresholding , in which the smallest or least significant entries in the covariance matrix are simply dropped and replaced with the value 0. This suggests an obvious way to address the computational difficulty as well: First, compute the identities of the K entries in the covariance matrix that are actually important in the sense that they will not be removed during thresholding, and then in a second step, compute the values of those entries. This can be done in O ( Kn ) time. If K ≪ p 2 and the identities of the important entries can be computed in reasonable time, then this is a big win. The key technical contribution of this article is the design and implementation of two different distributed algorithms for approximating the identities of the important entries quickly, using sampling. We have implemented these methods and tested them using an 800-core compute cluster. Experiments have been run using real datasets having millions of data points and up to 40, 000 dimensions. These experiments show that the proposed methods are both accurate and efficient. Zekai J. Gao, Chris Jermaine |
ACM Trans. Knowl. Discov. Data | 1 |
| 2014 | A comparison of platforms for implementing and running very large scale machine learning algorithmsabstractWe describe an extensive benchmark of platforms available to a user who wants to run a machine learning (ML) inference algorithm over a very large data set, but cannot find an existing implementation and thus must "roll her own" ML code. We have carefully chosen a set of five ML implementation tasks that involve learning relatively complex, hierarchical models. We completed those tasks on four different computational platforms, and using 70,000 hours of Amazon EC2 compute time, we carefully compared running times, tuning requirements, and ease-of-programming of each. Zhuhua Cai, Zekai J. Gao, Shangyu Luo, Luis Leopoldo Perez, Zografoula Vagena, Chris Jermaine |
SIGMOD Conference | 2 |