VLDB 2026 Research / reviewers in the wild / expert
Zunyao Mao
dblp:322/1503
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2024
0009-0006-8516-3723ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Query processing and optimization · 57% Machine learning and data management · 22% Information retrieval · 17% | |
| Artificial intelligence
2 papers |
Graph learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
GPUs and heterogeneous computing · 100% |
Topics — the 9 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs · NeurIPS 2024 |
Information retrieval › evaluation
benchmark |
0.8 | 1 | 2024 | 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBs · NeurIPS 2024 |
Machine learning › Graph learning
graph sampling |
0.7 | 1 | 2023 | gSampler: General and Efficient GPU-based Graph Sampling for Graph Learning · SOSP 2023 |
Query processing and optimization
cardinality estimation |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
Query processing and optimization › cardinality estimation
learned cardinality estimation |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
Query processing and optimization › adaptive query processing
query re-optimization |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
GPUs and heterogeneous computing
GPU graph processing |
0.7 | 1 | 2023 | gSampler: General and Efficient GPU-based Graph Sampling for Graph Learning · SOSP 2023 |
Query processing and optimization › query execution › hardware-accelerated query processing
GPU-accelerated query processing |
0.6 | 1 | 2022 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive · SIGMOD Conference 2022 |
Machine learning and data management
data management for machine learning |
0.2 | 1 | 2023 | gSampler: General and Efficient GPU-based Graph Sampling for Graph Learning · SOSP 2023 |
Methods — techniques the papers use, named apart from their topics
matrix-centric API · 2.0extract-compute-select-finalize model · 2.0data-flow intermediate representation · 2.0subsampling · 1.5graph neural network · 1.5refinement model · 0.7query-driven estimation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 4DBInfer: A 4D Benchmarking Toolbox for Graph-Centric Predictive Modeling on RDBsabstractGiven a relational database (RDB), how can we predict missing column values in some target table of interest? Although RDBs store vast amounts of rich, informative data spread across interconnected tables, the progress of predictive machine learning models as applied to such tasks arguably falls well behind advances in other domains such as computer vision or natural language processing. This deficit stems, at least in part, from the lack of established/public RDB benchmarks as needed for training and evaluation purposes. As a result, related model development thus far often defaults to tabular approaches trained on ubiquitous single-table benchmarks, or on the relational side, graph-based alternatives such as GNNs applied to a completely different set of graph datasets devoid of tabular characteristics. To more precisely target RDBs lying at the nexus of these two complementary regimes, we explore a broad class of baseline models predicated on: (i) converting multi-table datasets into graphs using various strategies equipped with efficient subsampling, while preserving tabular characteristics; and (ii) trainable models with well-matched inductive biases that output predictions based on these input subgraphs. Then, to address the dearth of suitable public benchmarks and reduce siloed comparisons, we assemble a diverse collection of (i) large-scale RDB datasets and (ii) coincident predictive tasks. From a delivery standpoint, we operationalize the above four dimensions (4D) of exploration within a unified, scalable open-source toolbox called 4DBInfer; please see https://github.com/awslabs/multi-table-benchmark . David P. Wipf, Zheng Zhang 0001, Christos Faloutsos, Weinan Zhang 0001, Muhan Zhang, Zhenkun Cai, Jiahang Li 0002, Zunyao Mao, Yakun Song, Yanlin Zhang, Chuan Lei, Xiao Qin 0003, Ning Li 0029, Han Zhang 0057 |
NeurIPS | 10 |
| 2023 | gSampler: General and Efficient GPU-based Graph Sampling for Graph LearningabstractGraph sampling prepares training samples for graph learning and can dominate the training time. Due to the increasing algorithm diversity and complexity, existing sampling frameworks are insufficient in the generality of expression and the efficiency of execution. To close this gap, we conduct a comprehensive study on 15 popular graph sampling algorithms to motivate the design of gSampler, a general and efficient GPU-based graph sampling framework. gSampler models graph sampling using a general 4-step Extract-Compute-Select-Finalize (ECSF) programming model, proposes a set of matrix-centric APIs that allow to easily express complex graph sampling algorithms, and incorporates a data-flow intermediate representation (IR) that translates high-level API codes for efficient GPU execution. We demonstrate that implementing graph sampling algorithms with gSampler is easy and intuitive. We also conduct extensive experiments with 7 algorithms, 4 graph datasets, and 2 hardware configurations. The results show that gSampler introduces sampling speedups of 1.14--32.7× and an average speedup of 6.54×, compared to state-of-the-art GPU-based graph sampling systems such as DGL, which translates into an overall time reduction of over 40% for graph learning. gSampler is open-source at https://tinyurl.com/29twthd4. Ping Gong 0009, Renjie Liu 0001, Zunyao Mao, Zhenkun Cai, Xiao Yan 0002, Cheng Li 0001, Zhuozhao Li |
SOSP | 3 |
| 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality EstimationabstractFast query execution requires learning-based cardinality estimators to have short inference time (as model inference time adds to end-to-end query execution time) and high estimation accuracy (which is crucial for finding good execution plan). However, existing estimators cannot meet both requirements due to the inherent tension between model complexity and estimation accuracy. We propose a novel Learning-based Progressive Cardinality Estimator (LPCE), which adopts a query re-optimization methodology. In particular, LPCE consists of an initial model (LPCE-I), which estimates cardinality before query execution, and a refinement model (LPCE-R), which progressively refines the cardinality estimations using the actual cardinalities of the executed operators. During query execution, re-optimization is triggered if the estimations of LPCE-I are found to have large errors, and more efficient execution plans are selected for the remaining operators using the refined estimations provided by LPCE-R. Both LPCE-I and LPCE-R are light-weight query-driven estimators but they achieve both good efficiency and high accuracy when used jointly. Besides designing the models for LPCE-I and LPCE-R, we also integrate re-optimization and LPCE into PostgreSQL, a popular database engine. Extensive experiments show that LPCE yields shorter end-to-end query execution time than state-of-the-art learning-based estimators. Fang Wang 0012, Xiao Yan 0002, Man Lung Yiu, Shuai Li 0014, Zunyao Mao, Bo Tang 0016 |
Proc. ACM Manag. Data | 5 |
| 2022 | GHive: accelerating analytical query processing in apache hive via CPU-GPU heterogeneous computingabstractAs a popular distributed data warehouse system, Apache Hive has been widely used for big data analytics in many organizations. Meanwhile, exploiting the massive parallelism of GPU to accelerate online analytical processing (OLAP) has been extensively explored in the database community. In this paper, we present GHive, which enhances CPU-based Hive via CPU-GPU heterogeneous computing. GHive is designed for the business intelligence applications and provides the same API as Hive for compatibility. To run SQL queries jointly on both CPU and GPU, GHive comes with three key techniques: (i) a novel data model gTable, which is column-based and enables efficient data movement between CPU memory and GPU memory; (ii) a GPU-based operator library Panda, which provides a complete set of SQL operators with extensively optimized GPU implementations; (iii) a hardware-aware MapReduce job placement scheme, which puts jobs judiciously on either GPU or CPU via a cost-based approach. In the experiments, we observe that GHive outperforms Hive in both query processing speed and operating expense on the Star Schema Benchmark (SSB). Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xiao Yan 0002, Xinying Zheng, Qiaomu Shen, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo |
SoCC | 9 |
| 2022 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache HiveabstractAs a distributed, fault-tolerant data warehouse system for large-scale data analytics, Apache Hive has been used for various applications in many organizations (e.g., Facebook, Amazon, and Huawei). Exploiting the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database system is a common practice in the industry. Meanwhile, it is a common practice to exploit the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database systems. This demo presents GHive, which enables Apache Hive to accelerate OLAP queries by jointly utilizing CPU and GPU in intelligent and efficient ways. The takeaways for SIGMOD attendees include: (1) the superior performance of GHive compared with vanilla Hive that only uses CPU; (2) intuitive visualizations of execution statistics for Hive and GHive to understand where the acceleration of GHive comes from; (3) detailed profiling of the time taken by each operator on CPU and GPU to show the advantages of GPU execution. Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xinying Zheng, Qiaomu Shen, Xiao Yan 0002, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo |
SIGMOD Conference | 9 |