EDBT 2026 Demo / reviewers in the wild / expert
Zezhou Huang
dblp:287/9765
· DBLP profile ↗
in reviewer pool
← Back
8ranked-venue papers in the field
7as first author
8since 2021 · last 2026
0009-0002-6613-0337ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 8 (7 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A decade of systems for human data interaction
Eugene Wu 0002, Haneen Mohammed, Zezhou Huang |
Inf. Syst. | 4 |
| 2025 | GPU Acceleration of SQL Analytics on Compressed Data
Zezhou Huang, Krystian Sakowski, Hans Lehnert, Carlo Curino, Matteo Interlandi, Marius Dumitru, Rathijit Sen |
Proc. VLDB Endow. | 1 |
| 2024 | The Fast and the Private: Task-based Dataset Search
Zezhou Huang, Eugene Wu 0002 |
CIDR | 1 |
| 2023 | Random Forests over normalized data in CPU-GPU DBMSesabstractThis short paper studies query execution based on message passing on CPU-GPU systems, using random forests training as the workload. We investigate different data placement and query execution strategies and find that the unique properties of training ML models using message passing necessitates different design decisions. We show that with proper data placement and CPU-GPU co-execution, training random forest models using pure SQL can outperform the leading LightGBM ML library by 1.5 × on SSB SF=10. Zezhou Huang, Pavan Kalyan Damalapati, Rathijit Sen, Eugene Wu 0002 |
DaMoN | 1 |
| 2023 | Lightweight Materialization for Fast Dashboards Over JoinsabstractDashboards are vital in modern business intelligence tools, providing non-technical users with an interface to access comprehensive business data. With the rise of cloud technology, there is an increased number of data sources to provide enriched contexts for various analytical tasks, leading to a demand for interactive dashboards over a large number of joins. Nevertheless, joins are among the most expensive operations in DBMSes, making the support of interactive dashboards over joins challenging. In this paper, we present Treant, a dashboard accelerator for queries over large joins. Treant uses factorized query execution to handle aggregation queries over large joins, which alone is still insufficient for interactive speeds. To address this, we exploit the incremental nature of user interactions using Calibrated Junction Hypertree (CJT), a novel data structure that applies lightweight materialization of the intermediates during factorized execution. CJT ensures that the work needed to compute a query is proportional to how different it is from the previous query, rather than the overall complexity. Treant manages CJTs to share work between queries and performs materialization offline or during user "think-times." Implemented as a middleware that rewrites SQL, Treant is portable to any SQL-based DBMS. Our experiments on single node and cloud DBMSes show that Treant improves dashboard interactions by two orders of magnitude, and provides 10x improvement for ML augmentation compared to SOTA factorized ML system. Zezhou Huang, Eugene Wu 0002 |
Proc. ACM Manag. Data | 1 |
| 2023 | Saibot: A Differentially Private Data Search PlatformabstractRecent data search platforms use ML task-based utility measures rather than metadata-based keywords, to search large dataset corpora. Requesters submit a training dataset, and these platforms search for augmentations ---join or union-compatible datasets---that, when used to augment the requester's dataset, most improve model (e.g., linear regression) performance. Although effective, providers that manage personally identifiable data demand differential privacy (DP) guarantees before granting these platforms data access. Unfortunately, making data search differentially private is nontrivial, as a single search can involve training and evaluating datasets hundreds or thousands of times, quickly depleting privacy budgets. We present Saibot , a differentially private data search platform that employs Factorized Privacy Mechanism (FPM), a novel DP mechanism, to calculate sufficient semi-ring statistics for ML over different combinations of datasets. These statistics are privatized once, and can be freely reused for the search. This allows Saibot to scale to arbitrary numbers of datasets and requests, while minimizing the amount that DP noise affects search results. We optimize the sensitivity of FPM for common augmentation operations, and analyze its properties with respect to linear regression. Specifically, we develop an unbiased estimator for many-to-many joins, prove its bounds, and develop an optimization to redistribute DP noise to minimize the impact on the model. Our evaluation on a real-world dataset corpus of 329 datasets demonstrates that Saibot can return augmentations that achieve model accuracy within 50--90% of non-private search, while the leading alternative DP mechanisms (TPM, APM, shuffling) are several orders of magnitude worse. Zezhou Huang, Daniel Alabi, Raul Castro Fernandez, Eugene Wu 0002 |
Proc. VLDB Endow. | 1 |
| 2023 | JoinBoost: Grow Trees Over Normalized Data Using Only SQLabstractAlthough dominant for tabular data, ML libraries that train tree models over normalized databases (e.g., LightGBM, XGBoost) require the data to be denormalized as a single table, materialized, and exported. This process is not scalable, slow, and poses security risks. In-DB ML aims to train models within DBMSes to avoid data movement and provide data governance. Rather than modify a DBMS to support In-DB ML, is it possible to offer competitive tree training performance to specialized ML libraries...with only SQL? We present JoinBoost, a Python library that rewrites tree training algorithms over normalized databases into pure SQL. It is portable to any DBMS, offers performance competitive with specialized ML libraries, and scales with the underlying DBMS capabilities. JoinBoost extends prior work from both algorithmic and systems perspectives. Algorithmically, we support factorized gradient boosting, by updating the Y variable to the residual in the non-materialized join result. Although this view update problem is generally ambiguous, we identify addition-to-multiplication preserving , the key property of variance semi-ring to support rmse the most widely used criterion. System-wise, we identify residual updates as a performance bottleneck. Such overhead can be natively minimized on columnar DBMSes by creating a new column of residual values and adding it as a projection. We validate this with two implementations on DuckDB, with no or minimal modifications to its internals for portability. Our experiment shows that JoinBoost is 3× (1.1×) faster for random forests (gradient boosting) compared to LightGBM, and over an order of magnitude faster than state-of-the-art In-DB ML systems. Further, JoinBoost scales well beyond LightGBM in terms of the # features, DB size (TPC-DS SF=1000), and join graph complexity (galaxy schemas). Zezhou Huang, Rathijit Sen, Eugene Wu 0002 |
Proc. VLDB Endow. | 1 |
| 2022 | Reptile: Aggregation-level Explanations for Hierarchical DataabstractUsers often can see from overview-level statistics that some results look "off", but are rarely able to characterize even the type of error. Reptile is an iterative human-in-the-loop explanation and cleaning system for errors in hierarchical data. Users specify an anomalous distributive aggregation result (a complaint), and Reptile recommends drill-down operations to help the user "zoom-in" on the underlying errors. Unlike prior explanation systems that intervene on raw records, Reptile intervenes by learning a group's expected statistics, and ranks drill-down sub-groups by how much the intervention fixes the complaint. This group-level formulation supports a wide range of error types (missing, duplicates, value errors) and uniquely leverages the distributive properties of the user complaint. Further, the learning-based intervention lets users provide domain expertise that Reptile learns from. Zezhou Huang, Eugene Wu 0002 |
SIGMOD Conference | 1 |