EDBT 2026 Demo / reviewers in the wild / expert
Xiening Dai
dblp:313/0146
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Query processing and optimization · 49% Data stream processing · 32% Machine learning and data management · 14% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › cardinality estimation
distinct element counting |
1.1 | 2 | 2022 | Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022 Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021 |
Query processing and optimization
cardinality estimation |
0.7 | 2 | 2022 | Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021 Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022 |
Data stream processing
frequency estimation |
0.6 | 1 | 2022 | Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022 |
Data stream processing › sketch
sketch-based estimation |
0.6 | 1 | 2022 | Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022 |
Machine learning and data management
learned database components |
0.5 | 1 | 2021 | Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021 |
Data integration and cleaning
data profiling |
0.1 | 1 | 2021 | Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021 |
Methods — techniques the papers use, named apart from their topics
sketching · 0.6sampling · 0.6supervised learning · 0.5maximum likelihood estimation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Sampling-based Estimation of the Number of Distinct Values in Distributed EnvironmentabstractIn data mining, estimating the number of distinct values (NDV) is a fundamental problem with various applications. Existing methods for estimating NDV can be broadly classified into two categories: i) scanning-based methods, which scan the entire data and maintain a sketch to approximate NDV; and ii) sampling-based methods, which estimate NDV using sampling data rather than accessing the entire data warehouse. Scanning-based methods achieve a lower approximation error at the cost of higher I/O and more time. Sampling-based estimation is preferable in applications with a large data volume and a permissible error restriction due to its higher scalability. However, while the sampling-based method is more effective on a single machine, it is less practical in a distributed environment with massive data volumes. For obtaining the final NDV estimators, the entire sample must be transferred throughout the distributed system, incurring a prohibitive communication cost when the sample rate is significant. This paper proposes a novel sketch-based distributed method that achieves sub-linear communication costs for distributed sampling-based NDV estimation under mild assumptions. Our method leverages a sketch-based algorithm to estimate the sample's frequency of frequency in the distributed streaming model, which is compatible with most classical sampling-based NDV estimators. Additionally, we provide theoretical evidence for our method's ability to minimize communication costs in the worst-case scenario. Extensive experiments show that our method saves orders of magnitude in communication costs compared to existing sampling- and sketch-based methods. Zhewei Wei, Bolin Ding, Xiening Dai, Jingren Zhou 0001 |
KDD | 4 |
| 2021 | Learning to be a Statistician: Learned Estimator for Number of Distinct ValuesabstractEstimating the number of distinct values (NDV) in a column is useful for many tasks in database systems, such as columnstore compression and data profiling. In this work, we focus on how to derive accurate NDV estimations from random (online/offline) samples. Such efficient estimation is critical for tasks where it is prohibitive to scan the data even once. Existing sample-based estimators typically rely on heuristics or assumptions and do not have robust performance across different datasets as the assumptions on data can easily break. On the other hand, deriving an estimator from a principled formulation such as maximum likelihood estimation is very challenging due to the complex structure of the formulation. We propose to formulate the NDV estimation task in a supervised learning framework, and aim to learn a model as the estimator. To this end, we need to answer several questions: i) how to make the learned model workload agnostic; ii) how to obtain training data; iii) how to perform model training. We derive conditions of the learning framework under which the learned model isworkload agnostic, in the sense that the model/estimator can be trained with synthetically generated training data, and then deployed into any data warehouse simply as,e.g., user-defined functions (UDFs), to offer efficient (within microseconds on CPU) and accurate NDV estimations forunseen tables and workloads.We compare the learned estimator with the state-of-the-art sample-based estimators on nine real-world datasets to demonstrate its superior estimation accuracy. We publish our code for training data generation, model training, and the learned estimator online for reproducibility. Renzhi Wu, Bolin Ding, Xu Chu 0002, Zhewei Wei, Xiening Dai, Jingren Zhou 0001 |
Proc. VLDB Endow. | 5 |