Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xiening Dai

dblp:313/0146 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 49% Data stream processing · 32% Machine learning and data management · 14%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › cardinality estimation
distinct element counting
1.122022
Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022
Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021
Query processing and optimization
cardinality estimation
0.722022
Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021
Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022
Data stream processing
frequency estimation
0.612022
Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022
Data stream processing › sketch
sketch-based estimation
0.612022
Sampling-based Estimation of the Number of Distinct Values in Distributed Environment · KDD 2022
Machine learning and data management
learned database components
0.512021
Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021
Data integration and cleaning
data profiling
0.112021
Learning to be a Statistician: Learned Estimator for Number of Distinct Values · Proc. VLDB Endow. 2021

Methods — techniques the papers use, named apart from their topics

sketching · 0.6sampling · 0.6supervised learning · 0.5maximum likelihood estimation · 0.5
YearPublicationVenuePosition
2022 Sampling-based Estimation of the Number of Distinct Values in Distributed Environment
abstract
In data mining, estimating the number of distinct values (NDV) is a fundamental problem with various applications. Existing methods for estimating NDV can be broadly classified into two categories: i) scanning-based methods, which scan the entire data and maintain a sketch to approximate NDV; and ii) sampling-based methods, which estimate NDV using sampling data rather than accessing the entire data warehouse. Scanning-based methods achieve a lower approximation error at the cost of higher I/O and more time. Sampling-based estimation is preferable in applications with a large data volume and a permissible error restriction due to its higher scalability. However, while the sampling-based method is more effective on a single machine, it is less practical in a distributed environment with massive data volumes. For obtaining the final NDV estimators, the entire sample must be transferred throughout the distributed system, incurring a prohibitive communication cost when the sample rate is significant. This paper proposes a novel sketch-based distributed method that achieves sub-linear communication costs for distributed sampling-based NDV estimation under mild assumptions. Our method leverages a sketch-based algorithm to estimate the sample's frequency of frequency in the distributed streaming model, which is compatible with most classical sampling-based NDV estimators. Additionally, we provide theoretical evidence for our method's ability to minimize communication costs in the worst-case scenario. Extensive experiments show that our method saves orders of magnitude in communication costs compared to existing sampling- and sketch-based methods.
Zhewei Wei, Bolin Ding, Xiening Dai, Jingren Zhou 0001
KDD4
2021 Learning to be a Statistician: Learned Estimator for Number of Distinct Values
abstract
Estimating the number of distinct values (NDV) in a column is useful for many tasks in database systems, such as columnstore compression and data profiling. In this work, we focus on how to derive accurate NDV estimations from random (online/offline) samples. Such efficient estimation is critical for tasks where it is prohibitive to scan the data even once. Existing sample-based estimators typically rely on heuristics or assumptions and do not have robust performance across different datasets as the assumptions on data can easily break. On the other hand, deriving an estimator from a principled formulation such as maximum likelihood estimation is very challenging due to the complex structure of the formulation. We propose to formulate the NDV estimation task in a supervised learning framework, and aim to learn a model as the estimator. To this end, we need to answer several questions: i) how to make the learned model workload agnostic; ii) how to obtain training data; iii) how to perform model training. We derive conditions of the learning framework under which the learned model isworkload agnostic, in the sense that the model/estimator can be trained with synthetically generated training data, and then deployed into any data warehouse simply as,e.g., user-defined functions (UDFs), to offer efficient (within microseconds on CPU) and accurate NDV estimations forunseen tables and workloads.We compare the learned estimator with the state-of-the-art sample-based estimators on nine real-world datasets to demonstrate its superior estimation accuracy. We publish our code for training data generation, model training, and the learned estimator online for reproducibility.
Renzhi Wu, Bolin Ding, Xu Chu 0002, Zhewei Wei, Xiening Dai, Jingren Zhou 0001
Proc. VLDB Endow.5