Zizhong Meng

dblp:227/6807 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2023
0009-0007-2702-9304ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 87% Machine learning and data management · 13%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
cardinality estimation
0.712023
Selectivity Estimation for Queries Containing Predicates over Set-Valued Attributes · Proc. ACM Manag. Data 2023
Query processing and optimization
selectivity estimation
0.712023
Selectivity Estimation for Queries Containing Predicates over Set-Valued Attributes · Proc. ACM Manag. Data 2023
Machine learning and data management
learned database components
0.212023
Selectivity Estimation for Queries Containing Predicates over Set-Valued Attributes · Proc. ACM Manag. Data 2023

Methods — techniques the papers use, named apart from their topics

query conversion · 0.7column factorization · 0.7
YearPublicationVenuePosition
2023 Selectivity Estimation for Queries Containing Predicates over Set-Valued Attributes
abstract
Selectivity estimation aims to estimate the size of query results size accurately and efficiently. Despite being a well-researched area for decades, most existing estimators are designed to handle comparison predicates over numeric and categorical data. Nevertheless, Set-valued data are ubiquitous in many applications such as information retrieval and recommendation systems. However, these estimators may not be effective for handling predicates over set-valued data. In this work, we presents novel techniques for selectivity estimation on queries involving predicates over set-valued attributes. We first propose the set-valued column factorization problem, whereby each each set-valued column is converted to multiple numeric subcolumns, and set containment predicates are converted to numeric comparison predicates. This enables us to leverage any existing estimator to perform selectivity estimation. We then develop two methods for column factorization and query conversion, namely ST and STH. We integrate ST and STH into three estimators, Postgres, Neurocard, and DeepDB. We then conduct a comprehensive empirical analysis by comparing our approach against three baselines across three different datasets. The experimental results demonstrate that our methods exhibit superior estimation accuracy while maintaining high efficiency compared to the baseline techniques.
Zizhong Meng, Xin Cao 0001, Gao Cong
Proc. ACM Manag. Data1
2022 Unsupervised Selectivity Estimation by Integrating Gaussian Mixture Models and an Autoregressive Model
Zizhong Meng, Peizhi Wu, Gao Cong, Shuai Ma 0001
EDBT1