Dake Zhong

dblp:358/4613 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 56% Data mining · 44%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
cost estimation
0.812024
QCFE: An Efficient Feature Engineering for Query Cost Estimation · ICDE 2024
Data mining
feature engineering
0.812024
QCFE: An Efficient Feature Engineering for Query Cost Estimation · ICDE 2024
Query processing and optimization › query planning
query plan representation
0.212024
QCFE: An Efficient Feature Engineering for Query Cost Estimation · ICDE 2024

Methods — techniques the papers use, named apart from their topics

feature reduction · 0.8difference propagation · 0.8
YearPublicationVenuePosition
2024 QCFE: An Efficient Feature Engineering for Query Cost Estimation
abstract
Query cost estimation is a classical task for database management. Recently, researchers have applied AI-driven methods to implement query cost estimation for achieving high accuracy. However, two defects of the feature design lead to poor time-accuracy efficiency in the query cost estimation task. On the one hand, existing works only encode the query plan and data statistics while ignoring some important variables, like storage structure, hardware, database knobs, etc. These variables also have a significant impact on the query cost. On the other hand, existing works suffer the heavy model training and model inference due to inefficient features, such as the index encoding of write-only workloads. To address the above two problems, we first propose an efficient feature engineering for query cost estimation, called QCFE, consisting of the feature snapshot and feature reduction algorithm. (1) We design a novel concept called feature snapshot to efficiently integrate the influences of the missing variables. (2) We propose a difference-propagation feature reduction method for query cost estimation to filter the ineffective features. Compared to state-of-the-art methods, QCFE demonstrates significant improvements in various aspects with well-known benchmarks. QCFE saves up to 50% time consumption for model training, resulting in more efficient and faster training processes. QCFE also optimizes the mean q-error by 19.8% in TPCH, leading to more precise query cost estimation. QCFE offers up to an impressive 8 times inference speedup in query inference throughput.
Hongzhi Wang 0001, Junfang Huang, Dake Zhong
ICDE4