Shuhuan Fan

dblp:299/2523 · also Shu Huan Fan · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0003-9191-1968ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Faper: Join Tree with Uncertainty Awareness for Faster, More Precise and Robust Cardinality Estimation
Junxin Zhu, Jincan Xiong, Shuhuan Fan, Mengshu Hou
PAKDD (1)5
2024 Precision Meets Resilience: Cross-Database Generalization with Uncertainty Quantification for Robust Cost Estimation
abstract
Learning-based models have shown promise in addressing query optimization challenges in the database field, where the learned cost model plays a central role. While these models outperform traditional optimizers on static datasets, their resilience and reliability in real-world applications remain a concern, limiting their widespread adoption. In this paper, we take a step towards a practical cost estimation model, named Tosure, which can quantify the uncerT ainty for cost estimation and generalizes to unseen databases accurately and efficiently. It consists primarily of two modules: a Cross-Database Representation (CDR) module and a Cost Estimation with Uncertainty (CEU) module. The CDR module captures the transferable features by focusing the minimal set based on deep-learning network, thereby enhancing the model's generalization capabilities. The CEU module introduces a novel Neural Network Gaussian Process (NNGP) to quantify the uncertainty in cost estimation, ensuring more robust estimations with an upper bound. To improve the model's performance, we perform pre-training on diverse large-scale datasets. Furthermore, we implement the model and integrate it with traditional query optimizer to validate its usability and effectiveness in real-world scenarios. Extensive experimentation demonstrates that Tosure outperforms state-of-the-art methods, achieving a 20% improvement in cost estimation accuracy and twice of the robustness.
Shuhuan Fan, Mengshu Hou, Wenwen Ma
CIKM1
2024 MLETune: Streamlining Database Knob Tuning via Multi-LLMs Experts Guided Deep Reinforcement Learning
abstract
Automatic knob tuning has emerged as a critical field of study within database optimization, focusing on simplifying the configuration of database parameters to boost performance, particularly in the realm of advanced modern database management systems with myriad adjustable knobs. The primary challenge revolves around identifying the ideal knob configurations that can markedly enhance system efficiency. Various machine learning techniques have been devised to automate this tuning process. Nevertheless, these methods frequently entail running extensive workloads, resulting in significant time and resource consumption. This inefficiency arises from their reliance on runtime feedback or the limited exploitation of domain knowledge.To overcome these limitations, we propose MLETune, a novel deep reinforcement learning-based approach guided by multi-Large Language Models (LLMs) experts. Our method leverages a remix retrieval-augmented generation algorithm to harness knowledge and distill expert guidance effectively. Additionally, we utilize a genetic algorithm for coarse-grained exploration based on system and query-level knob knowledge to expedite the cold start process in deep reinforcement learning. By classifying and compressing metrics and optimizing tuning knobs based on workload and knob-level insights, we aim to reduce the search space efficiently. In addition, adopting a delayed update strategy helps mitigate the training time required for the deep reinforcement learning model. Our extensive experiments demonstrate that MLETune outperforms existing methods by identifying superior configurations in significantly less time, showing an average improvement of ${6x}$, along with achieving up to a $26 \%$ performance improvement.
Wenlong Dong, Wei Liu 0279, Mengshu Hou, Shuhuan Fan
ICPADS5
2023 Graph-Attention-Network-Based Cost Estimation Model in Materialized View Environment
abstract
In database systems, materialized views (MV) pre-emptively materialize the common portion of query workloads to reduce redundant computations through query rewriting. However, the utilization of these rewritten queries depends on the accuracy of cost estimation models. Despite the promising performance of learning-based cost estimation models, they still exhibit limitations. Firstly, they are unable to capture the relationships between cross-node dependencies and node hierarchy across physical execution plan trees, hindering accuracy improvements. Secondly, they cannot simultaneously support original queries and rewritten queries, thereby limiting compatibility enhancements. In this paper, we introduce TGAE, a cost estimation model employing Graph Attention Network (GAT) to learn cross-node dependencies among physical execution plans. TGAE first utilizes learning embeddings instead of one-hot encoding and then introduces an efficient node feature encoding to facilitate the dynamic creation of base tables tailored to meet the requirements of MV environments. To demonstrate the effectiveness of TGAE, we design and implement AGatMv, a system with view design and exploitation capabilities. Experimental results on two query workloads from the real-world IMDb dataset show significant improvements in cost estimation accuracy and rewrite evaluation correctness compared to PostgreSQL.
Daobing Zhu, Shuhuan Fan, Xiaoyang Zeng, Mengshu Hou
ICPADS2
2021 GACE: Graph-Attention-Network-Based Cardinality Estimator
Daobing Zhu, Dongsheng He, Shuhuan Fan, Jianming Liao, Mengshu Hou
DEXA (2)3