VLDB 2026 Research / reviewers in the wild / expert
Fang Wang 0012
dblp:35/5625-12
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2023
0000-0002-0339-6552ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 68% Data mining · 27% Distributed and cloud data management · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Memory systems · 64% GPUs and heterogeneous computing · 36% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
cardinality estimation |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
Query processing and optimization › cardinality estimation
learned cardinality estimation |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
Query processing and optimization › adaptive query processing
query re-optimization |
0.7 | 1 | 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality Estimation · Proc. ACM Manag. Data 2023 |
Query processing and optimization › query execution › hardware-accelerated query processing
GPU-accelerated query processing |
0.6 | 1 | 2022 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache Hive · SIGMOD Conference 2022 |
Data mining › high-dimensional data analysis
high-dimensional data mining |
0.5 | 1 | 2021 | Accelerating Similarity-based Mining Tasks on High-dimensional Data by Processing-in-memory · ICDE 2021 |
Data mining
similarity computation |
0.5 | 1 | 2021 | Accelerating Similarity-based Mining Tasks on High-dimensional Data by Processing-in-memory · ICDE 2021 |
Memory systems
non-volatile memory |
0.5 | 1 | 2021 | Accelerating Similarity-based Mining Tasks on High-dimensional Data by Processing-in-memory · ICDE 2021 |
Memory systems
processing-in-memory |
0.5 | 1 | 2021 | Accelerating Similarity-based Mining Tasks on High-dimensional Data by Processing-in-memory · ICDE 2021 |
Methods — techniques the papers use, named apart from their topics
processing-in-memory · 1.0refinement model · 0.7query-driven estimation · 0.7dot-product · 0.5dot product · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Speeding Up End-to-end Query Execution via Learning-based Progressive Cardinality EstimationabstractFast query execution requires learning-based cardinality estimators to have short inference time (as model inference time adds to end-to-end query execution time) and high estimation accuracy (which is crucial for finding good execution plan). However, existing estimators cannot meet both requirements due to the inherent tension between model complexity and estimation accuracy. We propose a novel Learning-based Progressive Cardinality Estimator (LPCE), which adopts a query re-optimization methodology. In particular, LPCE consists of an initial model (LPCE-I), which estimates cardinality before query execution, and a refinement model (LPCE-R), which progressively refines the cardinality estimations using the actual cardinalities of the executed operators. During query execution, re-optimization is triggered if the estimations of LPCE-I are found to have large errors, and more efficient execution plans are selected for the remaining operators using the refined estimations provided by LPCE-R. Both LPCE-I and LPCE-R are light-weight query-driven estimators but they achieve both good efficiency and high accuracy when used jointly. Besides designing the models for LPCE-I and LPCE-R, we also integrate re-optimization and LPCE into PostgreSQL, a popular database engine. Extensive experiments show that LPCE yields shorter end-to-end query execution time than state-of-the-art learning-based estimators. Fang Wang 0012, Xiao Yan 0002, Man Lung Yiu, Shuai Li 0014, Zunyao Mao, Bo Tang 0016 |
Proc. ACM Manag. Data | 1 |
| 2022 | GHive: accelerating analytical query processing in apache hive via CPU-GPU heterogeneous computingabstractAs a popular distributed data warehouse system, Apache Hive has been widely used for big data analytics in many organizations. Meanwhile, exploiting the massive parallelism of GPU to accelerate online analytical processing (OLAP) has been extensively explored in the database community. In this paper, we present GHive, which enhances CPU-based Hive via CPU-GPU heterogeneous computing. GHive is designed for the business intelligence applications and provides the same API as Hive for compatibility. To run SQL queries jointly on both CPU and GPU, GHive comes with three key techniques: (i) a novel data model gTable, which is column-based and enables efficient data movement between CPU memory and GPU memory; (ii) a GPU-based operator library Panda, which provides a complete set of SQL operators with extensively optimized GPU implementations; (iii) a hardware-aware MapReduce job placement scheme, which puts jobs judiciously on either GPU or CPU via a cost-based approach. In the experiments, we observe that GHive outperforms Hive in both query processing speed and operating expense on the Star Schema Benchmark (SSB). Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xiao Yan 0002, Xinying Zheng, Qiaomu Shen, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo |
SoCC | 14 |
| 2022 | GHive: A Demonstration of GPU-Accelerated Query Processing in Apache HiveabstractAs a distributed, fault-tolerant data warehouse system for large-scale data analytics, Apache Hive has been used for various applications in many organizations (e.g., Facebook, Amazon, and Huawei). Exploiting the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database system is a common practice in the industry. Meanwhile, it is a common practice to exploit the large degrees of parallelism of GPU to improve the performance of online analytical processing (OLAP) in database systems. This demo presents GHive, which enables Apache Hive to accelerate OLAP queries by jointly utilizing CPU and GPU in intelligent and efficient ways. The takeaways for SIGMOD attendees include: (1) the superior performance of GHive compared with vanilla Hive that only uses CPU; (2) intuitive visualizations of execution statistics for Hive and GHive to understand where the acceleration of GHive comes from; (3) detailed profiling of the time taken by each operator on CPU and GPU to show the advantages of GPU execution. Bo Tang 0016, Jiashu Zhang, Yangshen Deng, Xinying Zheng, Qiaomu Shen, Xiao Yan 0002, Dan Zeng 0002, Zunyao Mao, Chaozu Zhang, Zhengxin You, Runzhe Jiang, Fang Wang 0012, Man Lung Yiu, Huan Li 0003, Mingji Han, Zhenghai Luo |
SIGMOD Conference | 14 |
| 2021 | Accelerating Similarity-based Mining Tasks on High-dimensional Data by Processing-in-memoryabstractSimilarity computation is a core subroutine of many mining tasks on multi-dimensional data, which are often massive datasets at high dimensionality. In these mining tasks, the performance bottleneck is caused by the `memorywall' problem as substantial amount of data needs to be transferred from memory to processors. Recent advances in non-volatile memory (NVM) enable processing-in-memory (PIM), which reduces data transfer and thus alleviates the performance bottleneck. Nevertheless, NVM PIM supports specific operations only (e.g., dot-product on non-negative integer vectors) but not arbitrary operations. In this paper, we tackle the above challenge and carefully exploit NVM PIM to accelerate similarity-based mining tasks on multi-dimensional data without compromising the accuracy of results. Experimental results on real datasets show that our proposed method achieves up to 10.5x and 8.5x speedup for state-of-artkNN classification andk-means clustering algorithms, respectively. Fang Wang 0012, Man Lung Yiu, Zili Shao |
ICDE | 1 |
| 2019 | ReRAM-based processing-in-memory architecture for blockchain platformsabstractBlockchain's decentralized and consensus mechanism has attracted lots of applications, such as IoT devices. Blockchain maintains a linked list of blocks and grows by mining new blocks. However, the Blockchain mining consumes huge computation resource and energy, which is unacceptable for resource-limited embedded devices. This paper for the first time presents a ReRAM-based processing-in-memory architecture for Blockchain mining, called Re-Mining. Re-Mining includes a message schedule module and a SHA computation module. The modules are composed of several basic ReRAM-based logic operations units, such as ROR, RSF and XOR. Re-Mining further designs intra-transaction and inter-transaction parallel mechanisms to accelerate the Blockchain mining. Simulation results show that the proposed Re-Mining architecture outperforms CPU-based and GPU-based implementations significantly. Fang Wang 0012, Zhaoyan Shen, Zili Shao |
ASP-DAC | 1 |