EDBT 2026 Demo / reviewers in the wild / expert
Zhuhe Fang
dblp:228/6022
· DBLP profile ↗
4ranked-venue papers
3as first author
0since 2021 · last 2020
0000-0001-6554-0014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 66% Distributed and cloud data management · 13% Database system architecture and tuning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 47% Processor architecture and microarchitecture · 20% Parallel and multicore computing · 20% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization › query scheduling
dynamic scheduling |
0.8 | 2 | 2020 | Scheduling Resources to Multiple Pipelines of One Query in a Main Memory Database Cluster · IEEE Trans. Knowl. Data Eng. 2020 Parallelizing Multiple Pipelines of One Query in a Main Memory Database Cluster · ICDE 2018 |
Query processing and optimization
parallel query processing |
0.8 | 2 | 2020 | Scheduling Resources to Multiple Pipelines of One Query in a Main Memory Database Cluster · IEEE Trans. Knowl. Data Eng. 2020 Parallelizing Multiple Pipelines of One Query in a Main Memory Database Cluster · ICDE 2018 |
Query processing and optimization
query scheduling |
0.8 | 2 | 2020 | Scheduling Resources to Multiple Pipelines of One Query in a Main Memory Database Cluster · IEEE Trans. Knowl. Data Eng. 2020 Parallelizing Multiple Pipelines of One Query in a Main Memory Database Cluster · ICDE 2018 |
Database system architecture and tuning
hybrid transactional and analytical processing |
0.4 | 1 | 2020 | TiDB: A Raft-based HTAP Database · Proc. VLDB Endow. 2020 |
Distributed systems
consensus |
0.4 | 1 | 2020 | TiDB: A Raft-based HTAP Database · Proc. VLDB Endow. 2020 |
Distributed systems › consensus › leader-based consensus
raft |
0.4 | 1 | 2020 | TiDB: A Raft-based HTAP Database · Proc. VLDB Endow. 2020 |
Processor architecture and microarchitecture
instruction set architecture |
0.4 | 1 | 2019 | Interleaved Multi-Vectorizing · Proc. VLDB Endow. 2019 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.4 | 1 | 2019 | Interleaved Multi-Vectorizing · Proc. VLDB Endow. 2019 |
Indexing and storage engines
column store |
0.1 | 1 | 2020 | TiDB: A Raft-based HTAP Database · Proc. VLDB Endow. 2020 |
Transaction processing and concurrency control
distributed transaction processing |
0.1 | 1 | 2020 | TiDB: A Raft-based HTAP Database · Proc. VLDB Endow. 2020 |
Memory systems
cache |
0.1 | 1 | 2019 | Interleaved Multi-Vectorizing · Proc. VLDB Endow. 2019 |
Memory systems › cache
cache miss reduction |
0.1 | 1 | 2019 | Interleaved Multi-Vectorizing · Proc. VLDB Endow. 2019 |
Methods — techniques the papers use, named apart from their topics
replicated state machine · 0.9multi-raft · 0.9preemption · 0.4list scheduling · 0.4adaptive filling · 0.4vectorization · 0.4prefetching · 0.4list with filling and preemption · 0.3cost-based preemption · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | TiDB: A Raft-based HTAP DatabaseabstractHybrid Transactional and Analytical Processing (HTAP) databases require processing transactional and analytical queries in isolation to remove the interference between them. To achieve this, it is necessary to maintain different replicas of data specified for the two types of queries. However, it is challenging to provide a consistent view for distributed replicas within a storage system, where analytical requests can efficiently read consistent and fresh data from transactional workloads at scale and with high availability. To meet this challenge, we propose extending replicated state machine-based consensus algorithms to provide consistent replicas for HTAP workloads. Based on this novel idea, we present a Raft-based HTAP database: TiDB. In the database, we design a multi-Raft storage system which consists of a row store and a column store. The row store is built based on the Raft algorithm. It is scalable to materialize updates from transactional requests with high availability. In particular, it asynchronously replicates Raft logs to learners which transform row format to column format for tuples, forming a real-time updatable column store. This column store allows analytical queries to efficiently read fresh and consistent data with strong isolation from transactions on the row store. Based on this storage system, we build an SQL engine to process large-scale distributed transactions and expensive analytical queries. The SQL engine optimally accesses row-format and column-format replicas of data. We also include a powerful analysis engine, TiSpark, to help TiDB connect to the Hadoop ecosystem. Comprehensive experiments show that TiDB achieves isolated high performance under CH-benCHmark, a benchmark focusing on HTAP workloads. Dongxu Huang, Qiu Cui, Zhuhe Fang, Yuxing Zhou, Menglong Huang, Wan Wei, Xuelian Wu, Lingyu Song, Ruoxi Sun 0005, Shuaipeng Yu, Nicholas Cameron 0001, Liquan Pei |
Proc. VLDB Endow. | 4 |
| 2020 | Scheduling Resources to Multiple Pipelines of One Query in a Main Memory Database ClusterabstractTo fully utilize the resources of a main memory database cluster, we additionally take the independent parallelism into account to parallelize multiple pipelines of one query. However, scheduling resources to multiple pipelines is an intractable problem. Traditional static approaches to this problem may lead to a serious waste of resources and suboptimal execution order of pipelines, because it is hard to predict the actual data distribution and fluctuating workloads at compile time. In response, we propose a dynamic scheduling algorithm, List with Filling and Preemption (LFPS), based on two novel techniques. (1) Adaptive filling improves resource utilization by issuing more extra pipelines to adaptively fill idle resource “holes” during execution. (2) Rank-based preemption strictly guarantees scheduling the pipelines on the critical path first at run time. Interestingly, the latter facilitates the former filling idle “holes” with best efforts to finish multiple pipelines as soon as possible. We implement LFPS in our prototype database system. Under the workloads of TPC-H, experiments show our work improves the finish time of parallelizable pipelines from one query up to 2.5X than a static approach and 2.1X than a serialized execution. Zhuhe Fang, Chuliang Weng, Huiqi Hu, Aoying Zhou |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Interleaved Multi-VectorizingabstractSIMD is an instruction set in mainstream processors, which provides the data level parallelism to accelerate the performance of applications. However, its advantages diminish when applications suffer from heavy cache misses. To eliminate cache misses in SIMD vectorization, we present interleaved multi-vectorizing (IMV) in this paper. It interleaves multiple execution instances of vectorized code to hide memory access latency with more computation. We also propose residual vectorized states to solve the control flow divergence in vectorization. IMV can make full use of the data parallelism in SIMD and the memory level parallelism through prefetching. It reduces cache misses, branch misses and computation overhead to significantly speed up the performance of pointer-chasing applications, and it can be applied to executing entire query pipelines. As experimental results show, IMV achieves up to 4.23X and 3.17X better performance compared with the pure scalar implementation and the pure SIMD vectorization, respectively. Zhuhe Fang, Beilei Zheng, Chuliang Weng |
Proc. VLDB Endow. | 1 |
| 2018 | Parallelizing Multiple Pipelines of One Query in a Main Memory Database ClusterabstractTo fully use the advanced resources of a main memory database cluster, we take independent parallelism into account to parallelize multiple pipelines of one query. However, scheduling resources to multiple pipelines is an intractable problem. Traditional static approaches to this problem may lead to a serious waste of resources and suboptimal execution order of pipelines, because it is hard to predict the actual data distribution and fluctuating workloads at compile time. In response, we propose a dynamic scheduling algorithm, List with Filling and Preemption (LFPS), based on two techniques. (1) Adaptive filling improves resource utilization by issuing more extra pipelines to adaptively fill idle resource "holes" during execution. (2) Cost-based preemption strictly guarantees scheduling the pipelines on a critical path first at run time. We implement LFPS in our prototype database system. Under the workloads of TPC-H, experiments show our work improves the finish time of parallelizable pipelines from one query up to 2.3X than a static approach and 1.7X than a serialized execution. Zhuhe Fang, Chuliang Weng, Aoying Zhou |
ICDE | 1 |