VLDB 2026 Research / reviewers in the wild / expert
Daokun Hu
dblp:250/0478
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0002-5538-0517ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LCL+: a Lock Chain Length-based Distributed Deadlock Detection and Resolution Service Built for OceanBaseabstractThe problem of deadlock detection and resolution in database systems has been studied for decades. Although it has long been a mature feature of classical centralized database systems for many years, its use in distributed database systems remains in its infancy. A simple and fully distributed deadlock detection and resolution algorithm was proposed by Don P. Mitchell and Michael J. Merritt ( M&M ), but its assumption that each process waits for only one resource at a time prevents it from being generally applicable. The distributed deadlock detection algorithm based on Lock Chain Length (LCL) surpasses the limitations of the M&M algorithm. However, it is less effective in quickly detecting deadlocks that encompass both distributed and local deadlocks. In this article, we introduce LCL + , an advanced and universally applicable algorithm specifically designed for the detection and resolution of resource deadlocks in distributed environments. This algorithm improves the efficiency of identifying distributed deadlocks by accelerating the detection of hybrid deadlocks. Our extensive experiments demonstrate that LCL + significantly outperforms its predecessor, LCL, in efficiency. In addition, it has been successfully implemented in the OceanBase distributed relational database system. Detailed analyses from multiple perspectives within OceanBase confirm that LCL + significantly improves the system’s scalability and ensures the provision of high-quality service. Xuwang Teng, Fanyu Kong 0004, Fusheng Han, Quanqing Xu, Daokun Hu |
ACM Trans. Comput. Syst. | 7 |
| 2024 | ZBTree: A Fast and Scalable B$^+$+-Tree for Persistent MemoryabstractIn this paper, we present the design and implementation of ZBTree, a hotness-aware B$^+$-Tree for persistent memory (PMem). ZBTree leverages the PMem+DRAM architecture, which is featured with a volatile operation layer to accelerate data access and an order-preserving persistent layer to achieve fast recovery and low-overhead consistency and persistence guarantees. The operation layer contains inner nodes for indexing and compacted leaf nodes (DLeaves) that hold metadata. Based on leaf node compaction, we present a data lodging method, which supports to load hot data into fast DRAM dynamically, avoiding PMem accesses for subsequent reads of hot data and achieving improved read performance without incurring extra DRAM usage. In addition, we present a lightweight node splitting mechanism with constant persistence overhead that does not vary with node size. Our extensive evaluations show that ZBTree achieves higher throughput by a factor of 1.4x-6.3x compared to state-of-the-art tree indexes under a wide range of workloads. Meanwhile, ZBTree achieves comparable or faster recovery speed compared to existing designs. Wenkui Che, Zhiwen Chen 0006, Daokun Hu, Jianhua Sun 0002, Hao Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | A quantitative evaluation of persistent memory hash indexes
Zhiwen Chen 0006, Daokun Hu, Wenkui Che, Jianhua Sun 0002, Hao Chen 0002 |
VLDB J. | 2 |
| 2023 | On the Performance Intricacies of Persistent Memory Aware Storage EnginesabstractAs key components of DBMSs, various storage engines and index structures have been proposed based on incorrect assumptions before PMem hardware is publicly available. Recent studies reveal that there is a significant performance gap in evaluating index structures on real PMem platforms as compared to DRAM-based emulators. However, a comprehensive evaluation for those PMem-aware database storage engines on real PMem hardware is still missing. Meanwhile, dynamic memory management is more important on PMem systems because PMem is slower than DRAM and unfriendly to random small-writes, and ensuring crash-consistency for the metadata of PMem allocators introduces extra overhead. Therefore, it is essential to understand the performance intricacies of PMem-aware database storage engines from the perspective of PMem allocators. This paper presents a systematic evaluation of three PMem-aware database storage engines using representative workloads and a unified benchmarking framework that is integrated with four PMem allocators. Besides the commonly used metrics, the impact of different hardware configurations (such as NUMA and eADR) on performance is also considered. Through in-depth analysis, we reveal caveats and pitfalls on using or designing PMem-aware storage engines and important insights that can serve as guidelines for future development of PMem allocators and other related components. Zhiwen Chen 0006, Wenkui Che, Daokun Hu, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Halo: A Hybrid PMem-DRAM Persistent Hash Index with Fast RecoveryabstractHash index, a fundamental component in many data management systems, can benefit from the emerging persistent memory (PMem) to achieve high performance and instant recovery. However, existing persistent hash indexes are suboptimal in at least three aspects. First, their performance suffers from the mismatch between small random write and access granularity of PMem hardware. Second, none of them are aware of the significance of write amplification caused by memory allocators and synchronization primitives. Third, hybrid designs (PMem+DRAM) focus on improving throughput at the cost of extremely long recovery time. Daokun Hu, Zhiwen Chen 0006, Wenkui Che, Jianhua Sun 0002, Hao Chen 0002 |
SIGMOD Conference | 1 |
| 2022 | TLB-pilot: Mitigating TLB Contention Attack on GPUs with Microarchitecture-Aware SchedulingabstractCo-running GPU kernels on a single GPU can provide high system throughput and improve hardware utilization, but this raises concerns on application security. We reveal that translation lookaside buffer (TLB) attack, one of the common attacks on CPU, can happen on GPU when multiple GPU kernels co-run. We investigate conditions or principles under which a TLB attack can take effect, including the awareness of GPU TLB microarchitecture, being lightweight, and bypassing existing software and hardware mechanisms. This TLB-based attack can be leveraged to conduct Denial-of-Service (or Degradation-of-Service) attacks. Furthermore, we propose a solution to mitigate TLB attacks. In particular, based on the microarchitecture properties of GPU, we introduce a software-based system, TLB-pilot, that binds thread blocks of different kernels to different groups of streaming multiprocessors by considering hardware isolation of last-level TLBs and the application’s resource requirement. TLB-pilot employs lightweight online profiling to collect kernel information before kernel launches. By coordinating software- and hardware-based scheduling and employing a kernel splitting scheme to reduce load imbalance, TLB-pilot effectively mitigates TLB attacks. The result shows that when under TLB attack, TLB-pilot mitigates the attack and provides on average 56.2% and 60.6% improvement in average normalized turnaround times and overall system throughput, respectively, compared to the traditional Multi-Process Service based co-running solution. When under TLB attack, TLB-pilot also provides up to 47.3% and 64.3% improvement (41% and 42.9% on average) in average normalized turnaround times and overall system throughput, respectively, compared to a state-of-the-art co-running solution for efficiently scheduling of thread blocks. Bang Di, Daokun Hu, Jianhua Sun 0002, Hao Chen 0002, Jinkui Ren, Dong Li 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2021 | Persistent Memory Hash Indexes: An Experimental EvaluationabstractPersistent memory (PM) is increasingly being leveraged to build hash-based indexing structures featuring cheap persistence, high performance, and instant recovery, especially with the recent release of Intel Optane DC Persistent Memory Modules. However, most of them are evaluated on DRAM-based emulators with unreal assumptions, or focus on the evaluation of specific metrics with important properties sidestepped. Thus, it is essential to understand how well the proposed hash indexes perform on real PM and how they differentiate from each other if a wider range of performance metrics are considered. To this end, this paper provides a comprehensive evaluation of persistent hash tables. In particular, we focus on the evaluation of six state-of-the-art hash tables including Level hashing, CCEH, Dash, PCLHT, Clevel, and SOFT, with real PM hardware. Our evaluation was conducted using a unified benchmarking framework and representative workloads. Besides characterizing common performance properties, we also explore how hardware configurations (such as PM bandwidth, CPU instructions, and NUMA) affect the performance of PM-based hash tables. With our in-depth analysis, we identify design trade-offs and good paradigms in prior arts, and suggest desirable optimizations and directions for the future development of PM-based hash tables. Daokun Hu, Zhiwen Chen 0006, Jianbing Wu, Jianhua Sun 0002, Hao Chen 0002 |
Proc. VLDB Endow. | 1 |