VLDB 2026 Research / reviewers in the wild / expert
Zhiwen Chen 0006
dblp:133/3514-6
· DBLP profile ↗
14ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-1897-7840ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | kShield: An eBPF runtime defence framework for linux kernel privilege escalation attacks
Guoyun Duan, Boying Chen 0001, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002 |
Inf. Softw. Technol. | 3 |
| 2024 | ZBTree: A Fast and Scalable B$^+$+-Tree for Persistent MemoryabstractIn this paper, we present the design and implementation of ZBTree, a hotness-aware B$^+$-Tree for persistent memory (PMem). ZBTree leverages the PMem+DRAM architecture, which is featured with a volatile operation layer to accelerate data access and an order-preserving persistent layer to achieve fast recovery and low-overhead consistency and persistence guarantees. The operation layer contains inner nodes for indexing and compacted leaf nodes (DLeaves) that hold metadata. Based on leaf node compaction, we present a data lodging method, which supports to load hot data into fast DRAM dynamically, avoiding PMem accesses for subsequent reads of hot data and achieving improved read performance without incurring extra DRAM usage. In addition, we present a lightweight node splitting mechanism with constant persistence overhead that does not vary with node size. Our extensive evaluations show that ZBTree achieves higher throughput by a factor of 1.4x-6.3x compared to state-of-the-art tree indexes under a wide range of workloads. Meanwhile, ZBTree achieves comparable or faster recovery speed compared to existing designs. Wenkui Che, Zhiwen Chen 0006, Daokun Hu, Jianhua Sun 0002, Hao Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | A quantitative evaluation of persistent memory hash indexes
Zhiwen Chen 0006, Daokun Hu, Wenkui Che, Jianhua Sun 0002, Hao Chen 0002 |
VLDB J. | 1 |
| 2023 | TEEFuzzer: A fuzzing framework for trusted execution environments with heuristic seed mutation
Guoyun Duan, Yuanzhi Fu, Peiyao Deng, Jianhua Sun 0002, Hao Chen 0002, Zhiwen Chen 0006 |
Future Gener. Comput. Syst. | 7 |
| 2023 | On the Performance Intricacies of Persistent Memory Aware Storage EnginesabstractAs key components of DBMSs, various storage engines and index structures have been proposed based on incorrect assumptions before PMem hardware is publicly available. Recent studies reveal that there is a significant performance gap in evaluating index structures on real PMem platforms as compared to DRAM-based emulators. However, a comprehensive evaluation for those PMem-aware database storage engines on real PMem hardware is still missing. Meanwhile, dynamic memory management is more important on PMem systems because PMem is slower than DRAM and unfriendly to random small-writes, and ensuring crash-consistency for the metadata of PMem allocators introduces extra overhead. Therefore, it is essential to understand the performance intricacies of PMem-aware database storage engines from the perspective of PMem allocators. This paper presents a systematic evaluation of three PMem-aware database storage engines using representative workloads and a unified benchmarking framework that is integrated with four PMem allocators. Besides the commonly used metrics, the impact of different hardware configurations (such as NUMA and eADR) on performance is also considered. Through in-depth analysis, we reveal caveats and pitfalls on using or designing PMem-aware storage engines and important insights that can serve as guidelines for future development of PMem allocators and other related components. Zhiwen Chen 0006, Wenkui Che, Daokun Hu, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Halo: A Hybrid PMem-DRAM Persistent Hash Index with Fast RecoveryabstractHash index, a fundamental component in many data management systems, can benefit from the emerging persistent memory (PMem) to achieve high performance and instant recovery. However, existing persistent hash indexes are suboptimal in at least three aspects. First, their performance suffers from the mismatch between small random write and access granularity of PMem hardware. Second, none of them are aware of the significance of write amplification caused by memory allocators and synchronization primitives. Third, hybrid designs (PMem+DRAM) focus on improving throughput at the cost of extremely long recovery time. Daokun Hu, Zhiwen Chen 0006, Wenkui Che, Jianhua Sun 0002, Hao Chen 0002 |
SIGMOD Conference | 2 |
| 2022 | CVFuzz: Detecting complexity vulnerabilities in OpenCL kernels via automated pathological input generation
Zhiwen Chen 0006, Xin He 0054, Guoyun Duan, Jianhua Sun 0002, Hao Chen 0002 |
Future Gener. Comput. Syst. | 2 |
| 2021 | Efficient parallel A* search on multi-GPU system
Xin He 0054, Yapeng Yao, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002 |
Future Gener. Comput. Syst. | 3 |
| 2021 | Persistent Memory Hash Indexes: An Experimental EvaluationabstractPersistent memory (PM) is increasingly being leveraged to build hash-based indexing structures featuring cheap persistence, high performance, and instant recovery, especially with the recent release of Intel Optane DC Persistent Memory Modules. However, most of them are evaluated on DRAM-based emulators with unreal assumptions, or focus on the evaluation of specific metrics with important properties sidestepped. Thus, it is essential to understand how well the proposed hash indexes perform on real PM and how they differentiate from each other if a wider range of performance metrics are considered. To this end, this paper provides a comprehensive evaluation of persistent hash tables. In particular, we focus on the evaluation of six state-of-the-art hash tables including Level hashing, CCEH, Dash, PCLHT, Clevel, and SOFT, with real PM hardware. Our evaluation was conducted using a unified benchmarking framework and representative workloads. Besides characterizing common performance properties, we also explore how hardware configurations (such as PM bandwidth, CPU instructions, and NUMA) affect the performance of PM-based hash tables. With our in-depth analysis, we identify design trade-offs and good paradigms in prior arts, and suggest desirable optimizations and directions for the future development of PM-based hash tables. Daokun Hu, Zhiwen Chen 0006, Jianbing Wu, Jianhua Sun 0002, Hao Chen 0002 |
Proc. VLDB Endow. | 2 |
| 2018 | Concurrent hash tables on multicore machines: Comparison, evaluation and implications
Zhiwen Chen 0006, Xin He 0054, Jianhua Sun 0002, Hao Chen 0002, Ligang He |
Future Gener. Comput. Syst. | 1 |
| 2017 | Exploring Synchronization in Cache Coherent Manycore Systems: A Case Study with Xeon PhiabstractIntel Xeon Phi is a many-core architecture, featuring more than 50 cores and 200 hardware threads. Given this scale and its other distinctive architectural features, highly-concurrent applications on Xeon Phi may behave differently than on tradi- tional multi-core systems. Yet, concurrency issues especially for synchronization intensive applications on this platform have not been thoroughly analyzed. In this paper, we conduct an extensive analysis at multiple layers, from the underlying hardware cache- coherence protocol up to the user-level applications, aiming to present the most exhaustive study of synchronization on Xeon Phi. Through a range of benchmarks, we testify the feasibility and advantage of accelerating concurrent applications with Xeon Phi. Meanwhile, we identify severe scalability issues relevant to synchronization, and solutions to these issues are discussed. We believe this work can be used as guidelines both for designing better synchronization mechanisms and in optimizing concurrent applications in order to fully exploit the capability of Xeon Phi. Xin He 0054, Zhiwen Chen 0006, Jianhua Sun 0002, Hao Chen 0002, Dong Li 0001, Zhe Quan |
ICPADS | 2 |
| 2017 | Optimizing Graph Processing on GPUsabstractDistributed vertex-centric model has been recently proposed for large-scale graph processing. Due to the simple but efficient programming abstraction, similar graph computing frameworks based on GPUs are gaining more and more attention. However, prior works of GPU-based graph processing suffer from load imbalance and irregular memory access because of the inherent characteristics of graph applications. In this paper, we propose a generalized graph computing framework for GPUs to simplify existing models but with higher performance. In particular, two novel algorithmic optimizations, lightweight approximate sorting and data layout transformation, are proposed to tackle the performance issues of current systems. With extensive experimental evaluation under a wide range of real world and synthetic workloads, we show that our system can achieve 1.6× to 4.5× speedups over the state-of-the-art. Wenyong Zhong, Jianhua Sun 0002, Hao Chen 0002, Zhiwen Chen 0006, Xuanhua Shi |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2016 | Automatically identifying apps in mobile trafficabstractSummary With the rapid development of smartphones in recent years, we have witnessed an exponential growth of the number of mobile apps. Considering the security and management issues, network operators need to have a clear visibility into the apps running in the network. To this end, this paper presents a novel approach to generating the fingerprints for mobile apps from network traffic. The fingerprints that characterize the unique behaviors of specific mobile apps can be used to identify mobile apps from the real network traffic. In order to handle the large volume of traffic efficiently, we use non‐negative matrix factorization (NMF) to perform traffic analysis to cluster similar network traffic into groups. Then, access patterns of individual apps that are extracted from each group can be used as fingerprints distinguishing apps from others uniquely. The experimental evaluations show that the proposed approach can identify the mobile apps from random and mixed network traffic with high precision. Copyright © 2015 John Wiley & Sons, Ltd. Jianhua Sun 0002, Lingjun She, Hao Chen 0002, Wenyong Zhong, Zhiwen Chen 0006, Shuna Yao |
Concurr. Comput. Pract. Exp. | 6 |
| 2015 | GPSA: A Graph Processing System with ActorsabstractDue to the increasing need to process the fast growing graph-structured data (e.g. Social networks and Web graphs), designing high performance graph processing systems becomes one of the most urgent problems facing systems researchers. In this paper, we introduce GPSA, a single-machine graph processing system based on an actor computation model inspired by the Bulk Synchronous Parallel(BSP) computation model. GPSA takes advantage of actors to improve the concurrency on a single machine with limited resource. GPSA improves the conventional BSP computation model to fit in the actor programming paradigm by decoupling the message dispatching from the computation. Furthermore, we exploit memory mapping to avoid explicit data management to improve I/O performance. Experimental evaluation shows that our system outperforms existing systems by 2x-6x in processing large-scale graphs on a single system. Jianhua Sun 0002, Dongwei Zhou, Hao Chen 0002, Zhiwen Chen 0006, Ligang He |
ICPP | 5 |