EDBT 2026 Demo / reviewers in the wild / expert
Zhulin Ma
dblp:214/1940
· DBLP profile ↗
10ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0003-3439-3570ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CXL shared coherent memory simulation and cross-host synchronization mechanisms design for data sharing
Ting Wu 0012, Qingyuan Song, Xihong Huang, Linbo Long, Zhulin Ma, Weichen Liu 0001 |
J. Syst. Archit. | 5 |
| 2026 | WPAlloc: An Efficient Wear-Leveling-Aware Parallel Allocator for Persistent Memory File SystemsabstractInternet and IoT applications have generated increasing amounts of data that require efficient storage. Many persistent memory file systems have been designed to handle high-performance storage demands by fully exploiting the senior features of persistent memory (PM). However, PMs suffer from limited write endurance. Existing PM file systems achieve PM wear-leveling by designing wear-leveling-aware allocators. These allocators focus on providing higher-balanced writes to PMs while neglecting the overhead. Moreover, they cause serious request conflicts in parallel block requests by multiple threads in modern multiprocessor computer systems. In this paper, we propose an efficient wear-leveling-aware parallel allocator, WPAlloc, for persistent memory file systems to achieve wear-leveling of PM and high parallel performance. The essential idea of WPAlloc is to allocate blocks with lower write counters for each allocation request and to provide parallel block allocation and deallocation for multiple threads via one free list per logical processor. WPAlloc consists of two key techniques: the bucket sort-based range management scheme (BRMS) and the parallel allocation and deallocation scheme (PADS). First, we design the BRMS to obtain the less worn free blocks. Based on BRMS, an online wear range adjustment algorithm has been designed to adjust the wear range dynamically. Then, we present the PADS to avoid request conflicts by multiple threads. We implement WPAlloc based on PMFS. Experimental results show that WPAlloc can reduce the maximum write by 54.2%, 10.2%, and 55.7%, while achieving average performance improvements of 5.68%, 54.18%, and 11.28% compared to PMFS, DWARM, and WASA, respectively. Ting Wu 0012, Linbo Long, Zhulin Ma, Duo Liu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | CMCache: An Adaptive Cross-Level Data Placement Method for Multilevel CacheabstractMultilevel cache systems enhance I/O performance by optimizing data placement across various cache levels from a global perspective. However, existing methods often struggle to place data at the optimal cache level promptly due to their reliance on historical access patterns and inflexible placement strategies. These methods face two main challenges: 1) for already cached data with sufficient access history, existing approaches only optimize movement between adjacent cache levels, potentially delaying data arrival at its globally optimal cache level and leading to unnecessary bandwidth consumption and increased latency and 2) for newly entered data without access history, current methods cannot accurately predict their future hotness and simply place them at a fixed cache level (i.e., first or final level), overlooking future accesses of new data and potentially resulting in high cache miss rates or cache pollution. To address these issues, we propose CMCache, an adaptive cross-level data placement method for multilevel cache. CMCache applies distinct placement strategies for cached and new data to reach the optimal level timely, considering their different characteristics. It also logically divides cache space into two sections to manage cached and new data separately, dynamically adjusting section sizes based on access patterns. This approach significantly improves data placement efficiency, achieving up to an 89% reduction in miss rates and a 79% decrease in average response times compared to existing methods. Zhaoyang Zeng, Yujuan Tan, Zhulin Ma, Sanle Zhao, Duo Liu 0002, Xianzhang Chen, Ao Ren |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | SAPredictor: a simple and accurate self-adaptive predictor for hierarchical hybrid memory systemabstractIn a hybrid memory system using DRAM as the NVM cache, DRAM and NVM can be accessed in serial or parallel mode. However, we found that using either mode alone will bring access latency and bandwidth problems. In this paper, we integrate these two access modes and design a simple but accurate predictor (called SAPredictor) to help choose the appropriate access mode, thereby avoiding long access latency and bandwidth problems to improve memory performance. Our experiments show that SAPredictor achieves an accuracy rate of up to 97.1% and helps reduce access latency by up to 35.6% at fairly low costs. Yujuan Tan, Wei Chen 0101, Zhulin Ma, Dan Xiao, Zhichao Yan 0001, Duo Liu 0002, Xianzhang Chen |
DAC | 3 |
| 2022 | GATLB: A Granularity-Aware TLB to Support Multi-Granularity Pages in Hybrid Memory SystemabstractThe parallel hybrid memory system that combines Non-volatile Memory (NVM) and DRAM can effectively expand the memory capacity. But it puts lots of pressure on TLB due to a limited TLB capacity. The superpage technology that manages pages with a large granularity (e.g., 2MB) is usually used to improve the TLB performance. However, its coarse-grained granularity conflicts with the fine-grained page migration in the hybrid memory system, resulting in serious invalid migration and page fragmentation problems. To solve these problems, we propose to maintain the coexistence of multi-granularity pages, and design a smart TLB called GATLB to support multi-granularity page management, coalesce consecutive pages and adapt to various changes in page size. Compared with the existing TLB technologies, GATLB can not only perceive page granularity to effectively expand the TLB coverage and reduce miss rate, but also provide faster address translation with a much lower overhead. Our experimental evaluations show that GATLB can expand the TLB coverage by 7.09x, reduce the TLB miss rate by 91.1%, and shorten the address translation cycle by 49.41%. Yujuan Tan, Yujie Xie, Zhulin Ma, Zhichao Yan 0001, Duo Liu 0002, Xianzhang Chen |
DATE | 3 |
| 2022 | Towards the Design of Efficient TCN-bascd Prefetcher for Hybrid NVM-DRAM MemoryabstractThe hybrid memory system has been widely studied, comprised of Non-volatile Memory (NVM) and DRAM, due to its larger capacity and lower power consumption than DRAM. As a data placement scheme, prefetching plays a vital role in the performance of the hybrid memory scenario. However, existing prefetchers fail to satisfy both high prediction accuracy and fast processing simultaneously: the hardware-based prefetcher becomes impractical due to the exploding prediction table size, as the application complexity increases; the LSTM-based prefetcher, as a promising software-based prefetcher, suffers from inefficient timeliness, unstable structures, and excessive memory consumption. In this paper, we demonstrate the potential of a temporal convolutional network (TCN) in prefetching because of its parallelizable convolution operations for acceleration, a more stable structure compared with RNNs, and adequately long history window size. However, using TCN directly in the prefetching brings some challenges: the design of the TCN structure requires consideration of the trade-off between the model size and its prediction accuracy; it is hard for TCN layers to learn the correlation between memory accesses comprehensively. Therefore, we propose a novel TCN-based memory prefetcher (TMP), which uses an appropriate number of dilated convolution layers to satisfy a sufficiently large receptive field while maintaining a relatively small model size. In addition, we use the attention mechanism to fully exploit the correlation between memory accesses and improve prefetching effectiveness. Our TMP model comprises an input module for dimensionality reduction, a TCN-Attention module for memory access pattern learning, and an output module for future access prediction. Compared to the state-of-the-art LSTM-based prefetcher, TMP is 1.6x and 4.9x faster in training and inference speed, respectively, meanwhile achieving as high as 84.1% accuracy on average on SPEC CPU 2017. Yujuan Tan, Zhulin Ma, Duo Liu 0002 |
IJCNN | 3 |
| 2021 | DFShards: effective construction of MRCs online for non-stack algorithmsabstractThe Miss Ratio Curve (MRC) describes the cache miss ratio as a function of the cache size. It has various shapes that represent the data access behaviors of workloads in the cache. MRC is an effective tool to guide cache partitioning, but its real-time construction is challenging. Miniature Simulation is a novel approach that constructs MRCs for non-stack algorithms in real time, via feeding a small number of sample references to multiple mini caches simultaneously to get the miss ratios. However, while using the Miniature Simulation, the size and number of mini-caches are difficult to set before the program runs. First, it may set too many mini-caches and cause repeated simulations. Second, it may miss some important cache sizes and consequently construct a less precise shape of MRC and result in incorrect cache partitioning. Ailing Yu, Yujuan Tan, Congcong Xu, Zhulin Ma, Duo Liu 0002, Xianzhang Chen |
CF | 4 |
| 2020 | Unified-TP: A Unified TLB and Page Table Cache Structure for Efficient Address TranslationabstractTo improve the performance of address translation in applications with large memory footprints, techniques, such as hugepages and HW coalescing, are proposed to increase the coverage of limited hardware translation entries by exploiting the contiguous memory allocation to lower Tanslation Lookaside Buffer (TLB) miss rate. Furthermore, Page Table Caches (PTCs) are proposed to store the upper-level page table entries to reduce the TLB miss handling latency. Both increasing TLB coverage and reducing TLB miss handling latency have proved to be effective in speeding up address translation, to a certain extent. Nevertheless, our preliminary studies suggest that the structural separation between TLBs and PTCs in existing computer systems makes these two methods less effective because they are exclusively used in TLBs and PTCs respectively. In particular, the separate structures cannot dynamically adjust their sizes according to the workloads, resulting in low resource utilization and inefficient address translation. To address these issues, we propose a unified structure, called Unified - Tp,which stores PTC and TLB entries together. Besides, Our modified LRU algorithm helps identify the cold TLB and PTC entries and dynamically adjust the numbers of TLB and PTC entries to adapt to different workloads. Furthermore, we introduce a scheme of parallel search when receiving memory access requests. Our experimental results show that Unified-TP can reduce the numbers of TLB misses by an average of 35.69 % and improve the performance by an average of 11.12% compared with separately structured TLBs and PTCs. Zhulin Ma, Yujuan Tan, Hong Jiang 0001, Zhichao Yan 0001, Duo Liu 0002, Xianzhang Chen, Qingfeng Zhuge, Edwin H.-M. Sha, Chengliang Wang 0002 |
ICCD | 1 |
| 2020 | Towards the design of efficient hash-based indexing scheme for growing databases on non-volatile memory
Zhulin Ma, Edwin H.-M. Sha, Qingfeng Zhuge, Weiwen Jiang, Runyu Zhang 0002, Shouzhen Gu |
Future Gener. Comput. Syst. | 1 |
| 2018 | Towards the Design of Efficient and Consistent Index Structure with Minimal Write Activities for Non-Volatile MemoryabstractIndex structures can significantly accelerate the data retrieval operations in data intensive systems, such as databases. Tree structures, such as B+-tree alike, are commonly employed as index structures; however, we found that the tree structure may not be appropriate for Non-Volatile Memory (NVM) in terms of the requirements for high-performance and high-endurance. This paper studies what is the best index structure for NVM-based systems and how to design such index structures. The design of an NVM-friendly index structure faces a lot of challenges. First, in order to prolong the lifetime of NVM, the write activities on NVM should be minimized. To this end, the index structure should be as simple as possible. The index proposed in this paper is based on the simplest data structure, i.e., linked list. Second, the simple structure brings challenges to achieve high-performance data retrieval operations. To overcome this challenge, we design a novel technique by explicitly building up a contiguous virtual address space on the linked list, such that efficient search algorithms can be performed. Third, we need to carefully consider data consistency issues in NVM-based systems, because the order of memory writes may be changed and the data content in NVM may be inconsistent due to write-back effects of CPU cache. This paper devises a novel indexing scheme, called “Virtual Linear Addressable Buckets” (VLAB). We implement VLAB in a storage engine and plug it into MySQL. Evaluations are conducted on an NVDIMM workstation using YCSB workloads and real-world traces. Results show that write activities of the state-of-the-art indexes are 6.98 times more than ours; meanwhile, VLAB achieves 2.53 times speedup. Edwin H.-M. Sha, Weiwen Jiang, Hailiang Dong, Zhulin Ma, Runyu Zhang 0002, Xianzhang Chen, Qingfeng Zhuge |
IEEE Trans. Computers | 4 |