Guanglei Xu

dblp:333/9702 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0005-1415-1592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2025 GroupRS: Node-Grouping-Based Data Placement Strategy in Erasure-Coded Data Center Storage for High Data Reliability
abstract
Data center storage systems commonly use erasure codes instead of replication to ensure data reliability at a lower cost, striping data into blocks and placing them randomly across nodes. Studies on replication have shown that random block placement can lead to data loss during multi-node failures, and grouping nodes and placing replicas within these groups is used to enhance data reliability. However, random grouping block placement in erasure-coded storage systems often falls short of the required repair parallelism across nodes. Additionally, creating a globally optimal grouping scheme that meets the requirements incurs prohibitively high time complexity. In this paper, we analyze the data reliability of erasure-coded storage systems using node-grouping-based data placement, revealing a trade-off between fault tolerance and repair parallelism. Based on the above analysis, we propose GroupRS, a node-grouping-based data placement strategy that utilizes a greedy heuristic to group nodes for greater node repair parallelism while maintaining the fault tolerance, thereby enhancing the reliability of the data. In addition, for rack failures, we propose GroupRSR atop GroupRS, a simulated annealing-based optimization strategy to reduce the probability of data loss caused by rack failures. Simulation results show that GroupRS improves system reliability by 18% over Copyset with the same fault tolerance, and by 1 × with the same repair parallelism. Cloud tests reveal that GroupRS reduces the average repair time by 47% with the same fault tolerance and increases MTTDL by 2.5 times. In racklevel fault scenarios, GroupRS-R further reduces the repair time by 10%.
Junyuan Huang, Yuchong Hu, Guanglei Xu
ICPADS3
2025 Accelerating Erasure Coding on Persistent Memory via Adaptive Prefetcher Scheduling
abstract
Compared to DRAM, persistent memory (PM) offers higher density and persistence but encounters more severe reliability challenges. Erasure coding is widely adopted to enhance reliability with minimal space overhead. Unfortunately, applying erasure coding to PM introduces significant additional latency. Previous work to mitigate coding latency has primarily focused on optimizing computational efficiency. Instead, we reveal that the main performance bottleneck is high memory latency due to inefficient hardware prefetchers, rather than computation. We further observe that the prefetching inefficiency mainly results from: (i) too wide or narrow coding stripes, (ii) small block sizes, and (iii) high concurrency.
Guanglei Xu, Hai Zhou 0002, Yuchong Hu, Dan Feng 0001, Renzhi Xiao
ICPP1
2025 CodePM: Parity-Based Crash Consistency for Log-Free Persistent Transactional Memory
abstract
Emerging persistent memory (PM) can provide large persistent capacity with performance comparable to DRAM in modern memory systems. Persistent transactional memory (PTM) needs to ensure data consistency after unexpected power loss or crashes. Therefore, crash consistency strategies, such as persistent logging, are still required. However, the additional overhead introduced by these strategies, such as significant extra writes on PM, can lead to system performance degradation. In this article, we propose CodePM, a fault-tolerant PM transactional library that utilizes parity-based crash consistency to remove logging overhead while guaranteeing the correct state of data. CodePM reuses the decoding capability of parity to detect and recover inconsistent objects. To ensure consistency without logs when updating, CodePM employs fine-grained memory fences to carefully align potential inconsistency with the repairability of parity. To detect inconsistency without logs when recovering, CodePM utilizes optimistic speculative scanning recovery by reusing checksum and parity, which supports instant recovery with transient degraded reliability. Moreover, we study the memory fence blocking effects and further augment CodePM with pipelined encoding and persistent writing to hide update latency. We implemented CodePM on Pangolin, the state-of-the-art parity-based PTM for fault-tolerance. Evaluation results with real-world workloads on Intel Optane DCPMM show that CodePM can achieve up to$3.4\times $higher throughput than Pangolin.
Guanglei Xu, Yuchong Hu, Dan Feng 0001, Wenpeng He, Junyuan Huang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Read-Optimized Persistent Hash Index for Query Acceleration through Fingerprint Filtering and Lock-Free Prefetching
abstract
Hash indexes are widely used in key-value storage systems due to their ability to perform rapid single-point queries. The persistent memory (PM) technology has received significant attention in both academia and industry due to its high performance, non-volatility, and large capacity characteristics. Currently, hash indexes tailored for persistent memories have been extensively researched. However, through an in-depth experimental study, we have discovered that existing persistent hash indexes suffer from low query performance. This is primarily due to persistent memory's higher read latency than DRAM's, which reduces the performance of both positive and negative queries in persistent hash indexes. Additionally, the former's higher read lock overhead further diminishes query performance. To address the above problems, we propose in this paper a Read-Optimized Persistent Hash Index, referred to as ROPHI, based on fingerprint filtering and lock-free prefetching. By employing a fingerprint filtering method, ROPHI introduces a DRAM-based Cuckoo filter to store fingerprints of keys on top of the PM-based hash table, effectively mitigating the time-consuming access overhead of persistent memory hash tables by accessing only the DRAM-based filter. Additionally, ROPHI employs lock-free prefetching for positive query acceleration, utilizing lock-free optimistic concurrent read techniques to avoid read lock overhead and high-speed cache prefetching techniques to reduce access overhead to persistent memory. Experimental results on the Intel Optane DC Persistent Memory Module (DCPMM) platform demonstrate that ROPHI significantly improves query performance over existing persistent hash index schemes. Specifically, ROPHI achieves an improvement of 2.67×-13.59× in negative query performance and 1.72x-7.86x in positive query performance. ROPHI outperforms the state-of-the-art SmartHT in positive query throughput by 34.5%, and in insertion and deletion throughput by 9.20% and 19.87% respectively, while sacrificing only 1.93% of negative query throughput. Additionally, it achieves a 5.07x improvement in recovery efficiency.
Renzhi Xiao, Dan Feng 0001, Yuchong Hu, Hong Jiang 0001, Lanlan Cui, Guanglei Xu, Fang Wang 0001
ICCD8