EDBT 2026 Demo / reviewers in the wild / expert
Xueliang Wei
dblp:07/7742
· DBLP profile ↗
27ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0003-3571-1702ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 5 first-author · 19 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An IR drop-robust Mapping Method for Reliable Memristive AcceleratorsabstractMemristive accelerators (MAs) facilitate efficient matrix-vector multiplication (MVM) by performing in situ computation within memory crossbar arrays, thereby ensuring a fast and energy-efficient application acceleration. A significant challenge associated with the MA lies in the limited computing accuracy caused by the IR drop effect. However, existing IR drop mitigation works provide an approximate compensation, resulting in less accurate results. In this paper, we propose an IR drop-robust mapping method for reliable memristive accelerators. Firstly, the IR drop-robust mapping (IRM) method exploits the residuals between the equivalent matrix after the IR drop effect and the original matrix, and iteratively maps them to the crossbars for IR drop compensation. Based on the IRM method, a novel mechanism of the matrix-vector multiplication (MVM) operation is derived, ensuring that MVM is computed correctly. Secondly, the Calibrate-Shift-Reflect (CSR) strategy is developed to significantly reduce the number of arrays required by the IRM method to map the residuals. Thirdly, the hardware support for the IRM method is designed, and the overhead is reduced by sharing drivers/selectors between neighboring arrays. The experimental results indicate that the IRM-CSR method can effectively mitigate the IR drop effect, restoring inference accuracy by at most 80% (for neural network applications), and achieving a reduction in the relative root-mean-squared error by 103×~1010× (for scientific computing), compared with the state-of-the-art methods. Shiyi Song, Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
DATE | 6 |
| 2026 | Secret Caching Sauce for High-Performance Secure Memory
Xu Jiang 0005, Xueliang Wei, Yifei Qu, Dan Feng 0001, Yulai Xie 0002, Wei Tong 0001 |
HPCA | 2 |
| 2026 | pTree: Building Efficient B${}^{+}$+-Tree on Non-Volatile Memory With Processing-in-MemoryabstractB+-Trees are widely used in storage systems and diverse applications. However, their performance is hindered by frequent memory accesses required for key comparisons and structural maintenance. The limited parallelism of CPUs, along with their sensitivity to data order and volume, further restricts the efficiency of B+-Trees. Processing-in-memory (PIM) offers a promising alternative with its in-situ parallel computing capability. In this paper, we propose pTree, a novel set of PIM techniques and architecture specifically designed to accelerate B+-Tree operations. pTree introduces in-situ parallel comparison mechanisms that significantly reduce costly memory accesses and inefficient CPU-side comparisons during tree traversal. This parallelism eliminates the need to maintain intra-node key order and, together with our in-situ parallel node bisecting techniques, greatly minimizes tree structure maintenance overhead during insertions and deletions. Additionally, pTree decouples both comparison and modification efficiency from node size, enabling the use of larger nodes to reduce tree height and traversal complexity. Evaluation shows that pTree reduces the average latency to 35%, 26%, 52%, and 32% compared to state-of-theart B+-Trees forInsert, Search, Update,andDeleteoperations. Bing Wu 0001, Shiyi Song, Xueliang Wei, Huan Cheng, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Computers | 4 |
| 2025 | MPFS: A Scalable User-Space Persistent Memory File System for Multiple Processesabstract11This work was supported by the Young Scientists Fund of the National Natural Science Foundation of China under Grant 62302182.Persistent memory (PM) leveraging memory-mapped I/O(MMIO) delivers superior I/O performance, leading to the development of user-space PM file systems based on MMIO. While effective in single-process scenarios, these systems encounter challenges in multi-process environments, such as performance degradation due to repeated page faults and cross-process synchronizations, as well as a large memory footprint from duplicated paging structures. To address these problems, we propose a Multi-process PM File System (MPFS). MPFS builds a shareable page table and shares it among processes, avoiding building duplicate paging structures for distinct processes, thereby significantly reducing the software overhead and memory footprint caused by repeated page faults. MPFS further proposes a PGD-aligned (512GB) mapping method to accelerate page table sharing. Furthermore, MPFS provides a cross-process memory protection mechanism based on the PGD-aligned mapping, ensuring multi-process data reliability with negligible overheads. The experimental results show that MPFS outperforms existing user-space PM file systems by 1560% in multi-process scenarios. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Yuchong Hu, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
DATE | 8 |
| 2025 | COVER: Alleviating Crash-Consistency Error Amplification in Secure Persistent Memory SystemsabstractData security (including confidentiality, integrity, and availability) and crash consistency guarantees are essential for building trusted persistent memory (PM) systems. Security and consistency metadata are added to enable the guarantees. Recent studies show that errors in security metadata have the amplified effect, which significantly affects data availability. However, the impact of consistency metadata errors on data availability has rarely been discussed. We identify the crash-consistency error amplification (CCEA) problem, several errors in consistency metadata can make a large portion of data in PM possibly inconsistent. The error sensitivity of consistency metadata is higher than data and security metadata, thus requiring special attention. It is inefficient to address this problem by using the methods that are proposed to alleviate the amplified effect of security metadata errors, because security metadata are generally designed for a single purpose (e.g., integrity verification), while consistency metadata are designed for multiple purposes, including inconsistency locating and recovery. To effectively and efficiently alleviate the CCEA problem, we propose a c rash c o nsistency ver ification approach (COVER) that decouples inconsistency locating and recovery. COVER provides three design options that support different tradeoffs between effectiveness and efficiency. Experimental results show that COVER effectively alleviates the problem with only about 1.0% performance degradation on average compared with the state-of-the-art secure PM design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Bing Wu 0001, Xu Jiang 0005 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | SEED: Speculative Security Metadata Updates for Low-Latency Secure MemoryabstractSecuring systems’ main memory is important for building trusted data centers. To ensure memory security, encryption and integrity verification techniques update the security metadata (e.g., encryption counters and integrity trees) during memory data writes. Existing studies are optimistic about the effect of data writes on system performance since they regard all data writes as background operations. However, we show that security metadata updates significantly increase data write latency. High-latency data writes frequently fill up write buffers in the system, forcing the system to perform the writes in the critical path. As a result, performance-critical data reads need to wait for the execution of these writes, which increases data read latency and degrades system performance. In this paper, we propose SEED that improves the performance of secure memory systems by speculatively updating security metadata in the background before data writes arrive. To enable speculative updates, SEED predicts which dirty cache lines will be written to memory through natural evictions. We find that cache evictions depend on multiple factors. To decouple the dependencies for accurate predictions, we devise a two-step eviction prediction method based on our observation that the next eviction victim rarely changes in a set. The first step predicts which cache sets will evict cache lines, while the second step predicts which cache lines will be evicted by finding the next eviction victims in the sets. For predicted evictions, we develop a speculative updater to perform speculative updates. We analyze the invariants that must be followed by the updater to ensure the correctness of speculative updates. The updater rolls back the speculatively updated security metadata of inaccurate predictions. To reduce the rollback overhead, we devise a rollback batching and an update pausing optimization for the updater. Experimental results show that SEED reduces data write latency by 39.8%, data read latency by 44.9%, and improves performance by 40.0% on average compared with the state-of-the-art secure memory design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Bing Wu 0001, Xu Jiang 0005 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | An Efficient Independent Read Scheme for Contemporary QLC SSDsabstractQLC solid-state disks (SSDs) are increasingly deployed in large-scale storage systems. While achieving remarkable storage density and cost-effectiveness, QLC NAND exhibits degraded performance. To alleviate the issue, Independent Multi-Plane (IMP) read has been proposed to leverage the plane-level parallelism under random read workloads. However, compared to the previous generations of chips, the variation in read latency of QLC chip has widened significantly, and the number of planes in a QLC chip has increased. As a result, the idle time in IMP commands has escalated dramatically. Moreover, the conventional layout of data and parities in redundant array of independent NAND (RAIN) increases the probability of high latency read occureneces, exacerbating the contribution to idle time. Consequently, the incorporation of IMP is inherently inefficient in contemporary QLC NAND flash chips. In this paper, we propose an efficient independent read (EFFIR) scheme to tackle this challenge. EFFIR features a read latency variation aware transaction service that properly combines read transactions to minimize idle time and proactively transfers read data to eliminate unnecessary delays. Moreover, EFFIR incorporates a read latency variation aware RAIN that reorganizes the layout of data and parity to mitigate the impact of high-latency data access on idle time. Our comprehensive experimental results demonstrate and elucidate how EFFIR significantly enhances SSD responsiveness while consistently delivering favorable performance across a diverse range of read-intensive workloads. Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001, Feng Zhu 0024, Maojun Yuan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | FADESIM: Enable Fast and Accurate Design Exploration for Memristive Accelerators Considering NonidealitiesabstractMemristive accelerators (MAs), built with memristive crossbar arrays (MCAs), have gained significant attention for their ability to efficiently perform matrix-vector multiplication in diverse applications. Modeling and simulation are indispensable tools for exploring and evaluating architectural design. Specifically, for MAs, maintaining high computation accuracy has been challenging because of realistic nonidealities like wire resistance (i.e., IR drop), I–V nonlinearity, program variation, and so on. Thus, for the system architects, fast and exact analysis of the effects of nonidealities is highly desirable, especially during the vast early design space exploration of the architecture. However, the SPICE model and existing MCA compact models (CMs) do not offer acceptable speeds for accurate simulation purposes, making them less practical. Additionally, the existing MCA simplified circuit models and predictive models fail to produce effective results due to the complexity of IR drop, let alone the coupling of multiple nonidealities. To enable fast and accurate design exploration for MAs with nonidealities considered, in this article, we propose FADESIM which includes fast IR drop simulation methods and a processing chain for the joint simulation of multiple nonidealities. Starting with the analysis of the accurate CM for IR drop, we explore the special properties of the model to enable fast iterative methods as well as a method to skip invalid calculations. This significantly reduces the time complexity of the simulation from naive$O(n^{6})$to$O(n^{3})$. For less severe IR drop cases, a custom iterative update algorithm is presented for faster simulations as a supplement, specifically with a time complexity of near$O(n^{2})$and proven applicable conditions. To simulate multiple nonidealities, we introduce a processing chain to inject corresponding processing functions before, during, and after the proposed fast IR drop simulation process according to the stages at which nonidealities take effect. The array-level experimental results show that our method achieves accurate simulations, with$19.8 \times - 884.9 \times $faster and$276.7 \times - 8018.8 \times $reduced memory usage compared to SPICE simulation. Further experiments at the algorithm-level demonstrate the effectiveness of our method in assisting architects with evaluating their designs. Bing Wu 0001, Huan Cheng, Xueliang Wei, Wei Tong 0001, Dan Feng 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | CMD: A Cache-Assisted GPU Memory Deduplication ArchitectureabstractMassive off-chip accesses in graphics processing units (GPUs) are the main performance bottleneck. We find that many writes are duplicate, and the duplication can beinter-dupandintra-dup. Whileinter-dupmeans different memory blocks are identical, andintra-dupmeans all the 4B elements in a line are the same. In this work, we propose a cache-assisted GPU memory deduplication architecture named cache-assisted GPU memory deduplicated (CMD) to reduce the off-chip accesses via utilizing the data duplication in GPU applications. CMD includes three key design contributions which aim to reduce the three kinds of accesses: 1) a novel GPU memory deduplication architecture that removes theintra-dupandinter-duplines. We design several techniques to manage duplicate blocks, reducing massive off-chip writes; 2) we propose a cache-assisted read scheme to reduce the reads to duplicate data. When an L2 cache miss wants to read the duplicate block, if the reference block has been fetched to L2 and it is clean, we can copy it to the L2 missed block without accessing off-chip DRAM. As for the reads tointra-dupdata, CMD uses the on-chip metadata cache to get the data; and 3) when a cache line is evicted, the clean sectors in the line are invalidated while the dirty sectors are written back. However, most read-only victims are rereferenced from DRAM more than twice. Therefore, we add a full-associate FIFO to accommodate the read-only (it is also clean) victims to reduce the rereference counts. Experiments show that CMD can decrease the off-chip accesses by 31.01%, reduce the energy by 32.78% and improve performance by 42.53%. Besides, CMD can improve the performance of memory-intensive workloads by 57.56%. Wei Zhao 0034, Dan Feng 0001, Wei Tong 0001, Xueliang Wei, Bing Wu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | A Read Latency Variation Aware Independent Read Scheme for QLC SSDsabstractQLC flash-based SSDs has attracted growing interest and is expected to fit in read-intensive scenarios owing to its higher cost-effectiveness and shorter write endurance compared with the Triple-Level Cell (TLC) SSDs. Recently, new commands supporting independent reads such as Single Operation Multiple Locations (SOML) and Independent multi-plane (IMP) read are proposed to improve read performance. Unfortuantely, while independent read exhibits a significant performance improvement, we show in this paper that the exisiting approach fails to fully exploit its potential due to the larger read latency variation and more planes per die for current SSD architecture. Through a set of experiments, we demonstrate that the lack of a read latency variation aware machanism leads to low performance of independent read among a wide variety of workloads. To alleviate this issue, we propose LITA, which key idea is to combine read transactions with similar latency into one command. LITA includes (1) LIT, a latency variation aware transaction combination, (2) and TASAP, a latency variation aware transaction completion service. The experimental results show LITA can reduce read latency by 20.4% and 9.6% on average for 4-planes QLC SSDs under IMP and SOML, respectively. Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001 |
DATE | 6 |
| 2024 | Enabling Reliable Memory-Mapped I/O With Auto-Snapshot for Persistent Memory SystemsabstractPersistent memory (PM) is promising to be the next-generation storage device with better I/O performance. Since the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of being corrupted by scribbles and non-atomic I/O interfaces, causing serious reliability issues. To address these issues, we propose RMMIO, an efficient user-space library that provides reliable MMIO for PM systems. RMMIO provides atomic I/O interfaces and lightweight snapshots to ensure the reliability of MMIO. Compared with existing schemes, RMMIO mitigates additional writes and extra software overheads caused by reliability guarantees, thus achieving MMIO-like performance. In addition, we also propose an automatic snapshot with efficient memory management for RMMIO to minimize data loss incurred by reliability issues. The experimental results of microbenchmarks show that RMMIO achieves 8.49x and 2.31x higher throughput than ext4-DAX and the state-of-the-art MMIO-based scheme, respectively, while ensuring data reliability. The real-world application accelerated by RMMIO achieves at most 7.06x higher throughput than that of ext4-DAX. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
IEEE Trans. Computers | 5 |
| 2023 | ODLPIM: A Write-Optimized and Long-Lifetime ReRAM-Based Accelerator for Online Deep LearningabstractReRAM-based Processing-In-Memory (PIM) architectures have demonstrated high energy efficiency and performance in deep neural network (DNN) acceleration. Most of the existing PIM accelerators for DNN focus on offline batch learning (OBL) which requires the whole dataset to be available before training. However, in the real world, data instances arrive in sequential settings, and even the data pattern may change, which calls concept drift. OBL requires expensive retraining to solve concept drift, whereas online deep learning (ODL) is evidenced to be a better solution to keep the model evolving over streaming data. Unfortunately, when ODL optimizes models over a large-scale data stream in the PIM system, unbalanced writes are more severe than OBL, due to the heavier weight updates, resulting in the amplification of unbalanced writes and lifetime deterioration. In this work, we propose ODLPIM, an online deep learning PIM accelerator that extends the system lifetime through algorithm-hardware co-optimization. ODLPIM adopts a novel write-optimized parameter update (WARP) scheme that reduces the non-critical weight updates in hidden layers. Besides, a table-based inter-crossbar wear-leveling (TIWL) scheme is proposed and applied to the hardware controller to achieve wear-leveling between crossbars for lifetime improvement. Experiments show that WARP reduces weight updates on average to 15.25% and up to 24% compared to that without WARP, and eventually prolongs system lifetime on average to 9.65% and up to 26.81%, with a negligible rise in cumulative error rate (up to 0.31%). By combining WARP with TIWL, the lifetime of ODLPIM is improved by an average of$\mathbf{12}.\mathbf{59}\times$and up to$\mathbf{17}.\mathbf{73}\times$. Bing Wu 0001, Huan Cheng, Wei Zhao 0034, Xueliang Wei, Dan Feng 0001, Wei Tong 0001 |
DATE | 5 |
| 2023 | LifetimeKV: Narrowing the Lifetime Gap of SSTs in LSMT-based KV Stores for ZNS SSDsabstractZone Namespace (ZNS) SSDs delegate data placement and garbage collection (GC) to the host, enabling applications on the host to perform efficient GC. The existing works on LSMT-based KV stores adopt ZenFS (a user-level file system) to manage ZNS SSDs. ZenFS assumes that SSTs within the same level have similar lifetimes and places SSTs with similar lifetimes into the same zone to minimize data migration in GC. However, we observe significant disparity in the lifetimes of SSTs within the same level, resulting in fragmented zones and huge data migration in GC. To reduce data migration in GC and improve performance, we present LifetimeKV, an LSMT-based KV store for ZNS SSDs. LifetimeKV introduces a range compaction algorithm to reduce short-lived SSTs and an overlap-ratio-lifetime victim SST selection algorithm to reduce long-lived SSTs, thereby reducing the lifetime disparity among SSTs within the same level. We evaluate LifetimeKV performance on a real ZNS SSD. The results demonstrate that LifetimeKV reduces data migration in GC by 63.23% and achieves a throughput improvement of 98.81% under write-intensive workload than state-of-the-art work. Biyong Liu, Yuan Xia, Xueliang Wei, Wei Tong 0001 |
ICCD | 3 |
| 2023 | Accelerating Persistent Hash Indexes via Reducing Negative SearchesabstractHashing is a widely used and efficient indexing mechanism for key-value storage. Persistent memory (PM) has attracted extensive attention in research due to its non-volatility and DRAM-like performance. Intel DCPMM, as a PM, can provide large capacity and low total cost of ownership, further promoting the research of PM-based hash index. However, based on real-world workloads, we found that negative searches of existing PM-based hash indexes significantly degrade system performance. A direct method to solve this problem is to use a PM-based Bloom filter to reduce negative searches, but at the cost of the decreased lifespan of PM due to extra PM writes. An alternative method is to use a DRAM-based Bloom filter, but it still faces increased multi-threaded insertion/deletion/positive-search scalability overhead as well as increased data consistency and recovery overhead.In this paper, we propose SmartHT, a small-size DRAM-based Bloom filter to accelerate hash table operations for PM while solving the aforementioned problems. SmartHT uses efficient merge write optimization with head insertion, lazy deletion, and shortened average chained length of head-bucket to provide high insertion/deletion/positive-search scalability, respectively. On the other hand, it utilizes a merged-flush mechanism based on an 8-byte failure-atomic write method to reduce flush instructions and extra PM writes to achieve low data consistency overhead. Experimental results on Intel Optane DCPMM show that, compared with the state-of-the-art persistent hash indexes, SmartHT improves multi-threaded negative queries under uniform and skewed distributions by 4.61x-13.86x and 2.76x-12.99x respectively, achieves high multi-threaded scalability and low data consistency overhead, at the modest cost of recovery time overhead. Renzhi Xiao, Hong Jiang 0001, Dan Feng 0001, Yuchong Hu, Wei Tong 0001, Kang Liu 0017, Xueliang Wei, Zhengtao Li |
ICCD | 8 |
| 2023 | SplitZNS: Towards an Efficient LSM-Tree on Zoned Namespace SSDsabstractThe Zoned Namespace (ZNS) Solid State Drive (SSD) is a nascent form of storage device that offers novel prospects for the Log Structured Merge Tree (LSM-tree). ZNS exposes erase blocks in SSD as append-only zones, enabling the LSM-tree to gain awareness of the physical layout of data. Nevertheless, LSM-tree on ZNS SSDs necessitates Garbage Collection (GC) owing to the mismatch between the gigantic zones and relatively small Sorted String Tables (SSTables). Through extensive experiments, we observe that a smaller zone size can reduce data migration in GC at the cost of a significant performance decline owing to inadequate parallelism exploitation. In this article, we present SplitZNS, which introduces small zones by tweaking the zone-to-chip mapping to maximize GC efficiency for LSM-tree on ZNS SSDs. Following the multi-level peculiarity of LSM-tree and the inherent parallel architecture of ZNS SSDs, we propose a number of techniques to leverage and accelerate small zones to alleviate the performance impact due to underutilized parallelism. (1) First, we use small zones selectively to prevent exacerbating write slowdowns and stalls due to their suboptimal performance. (2) Second, to enhance parallelism utilization, we propose SubZone Ring, which employs a per-chip FIFO buffer to imitate a large zone writing style; (3) Read Prefetcher, which prefetches data concurrently through multiple chips during compactions; (4) and Read Scheduler, which assigns query requests the highest priority. We build a prototype integrated with SplitZNS to validate its efficiency and efficacy. Experimental results demonstrate that SplitZNS achieves up to 2.77× performance and reduces data migration considerably compared to the lifetime-based data placement. 1 Dan Feng 0001, Bo Ding 0002, Wei Zhao 0034, Xueliang Wei, Wei Tong 0001 |
ACM Trans. Archit. Code Optim. | 6 |
| 2023 | A Low-Latency and High-Endurance MLC STT-MRAM-Based Cache SystemabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising cache memory candidate due to its high density, low leakage power, and nonvolatility. Multilevel cell (MLC) STT-MRAM can further increase density by storing 2 bits in one cell’s hard and soft domain, respectively. However, MLC STT-MRAM suffers two-step write, leading to high write energy, long latency, and severe lifetime degradation. Current encoding techniques propose to encode the new data to reduce the two-step data writes. However, they have two weaknesses: 1) high area overhead, e.g., recent work TSE (Hsieh et al., 2020) needs extra 37.5% MLCs and 2) prolong the write latency due to an extra read. Therefore, we propose enhanced one-step write (EOSwrite) to write data in one step. EOSwrite includes line bypassing and four intraline encoding techniques. Line bypassing schemes can bypass the writes to zero or clean lines, leading to low write/read latency. As for the intraline techniques, we propose four write modes. They utilize the data patterns and the clean data in cache lines to write data in one step, therefore reducing the data write latency. The key idea of one-step write is to write as much data as possible in the soft domain of MLC STT-MRAM. EOSwrite can greatly relieve the weaknesses of the current encoding schemes. Evaluation results show that EOSwrite can improve the lifetime of MLC STT-MRAM by 56.96%, reduce dynamic energy by 33.95%, reduce access latency by 36.95%, and improve system performance of MLC STT-MRAM by 4.30%, respectively. While the area overhead of EOSwrite is only 7.27%. Wei Zhao 0034, Jie Xu 0013, Xueliang Wei, Bing Wu 0001, Chengning Wang, Weilin Zhu, Wei Tong 0001, Dan Feng 0001, Jingning Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | RMMIO: Enabling Reliable Memory-Mapped I/O for Persistent Memory SystemsabstractThe byte-addressable persistent memory (PM) is coming to be the next-generation storage device for better I/O performance. As the traditional I/O path is too lengthy to drive PM featuring low latency and high bandwidth, prior works have proposed memory-mapped I/O (MMIO) to shorten the I/O path to PM. However, native MMIO directly maps files into the user address space, which puts files at risk of user-space scribbles and non-atomic I/O interfaces, termed reliability issues. Since existing reliability schemes cause significant extra overheads, we propose RMMIO, an efficient user-space library that provides reliable memory-mapped I/O interfaces for PM systems. RMMIO achieves a good balance between efficiency and reliability by introducing a memory-mapped cache layer upon kernel file systems. The cache layer accelerates I/O requests and carries the file system’s responsibility for data reliability by data isolation. In addition, RMMIO further employs lightweight snapshots and efficient atomic I/O interfaces to guarantee the integrity and consistency of the data in the cache layer at low costs. The experimental results show that RMMIO achieves 8.49x higher throughput than ext4-DAX and 2.31x higher throughput than state-of-the-art MMIO-based schemes for PM while ensuring data reliability. Bo Ding 0002, Wei Tong 0001, Yu Hua 0001, Zhangyu Chen, Xueliang Wei, Dan Feng 0001 |
ICCD | 5 |
| 2021 | Crash-Consistency-Aware Encryption for Non-Volatile MemoriesabstractData security is a very important issue in non-volatile memory (NVM) systems. Due to high security and low decryption latency, counter mode encryption (CME) is often used, which unfortunately suffers from extra efforts to guarantee the crash consistency of counters. Existing schemes fail to efficiently guarantee the crash consistency of counters and data, thus resulting in inefficient recovery and high costs. In this paper, we propose a Crash-Consistency-Aware Encryption scheme (CCAE) for NVM systems. CCAE is inspired by two key observations. First, we observe that all used log entries in a transaction have the same number of encryption times before being recycled due to the write appending feature. This motivates our shared counter optimization for log encryption. Second, we observe that the counters related to uncommitted data blocks only need to be restored to the latest or newer value before the crash to avoid OTP reuse. This motivates our delayed counter persistency scheme for data encryption. Evaluation results show that compared with the state-of-the-art design, CCAE reduces NVM write traffic caused by counters by 67%, improves system performance by 14%, and decreases NVM energy consumption by 35%. Mengya Lei, Fang Wang 0001, Dan Feng 0001, Xueliang Wei |
ICPP | 5 |
| 2021 | Multi-task prediction model based on ConvLSTM and encoder-decoderabstractThe energy load data in the micro-energy network are a time series with sequential and nonlinear characteristics. This paper proposes a model based on the encode-decode architecture and ConvLSTM for multi-scale prediction of multi-energy loads in the micro-energy network. We apply ConvLSTM, LSTM, attention mechanism and multi-task learning concepts to construct a model specifically for processing the energy load forecasting of the micro-energy network. In this paper, ConvLSTM is used to encode the input time series. The attention mechanism is used to assign different weights to the features, which are subsequently decoded by the decoder LSTM layer. Finally, the fully connected layer interprets the output. This model is applied to forecast the multi-energy load data of the micro-energy network in a certain area of Northwest China. The test results prove that our model is convergent, and the evaluation index value of the model is better than that of the multi-task FC-LSTM and the single-task FC-LSTM. In particular, the application of the attention mechanism makes the model converge faster and with higher precision. Xudong Cao, Xueliang Wei |
Intell. Data Anal. | 6 |
| 2021 | Improving Multilevel Writes on Vertical 3-D Cross-Point Resistive MemoryabstractResistive memory is promising to be constructed as a high-density storage-class memory. Multilevel cell, access-transistor-free cross-point array structure, and 3-D array integration are three approaches to scale up the density of resistive memory. However, composing the three approaches together strengthens the interactions between array-level and cell-level nonidealities (interconnect resistance-induced IR drop, sneak current, and device variability) of resistive memory arrays during write operations and significantly degrades write performance and reliability. In this article, we analyze the dynamic voltage-dividing effect along a selected write current path in 3-D cross-point memory arrays. We propose a nonideality-tolerant high-density resistive memory (HD-RRAM) architecture, that can weaken the interactions between nonidealities and mitigate their degradation effects on the performance and reliability of array multilevel write operations. HD-RRAM is equipped with a double-transistor array architecture with two-transistor- n-resistor (2TnR) cell organization along pillars to reduce the current driving requirement and the large undesired voltage drop across each vertical pillar access transistor. Moreover, multiside asymmetric bias improves the resistive switching velocity by leveraging current-dividing effects. Variability-aware multilevel state partition reduces the worst-case write error rate by leveraging target state dependency of variability. Proportional-control multilevel state tuning reduces the average number of required write-and-verify iterations by leveraging pulse amplitude dependency of variability. Multilevel cell parallel writing improves the cell-level parallelism by leveraging the pass-through feature of intermediate resistance states. The evaluations show that HD-RRAM reduces both memory access latency and energy consumption over an aggressive baseline. Chengning Wang, Dan Feng 0001, Wei Tong 0001, Yu Hua 0001, Jingning Liu, Bing Wu 0001, Wei Zhao 0034, Linghao Song, Yang Zhang 0051, Jie Xu 0013, Xueliang Wei, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 11 |
| 2020 | CCHL: Compression-Consolidation Hardware Logging for Efficient Failure-Atomic Persistent Memory UpdatesabstractNon-volatile memory (NVM) is emerging as a fast byte-addressable persistent memory (PM) that promises data persistence at the main memory level. One of the common choices for providing failure-atomic updates in PM is the write-ahead logging (WAL) technique. To mitigate logging overhead, recent studies propose WAL-based hardware logging designs that overlap log writes with transaction execution. However, existing hardware logging designs incur a large number of unnecessary log writes. Many log writes are still performed in the critical path, which causes high performance overhead, particularly for the multi-core systems with many threads. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Chengning Wang, Liuqing Ye |
ICPP | 1 |
| 2020 | MorLog: Morphable Hardware Logging for Atomic Persistence in Non-Volatile Main MemoryabstractByte-addressable non-volatile memory (NVM) is emerging as an alternative for main memory. Non-volatile main memory (NVMM) systems are required to support atomic persistence and deal with the high overhead of programming NVM cells. To this end, recent studies propose hardware logging and data encoding designs for NVMM systems. However, prior hardware logging designs incur either extra ordering constraints or redundant log data. Moreover, existing data encoding designs are unaware of the characteristics of log data, resulting in writing unnecessary log bits.In this paper, we propose a morphable hardware logging design (MorLog) that only logs the data necessary for recovery and dynamically selects encoding methods with least write overhead. We observe that (1) only the oldest undo and the newest redo data in each transaction are necessary for recovery, and (2) the log data for clean bits are clean. The first motivates our morphable logging mechanism. This mechanism logs both undo and redo data for the first update to the data in a transaction, and then logs only redo data. Undo data are eagerly written to NVMM to ensure atomicity, while redo data are buffered in a volatile log buffer and L1 caches to write only the newest redo data to NVMM. The second motivates our selective log data encoding mechanism. This mechanism simultaneously encodes log data with different methods, and writes the encoded log data with the least write cost to NVMM. We devise a differential log data compression method to exploit the characteristics of log data. This method directly discards clean bits from log data and compresses remained dirty bits. Our evaluation shows that MorLog improves performance by 72.5%, reduces NVMM write traffic by 41.1%, and decreases NVMM write energy by 49.9% compared with the state-of-the-art design. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Liuqing Ye |
ISCA | 1 |
| 2020 | STC: Sub-packetization tunable codes for fast recovery
Liuqing Ye, Dan Feng 0001, Yuchong Hu, Xueliang Wei |
J. Syst. Archit. | 4 |
| 2020 | Hybrid Codes: Flexible Erasure Codes with Optimized Recovery PerformanceabstractErasure codes are being extensively deployed in practical storage systems to prevent data loss with low redundancy. However, these codes require excessive disk I/Os and network traffic for recovering unavailable data. Among all erasure codes, Minimum Storage Regenerating (MSR) codes can achieve optimal repair bandwidth under the minimum storage during recovery, but some open issues remain to be addressed before applying them in real systems. Facing with the huge burden during recovery, erasure-coded storage systems need to be developed with high repair efficiency. Aiming at this goal, a new class of coding scheme is introduced—Hybrid Regenerating Codes (Hybrid-RC). The codes utilize the superiority of MSR codes to compute a subset of data blocks while some other parity blocks are used for reliability maintenance. As a result, our design is near-optimal with respect to storage and network traffic and shows great improvements in recovery performance. Liuqing Ye, Dan Feng 0001, Yuchong Hu, Xueliang Wei |
ACM Trans. Storage | 4 |
| 2019 | A Generic Construction for All Parameters in Minimum Storage Regenerating CodesabstractMinimum-Storage Regenerating (MSR) codes have become superior alternatives to traditional erasure codes as they can provide optimal repair bandwidth while the reliability and the storage overhead are still optimal. So far, the state-of-the-art MSR codes mainly focus on connecting to all the remaining nodes to repair a single failure. Since the recovery latency may be bottlenecked by the time taken to retrieve the slowest or straggling block, it is however impractical to have the highest connectivity in MSR codes. In this paper, we introduce a generic construction for all parameters in MSR codes, which allows for bandwidth-efficient repair of a single data node failure with an arbitrary (but fixed) number of accessed nodes d. Our method provides explicit and generic encoding and repairing processes, and such codes were not previously known to exist. When d = n - 1, we show that our codes are not inferior to the other MSR codes and retain the same recovery optimality. Furthermore, the arbitrary number of d allows our codes to adapt to the late binding strategy to avoid stragglers (overloaded sites) for efficient recovery performance. As a result, our codes outperform both traditional erasure codes and the state-of-the-art MSR codes on the aspect of the response time when the system is subjected to an imbalanced load. Liuqing Ye, Dan Feng 0001, Yuchong Hu, Xueliang Wei |
SRDS | 4 |
| 2019 | NICO: Reducing Software-Transparent Crash Consistency Cost for Persistent MemoryabstractEmerging non-volatile byte-addressable memory (NVM) introduces many opportunities and challenges to memory system designs. As data become persistent at main memory level, persistent memory systems need to guarantee the consistent state of data in the event of system failures (i.e., crash consistency). Existing studies propose persistent memory designs with software-transparent crash consistency guarantee to reduce programmers' manual effort when taking advantage of persistent memory. However, these designs are suboptimal due to their performance overhead caused by creating checkpoints. In this paper, we propose a Non-Intrusive memory COntroller design (NICO) that uses backend operations for achieving software-transparent crash consistency with minimized checkpointing overhead. By moving data persist operations to the background, NICO fully decouples data persist operations from volatile execution and cache management. To efficiently enforce crash consistency, we design a lightweight checkpointing scheme which only needs to flush and modify a very small amount of data when creating a consistent snapshot of persistent memory data. Our results show that NICO reduces the percent of time spent on checkpointing to within 0.9 percent across different benchmarks, and improves performance by 2.04× compared with existing checkpoint-based designs on average. Xueliang Wei, Dan Feng 0001, Wei Tong 0001, Jingning Liu, Liuqing Ye |
IEEE Trans. Computers | 1 |
| 2010 | A signal-noise model for significance analysis of ChIP-seq with negative controlabstractMOTIVATION: ChIP-seq is becoming the main approach to the genome-wide study of protein-DNA interactions and histone modifications. Existing informatics tools perform well to extract strong ChIP-enriched sites. However, two questions remain to be answered: (i) to which extent is a ChIP-seq experiment able to reveal the weak ChIP-enriched sites? (ii) are the weak sites biologically meaningful? To answer these questions, it is necessary to identify the weak ChIP signals from background noise. RESULTS: We propose a linear signal-noise model, in which a noise rate was introduced to represent the fraction of noise in a ChIP library. We developed an iterative algorithm to estimate the noise rate using a control library, and derived a library-swapping strategy for the false discovery rate estimation. These approaches were integrated in a general-purpose framework, named CCAT (Control-based ChIP-seq Analysis Tool), for the significance analysis of ChIP-seq. Applications to H3K4me3 and H3K36me3 datasets showed that CCAT predicted significantly more ChIP-enriched sites that the previous methods did. With the high sensitivity of CCAT prediction, we revealed distinct chromatin features associated to the strong and weak H3K4me3 sites. AVAILABILITY: http://cmb.gis.a-star.edu.sg/ChIPSeq/tools.htm. Han Xu 0013, Lusy Handoko, Xueliang Wei, Chaopeng Ye, Jianpeng Sheng, Chia-Lin Wei, Wing-Kin Sung |
Bioinform. | 3 |