EDBT 2026 Demo / reviewers in the wild / expert
Yajuan Du
dblp:202/9581
· DBLP profile ↗
31ranked-venue papers
6as first author
17since 2021 · last 2025
0000-0002-8937-8055ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-author · 14 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Novel Computational Approach to Transcriptome Mapping in Polyploid SpeciesabstractAccurate transcriptome mapping in polyploid species remains a major computational challenge due to high sequence similarity among homeologous chromosome groups. Misassignment of multi-mapped reads can severely distort gene expression estimates and downstream analyses such as allelespecific expression (ASE) or subgenome dominance studies. We propose a novel computational approach for improving transcriptome mapping accuracy in polyploid species, tested on Fragaria orientalis, a tetraploid strawberry species. We first conducted controlled analysis using two highly similar sequences with SNP-containing genes to examine the effect of SNP-induced divergence on alignment accuracy. While the presence of SNPs alone under standard alignment scoring yielded approximately 89% overall accuracy, a small subset of genes (3.47%) with extremely high sequence similarity still suffered from substantial misassignment, with over 97% of misaligned reads lacking distinguishing SNP information. Based on these limitations, we developed a novel computational approach that leverages unique read density within gene regions to inform multi-mapping decisions. Validation on tetraploid strawberry data demonstrated that our approach significantly improves assignment accuracy from 50% baseline to 70.26%, representing a 40% relative improvement. Yajuan Du, Changxiang Mao, Tianxi Wu, XiaoHui Yuan |
BIBM | 3 |
| 2025 | RemapCom: Optimizing Compaction Performance of LSM Trees via Data Block Remapping in SSDsabstractIn LSM-based KV stores, typically deployed on systems with DRAM-SSD storage, compaction degrades write performance and SSD endurance due to significant write amplification. To address this issue, recent proposals have mostly focused on redesigning the structure of LSM trees. In this paper, we observe the prevalence of data blocks that are are simply read and written back without being altered during the LSM-tree compaction process, which we refer to as Unchanged Data Blocks (UDBs). These UDBs are source of unnecessary write amplification leading to performance degradation and shortening of SSD lifetime. To address this duplication issue, we propose a remapping-based compaction method, which we call RemapCom. RemapCom considers the identification and retention by designing a lightweight state machine to track the status of the KV items in each data block as well as designing a UDB retention strategy to prevent data blocks from being split due to adjacent intersecting blocks. We implement a prototype of RemapCom on LevelDB by providing two primitives for the remapping. Compared to the state-of-the-art, evaluation results demonstrate that RemapCom can reduce the write amplification by up to 53%. Yajuan Du, Sam H. Noh |
DATE | 2 |
| 2025 | Eliminating duplicate writes of logging via no-logging flash translation layer in SSDs
Zhenghao Yin, Yajuan Du, Sam H. Noh |
J. Syst. Archit. | 2 |
| 2025 | Select Edges Wisely: Monotonic Path Aware Graph Layout Optimization for Disk-based ANN Search
Ziyang Yue, Bolong Zheng, Kanru Xu, Shuhao Zhang 0001, Yajuan Du, Yunjun Gao, Xiaofang Zhou 0001, Christian S. Jensen |
Proc. VLDB Endow. | 6 |
| 2024 | Extending SSD Lifetime via Balancing Layer Endurance in 3D NAND Flash MemoryabstractBy stacking layers vertically, 3D flash memory enables continuous growth in capacity. In this paper, we study the layer variation in 3D flash blocks and find that bottom layer pages exhibit the lowest endurance, whereas middle layer pages demonstrate the highest endurance. The imbalanced endurance across different layers will diminish the overall SSD lifetime. To address this issue, we introduce a novel layer-aware write strategy, named LA-Write. It performs write-skip operations with layer-specific probabilities. The endurance of bottom layer pages with the highest probability would be noticeably improved, which could balance the layer endurance. Experiment results show that LA-Write can improve SSD lifetime by 29%. Yajuan Du, Cheng Ji 0002 |
DATE | 2 |
| 2024 | LA-Write: Balancing Endurance of Inter-Layer for Prolonging 3D NAND Flash Memory LifetimeabstractWith vertical stacking, 3D NAND flash memory can achieve continuous capacity growth. However, as the number of stacked layers in a flash block increases, the endurance variation between the stacked layers becomes more and more significant due to process variation, which will seriously affect the lifetime of 3D NAND flash memory. We investigated the endurance variation characteristics between layers and divided the stacked layers into top, middle, and bottom layers according to the endurance characteristics. We found that the endurance of the bottom layer pages is much weaker than that of the other two layers, in response to this endurance variation feature, we proposed a new layer-aware write strategy, called LA-Write. First of all, the write-skip unit in LA-Write will reduce the wear pressure of the pages through write-skip operations. Secondly, LA-Write maintains a layer-aware table, which stores the probability of pages in different layers performing write-skip operation. Setting the probability of the bottom pages to the highest will result in more write-skip operations on the bottom layers, mitigating endurance variations between layers. Experimental results show that LA-Write can increase SSD lifetime by an average of 31%. Yajuan Du |
NAS | 2 |
| 2024 | Characterizing and Optimizing LDPC Performance on 3D NAND Flash MemoriesabstractWith the development of NAND flash memories’ bit density and stacking technologies, while storage capacity keeps increasing, the issue of reliability becomes increasingly prominent. Low-density parity check (LDPC) code, as a robust error-correcting code, is extensively employed in flash memory. However, when the RBER is prohibitively high, LDPC decoding would introduce long latency. To study how LDPC performs on the latest 3D NAND flash memory, we conduct a comprehensive analysis of LDPC decoding performance using both the theoretically derived threshold voltage distribution model obtained through modeling (Modeling-based method) and the actual voltage distribution collected from on-chip data through testing (Ideal case). Based on LDPC decoding results under various interference conditions, we summarize four findings that can help us gain a better understanding of the characteristics of LDPC decoding in 3D NAND flash memory. Following our characterization, we identify the differences in LDPC decoding performance between the Modeling-based method and the Ideal case. Due to the accuracy of initial probability information, the threshold voltage distribution derived through modeling deviates by certain degrees from the actual threshold voltage distribution. This leads to a performance gap between using the threshold voltage distribution derived from the Modeling-based method and the actual distribution. By observing the abnormal behaviors in the decoding with the Modeling-based method, we introduce an Offsetted Read Voltage (ΔRV) method for optimizing LDPC decoding performance by offsetting the reading voltage in each layer of a flash block. The evaluation results show that our ΔRV method enhances the decoding performance of LDPC on the Modeling-based method by reducing the total number of sensing levels needed for LDPC decoding by 0.67% to 18.92% for different interference conditions on average, under the P/E cycles from 3,000 to 7,000. Qiao Li 0001, Guanyu Wu, Yajuan Du, Xinbiao Gan, Jie Zhang 0048, Zhirong Shen, Jiwu Shu, Chun Jason Xue |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | GPU Performance Optimization via Intergroup Cache CooperationabstractModern GPUs have integrated multilevel cache hierarchy to provide high bandwidth and mitigate the memory wall problem. However, the benefit of on-chip cache is far from achieving optimal performance. In this article, we investigate existing cache architecture and find that the cache utilization is imbalanced and there exists serious data duplication among L1 cache groups.In order to exploit the duplicate data, we propose an intergroup cache cooperation (ICC) method to establish the cooperation across L1 cache groups. According the cooperation scope, we design two schemes of the adjacent cache cooperation (ICC-AGC) and the multiple cache cooperation (ICC-MGC). In ICC-AGC, we design an adjacent cooperative directory table to realize the perception of duplicate data and integrate a lightweight network for communication. In ICC-MGC, a ring bi-directional network is designed to realize the connection among multiple groups. And we present a two-way sending mechanism and a dynamic sending mechanism to balance the overhead and efficiency involved in request probing and sending.Evaluation results show that the proposed two ICC methods can reduce the average traffic to L2 cache by 10% and 20%, respectively, and improve overall GPU performance by 19% and 49% on average, respectively, compared with the existing work. Yajuan Du |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Transparent File Deduplication with Reduced Update Cost on Encryption Enabled Mobile DevicesabstractData deduplication has been long studied to achieve data reduction. However, deploying deduplication on encryption enabled mobile systems might consume much memory footprint and computation time for hash calculations. Moreover, frequent file updates on deduplicated files could badly degrade the deduplication efficacy due to the increased file-system metadata penalty. Considering the characteristics of mobile devices, an efficient data deduplication method is proposed in this paper. First, it separates the hash calculation into foreground and background stages. The background stage calculates the hash values of potentially duplicate files while the foreground stage quickly hashes the file which is being written using a lightweight hash algorithm. Second, a dual-level node structure is proposed to improve the file update efficacy for deduplicated files, saving more storage space against the file re-splitting. Besides, we implement a superlink call to make deduplication process compatible with file-based encryption. These methods are combined to realize a transparent file deduplication (TFDedup) approach, which eliminates redundant data and reduces the associated cost of file update. Experimental results show that TFDedup succeeds to lower the space consumption when serving file updates by 55.6% and accelerate the deduplication process by 50.3%. Junbin Ren, Cheng Ji 0002, Weiwei Jin, Weichao Guo, Yajuan Du, Zongwei Zhu |
ICPADS | 6 |
| 2023 | GPU Performance Acceleration via Intra-Group Sharing TLBabstractUnified virtual memory greatly simplifies GPU programming, but it introduces huge address translation overhead. To reduce this overhead, modern GPUs utilize the translation lookaside buffer (TLB) to accelerate the address translation process. However, the benefit of TLB is far from achieving optimal performance. In this work, we find that GPU performance deficiency mainly stems from the private property of L1 TLBs. First, there exist a lot of duplicate page table entries among L1 TLBs, which induces insufficient space utilization. Second, the miss rate of L2 TLB is high due to the massive number of requests from L1 TLB miss, which leads to a significant GPU performance degradation. To reduce L1 TLB miss and improve the address translation performance of GPU, we propose a hardware scheme by exploiting an Intra-Group Sharing approach, named IGS-TLB. In IGS-TLB, L1 TLBs are decoupled from the compute units and aggregated into groups. Specifically, there only exist shared L1 TLB entries inside TLB groups that are responsible for non-overlapping address ranges. This greatly eliminates duplicate page table entries in L1 TLBs and significantly reduces the request number of L1 TLB misses. Our evaluation on a wide set of GPU workloads shows that IGS-TLB can effectively reduce L1 TLB miss rate and the L2 TLB traffic, speeding up the GPU performance by 20.5% on average. Yajuan Du |
ICPP | 2 |
| 2023 | LDPC Level Prediction Toward Read Performance of High-Density Flash MemoriesabstractHigh-density NAND flash memories have been prevailing in storage systems to achieve large capacities for explosive data. However, they suffer from more severe reliability degradation due to the narrowed margins between threshold voltage states. Low-density parity-check (LDPC) codes have been widely applied in high-density flash memories to ensure data reliability. Due to the increased number of cell states, more read voltages are required in reading a flash page correctly. This induces more soft levels to read pages with high-bit error rates in LDPC decoding. Read latency is significantly increased in high-density flash memories. To enhance the read performance of high-density flash memories, this article proposes PreLDPC, an LDPC-level prediction approach with fine-grained LDPC reading. The key idea of PreLDPC is to predict the final read level during the early read iteration, thus, avoiding unnecessary read-retry latency. From a preliminary study, we observe that after decoding in the first two iterations, the ratio of cells that lie in the error-prone area (i.e., adjacent area of two cell states) can be obtained. The ratio is closely related to the final read level for a successful decoding. By exploiting this observation, PreLDPC directly uses the predicted read level for LDPC reading, which could eliminate the excessive number of read retries. Furthermore, by exploiting the benefit of fine-grained LDPC reading, this article further divides the existing integer level (called i-level, e.g., level-1 and level-2) into a finer decimal level (called d-level, e.g., level-1.25 and level-1.5), and proposes a fine-grained read method. By combining the prediction method and fine-grained method together, PreLDPC can first estimate the i-level and then perform the read-retry iteration with d-levels to eliminate unnecessary read latency as much as possible. From experimental results of real-world workloads on Disksim with SSD extensions, it is verified that PreLDPC can effectively reduce read latency in high-density flash memories. Yajuan Du, Qiao Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Towards LDPC Read Performance of 3D Flash Memories with Layer-induced Error Characteristicsabstract3D flash memories have been widely developed to further increase the storage capacity of SSDs by vertically stacking multiple layers. However, this special physical structure brings new error characteristics. Existing studies have discovered that there exist significant Raw Bit Error Rate (RBER) variations among different layers and RBER similarity inside the same layer due to the manufacturing process. These error characteristics would introduce a new data reliability issue. Currently, Low-Density Parity-Check (LDPC) code has been widely used to ensure the data reliability of flash memories. It can provide stronger error correction capability for high RBERs by trading with longer read latency. Traditional LDPC codes designed for planar flash memories do not consider the layer RBER characteristics of 3D flash memories, which may induce sub-optimal read performance. This article first investigates the effect of RBER characteristics of 3D flash memories on read performance and then obtains two observations. On one hand, we observe that LDPC read latencies are largely diverse in different flash layers and increase in diverse speeds along with data retention. This phenomenon is caused by the inter-layer RBER variation. On the other hand, we also compare RBERs between different pages of the same flash layer and observe that read latencies with LDPC codes are quite similar, which is caused by the intra-layer RBER similarity. Then, by exploiting these two observation results, this article proposes a Multi-Granularity LDPC (MG-LDPC) read method to adapt read latency increase characteristics across 3D flash layers. In detail, we design five LDPC decoding engines with varied read level increase granularity (higher level induces higher latency) and assign these engines to each layer dynamically according to prior information, or in a fixed way. A series of experimental results demonstrate that the fixed and dynamic MG-LDPC methods can reduce SSD read response time by 21% and 51% on average, respectively. Yajuan Du, Yao Zhou 0012, Qiao Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2022 | Efficient Atomic Durability on eADR-Enabled Persistent MemoryabstractApplications atop persistent memory (PM) require atomic durability to ensure crash consistency. However, existing atomic durability techniques designed for PM systems are based on volatile cache and incur non-negligible performance overhead. Recently, Intel introduces a new feature called eADR (enhanced Asynchronous DRAM Refresh) for Optane PM, which brings an opportunity to build a much more efficient atomic durability system for PM. Taiyu Zhou, Yajuan Du, Fan Yang 0134, Xiaojian Liao, Youyou Lu |
PACT | 2 |
| 2022 | Work-in-Progress: Prediction-based Fine-Grained LDPC Reading to Enhance High-Density Flash Read PerformanceabstractLDPC codes have been widely applied in high-density flash memories, e.g., TLC flash and QLC flash, to ensure data reliability. In order to reduce the read latency of high-density flash memories, this paper proposes a prediction-based fine-grained LDPC reading method, named as PreLDPC. From a preliminary study, we observe that the ratio of cells that lie in error-prone areas (i.e., the areas between two adjacent cell states) is closely related to the final read level for successful decoding. Based on this observation, PreLDPC predicts the read level for LDPC reading, which could avoid excessive unnecessary read-retries. Furthermore, a fine-grained read method with fine sub-levels is used in the read-retry iteration for read latency reduction. From experimental results over real-world workloads on Disksim with SSD extensions, the effectiveness of PreLDPC on reducing read latency is verified in high-density flash memories. Yajuan Du, Qiao Li 0001 |
CASES | 1 |
| 2022 | Enhancing GPU Performance via Neighboring Directory Table Based Inter-TLB SharingabstractModern discrete GPUs support Unified Virtual Memory (UVM), simplifying GPU programming. However, UVM entails address translation on each memory access, which introduces expensive performance overhead during address translation. In this work, we select various workloads and conduct experiments on GPU performance. Our investigation shows that many workloads have low L1 TLB hit ratios of less than 40% on average. Even for a particular workload, the hit ratio is as low as 15%, which leads to significant performance degradation. Through further analysis, we find that a lot of common entries exist between neighboring private L1 TLBs, showing clear inter-TLB sharing behavior. To leverage the sharing, we propose a Neighboring Directory table based hardware scheme, named NeiDty. In NeiDty, L1 TLBs can probe physical addresses from neighboring L1 TLBs through a lightweight interconnect network. And NeiDty uses neighboring directory tables to keep track of the shared entries among neighboring L1-TLBs. In addition, we find it better to update address translation after two consecutive neighboring TLB hits than one hit. We run eight typical workloads with Gem5-GPU, and the results show that NeiDty increases the average hit ratio of L1 TLB TLB by 14% and improves the average performance by 10%. Yajuan Du, Xulong Tang |
ICCD | 1 |
| 2021 | Stereo Superpixel Segmentation Via Dual-Attention Fusion NetworksabstractStereo image pairs can improve performance of many tasks benefiting from the additional information obtained from a second viewpoint when compared with single images. Existing superpixel segmentation algorithms for stereo images mostly adopt single images as input, and neglect the correspondence between the left and right views. In this work, we consider to exploit the depth information between stereo image pairs, and propose an end-to-end dual-attention fusion network for stereo images to generate parallax-consistency superpixels. We first utilize a deep convolution network to extract the deep features of stereo images. Then, to effectively utilize the additional information from the other view, features of the left and right views is integrated by a parallax attention and channel attention mechanism. Finally, the stereo superpixels are generated by a differentiable clustering algorithm, which is end-to-end trainable with deep learning networks. Comprehensive experimental results demonstrate that our method can outperform the state-of-the-art performance on the KITTI2015 and Cityscapes dataset. Yajuan Du, Hua Li 0012, Yucong Dai |
ICME | 2 |
| 2021 | Read-Ahead Efficiency on Mobile Devices: Observation, Characterization, and OptimizationabstractRead-ahead schemes have been widely used in page cache to improve read performance of Linux systems. As the Android system inherits the Linux kernel, the traditional read-ahead scheme is directly transplanted to mobile devices. However, request sizes and page cache sizes on mobile devices are much smaller, which may degrade read-ahead efficiency and therefore hurt user experience. This article first observes that many pages pre-fetched by read-ahead are unused, which causes frequent page cache eviction. And these evict operations could induce extra access latency, especially when write-back is conducting. Then, this article proposes a new analysis model to characterize the factors that closely relate to the access latency. It is found that there exists a trade-off between read-ahead size and access latency. Finally, this article proposes two optimized read-ahead schemes to exploit this trade-off under different situations. Size-tuning scheme aims to find the proper maximum size of read-ahead according to the characteristics of mobile devices. While MobiRA scheme improves the read-ahead efficiency by dynamically tuning read-ahead size and stop-settings. Experimental results on real mobile devices show that the proposed schemes can increase the efficiency of read-ahead scheme and improve the overall performance of mobile devices. Yu Liang 0004, Riwei Pan, Yajuan Du, Chenchen Fu, Liang Shi 0001, Tei-Wei Kuo, Chun Jason Xue |
IEEE Trans. Computers | 3 |
| 2020 | Protecting the Intellectual Property of Deep Neural Networks with Watermarking: The Frequency Domain ApproachabstractSimilar to other digital assets, deep neural network (DNN) models could suffer from piracy threat initiated by insider and/or outsider adversaries due to their inherent commercial value. DNN watermarking is a promising technique to mitigate this threat to intellectual property. This work focuses on black-box DNN watermarking, with which an owner can only verify his ownership by issuing special trigger queries to a remote suspicious model. However, informed attackers, who are aware of the watermark and somehow obtain the triggers, could forge fake triggers to claim their ownerships since the poor robustness of triggers and the lack of correlation between the model and the owner identity. This consideration calls for new watermarking methods that can achieve better trade-off for addressing the discrepancy. In this paper, we exploit frequency domain image watermarking to generate triggers and build our DNN watermarking algorithm accordingly. Since watermarking in the frequency domain is high concealment and robust to signal processing operation, the proposed algorithm is superior to existing schemes in resisting fraudulent claim attack. Besides, extensive experimental results on 3 datasets and 8 neural networks demonstrate that the proposed DNN watermarking algorithm achieves similar performance on functionality metrics and better performance on security metrics when compared with existing algorithms. Leo Yu Zhang, Yajuan Du, Jun Zhang 0010, Yong Xiang 0001 |
TrustCom | 4 |
| 2020 | Static detection of real-world buffer overflow induced by loop
Deqing Zou, Yajuan Du, Hai Jin 0001, Changming Liu, Jinan Shen |
Comput. Secur. | 3 |
| 2020 | Using Error Modes Aware LDPC to Improve Decoding Performance of 3-D TLC NAND Flashabstract3-D triple-level cell (3-D TLC) NAND flash has high storage density and capacity, but degrading data reliability due to high raw bit error rates induced by a certain number of program/erase cycles. To guarantee data reliability, low-density parity-check (LDPC) codes are selected as the error correction codes in modern flash memories because of strong error correction capability. However, directly adopting LDPC codes induces high decoding latency due to iterative updating of log-likelihood ratio (LLR) information in the decoding process. Increasing LLR information accuracy can greatly improve decoding performance. In this paper, we propose EMAL: using error modes aware LDPC codes for further enhancing the decoding performance of 3-D TLC NAND flash. We first obtain 3-D TLC error modes based on an FPGA testing platform, and then exploit the error modes to optimize LLR information and enable the decoding to converge at a high speed. The simulation results show that the decoding performance is significantly improved, resulting in reduced bit error rates and decoding latency. Fei Wu 0005, Meng Zhang 0014, Yajuan Du, Zuo Lu, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | TLFW: A Three-Layer Framework in Wireless Rechargeable Sensor Network with a Mobile Base StationabstractWireless sensor networks as the base support for the Internet of things have been a large number of popularity and application. Such as intelligent agriculture, we have to use the sensor network to obtain the growing environment data of crops and others. However, the difficulty of power supply of wireless nodes has seriously hindered the application and development of Internet of things. In order to solve this problem, people use low-power sleep scheduling and other energy-saving methods on the nodes. Although these methods can prolong the working time of nodes, they will eventually become invalid because of the exhaustion of energy. The use of solar energy, wind energy, and wireless signals in the environment to obtain energy is another way to solve the energy problem of nodes. However, these methods are affected by weather, environment, and other factors, and they are unstable. Thus, the discontinuity work of the node is caused. In recent years, the development of wireless power transfer (WPT) has brought another solution to this problem. In this paper, a three-layer framework is proposed for mobile station data collection in rechargeable wireless sensor networks to keep the node running forever, named TLFW which includes the sensor layer, cluster head layer, and mobile station layer. And the framework can minimize the total energy consumption of the system. The simulation results show that the scheme can reduce the energy consumption of the entire system, compared with a Mobile Station in a Rechargeable Sensor Network (MSiRSN). Anwen Wang, Xianjia Meng, Lvju Wang, Baoying Liu, Feng Chen 0002, Yajuan Du, Guangcheng Yin |
Wirel. Commun. Mob. Comput. | 8 |
| 2019 | Adapting Layer RBERs Variations of 3D Flash Memories via Multi-granularity Progressive LDPC ReadingabstractExisting studies have uncovered that there exist significant Raw Bit Error Rates (RBERs) variations among different layers of 3D flash memories due to manufacture process variation. These RBER variations would cause significantly diversed read latencies when reading data with traditional Low-Density Parity-Check (LDPC) codes designed for planar flash memories, which induces sub-optimal read performance of flash-based Solid-State Drives (SSDs). Yajuan Du, Meng Zhang 0014, Shengwu Xiong 0001 |
DAC | 1 |
| 2019 | Pair-Bit Errors Aware LDPC Decoding in MLC NAND Flash MemoryabstractBy storing multibit per cell, multilevel cell (MLC) NAND flash memory achieves high storage capacity, but sacrificing data reliability. Error correction codes, such as Bose–Chaudhuri–Hocquenghem (BCH) codes, are widely used to ensure data reliability. However, high raw bit error rates induced by interference noises make BCH codes become insufficient to guarantee data reliability. Low-density parity-check (LDPC) codes are considered as the replacement due to the stronger error correction capability. Nevertheless, directly exploiting LDPC codes introduces a concern about decoding latency because of their iterative decoding in the soft decision process. To develop effective LDPC decoding algorithms, it is necessary to have a more profound understanding on flash failure patterns. This paper first observes the pair-bit errors (PBEs) characteristic of MLC NAND flash memory on a real field-programmable gate array testing platform, then proposes a PBE-aware LDPC (PAL) decoding scheme-based upon this observation, in which PBE provides the promotion information for LDPC decoding to reduce decoding latency. Simulation results show that the decoding latency can be reduced by up to 54%, compared with the conventional LDPC codes. Meng Zhang 0014, Fei Wu 0005, Yajuan Du, Changsheng Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | FastGC: accelerate garbage collection via an efficient copyback-based data migration in SSDsabstractCopyback is an advanced command contributing to accelerating data migration in garbage collection (GC). Unfortunately, detecting copyback feasibility (whether copyback can be carried out with assurable reliability) against data corruption in the traditional copyback-based GC causes an expensive performance penalty. This paper first explores copyback error characteristics on real NAND flash chips, then proposes a fast garbage collection scheme called FastGC. It utilizes copyback error characteristics to efficiently detect copyback feasibility of data instead of transferring out all valid data for detecting. Experiment results in the SSDsim show the proposed FastGC greatly promotes write response time and read response time by up to 44.2% and 66.3% respectively, compared to the traditional copyback-based GC. Fei Wu 0005, Jiaona Zhou, Shunzhuo Wang, Yajuan Du, Chengmo Yang, Changsheng Xie 0001 |
DAC | 4 |
| 2018 | DigHR: precise dynamic detection of hidden races with weak causal relation analysis
Deqing Zou, Hai Jin 0001, Yajuan Du, Long Zheng 0003, Jinan Shen |
J. Supercomput. | 4 |
| 2017 | Reducing LDPC Soft Sensing Latency by Lightweight Data Refresh for Flash Read Performance ImprovementabstractIn order to relieve reliability problem caused by technology scaling, LDPC codes have been widely applied in flash memories to provide high error correction capability. However, LDPC read performance slowdown along with data retention largely weakens the access speed advantage of flash memories. This paper considers to apply the concept of refresh, that were used for flash lifetime improvement, to optimize flash read performance. Exploiting data read characteristics, this paper proposes LDR, a lightweight data refresh method, that aggressively corrects errors in read-hot pages with long read latency and reprograms error-free data into new pages. Experimental results show that LDR can achieve 29% read performance improvement with only 0.2% extra P/E cycles on average, which causes negligible overhead on flash lifetime. Yajuan Du, Qiao Li 0001, Liang Shi 0001, Deqing Zou, Hai Jin 0001, Chun Jason Xue |
DAC | 1 |
| 2017 | Exploiting Process Variation for Read Performance Improvement on LDPC Based Flash Memory Storage SystemsabstractWith the development of bit density and technology scaling, the process variation (PV) has become much severe on NAND flash memory. As PV presents reliability among flash blocks, which causes read performance variation to read data on different blocks. This paper proposes to improve read performance of LDPC based flash memory by exploiting the reliability characteristics of PV. First, a block grouping approach is proposed to classify the flash blocks based on their reliability. Then, a read data placement scheme is proposed, which is designed to place read-hot data on flash blocks with high reliability and move read-cold data to blocks with low reliability. Experiment results show that, with negligible overhead, the proposed scheme is able to significantly improve the read performance. Qiao Li 0001, Liang Shi 0001, Yejia Di, Yajuan Du, Chun Jason Xue, Edwin H.-M. Sha |
ICCD | 4 |
| 2017 | CooECC: A Cooperative Error Correction Scheme to Reduce LDPC Decoding Latency in NAND FlashabstractThe storage capacity of NAND Flash has increased by scaling down to smaller cell size and using multi-level storage technology, but data reliability is degraded by severer retention errors. To ensure data reliability, error correction codes (ECC) are adopted, such as BCH and low-density parity check (LDPC) codes. However, BCH codes are insufficient when raw bit error rates (RBER) caused by retention errors are high. As a result, BCH codes are inevitably replaced with LDPC codes with stronger error correction capability. Traditional LDPC codes are used to independently correct bit errors in the LSB and MSB pages. Unfortunately, decoding latency in such two pages is significantly unbalanced, MSB pages take much higher latency due to higher RBER, leading to suboptimal flash read performance. This paper proposes a cooperative error correction scheme, called CooECC, to reduce LDPC decoding latency of the MSB page in NAND Flash. By exploiting data error characteristics introduced by retention errors, CooECC integrates the decoding result of the LSB page into the initial information of LDPC decoding for the MSB page, making it more accurate. This in turn enables decoding to converge at a higher rate. Simulation results show that for LDPC schemes with information lengths of 2KB and 4KB, the decoding latency can be reduced by up to 87% and 84%, respectively, when RBER is as high as 8.0 × 10^-3. Meng Zhang 0014, Fei Wu 0005, Yajuan Du, Chengmo Yang, Changsheng Xie 0001, Jiguang Wan 0001 |
ICCD | 3 |
| 2017 | An empirical study of F2FS on mobile devicesabstractFlash Friendly File System (F2FS) is getting popular among mobile devices. However, lack of empirical and comprehensive analysis for characteristics of F2FS prohibits better application of F2FS. In this paper, we present a set of comprehensive experimental studies on mobile devices and show several counterintuitive observations on F2FS, including imprecise hot/cold data separation, unexpected trigger condition of background GC, impact of fragmentation on read performance and impact of readahead by fragments and available space. Based on these observations, we further provide several pilot solutions to improve the performance of these mobile devices. The objective is to inspire researchers and users to pay attention to F2FS characteristics, and further optimize its performance. Yu Liang 0004, Chenchen Fu, Yajuan Du, Aosong Deng, Mengying Zhao, Liang Shi 0001, Chun Jason Xue |
RTCSA | 3 |
| 2017 | A Program Interference Error Aware LDPC Scheme for Improving NAND Flash Decoding PerformanceabstractBy scaling down to smaller cell size, NAND flash has significantly increased the storage capacity in order to lower the unit cost down. However, the reliability is sacrificed due to much higher raw bit error rates. As a result, conventional error correction codes (ECCs), such as BCH codes, are not sufficient. Low-density parity check (LDPC) codes with stronger error correction capability are adopted in NAND flash to guarantee data reliability. However, read performance using LDPC is poor because of its decoding complexity. It has been found that flash cells with fewer electrons are more prone to program interference errors. As a result, program interference errors show the characteristic of value dependence. This characteristic can be exploited and translated into extra information facilitating the decoding convergence. Motivated by this observation, we propose PEAL: a flash program interference error aware LDPC scheme to enhance the decoding performance. PEAL integrates the obtained extra information from the value dependence into the soft-to-hard decision process in LDPC decoding to decrease decoding iterations and improve the decoding convergence speed. Simulation results show that decoding iterations are reduced by up to 69.37% and the decoding convergence speed is improved by up to 2.5×, compared with the normalized min-sum (NMS) algorithm with 2KB information lengths at an approximate raw bit error rate of 11.5 × 10 −3 . Fei Wu 0005, Meng Zhang 0014, Yajuan Du, Xubin He, Ping Huang 0001, Changsheng Xie 0001, Jiguang Wan 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2017 | A dynamic predictive race detector for C/C++ programs
Deqing Zou, Hai Jin 0001, Yajuan Du, Jinan Shen |
J. Supercomput. | 4 |