VLDB 2026 Research / reviewers in the wild / expert
Yina Lv
dblp:249/4229
· DBLP profile ↗
34ranked-venue papers
8as first author
30since 2021 · last 2026
0000-0003-3971-3123ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 6 first-author · 26 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nemo: A Low-Write-Amplification Cache for Tiny Objects on Log-Structured Flash Devices
Xufeng Yang, Jingxin Hu, Congming Gao, Tianyang Jiang, Linbo Long, Yina Lv, Jiwu Shu |
ASPLOS (2) | 9 |
| 2026 | The Evolution of LSM-Tree Key-Value Stores: A Tutorial on State-Of-The-Art and Future Directions
Yina Lv, Qiao Li 0001, Quanqing Xu, Chun Jason Xue |
ICDE | 1 |
| 2026 | Tetris: Lightweight Hyperparameter Auto-Tuning for Mitigating Performance Spikes in LSM-KVS
Yina Lv, Qiao Li 0001, Quanqing Xu, Congming Gao, Chuanhui Yang, Xiaoli Wang 0002, Chun Jason Xue |
ICDE | 1 |
| 2026 | Breaking Barriers in Atomic Scaling: A Hardware-Software-Collaborated Framework to Deconstruct RDMA Atomic
Guangyang Deng, Qiangsheng Su, Zhirong Shen, Qing Wang 0031, Yina Lv, Ronglong Wu, Jiwu Shu |
ISCA | 5 |
| 2026 | LOONG: Utilizing Long-Stride Reprogramming to Enhance the Performance of SSDs
Congming Gao, Jiancong Zheng, Xufeng Yang, Qiao Li 0001, Yina Lv, Chun Jason Xue, Jiwu Shu |
ISCA | 8 |
| 2026 | Enhancing optimal read voltage prediction for three-dimensional NAND flash memory through data augmentation techniques
Xiangyu Yao, Guanyu Wu, Yina Lv, Jie Zhang 0048, Xinbiao Gan, Qiao Li 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | CrossFS: Improving Cross-Domain File System Performance with CRDT-Based Metadata SynchronizationabstractModern data-intensive applications increasingly demand efficient and scalable file systems that can operate across distributed and cross-domain environments. However, existing file systems are inefficient in metadata management, synchronization efficiency, and system scalability under high-concurrency and metadata-intensive workloads in cross-domain environments. To address these challenges, this article introduces CrossFS (CFS), a cross-domain distributed file system that enhances consistency guarantees and metadata indexing. Specifically, CFS leverages conflict-free replicated data types (CRDTs) to synchronize metadata, achieving strong eventual consistency with minimal synchronization overhead, even across network partitions. Furthermore, CFS employs a Hybrid Tree indexing structure, tailored for distributed environments, which optimizes metadata operations by reducing query latency by up to 33.4% and write amplification by 30.7%. Additionally, CFS achieves adaptive caching strategies and a hybrid synchronization model that effectively balances consistency latency with data availability. Extensive evaluations show that CFS outperforms CephFS and GlusterFS, achieving up to 33.9% higher metadata throughput, 36% lower latency, and 42% better data operation efficiency. Qiwen Ke, Yina Lv, Zhirong Shen, Yue Yu 0001, Zhenlong Song, Xinbiao Gan, Dongsheng Li 0001, Xin Yao 0008, Yiming Zhang 0003 |
ACM Trans. Storage | 2 |
| 2025 | DISS: A Novel Data Invalidation Scheme for Swap-Data on Flash Storage SystemsabstractStorage swapping has been a critical technique used to relieve memory pressure and improve user experience. However, it generates lots of data writes in flash storage, deteriorating the lifetime and performance. In this paper, inspired by empirical studies on swap data access characteristics, we propose a novel data invalidation scheme, namely DISS, which includes two methods. First, a cross-layer swap-data invalidation method is proposed to invalidate swapped-in data at a low cost. Second, a swap data separation method is proposed to schedule swap data and file-backed data into different places. Experimental results show that DISS achieves encouraging flash lifetime and performance optimization. Dingcui Yu, Longfei Luo, Han Wang 0051, Yina Lv, Liang Shi 0001 |
ASP-DAC | 4 |
| 2025 | Phoenix: A Refactored I/O Stack for GPU Direct Storage without Phony BuffersabstractGPU Direct Storage (GDS) plays a vital role in GPU-based training and inference systems, leveraging Peer-to-Peer Direct Memory Access (P2P-DMA) to establish a direct data transfer path between the GPU and the storage device. The direct I/O path reduces GPU storage access latency and CPU overhead, thus improving the efficiency of data transfer. Currently, however, GDS employs a phony buffer in the host memory to interact with the Linux kernel, which results in suboptimal I/O performance, extra resource consumption, and high deployment complexity. Jianqin Yan, Shi Qiu 0012, Yina Lv, Hao Chen 0080, Zhirong Shen, Xin Yao 0008, Renhai Chen, Jiwu Shu, Gong Zhang 0001, Yiming Zhang 0003 |
SC | 3 |
| 2025 | Revisiting Multiple ECC on High-Density NAND Flash memoryabstractThree-dimensionalnandflash memory using the advanced multibit-per-cell technique is widely adopted due to its high density. However, it faces the problem of deteriorating read performance and energy consumption due to decreased reliability. Low-density parity-check code (LDPC) is typically adopted as an error correction code (ECC) to encode data and provide fault tolerance. To reduce the cost, LDPC with a high code rate is always adopted. However, LDPC will lead to read retry operations when the accessed data are not successfully decoded, and such retry-induced performance degradation is serious, especially for modern high-density flash memory. In this work, a reliability-aware differential ECC (READECC) approach is proposed to reduce redundancy protection and storage cost of LDPC with a low code rate and optimize the read performance. The basic idea is to adopt LDPC with a suitable code rate considering both data access characteristics and flash reliability characteristics. First, hot reads are identified based on the frequency of being accessed. Second, based on the reliability variation characteristics, the life of flash memory is divided into three reliability periods. As the reliability period shifts, the code rate of the LDPC adjusts adaptively to minimize redundancy protection. Third, an adaptive-sized logical page approach is further proposed to support LDPC with strong error correction capability (a low code rate) with a low storage cost. Through careful design and evaluation on 3-D triple-level-cellnandflash memory, READECC achieves encouraging optimizations with a negligible cost. Yunpeng Song, Yina Lv, Wentong Li 0002, Liang Shi 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Achieving Near-Zero Read Retry for 3D NAND Flash MemoryabstractAs the flash-based storage devices age with program/erase (P/E) cycles, they require an increasing number of read retries for error correction, which in turn deteriorates their read performance. The design of read-retry methods is critical to flash read performance. Current flash chips embed pre-defined read retry tables (RRT) for retry, but these tables fail to consider the read granularity and error behaviors. We characterize different types of real flash chips, based on which we further develop models for the correlation among the optimal read offsets of read voltages required for reading each page. By leveraging characterization observations and the models, we propose a methodology to generate a tailored RRT for each flash model. We introduce a dynamic read retry procedure to pick up proper read voltages from the table, followed by a proximity-search method for fine-tuning the read offsets. Experiments on real flash chips show that the proposed methodology can achieve near-zero retries. It reduces the average number of read retries to below 0.003 for data with high retention time at 8K P/E cycles, whereas the state-of-the-art approaches incur over 3 read retries on average once the flash is aged to 3K P/E cycles. Qiao Li 0001, Yina Lv, Jie Zhang 0048, Daniel Wen, Tei-Wei Kuo, Chun Jason Xue |
ASPLOS (2) | 3 |
| 2024 | CPF: A Cross-Layer Prefetching Framework for High-Density Flash-Based StorageabstractThe pseudo-single-level-cell (pSLC) technique is widely adopted in high-density flash-based storage to mitigate the performance and endurance problem of high-density flash memory. Furthermore, prefetching schemes can compensate for performance differences among storage tiers. Existing prefetchers are implemented in the operating system (OS) or storage layers. However, OS layer prefetchers are conservative since it is a challenge to achieve both high accuracy and large coverage simultaneously. Storage layer prefetchers are sub-optimal due to the performance differences between pSLC and DRAM. In this paper, a cross-layer prefetching framework (CPF) is proposed to prefetch data selectively. The basic idea is that high-accuracy data will be prefetched to DRAM and large-coverage data will be prefetched to the pSLC flash in storage. To make it practical, an adaptive regulator is further designed to dynamically adjust the cross-layer prefetching to ensure accuracy and coverage. Evaluations show that CPF can improve read performance and reduce data transfer costs significantly. Longfei Luo, Han Wang 0051, Dingcui Yu, Yina Lv, Liang Shi 0001 |
DATE | 4 |
| 2024 | EEPC: Energy-Efficient Persistent Cache Scheme for Mobile Distributed File SystemsabstractFor mobile distributed file systems (MDFSs), files can be easily shared among multiple mobile devices. However, it requires the connected remote devices to be online all the time for timely file browsing, which incurs significant energy consumption. This is unacceptable for battery-powered mobile devices. To address this issue, we propose EEPC, an energy-efficient persistent cache scheme for MDFSs. It consists of several techniques. First, a proactive cache invalidation mechanism is designed to ensure optimistic access to the local persistent cache, which greatly reduces unnecessary read requests. Second, a lazy cache synchronization policy is designed to reorganize writeback requests, which ensures that remote devices remain in a low-power state for a long time. Finally, a cache admission and eviction scheme is proposed, which considers both file access frequency and recency, and an adaptable file prefetching scheme is adopted to quickly recover invalidated cache files. Evaluations on real devices show that EEPC maintains at least 60% of sleep time for remote devices and greatly extends the interval between two wake-ups, regardless of the frequency of remote file accesses. Compared with the state-of-the-art, the energy consumption of remote devices can be reduced by 33.6%, on average. Wentong Li 0002, Yina Lv, Liang Shi 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Access Characteristic-Guided Remote Swapping Across Mobile DevicesabstractMemory swapping ensures smooth application switching for mobile systems by caching applications in the background. To further play the role of memory swapping, remote swapping across mobile devices has been widely studied, which caches applications to nearby remote devices by remote paging. However, due to the massive remote I/Os and unguaranteed swap throughput, the current remote swapping is limited with an unsatisfactory user experience, especially under variable network conditions. This paper first studies the access characteristics of applications and clarifies the impact of various network traffic on remote swapping. Motivated by these, an efficient access characteristic-guided remote swapping framework (ACR-Swap + ) is proposed to optimize remote swapping across mobile devices with resilient remote paging. ACR-Swap + first performs selective remote paging based on the swap-in frequency of different processes and then prefetches data across devices based on the process running states. Finally, it conducts hierarchical remote paging to avoid the impact of network traffic on remote swapping. Evaluations on Google Pixel 6 show that ACR-Swap + reduces the application switching latency by 21.6% and achieves a negligible performance fluctuation under various network traffic compared to the state of the art. Wentong Li 0002, Yina Lv, Longfei Luo, Yunpeng Song, Liang Shi 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Critical Data Backup with Hybrid Flash-Based Consumer DevicesabstractHybrid flash-based storage constructed with high-density and low-cost flash memory has become increasingly popular in consumer devices in the last decade due to its low cost. However, its poor reliability is one of the major concerns. To protect critical data for guaranteeing user experience, some methods are proposed to improve the reliability of consumer devices with non-hybrid flash storage. However, with the widespread use of hybrid storage, these methods will result in severe problems, including significant performance and endurance degradation. This is caused by the fact that the different characteristics of flash memory in hybrid storage are not considered, e.g., performance, endurance, and access granularity. To address these problems, a critical data backup (CDB) design is proposed to ensure critical data reliability at a low cost. The basic idea is to accumulate two copies of critical data in the fast memory first to make full use of its performance and endurance. Then, one copy will be migrated to the slow memory in the stripe to avoid the write amplification caused by different access granularity between them. By respecting the different characteristics of flash memory in hybrid storage, CDB can achieve encouraging performance and endurance improvement compared with the state-of-the-art. Furthermore, to avoid performance and lifetime degradation caused by the backup data occupying too much space of fast memory, CDB Pro is designed. Two advanced schemes are integrated. One is making use of the pseudo-single-level-cell (pSLC) technique to make a part of slow memory become high-performance. By supplying some high-performance space, data will be fully updated before being evicted to slow memory. More invalid data are generated which reduces eviction costs. Another is to categorize data into three types according to their different life cycles. By putting the same type of data in a block, the eviction efficiency is improved. Therefore, both can improve device performance and lifetime based on CDB. Experiments are conducted to prove the efficiency of CDB and CDB Pro. Experimental results show that compared with the state-of-the-arts, CDB can ensure critical data reliability with lower device performance and lifetime loss whereas CDB Pro can diminish the loss further. Longfei Luo, Dingcui Yu, Yina Lv, Liang Shi 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Revisiting TRIM on High-Density Flash-Based Hybrid Storage SystemsabstractHybrid solid state drives (SSDs) that integrate high-performance and large-capacity flash are widely used due to their cost-effectiveness. The TRIM command, which is a popular command in normal SSDs to improve performance and endurance, is also recommended in hybrid SSDs. However, employing TRIM on hybrid SSDs as on normal SSDs will induce performance loss and sub-optimal endurance due to the different characteristics of flash in hybrid SSDs. To solve the problem, this paper first explores the critical factors of issuing TRIM commands to different flash. Then, this paper proposed a differential TRIM method (dTRIM), which suggests performing early TRIM on high-performance flash and lazy TRIM on high-capacity flash. Specifically, early TRIM will minimize garbage collection costs while lazy TRIM tries to avoid conflicting user requests. Experimental results demonstrate that dTRIM can significantly improve the performance and endurance of hybrid SSDs compared with the state-of-the-arts. Longfei Luo, Dingcui Yu, Yunpeng Song, Yina Lv, Edwin H.-M. Sha, Liang Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | Near-Free Lifetime Extension for 3-D nand Flash via Opportunistic Self-Healingabstract3-Dnandflash memories are the dominant storage media in modern data centers due to their high performance, large storage capacity, and low-power consumption. However, the lifetime of flash memory has decreased as technology scaling advances. Recent work has revealed that the number of achievable program/erase (P/E) cycles of flash blocks is related to the dwell time (DT) between two adjacent erase operations. A longer DT can lead to higher-achievable P/E cycles and, therefore, a longer lifetime for flash memories. This article found that the achievable P/E cycles would increase when flash blocks endure uneven DT distribution. Based on this observation, this article presents an opportunistic self-healing method to extend the lifetime of flash memory. By maintaining two groups with unequal block counts, namely, Active Group and Healing Group, the proposed method creates an imbalance in erase operation distribution. The Active Group undergoes more frequent erase operations, resulting in shorter DT, while the Healing Group experiences longer DT. Periodically, the roles of the two groups are switched based on the Active Group’s partitioning ratio. This role switching ensures that each block experiences both short and long DT periods, leading to an uneven DT distribution that magnifies the self-healing effect. The evaluation shows that the proposed method can improve the flash lifetime by 19.3% and 13.2% on average with near-free overheads, compared with the baseline and the related work, respectively. Qiao Li 0001, Yina Lv, Nan Guan, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Adaptive Differential Wearing for Read Performance Optimization on High-Density nand Flash MemoryabstractWith cost reduction and density optimization, high-density NAND flash memory has been widely deployed in data centers and consumer devices. However, this trend has significantly degraded the read performance and lifetime of high-density NAND flash memory during the last decade. Previous works proposed to optimize flash lifetime with wear leveling (WL) and optimize read performance with reliability improvement. Although WL can improve flash lifetime, it leads to the reliability of all blocks in 3-D NAND flash decreasing simultaneously. The reliability and read performance will be degraded with flash wearing. To solve this problem, an adaptive differential wearing (ADWR) scheme is proposed to optimize the read performance and lifetime in this work. The basic idea of ADWR is to determine the size of the high-reliability area to serve hot reads based on workload characteristics. Specifically, first, a differential wearing scheme is proposed to construct different reliability areas based on the characteristics of the data. Second, a lifetime model is constructed for the ADWR to clarify the lifetime impact. Based on this, a lifetime optimization scheme is proposed to improve the flash lifetime. Finally, a differential refresh scheme is proposed to reduce the impact of read disturbance on read performance. The experiments on real-life workloads show that ADWR achieves encouraging read performance optimization with negligible impacts on the lifetime of 3-D TLC NAND flash memory. Yunpeng Song, Yina Lv, Liang Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | DECC: Differential ECC for Read Performance Optimization on High-Density NAND Flash Memoryabstract3D NAND flash memory with advanced multi-level-cell technology has been widely adopted due to its high density, but with significantly degraded reliability. To solve the reliability issue, flash memory often adopts the low-density parity-check code (LDPC) as error correction code (ECC) to encode data and provide fault tolerance. For LDPC with a low code rate, it can provide a strong correction capability, but with a high energy cost. To avoid the cost, LDPC with a higher code rate is always adopted. When the accessed data is not successfully decoded, LDPC will rely on read retry operations to improve the error correction capability. However, the read retry operation will induce degraded read performance. In this work, a differential ECC (DECC) method is proposed to improve the read performance. The basic idea of DECC is to adopt LDPC with different code rates for data with different access characteristics. Specifically, when data is hot read and retried due to reliability, LDPC with a low code rate will be adopted to optimize performance. With this approach, the cost from LDPC with a low code rate is minimized and the performance is optimized. Through careful design and real-world workloads evaluation on a 3D triple-level-cell (TLC) NAND flash memory, DECC achieves encouraging read performance optimization. Yunpeng Song, Yina Lv, Liang Shi 0001 |
ASP-DAC | 2 |
| 2023 | When F2FS Meets Compression-Based SSD!abstractCompression-based schemes have been widely studied to improve the lifetime and performance of solid-state drives (SSDs). Recently, the most popular flash-friendly file system (F2FS) started supporting compression to maximize the lifetime of NAND flash-based storage. Also, compression-based computational SSDs (CSDs) are developed due to their high performance, transparency, and easy adoption. This paper will first study the compression of F2FS and CSD to understand their features. Then, cooperative compression (COCO) is proposed to optimize performance and power consumption based on the combination of F2FS and CSD. Experiments on real devices show that COCO has encouraged optimization. Yunpeng Song, Yiyang Huang 0001, Yina Lv, Liang Shi 0001 |
HotStorage | 3 |
| 2023 | MGC: Multiple-Gray-Code for 3D NAND Flash based High-Density SSDsabstractQLC (4-bit-per-cell) and more-bit-per-cell 3D NAND flash memories are increasingly adopted in large storage systems. While achieving significant cost reduction, these memories face degraded performance and reliability issues. The industry has adopted two-step programming (TSP), rather than one-step programming, to perform fine-granularity program control and choose gray-code encoding, as well as LDPC (Low-Density Parity-Check Code) for error correction. Different flash manufacturers often integrate different gray-codes in their products, which exhibit different performance and reliability characteristics. Unfortunately, a fixed gray-code encoding design lacks the ability to meet the dynamic read and program performance requirements at both application and device levels.In this paper, we propose MGC, a multiple-gray-code encoding strategy, that adaptively chooses the best gray-code to meet the optimization goals at runtime. In particular, MGC first extracts the performance and reliability requirements based on application-level access patterns and detects the reliability degree of SSD. It then determines the appropriate gray-code to encode the data, either from host/user application or due to garbage collection, before writing the pages to the flash memory. MGC is integrated in FTL (flash translation layer) and enhances the flash controller to enable runtime gray-code arbitration. We evaluate the proposed MGC scheme. The results show that MGC achieves better performance and lifetime guarantee compared with state-of-the-arts and introduces little overhead. Yina Lv, Liang Shi 0001, Qiao Li 0001, Congming Gao, Yunpeng Song, Longfei Luo, Youtao Zhang |
HPCA | 1 |
| 2023 | Performance and reliability optimization for high-density flash-based hybrid SSDs
Longfei Luo, Yina Lv, Liang Shi 0001 |
J. Syst. Archit. | 3 |
| 2023 | Access Characteristic Guided Partition for Nand Flash-Based High-Density SSDsabstractnand flash-based solid-state drives (SSDs) are a kind of widely adopted storage. However, state-of-the-art works presented that the SSD always suffers from significant read performance degradation. One of the most critical reasons is access interference between read and write operations. This is because the read and write latency gaps are more pronounced for the latest nand flash in SSDs. In this article, an interference reduction scheme is proposed to improve performance. This is motivated by the observation from several server workloads, where read and write operations can be easily separated based on access characteristics. Considering that SSDs are always organized with many parallel units (PUs), the basic idea of this work is to partition the PUs of the SSD into different areas and place data in the corresponding area according to access characteristics. Then, the interference can be optimized by issuing read and write requests to the different areas. To realize the above design, several approaches are proposed: first, an access characteristic-based data placement and migration method is proposed for read and write request separation. Second, to further adapt the parallel requirement for different workloads, a workload-based partitioning scheme is proposed to determine the number of PUs for read and write areas. Finally, based on partitioned SSD, a hot-data driven wear-leveling method is further proposed to balance the wearing of PUs in read and write areas. Experimental results show that partitioned SSD can significantly improve the read performance and wear leveling of partitioned SSD can guarantee performance and lifetime. Yina Lv, Liang Shi 0001, Yunpeng Song, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | DWR: Differential Wearing for Read Performance Optimization on High-Density NAND Flash MemoryabstractWith the cost reduction and density optimization, the read performance and lifetime of high-density NAND flash memory have been significantly degraded during the last decade. Previous works proposed to optimize lifetime with wear leveling and optimize read performance with reliability improvement. However, with wearing, the reliability and read performance will be degraded along with the life of the device. To solve this problem, a differential wearing scheme (DWR) is proposed to optimize the read performance. The basic idea of DWR is to partition the flash memory into two areas and wear them at different speeds. For the area with low wearing speed, read operations are scheduled for read performance optimization. For the area with high wearing speed, write operations are scheduled but designed to avoid generating bad blocks early. Through careful design and real workloads evaluation on 3D TLC NAND flash, DWR achieves encouraging read performance optimization with negligible impacts to the lifetime. Yunpeng Song, Qiao Li 0001, Yina Lv, Changlong Li 0006, Liang Shi 0001 |
DATE | 3 |
| 2022 | Read latency variation aware performance optimization on high-density NAND flash based storage systems
Liang Shi 0001, Yina Lv, Longfei Luo, Changlong Li 0006, Chun Jason Xue, Edwin H.-M. Sha |
CCF Trans. High Perform. Comput. | 2 |
| 2022 | Practical optimizations for lightweight distributed file system on consumer devices
Yuze Xu, Han Wang 0051, Ben Gu, Yina Lv, Longfei Luo, Changlong Li 0006, Liang Shi 0001 |
CCF Trans. High Perform. Comput. | 5 |
| 2022 | Tail Latency Optimization for LDPC-Based High-Density and Low-Cost Flash Memory DevicesabstractFlash memory has been developed with bit density improvement, technology scaling, and 3-D stacking. With this trend, its reliability has been significantly degraded. Error correction code (ECC), such as low-density parity code (LDPC), which has strong error correction capability, has been deployed to solve this problem. However, one of the critical issues of LDPC is that it would introduce a long decoding latency on devices with low reliability. In this case, tail latency would happen, which will significantly impact the quality of service. In this work, a set of smart refresh schemes is proposed to optimize the tail latency. The basic idea of the work is to refresh data when the accessed data have a long decoding latency. Two smart refresh schemes are proposed for this work. The first refresh scheme is designed to refresh data with a long access latency when they are accessed several times. The second refresh scheme is designed to periodically check data with an extremely long access latency and refresh them. To further optimize the refresh overhead caused by the above refresh schemes, a dual-ECC-based refresh scheme is proposed. Besides, a mathematical model for all proposed schemes is constructed to clarify the benefit of each scheme. The experimental results show that the proposed schemes can significantly improve the tail latency with acceptable overhead. What is more, the access performance is well maintained compared with the state-of-the-art work. Yina Lv, Liang Shi 0001, Longfei Luo, Changlong Li 0006, Chun Jason Xue, Edwin H.-M. Sha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | SAC: A Stream Aware Write Cache Scheme for Multi-Streamed Solid State DrivesabstractThis work found that the state-of-the-art multi-streamed SSDs are inefficiently used due to two issues. First, the write cache inside SSDs is not aware of data from different streams, which induce conflict among streams. Second, the current stream identification methods are not accurate, which should be optimized inside SSDs. This work proposed a novel write cache scheme to efficiently utilize and optimize the multiple streams. First, an inter-stream aware cache partitioning scheme is proposed to manage the data from different streams. Second, an intra-stream based active cache evicting scheme is proposed to evict data to block with more invalid pages in priority. Experiment results show that the proposed scheme significantly reduces the write amplification (WAF) of multi-streamed SSDs by up to 28% with negligible cost. Chuanming Ding, Yina Lv, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha, Liang Shi 0001 |
ASP-DAC | 3 |
| 2021 | Dynamic File Cache Optimization for Hybrid SSDs with High-Density and Low-Cost Flash MemoryabstractOver the last few years, hybrid solid-state drives (SSDs) have been widely adopted due to their high performance and high capacity. Devices equipped with hybrid SSDs can be utilized to cache files from the network for performance improvement. However, this paper finds an interesting observation, that is, the efficiency of hybrid SSDs is significantly degraded instead of improved when too much data is cached. This is because the internal mode switching between different types of flash memory is affected by the device utilization. This paper proposes a dynamic file cache optimization scheme for hybrid SSDs, DFCache, which optimizes the device’s efficiency and limits unreasonable space consumption. DFCache includes two key ideas, dynamic cache space management, and intelligent cache file sifting. DFCache is implemented in Linux kernel and tested under real hybrid SSDs. Experimental results show that the I/O performance outperforms the state-of-the-art by up to 3.7x. Ben Gu, Longfei Luo, Yina Lv, Changlong Li 0006, Liang Shi 0001 |
ICCD | 3 |
| 2021 | Understanding and Optimizing Hybrid SSD with High-Density and Low-Cost Flash MemoryabstractWith the development of NAND flash technology, hybrid SSDs with high-density and low-cost flash memory have become the mainstream of the existing SSD architecture. In this architecture, two flash modes can be dynamically switched, such as single-level cell (SLC) mode and quad-level cell (QLC) mode. Based on evaluations and analysis of multiple real devices, this paper presents two interesting findings. They demonstrate that the coordination between the two flash-modes is not well-designed in existing architectures. This paper proposes HyFlex, which redesigns the strategies of data placement and flash-mode management of hybrid SSDs in a flexible approach. Specifically, two novel optimization strategies are proposed: velocity-based I/O scheduling (VIS) and garbage collection (GC)-aware capacity tuning (GCT). Experimental results show that HyFlex achieves encouraging performance and endurance improvement. Liang Shi 0001, Longfei Luo, Yina Lv, Changlong Li 0006, Edwin H.-M. Sha |
ICCD | 3 |
| 2020 | Access Characteristic Guided Partition for Read Performance Improvement on Solid State DrivesabstractSolid state drives (SSDs) are now widely deployed due to the development of high-density and low-cost NAND flash memories. Previous works have identified that the read performance of SSDs is degrading along with the development. One of the most critical reasons is the access interference between reads and writes, as the latest NAND flash memories have significant latency gap between reads and writes. This paper addresses this issue with the assistance of access characteristic guided SSD partitioning. First, several server workloads are studied and it is shown that reads and writes can be separated based on their access characteristics. Second, a set of techniques is proposed to place data judiciously for requests separation. Finally, a workload based SSD partitioning scheme is proposed to improve the read performance. The experimental results show that the proposed solution can improve read performance by 36% on average compared with the state-of-the-art solutions. Yina Lv, Liang Shi 0001, Qiao Li 0001, Chun Jason Xue, Edwin H.-M. Sha |
DAC | 1 |
| 2020 | Latency Variation Aware Read Performance Optimization on 3D High Density NAND Flash MemoryabstractState-of-the-art high density NAND flash memory has been recommended as read intensive storage device due to their excellent read performance. However, recent studies and reports show that the read latency of high density NAND flash memory is increasing. The reason comes from at least two aspects: First, high density flash generally adopts multiple bits per cell technique, where the access latency of the most significant bits is largely increased. Second, due to the reliability variation among these bits, the access latency of the most significant bits is further increased. We introduce RLV, a read performance optimization scheme is proposed to exploit the read latency variation among the multiple bits. The basic idea is that firstly identify the hotness of read data and then move them to the places with corresponding read latency. Our evaluation shows that RLV incurs negligible overhead, while improving read performance by 14% on average compared with state-of-the-arts. Yina Lv, Liang Shi 0001, Chun Jason Xue, Qingfeng Zhuge, Edwin H.-M. Sha |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | An Empirical Study of Hybrid SSD with Optane and QLC FlashabstractEmerging non-volatile memory (NVM) technologies provide a new way to solve the I/O bottleneck problem. As one of the widely respected solutions, hybrid storage device performance in the real environment is worth studying. Previously, due to the delayed progress of NVM, most of the studies are proceeded on simulated devices. In this paper, an empirical study is presented on the state-of-the-art hybrid storage device - Intel Optane H10, which is designed with Optane Memory and Quad-Level Cell (QLC) NAND flash. Several interesting findings are concluded with the study, which should be well considered during the employment. Yina Lv, Changlong Li 0006, Shouzhen Gu, Liang Shi 0001 |
ICCD | 2 |
| 2019 | Optimizing Tail Latency of LDPC based Flash Memory Storage Systems Via Smart RefreshabstractFlash memory has been developed with bit density improvement, technology scaling, and 3D stacking. With this trend, its reliability has been degraded significantly. Error correction code, low density parity code (LDPC), which has strong error correction capability, has been employed to solve this issue. However, one of the critical issues of LDPC is that it would introduce a long decoding latency on devices with low reliability. In this case, tail latency would happen, which will significantly impact the quality of service (QoS). In this work, a set of smart refresh schemes is proposed to optimize the tail latency. The basic idea of the work is to refresh data when the accessed data has a long decoding latency. Two smart refresh schemes are proposed for this work: The first refresh scheme is designed to refresh long access latency data when it is accessed several times for access performance optimization; The second refresh scheme is designed to periodical detecting data with extremely long access latency and refreshing them for tail latency optimization. Experiment results show that the proposed schemes are able to significantly improve the tail latency and access performance with little overhead. Yina Lv, Liang Shi 0001, Qiao Li 0001, Congming Gao, Chun Jason Xue, Edwin H.-M. Sha |
NAS | 1 |