Chenlin Ma

dblp:148/8695 · DBLP profile ↗
← Back
35ranked-venue papers
13as first author
25since 2021 · last 2026
0000-0002-0497-123XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 31 · 13 first-author · 23 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Resolving Gray Code Dilemma With Bidirectional Programming for Efficient QLC SSDs
abstract
QLC NAND flash is widely adopted in modern storage systems. By trading off read/write performance for storage density through a “time-for-space" approach, QLC enables ultra-high storage capacity. To mitigate performance degradation, Gray code and the two-step programming (TSP) algorithm are used. However, Gray code also has limitations: multiple Gray codes incur circuit overhead, while a single Gray code causes extra I/O latency overhead. This paradox seems unsolvable at first glance, requiring an innovative solution that maintains I/O performance without additional circuit overhead. A promising solution lies in selecting an appropriate Gray code and preventing hot data placement on slow physical pages. This paper proposes BDP, a novel Bi-Directional Programming scheme that adopts a single Gray code to fit both traditional (forward) and reverse programming directions based on TSP. The objective of BDP is to resolve the inherent contradiction between I/O performance preservation and implementation overhead. BDP optimizes the system performance through hardware/software co-design. At the hardware level, a fixed Gray code is employed to avoid additional circuit complexity. At the software level, two strategies (i.e. hotness-aware data allocation and background data migration) are proposed to further mitigate the misplacement of hot data on slow pages in QLC SSDs. The experimental results demonstrate that BDP significantly reduces the allocation of hot data to slow pages and enhances overall I/O performance compared to representative schemes.
Yi Wang 0003, Shaoqi Li, Yongbiao Zhu, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 PCD-ORAM: A Path-Aware and Cross-Layer Design to Enhance Data Locality in Oblivious RAM
Yi Wang 0003, Zhencheng Wang, Weixuan 'Vincent' Chen, Xianhua Wang, Chenlin Ma, Tianyu Wang 0009, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 Registry: Enhancing Vertex Reusability for GCN Inference on Hybrid Stacked Memory
Zhaoyu Zhong, Jiaxian Chen, Yunhao Dong, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Unifying Two Operators with One PIM: Leveraging Hybrid Bonding for Efficient LLM Inference
Jiaxian Chen, Yuxuan Qi, Kaoyi Sun, Zhiliang Lin, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
APPT6
2025 Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language Models
abstract
Retrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of $273 \times 55 \times$, and $2.41 \times$ over CPUs, GPUs, and prior art accelerators, respectively.
Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
DAC6
2025 Anchor First, Accelerate Next: Revolutionizing GNNs with PIM by Harnessing Stationary Data
abstract
Substantial data movement caused by irregular graph topologies hinders the efficient processing of graph neural networks (GNNs). Although the emerging near-bank processing-in-memory (PIM) architecture offers a promising solution to reduce data transfer between memory and computing units, cross-bank communication remains a critical challenge, limiting the benefits of PIM architectures. Our findings indicate that only $35.6 \%$ of the data can stay stationary within PIM units on average, with the rest requiring movement due to graph dependencies. This situation worsens as the number of PIM units increases, reducing the ratio to $18.7 \%$. In this paper, we argue that to fully leverage PIM architectures, systems must maximize stationary data and minimize the movement of non-stationary data. Following this principle, we propose Anchor, a scalable PIM architecture that exploits stationary data for GNNs through a hardware-software co-design approach. To maximize stationary data, we introduce the graph partitioning algorithm Mastav, which carefully allocates vertices and edges to preserve data locality. To minimize the movement of non-stationary data, we employ a two-step strategy. First, a customized dataflow ensures that non-stationary data is accessed and distributed exactly once. Second, an optimized communication mechanism reduces redundant data transfers through critical paths. Our extensive experiments demonstrate that Anchor significantly reduces processing latency and data movement compared to representative schemes.
Jiaxian Chen, Yuxuan Qi, Yongbiao Zhu, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
DAC7
2025 MiniWear: Minimizing Flash Wear via Hybrid Persistent Cache for Extended EF-SMR Lifetime
abstract
As the huge discrepancy between traffic and capacity persists, the lifetime of flash in EF-SMR systems faces a grave issue. EF-SMR systems combine NAND flash with Shingled Magnetic Recording (SMR) disks to achieve both low cost and high performance. However, previous research has primarily focused on issues such as write amplification and tail-latency in EF-SMR disks, overlooking the critical issue of flash lifetime. Studying the durability of EF-SMR systems is essential for developing future high-performance, low-cost storage solutions.This paper presents MiniWear, a hybrid persistent cache (PC) design aimed at extending the lifetime of EF-SMR systems. MiniWear adopts a hybrid medium persistent cache and proposes a customized scheduling strategy to reduce flash wear without impacting the EF-SMR system performance. At the hardware level, the hybrid PC of EF-SMR, composed of flash and SMR disk, is organized into Flash-PC and SMR-PC. At the software level, a fine-grained scheduling strategy is proposed to better manage PC resources. Additionally, we introduce a proactive balancing strategy to address PC resource idleness. Experimental results show that, compared to existing methods, MiniWear can reduce flash wear by up to 66.67%.
Chenlin Ma, Kaoyi Sun, Yuxuan Qi, Jiaxian Chen, Xiaochuan Zheng, Tianyu Wang 0009, Yi Wang 0003
DAC1
2025 Dancer: Dynamic Compression and Quantization Architecture for Deep Graph Convolutional Network
abstract
Graph Convolutional Networks (GCNs) have been widely applied in fields such as social network analysis and recommendation systems. Recently, deep GCNs have emerged, enabling the exploration of deeper hidden information. Compared to traditional shallow GCNs, deep GCNs feature significantly more layers, leading to considerable computational and data movement challenges. Processing-In-Memory (PIM) offers a promising solution for efficiently handling GCNs by enabling near-data computation, thus reducing data transfer between processing units and memory. However, previous work mainly focused on shallow GCNs and has shown limited performance with deep GCNs. In this paper, we present Dancer, an innovative PIM-based GCN accelerator. Dancer optimizes data movement during the inference process, significantly improving efficiency and reducing energy consumption. Specifically, we introduce a novel compressed graph storage architecture and a dynamic quantization technique to minimize data transfers at each layer of the GCN. Additionally, through a detailed analysis of weight dynamics changes, we propose a sparsity propagation strategy to further alleviate the computational and data transfer burden between layers. Experimental results demonstrate that, compared to current state-of-the-art methods, Dancer achieves 3.7× speedup, 7.6× energy efficiency, and reduces of 9.6× DRAM access on average.
Yunhao Dong, Zhaoyu Zhong, Yi Wang 0003, Chenlin Ma, Tianyu Wang 0009
DATE4
2025 One Gray Code Fits All: Optimizing Access Time with Bi-Directional Programming for QLC SSDs
abstract
Gray code, a voltage-level-to-data-bit translation scheme, is widely used in QLC SSDs. However, it causes the four data bits in QLC to exhibit significantly different read and write performance with up to 8 × latency variation, severely impacting the worst-case performance of QLC SSDs. This paper presents BDP, a novel Bi-Directional Programming scheme. Based on a fixed Gray code, BDP combines both the normal (forward) and reverse programming directions to enable runtime programming direction arbitration. Experimental results show that BDP can effectively improve the read and write performance of SSD compared to representative schemes.
Shaoqi Li, Tianyu Wang 0009, Yongbiao Zhu, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
DATE4
2025 EF-IMR: Embedded Flash with Interlaced Magnetic Recording Technology
abstract
Interlaced Magnetic Recording (IMR), a technology that improves storage density through track overlap, introduces significant latency due to Read-Modify-Write (RMW) operations. Writing to overlapped tracks affects underlying tracks, requiring additional I/O operations to read, back up, and rewrite them, resulting in significant head movement latency. We propose EF-IMR, a new architecture that ensures crash consistency in IMR while minimizing RMW latency and head movement. EF-IMR reduces head movement during RMW operations and decreases redundant RMW operations. Evaluations under real-world, intensive I/O workloads show that EF-IMR reduces RMW latency by 20.11 % and head movement latency by 89.37% compared to existing methods.
Chenlin Ma, Xiaochuan Zheng, Kaoyi Sun, Tianyu Wang 0009, Yi Wang 0003
DATE1
2024 Leanor: A Learning-Based Accelerator for Efficient Approximate Nearest Neighbor Search via Reduced Memory Access
abstract
Approximate Nearest Neighbor Search (ANNS) is a classical problem in data science. ANNS is both computationally-intensive and memory-intensive. As a typical implementation of ANNS, Inverted File with Product Quantization (IVFPQ) has the properties of high precision and rapid processing. However, the traversal of non-nearest neighbor vectors in IVFPQ leads to redundant memory accesses. This significantly impacts retrieval efficiency. A promising approach involves the utilization of learned indexes, leveraging insights from data distribution to optimize search efficiency. Existing learned indexes are primarily customized for low-dimensional data. How to tackle ANNS in high-dimensional vectors is a challenging issue.
Yi Wang 0003, Jianan Yuan, Jiaxian Chen, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001
DAC6
2024 Rapper: A Parameter-Aware Repair-in-Memory Accelerator for Blockchain Storage Platform
abstract
Blockchain storage platforms reward storage nodes for keeping user-uploaded data for a certain amount of time. These storage nodes are unstable and can go online or offline unpredictably at any time, leading to potential data loss. To prevent data loss, blockchain storage platforms adopt erasure codes on user-uploaded encrypted data. Data repair processes will be performed to recover the lost data. However, the data repair processes heavily rely on time-consuming erasure coding algorithms, mainly consisting of vector-matrix multiplications. The emerging processing-in-memory technique can efficiently speed up the processing of vector-matrix multiplications. It can be integrated into blockchain storage platforms to solve the data repair issue. This paper presents Rapper, a parameter-aware repair-inmemory accelerator for blockchain storage platforms. Rapper utilizes the computing power of emerging processing-in-memory architecture so that data repair processes can be processed in a parallel manner and the overall efficiency can be improved significantly. Specifically, at the hardware level, the ReRAM memory is reorganized into our proposed double bank, XRU, XGroup, and ReRAM crossbars structure. At the software level, a parallel decoding/encoding strategy is proposed to fully exploit the internal parallelism of ReRAM. We also propose an adaptive parameter-aware mapping to handle various sizes of stripes. To demonstrate the viability of the proposed technique, a representative blockchain storage project Storj is adopted as the default storage infrastructure. Experimental results show that Rapper can achieve a 1.96 × speedup on average compared to the representative scheme.
Chenlin Ma, Yingping Wang, Fuwen Chen, Jing Liao 0008, Yi Wang 0003, Rui Mao 0001
HPCA1
2024 Boosting Write Performance of KV Stores: An NVM - Enabled Storage Collaboration Approach
abstract
As the most common data structure for key-value stores, LogStructured Merge Tree (LSM-tree) can eliminate random write operations and keep acceptable read performance. However, write stall and write amplification introduced by the leveled compaction of LSM-tree significantly degrade the system performance. The emerging non-volatile memory (NVM) provides byte-addressable access and low-latency data persistence. Integrating DIMM-interface NVM in the design of the LSM-tree can potentially alleviate the write stall and write amplification issue, as the access speed of NVM is several orders of magnitude faster than hard disk drives or flash memory-based solid-state drives. This hybrid storage should be carefully designed, requiring new architectural and key-value structural support. This paper presents ZigZagDB, an NVM-enabled data man-agement scheme for LSM-tree-based key-value stores. ZigZagDB adds additional layers of key-value stores and uses non-volatile memory as the storage media to hold these additional layers of data. The newly designed key-value stores alternately access the data from either SSD or NVM. This ‘ZigZag’ shape of storage collaboration and synchronization can benefit write efficiency and space utilization. By utilizing the NVM with very limited capacity, the redesigned organization of LSM-tree can effectively solve the write stall and write amplification issue. We demonstrate the viability of the proposed ZigZagDB using a set of extensive experiments. Experimental results show that ZigZagDB can significantly reduce the write amplification and boost the throughput in comparison with representative schemes.
Yi Wang 0003, Jiajian He, Kaoyi Sun, Yunhao Dong, Jiaxian Chen, Chenlin Ma, Amelie Chi Zhou, Rui Mao 0001
ICDE6
2024 LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV Table
abstract
Log-structured merge-tree (LSM-tree) is a storage architecture widely used in key-value (KV) stores. To enhance the read efficiency of LSM-tree, recent works utilize the learned index to learn the mapping between keys and locations. However, in existing learned-index-aided KV stores, inefficient design of the learned index and disk access significantly impact the read performance. How to design a learned KV store to improve index efficiency and minimize disk access remains a critical problem. This paper presents LeaderKV, a read-optimized LSM-tree-based KV store. LeaderKV employs decoupled KV tables (DK-Table) and efficient learned indexes for data retrieval. DKTables are storage files in Leader Kvbecause they avoid reading irrelevant data in collaboration with learned indexes during queries. A learned index called Leader is proposed to accelerate data retrieval within DKTable. Leader is composed of precise models and approximate models. A redirect mechanism is designed to reduce the cost of mispredictions in Leader. We integrate DKTable and Leader into LeaderKV and demonstrate its effectiveness using a variety of datasets and workloads. Experimental results show that LeaderKV significantly improves the read performance compared to representative schemes.
Yi Wang 0003, Jianan Yuan, Shangyu Wu, Jiaxian Chen, Chenlin Ma, Jianbin Qin
ICDE6
2024 NICE: A Nonintrusive In-Storage-Computing Framework for Embedded Applications
abstract
Embedded machine learning applications face challenges related to massive data movement and high computational intensity, exacerbated by the limited performance of mobile devices. Computational storage devices (CSDs) pose huge potential for accelerating both data-intensive and computation-intensive embedded machine learning tasks by effectively reducing data movement and leveraging built-in accelerators. However, existing in-storage-computing (ISC) frameworks either require invasive customization of existing host driver layers or necessitate complex device firmware modifications, hindering the widespread deployment of CSDs. In addition, the lack of file semantics and the constrained internal resources within CSD implicitly compromise system performance and impact normal read/write performance. In this article, we aim to provide a nonintrusive in-storage-computing framework for embedded applications, named NICE. This framework includes an easy-to-use ISC programming interface that bypasses the kernel stack and requires no modification to the host NVMe driver, which is achieved through a novel hyper-addressing-based programming library and a file-aware page data layout within the CSD. In addition, we incorporate a lightweight kernel with coroutine-based command scheduling and several FPGA-based accelerators within the storage device firmware to enhance the performance of embedded machine learning applications while ensuring that the normal I/O performance remains unaffected. NICE is implemented on real CSD hardware integrated with ARM and FPGA. Experimental results demonstrate that our NICE framework can achieve an average latency performance improvement of$43.5\times $($9.32\times $) compared to CPU-(GPU-) based embedded machine learning solutions using the state-of-the-art NVIDIA Jetson NX platform, with$27.5\times $($4.3\times $) higher energy efficiency. NICE also has$34.2\times $less software and I/O performance overheads than state-of-the-art ISC frameworks.
Tianyu Wang 0009, Yongbiao Zhu, Shaoqi Li, Jin Xue, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Lift: Exploiting Hybrid Stacked Memory for Energy-Efficient Processing of Graph Convolutional Networks
abstract
Graph Convolutional Networks (GCNs) are powerful learning approaches for graph-structured data. GCNs are both computing- and memory-intensive. The emerging 3D-stacked computation-in-memory (CIM) architecture provides a promising solution to process GCNs efficiently. The CIM architecture can provide near-data computing, thereby reducing data movement between computing logic and memory. However, previous works do not fully exploit the CIM architecture in both dataflow and mapping, leading to significant energy consumption.This paper presents Lift, an energy-efficient GCN accelerator based on 3D CIM architecture using software and hardware co-design. At the hardware level, Lift introduces a hybrid architecture to process vertices with different characteristics. Lift adopts near-bank processing units with a push-based dataflow to process vertices with strong re-usability. A dedicated unit is introduced to reduce massive data movement caused by high-degree vertices. At the software level, Lift adopts a hybrid mapping to further exploit data locality and fully utilize the hybrid computing resources. The experimental results show that the proposed scheme can significantly reduce data movement and energy consumption compared with representative schemes.
Jiaxian Chen, Zhaoyu Zhong, Kaoyi Sun, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
DAC4
2023 Meta-Block: Exploiting Cross-Layer and Direct Storage Access for Decentralized Blockchain Storage Systems
abstract
Decentralized storage systems such as blockchain storage applications adopt the distributed storage technology and use distributed storage nodes to store the persistent data. For each off-chain storage node, key-value (KV) stores are normally used to manage data. As the most common data structure for KV store, Log Structured Merge Tree (LSM-Tree) eliminates random write operations and keeps acceptable read performance. Although LSM-Tree-based decentralized storage system can provide a secure and reliable storage platform, the unique feature of blockchain applications is not fully exploited. In blockchain storage applications, the generation of keys for KV stores is based on the encrypted data, and the key determines the allocation of data. Since the granularity for a read/write request at the level of blockchain storage platform is much smaller than that at the level of LSM-Tree or flash memory, a physical block in a solid-state drive (SSD) could be filled with data from different system users. This mixture of workloads will lead to the inefficient usage of physical spaces in the SSD and cause extra compaction operations for LSM-Tree. This paper presentsMeta-Block, a cross-layer and efficient storage management strategy for decentralized blockchain storage applications. Meta-Block utilizes rich functionalities provided by the system infrastructure of open-channel SSD to provide direct storage accesses for the off-chain storage node. The objective is to capture the features of blockchain storage applications and reduce unnecessary read and write operations across different storage layers. As a cross-layer design, Meta-Block redesigns the organization of LSM-Tree, which can effectively reduce the write amplification. We also design a data prefetching strategy to speed up the indexing and enable direct storage access. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that Meta-Block can effectively reduce the write amplification and extend the lifetime of SSDs in comparison with representative schemes.
Yi Wang 0003, Jing Liao 0008, Jing Yang 0018, Zhengda Li, Chenlin Ma, Rui Mao 0001
IEEE Trans. Computers5
2023 Tidal-Tree-Mem: Toward Read-Intensive Key-Value Stores With Tidal Structure Based on LSM-Tree
abstract
The log-structured merge-tree (LSM-tree)-based key-value store has been widely adopted by many large-scale data storage applications for its excellent write performance. However, such write performance gains mainly come from scarifying read performance due to the leveled and log-structured intrinsic characteristics of the LSM-tree. Therefore, the critical challenge of the existing LSM-tree is how to improve the read efficiency by reducing read amplification. This article for the first time proposes Tidal-tree-Mem, a novel data structure where data flow inside the LSM-tree-like Tidal waves. First, a floating strategy is proposed to allow frequently accessed files at the bottom of the LSM-tree to move to higher positions, reducing read amplification. Second, a stretching strategy is proposed to vary the shape of the LSM-tree to adapt to workloads with different characteristics. To evaluate the performance of Tidal-tree-Mem, we conduct a series of experiments using standard benchmarks from YCSB. The experimental results show that Tidal-tree-Mem can effectively reduce read amplification and the overall latency by over 71.94% and 47.34%, respectively, compared with representative schemes.
Chenlin Ma, Shangyu Wu, Yi Wang 0003, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Work-in-Progress: Lark: A Learned Secondary Index Toward LSM-tree for Resource-Constrained Embedded Storage Systems
abstract
LSM-tree-based key-value stores are popular in embedded storage systems. With the growing demands of data analysis, the secondary index is created to support non-primary-key lookups. However, the lookup efficiency and space consumption of secondary index remain for further optimization. Inspired by the learned index, this paper presents Lark, a learned secondary index toward LSM-tree for resource-constrained embedded storage systems. Lark employs machine learning to speed up the non-primary-key queries and compress secondary indexes. Our preliminary evaluations show that, in comparison with traditional secondary index schemes, Lark achieves better lookup performance with less space consumption.
Jianan Yuan, Shangyu Wu, Yiquan Lin, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
CODES+ISSS6
2022 MU-RMW: Minimizing Unnecessary RMW Operations in the Embedded Flash with SMR Disk
abstract
Emerging Shingled Magnetic Recording (SMR) Disk can improve the storage capacity significantly by overlapping multiple tracks with the shingled direction. However, the shingled-like structure leads to severe write amplification caused by RMW operations inner SMR disks. As the mainstream solid-state storage technology, NAND flash has the advantages of tiny size, cost-effective, high performance, making it suitable and promising to be incorporated into SMR disks to boost the system performance. In this hybrid embedded storage system (i.e., the Embedded Flash with SMR disk (EF-SMR) system), we observe that physical flash blocks can contain a mixture of data associated with different SMR data bands; when garbage collecting such flash blocks, multiple RMW operations are triggered to rewrite the involved SMR bands and the performance is further exacerbated. Therefore, in this paper, we for the first time present MU-RMW to guarantee data from different SMR bands will not be mixed up within the flash blocks with an aim at minimizing unnecessary RMW operations. The effectiveness of MU-RMW was evaluated with realistic and intensive I/O workloads and the results are encouraging.
Chenlin Ma, Zhuokai Zhou, Yingping Wang, Yi Wang 0003, Rui Mao 0001
DATE1
2022 GCIM: Toward Efficient Processing of Graph Convolutional Networks in 3D-Stacked Memory
abstract
Graph convolutional networks (GCNs) have become a powerful deep learning approach for graph-structured data. Different from traditional neural networks such as convolutional neural networks, GCNs handle irregular input graph data, and GCNs are both computation-bound and memory-bound. How to efficiently utilize the underlying computation and memory resource becomes a critical issue. The emerging 3D-stacked computation-in-memory (CIM) architecture can reduce the data movement between computing logic and memory, thereby presenting a promising solution for the processing of GCNs. An unsolved key challenge is how to allocate GCNs to take advantage of fast near-data processing of the 3D-stacked CIM architecture. This article presents GCIM, a software–hardware co-design approach to exploit the efficient processing of GCNs on the CIM architecture. At the level of hardware design, GCIM integrates lightweight computing units near memory banks to fully exploit bank-level bandwidth and parallelism. At the level of software design, a locality-aware data mapping algorithm is proposed to partition the input graph and achieve workload balancing. GCIM is evaluated through a set of representative GCN models and standard graph datasets. The experimental results show that GCIM can significantly reduce the processing latency and data movement overhead compared with representative schemes.
Jiaxian Chen, Yiquan Lin, Kaoyi Sun, Jiexin Chen, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 Rebirth-FTL: Lifetime Optimization via Approximate Storage for NAND Flash Memory
abstract
The lifetime of NAND flash cells significantly degrades with feature-size reductions and multilevel cell technology. On the other hand, we have more and more approximate data, such as images and videos that are more error tolerant than regular data like text. In this article, we propose Rebirth-FTL, which reuses faulty blocks that contain uncorrectable errors to store approximate data for lifetime optimization. Rebirth-FTL effectively manages two spaces, namely, the approximate space and the normal space, with an efficient address translator, a coordinated garbage collection, and a differential wear leveler. In addition, we develop an migration times restriction (MTR) policy to restrict the movement of the approximate data in the approximate space. We also develop a scheme to pass approximate information from userland to kernel space in Linux. Finally, a lifetime model is presented for lifetime analysis. Our experimental results show that Rebirth-FTL can extend the lifetime by 41.63% on average.
Chenlin Ma, Zhuokai Zhou, Zhaoyan Shen, Yi Wang 0003, Renhai Chen, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 MAID-Q: Minimizing Tail Latency in Embedded Flash With SMR Disk via -Learning Model
abstract
As the mainstream solid-state storage technology, NAND flash has the advantages of tiny size, cost-effective, and high performance, which make it a promising candidate to be embedded into the shingled magnetic recording (SMR) disk to build a faster, denser, and cheaper storage system. However, such an embedded flash with SMR (EF-SMR) disk system suffers from lengthy tail-latency due to “reclamation issues” in both the NAND flash and the SMR disk. Our preliminary observations reveal that tremendous idle time intervals exist in real-world scenarios, and few prior works have focused on addressing the tail-latency issue in the EF-SMR disk. In this article, we propose a novel method termed MAID-Q to fully exploit the idle time intervals to minimize the lengthy tail-latency of the EF-SMR disk based on a lightweight reinforcement learning model (i.e., the$Q$-learning model). In addition, fine-grained block-level space management and a parallel reclamation strategy are proposed to improve the reclamation efficiency and hide the reclamation overheads. The effectiveness of our proposed design was evaluated with realistic I/O traces, and the results show that the proposed design can remedy the tail-latency by 88.31% and improve the overall performance by 79.03%.
Chenlin Ma, Zhuokai Zhou, Yingping Wang, Yi Wang 0003, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Tiler: An Autonomous Region-Based Scheme for SMR Storage
abstract
Shingled Magnetic Recording (SMR) Disks are adopted as a high-density, non-volatile media that significantly precedes conventional disks in both the storage capacity and cost. However, inefficient read-modify-writes (RMWs) greatly challenge the management of SMR disks. This article for the first time presents an approach called Tiler to manage SMR disks by dividing the physical space into small autonomous regions (ARs). Each AR can manage its space allocation, address mapping, and cleaning independently. By managing these ARs in a log-structured way, RMWs can be avoided; besides, ARs can also help update data when the adjacent tracks contain no valid data. Tiler is capable of partitioning a large-scale cleaning into self-contained-small-scale cleaning and thus, the data that need to be relocated are limited inside independent ARs, which further minimizes the performance overhead. Our experimental results show that Tiler can shorten the overall system response time by 50.21 percent and reduce the cleaning time by 90.24 percent on average.
Chenlin Ma, Zhaoyan Shen, Yi Wang 0003, Renhai Chen, Zili Shao
IEEE Trans. Computers1
2021 Leveraging the Interplay of RAID and SSD for Lifetime Optimization of Flash-Based SSD RAID
abstract
Flash-based SSD RAID arrays are increasingly being deployed in data centers. Compared with HDD arrays, SSD arrays drastically enhance I/O performance and density, and reduce power, cooling, and rack space. Nevertheless, SSDs suffer aging issues. Especially, an SSD has limited endurance and needs to be replaced when it reaches to the end of its lifetime. Although prior studies have been conducted to address this disadvantage, effective techniques of RAID/SSD controllers are urgently needed to extend the lifetime of SSD arrays. In this article, we propose a novel RAID architecture, called FreeRAID, to leverage the interplay of RAID and SSD controllers to optimize the lifespan of flash-based SSD arrays. FreeRAID adds a new exploitable phase to the life cycle of flash blocks. In FreeRAID, flash space is separated into normal space and exploitable space, and they are used to serve normal data and approximate data, respectively. We design a dual-space management scheme for RAID controllers to intelligently allocate SSD spaces based on their aging status. Inside an SSD, we propose an adaptive flash translation layer for the SSD controller to maintain the reliability and space efficiency of flash memories. We implemented a prototype of FreeRAID based on an SSD array simulator. Our experiments show that FreeRAID can significantly increase the lifetime by up to 3.07 × compared with conventional SSD-based RAID arrays.
Zhaoyan Shen, Chenlin Ma, Zhiping Jia, Tao Li 0006, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 KFR: Optimal Cache Management with K-Framed Reclamation for Drive-Managed SMR Disks
abstract
Shingled Magnetic Recording (SMR) disks have been proposed as a promising solution to satisfy the increasing capacity need in the big data era. Drive-Managed SMR (DM-SMR) disk which acts as a traditional block device is favored for providing high compatibility. However, DM-SMR disks suffer from high performance recovery time (PRT) due to the "SMR space reclamation" issue. This paper proposes an optimal cache management named K-Framed Reclamation (KFR) to minimize PRT within the DM-SMR disk. The effectiveness of our proposed design was evaluated with realistic and intensive I/O workloads and the results are encouraging.
Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
DAC1
2020 An Efficient Directory Entry Lookup Cache With Prefix-Awareness for Mobile Devices
abstract
Modern mobile devices, such as smartphones, maintain a directory cache (DCache) to accelerate directory and file accesses. However, the original DCache recursively walks through all the components of a path name for each directory entry (dentry) lookup operation, leading to low lookup efficiency. In this article, we first investigate intrinsic characteristics of dentry lookup operations in smartphones and make several interesting findings: 1) file path lookup operations are called frequently (up to 104times per second) by mobile applications; 2) file path lookup operations present high temporal and spatial locality; and 3) the latency of a file path lookup is linear to the depth of the path name. Based on our findings, we further propose an efficient directory entry lookup cache architecture, named dynamic skipping cache (DS-Cache), which adopts an ASCII-based hash table to simplify the path lookup complexity. In DS-Cache, a dynamic skip lookup algorithm is proposed to skip the common prefixes of different accessing paths. A rename algorithm and a delete algorithm are designed to promise the consistence of DS-Cache and the backend file system. We also design a prefix-aware cache replacement scheme to optimize the DS-Cache hit ratio. We have implemented and deployed DS-Cache on a Google Nexus 6P smartphone. The experimental results show that we can significantly reduce the latency of invoking system calls by up to 81%, and further reduce the completion time of real-world mobile applications by up to 67%.
Zhaoyan Shen, Renhai Chen, Chenlin Ma, Zhiping Jia, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 MNFTL: An Efficient Flash Translation Layer for MLC NAND Flash Memory
abstract
The write constraints of Multi-Level Cell (MLC) NAND flash memory make most of the existing flash translation layer (FTL) schemes inefficient or inapplicable. In this article, we solve several fundamental problems in the design of MLC flash translation layer. The objective is to reduce the garbage collection overhead to reduce the average system response time. We make the key observation that the valid pages copy is the essential garbage collection overhead. Based on this observation, we propose two approaches, namely, concentrated mapping and postponed reclamation, to effectively reduce the valid pages copy. Besides, we propose a progressive garbage collection that can well utilize the system idle time to reclaim more spaces. We conduct a series of experiments on an embedded developing board with a set of benchmarks. The experimental results show that our scheme can achieve a significant reduction in the average system response time compared with the previous work.
Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Renhai Chen, Zili Shao
ACM Trans. Design Autom. Electr. Syst.1
2019 Delay-based I/O request scheduling in SSDs
Renhai Chen, Qiming Guan, Chenlin Ma, Zhiyong Feng 0002
J. Syst. Archit.3
2019 FC: Built-in flash cache with fast cleaning for SMR storage systems
Chenlin Ma, Zhaoyan Shen, Renhai Chen, Zili Shao
J. Syst. Archit.1
2019 Alleviating Hot Data Write Back Effect for Shingled Magnetic Recording Storage Systems
abstract
Shingled magnetic recording (SMR) is a novel technology that can significantly increase the disk capacity and reduce the cost. However, this new technology leads to write constraint that prohibits random writes to SMR disks. In order to solve this, current solutions utilize a persistent cache (PC) to change random writes to sequential ones. By doing this, nevertheless, the hot-data write-back effect will be triggered which inevitably incurs extra read-modify-write (RMW) operations. This paper presents a cache management scheme calleddual-bufferto manage the overall SMR space, by which the PC is partitioned into the persistent buffer and the filter buffer. It leaves hot data in the filter buffer and moves cold data back to the SMR disk so that the hot data write-back effect can be alleviated. Since the PC may contain a large amount of valid data,dual-bufferalso presents a prediction-based dynamic configuration strategy so hot data can be cached as much as possible. We conducted a series of experiments with synthetic traces. Experimental results show that ourdual-bufferscheme can shorten the average response time by 51.66% on average and reduce the total number of RMW operations by 98.76% on average compared to the previous work.
Chenlin Ma, Zhaoyan Shen, Yi Wang 0003, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 RMW-F: A Design of RMW-Free Cache Using Built-in NAND-Flash for SMR Storage
abstract
Shingled Magnetic Recording (SMR) disks have been proposed as a high-density, non-volatile media and precede traditional hard disk drives in both storing capacity and cost. However, the intrinsic characteristics of SMR disks raise a major performance challenge named read-modify-write operations (RMWs) that are time-consuming and can significantly degrade the overall system performance. Current designs of SMR disks usually adopt a persistent cache to alleviate the negative effect brought by RMWs and the cache is used as a first-level cache to buffer all the incoming writes of the whole SMR storage system. In this paper, we propose to change the functionality of the cache, that is, the cache will no longer serve as a first-level cache like previous. Incoming data are distinguished according to their different write-back behavior and those data which will incur RMWs will be left in our built-in NAND flash cache called RMW-free Cache (RMW-F) to eliminate the need of RMWs. Besides, RMW-F improves the cleaning efficiency by a model that takes both write-back cost and data popularity into considerations. Our experimental results show that RMW-F can achieve both system performance and cleaning efficiency improvements.
Chenlin Ma, Zhaoyan Shen, Zili Shao
ACM Trans. Embed. Comput. Syst.1
2017 A Block-Level Log-Block Management Scheme for MLC NAND Flash Memory Storage Systems
abstract
NAND flash memory is the major storage media for both mobile storage cards and enterprise Solid-State Drives (SSDs). Log-block-based Flash Translation Layer (FTL) schemes have been widely used to manage NAND flash memory storage systems in industry. In log-block-based FTLs, a few physical blocks called log blocks are used to hold all page updates from a large amount of data blocks. Frequent page updates in log blocks introduce big overhead so log blocks become the system bottleneck. To address this problem, this paper presents BLog, a block-level log-block management scheme for MLC NAND flash memory storage system. In BLog, with block-level management, the update pages of a data block can be collected together and put into the same log block as much as possible; therefore, we can effectively reduce the associativities of log blocks so as to reduce the garbage collection overhead. We also propose a novel partial merge operation strategy called reduced-order merge by which we can effectively postpone the garbage collection of log blocks so as to maximally utilize valid pages and reduce unnecessary erase operations in log blocks. Based on BLog, we design an FTL called BLogFTL for Multi-Level Cell (MLC) NAND flash. We conduct a set of experiments on a real hardware platform. Both representative FTL schemes and the proposed BLogFTL have been implemented in the hardware evaluation board. The experimental results show that our scheme can effectively reduce the garbage collection operations and reduce the system response time compared to the previous log-block-based FTLs for MLC NAND flash.
Chenlin Ma, Renhai Chen, Yi Wang 0003, Zili Shao
IEEE Trans. Computers3
2016 NVMRA: utilizing NVM to improve the random write operations for NAND-flash-based mobile devices
abstract
NAND flash memory has become the major storage media in mobile devices, such as smartphones. However, the random write operations of NAND flash memory heavily affect the I/O performance, thus seriously degrading the application performance in mobile devices. The main reason for slow random write operations is the out-of-place update feature of NAND flash memory. Newly emerged non-volatile memory, such as phase-change memory, spin transfer torque, supports in-place updates and presents much better I/O performance than that of flash memory. All these good features make non-volatile memory (NVM) as a promising solution to improve the random write performance for NAND flash memory. In this paper, we propose a non-volatile memory for random access (NVMRA) scheme to utilize NVM to improve the I/O performance in mobile devices. NVMRA exploits the I/O behaviors of applications to improve the random write performance for each application. Based on different I/O behaviors, such as random write-dominant I/O behavior, NVMRA adopts different storing decisions. The scheme is evaluated on a real Android 4.2 platform. The experimental results show that the proposed scheme can effectively improve the I/O performance and reduce the I/O energy consumption for mobile devices. Copyright © 2015 John Wiley & Sons, Ltd.
Renhai Chen, Zhaoyan Shen, Chenlin Ma, Zili Shao
Softw. Pract. Exp.3
2014 Reference direction based immune clone algorithm for many-objective optimization
Ruochen Liu 0006, Chenlin Ma, Wenping Ma 0001, Licheng Jiao
Frontiers Comput. Sci.2