Yi Wang 0003

dblp:67/6649-3 · DBLP profile ↗
← Back
103ranked-venue papers
34as first author
35since 2021 · last 2026
0000-0002-5773-3817ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 86 · 27 first-author · 31 since 2021Software engineering, systems software and programming languages · 15 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 Pipelonk: Accelerating End-to-End Zero-Knowledge Proof Generation on GPUs for PLONK-Based Protocols
abstract
Zero-knowledge proofs (ZKPs) are cryptographic protocols that allow verification of statements without disclosing the underlying information. Among them, PLONK-based ZKPs are particularly notable for offering succinct, non-interactive proofs of knowledge with a universal trusted setup, leading to widespread adoption in blockchain and cryptocurrency applications. Nonetheless, their broader deployment is hindered by long proof-generation times and substantial memory demands. While GPUs can accelerate these computations, their limited memory capacity introduces significant challenges for efficient end-to-end proof generation.
Zhiyuan Zhang 0008, Yanxin Cai, Wenhao Yin, Xueyu Wu 0001, Yi Wang 0003, Lei Ju 0001, Zhuoran Ji
PPoPP5
2026 The design and formal verification of a merging protocol for autonomous vehicles
Rui Wang 0024, Fang Qi, Yixiao Yang, Yi Wang 0003
Future Gener. Comput. Syst.6
2026 PIRacle: A Fast and Scalable Private Information Retrieval System for Key-Value Stores
Zhaoyan Shen, Yi Wang 0003, Lei Ju 0001
IEEE Trans. Computers3
2026 Resolving Gray Code Dilemma With Bidirectional Programming for Efficient QLC SSDs
abstract
QLC NAND flash is widely adopted in modern storage systems. By trading off read/write performance for storage density through a “time-for-space" approach, QLC enables ultra-high storage capacity. To mitigate performance degradation, Gray code and the two-step programming (TSP) algorithm are used. However, Gray code also has limitations: multiple Gray codes incur circuit overhead, while a single Gray code causes extra I/O latency overhead. This paradox seems unsolvable at first glance, requiring an innovative solution that maintains I/O performance without additional circuit overhead. A promising solution lies in selecting an appropriate Gray code and preventing hot data placement on slow physical pages. This paper proposes BDP, a novel Bi-Directional Programming scheme that adopts a single Gray code to fit both traditional (forward) and reverse programming directions based on TSP. The objective of BDP is to resolve the inherent contradiction between I/O performance preservation and implementation overhead. BDP optimizes the system performance through hardware/software co-design. At the hardware level, a fixed Gray code is employed to avoid additional circuit complexity. At the software level, two strategies (i.e. hotness-aware data allocation and background data migration) are proposed to further mitigate the misplacement of hot data on slow pages in QLC SSDs. The experimental results demonstrate that BDP significantly reduces the allocation of hot data to slow pages and enhances overall I/O performance compared to representative schemes.
Yi Wang 0003, Shaoqi Li, Yongbiao Zhu, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 PCD-ORAM: A Path-Aware and Cross-Layer Design to Enhance Data Locality in Oblivious RAM
Yi Wang 0003, Zhencheng Wang, Weixuan 'Vincent' Chen, Xianhua Wang, Chenlin Ma, Tianyu Wang 0009, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2026 Registry: Enhancing Vertex Reusability for GCN Inference on Hybrid Stacked Memory
Zhaoyu Zhong, Jiaxian Chen, Yunhao Dong, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 Unifying Two Operators with One PIM: Leveraging Hybrid Bonding for Efficient LLM Inference
Jiaxian Chen, Yuxuan Qi, Kaoyi Sun, Zhiliang Lin, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
APPT7
2025 Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language Models
abstract
Retrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of $273 \times 55 \times$, and $2.41 \times$ over CPUs, GPUs, and prior art accelerators, respectively.
Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
DAC7
2025 Anchor First, Accelerate Next: Revolutionizing GNNs with PIM by Harnessing Stationary Data
abstract
Substantial data movement caused by irregular graph topologies hinders the efficient processing of graph neural networks (GNNs). Although the emerging near-bank processing-in-memory (PIM) architecture offers a promising solution to reduce data transfer between memory and computing units, cross-bank communication remains a critical challenge, limiting the benefits of PIM architectures. Our findings indicate that only $35.6 \%$ of the data can stay stationary within PIM units on average, with the rest requiring movement due to graph dependencies. This situation worsens as the number of PIM units increases, reducing the ratio to $18.7 \%$. In this paper, we argue that to fully leverage PIM architectures, systems must maximize stationary data and minimize the movement of non-stationary data. Following this principle, we propose Anchor, a scalable PIM architecture that exploits stationary data for GNNs through a hardware-software co-design approach. To maximize stationary data, we introduce the graph partitioning algorithm Mastav, which carefully allocates vertices and edges to preserve data locality. To minimize the movement of non-stationary data, we employ a two-step strategy. First, a customized dataflow ensures that non-stationary data is accessed and distributed exactly once. Second, an optimized communication mechanism reduces redundant data transfers through critical paths. Our extensive experiments demonstrate that Anchor significantly reduces processing latency and data movement compared to representative schemes.
Jiaxian Chen, Yuxuan Qi, Yongbiao Zhu, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003
DAC8
2025 MiniWear: Minimizing Flash Wear via Hybrid Persistent Cache for Extended EF-SMR Lifetime
abstract
As the huge discrepancy between traffic and capacity persists, the lifetime of flash in EF-SMR systems faces a grave issue. EF-SMR systems combine NAND flash with Shingled Magnetic Recording (SMR) disks to achieve both low cost and high performance. However, previous research has primarily focused on issues such as write amplification and tail-latency in EF-SMR disks, overlooking the critical issue of flash lifetime. Studying the durability of EF-SMR systems is essential for developing future high-performance, low-cost storage solutions.This paper presents MiniWear, a hybrid persistent cache (PC) design aimed at extending the lifetime of EF-SMR systems. MiniWear adopts a hybrid medium persistent cache and proposes a customized scheduling strategy to reduce flash wear without impacting the EF-SMR system performance. At the hardware level, the hybrid PC of EF-SMR, composed of flash and SMR disk, is organized into Flash-PC and SMR-PC. At the software level, a fine-grained scheduling strategy is proposed to better manage PC resources. Additionally, we introduce a proactive balancing strategy to address PC resource idleness. Experimental results show that, compared to existing methods, MiniWear can reduce flash wear by up to 66.67%.
Chenlin Ma, Kaoyi Sun, Yuxuan Qi, Jiaxian Chen, Xiaochuan Zheng, Tianyu Wang 0009, Yi Wang 0003
DAC7
2025 Dancer: Dynamic Compression and Quantization Architecture for Deep Graph Convolutional Network
abstract
Graph Convolutional Networks (GCNs) have been widely applied in fields such as social network analysis and recommendation systems. Recently, deep GCNs have emerged, enabling the exploration of deeper hidden information. Compared to traditional shallow GCNs, deep GCNs feature significantly more layers, leading to considerable computational and data movement challenges. Processing-In-Memory (PIM) offers a promising solution for efficiently handling GCNs by enabling near-data computation, thus reducing data transfer between processing units and memory. However, previous work mainly focused on shallow GCNs and has shown limited performance with deep GCNs. In this paper, we present Dancer, an innovative PIM-based GCN accelerator. Dancer optimizes data movement during the inference process, significantly improving efficiency and reducing energy consumption. Specifically, we introduce a novel compressed graph storage architecture and a dynamic quantization technique to minimize data transfers at each layer of the GCN. Additionally, through a detailed analysis of weight dynamics changes, we propose a sparsity propagation strategy to further alleviate the computational and data transfer burden between layers. Experimental results demonstrate that, compared to current state-of-the-art methods, Dancer achieves 3.7× speedup, 7.6× energy efficiency, and reduces of 9.6× DRAM access on average.
Yunhao Dong, Zhaoyu Zhong, Yi Wang 0003, Chenlin Ma, Tianyu Wang 0009
DATE3
2025 One Gray Code Fits All: Optimizing Access Time with Bi-Directional Programming for QLC SSDs
abstract
Gray code, a voltage-level-to-data-bit translation scheme, is widely used in QLC SSDs. However, it causes the four data bits in QLC to exhibit significantly different read and write performance with up to 8 × latency variation, severely impacting the worst-case performance of QLC SSDs. This paper presents BDP, a novel Bi-Directional Programming scheme. Based on a fixed Gray code, BDP combines both the normal (forward) and reverse programming directions to enable runtime programming direction arbitration. Experimental results show that BDP can effectively improve the read and write performance of SSD compared to representative schemes.
Shaoqi Li, Tianyu Wang 0009, Yongbiao Zhu, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
DATE5
2025 EF-IMR: Embedded Flash with Interlaced Magnetic Recording Technology
abstract
Interlaced Magnetic Recording (IMR), a technology that improves storage density through track overlap, introduces significant latency due to Read-Modify-Write (RMW) operations. Writing to overlapped tracks affects underlying tracks, requiring additional I/O operations to read, back up, and rewrite them, resulting in significant head movement latency. We propose EF-IMR, a new architecture that ensures crash consistency in IMR while minimizing RMW latency and head movement. EF-IMR reduces head movement during RMW operations and decreases redundant RMW operations. Evaluations under real-world, intensive I/O workloads show that EF-IMR reduces RMW latency by 20.11 % and head movement latency by 89.37% compared to existing methods.
Chenlin Ma, Xiaochuan Zheng, Kaoyi Sun, Tianyu Wang 0009, Yi Wang 0003
DATE5
2024 Leanor: A Learning-Based Accelerator for Efficient Approximate Nearest Neighbor Search via Reduced Memory Access
abstract
Approximate Nearest Neighbor Search (ANNS) is a classical problem in data science. ANNS is both computationally-intensive and memory-intensive. As a typical implementation of ANNS, Inverted File with Product Quantization (IVFPQ) has the properties of high precision and rapid processing. However, the traversal of non-nearest neighbor vectors in IVFPQ leads to redundant memory accesses. This significantly impacts retrieval efficiency. A promising approach involves the utilization of learned indexes, leveraging insights from data distribution to optimize search efficiency. Existing learned indexes are primarily customized for low-dimensional data. How to tackle ANNS in high-dimensional vectors is a challenging issue.
Yi Wang 0003, Jianan Yuan, Jiaxian Chen, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001
DAC1
2024 Rapper: A Parameter-Aware Repair-in-Memory Accelerator for Blockchain Storage Platform
abstract
Blockchain storage platforms reward storage nodes for keeping user-uploaded data for a certain amount of time. These storage nodes are unstable and can go online or offline unpredictably at any time, leading to potential data loss. To prevent data loss, blockchain storage platforms adopt erasure codes on user-uploaded encrypted data. Data repair processes will be performed to recover the lost data. However, the data repair processes heavily rely on time-consuming erasure coding algorithms, mainly consisting of vector-matrix multiplications. The emerging processing-in-memory technique can efficiently speed up the processing of vector-matrix multiplications. It can be integrated into blockchain storage platforms to solve the data repair issue. This paper presents Rapper, a parameter-aware repair-inmemory accelerator for blockchain storage platforms. Rapper utilizes the computing power of emerging processing-in-memory architecture so that data repair processes can be processed in a parallel manner and the overall efficiency can be improved significantly. Specifically, at the hardware level, the ReRAM memory is reorganized into our proposed double bank, XRU, XGroup, and ReRAM crossbars structure. At the software level, a parallel decoding/encoding strategy is proposed to fully exploit the internal parallelism of ReRAM. We also propose an adaptive parameter-aware mapping to handle various sizes of stripes. To demonstrate the viability of the proposed technique, a representative blockchain storage project Storj is adopted as the default storage infrastructure. Experimental results show that Rapper can achieve a 1.96 × speedup on average compared to the representative scheme.
Chenlin Ma, Yingping Wang, Fuwen Chen, Jing Liao 0008, Yi Wang 0003, Rui Mao 0001
HPCA5
2024 Boosting Write Performance of KV Stores: An NVM - Enabled Storage Collaboration Approach
abstract
As the most common data structure for key-value stores, LogStructured Merge Tree (LSM-tree) can eliminate random write operations and keep acceptable read performance. However, write stall and write amplification introduced by the leveled compaction of LSM-tree significantly degrade the system performance. The emerging non-volatile memory (NVM) provides byte-addressable access and low-latency data persistence. Integrating DIMM-interface NVM in the design of the LSM-tree can potentially alleviate the write stall and write amplification issue, as the access speed of NVM is several orders of magnitude faster than hard disk drives or flash memory-based solid-state drives. This hybrid storage should be carefully designed, requiring new architectural and key-value structural support. This paper presents ZigZagDB, an NVM-enabled data man-agement scheme for LSM-tree-based key-value stores. ZigZagDB adds additional layers of key-value stores and uses non-volatile memory as the storage media to hold these additional layers of data. The newly designed key-value stores alternately access the data from either SSD or NVM. This ‘ZigZag’ shape of storage collaboration and synchronization can benefit write efficiency and space utilization. By utilizing the NVM with very limited capacity, the redesigned organization of LSM-tree can effectively solve the write stall and write amplification issue. We demonstrate the viability of the proposed ZigZagDB using a set of extensive experiments. Experimental results show that ZigZagDB can significantly reduce the write amplification and boost the throughput in comparison with representative schemes.
Yi Wang 0003, Jiajian He, Kaoyi Sun, Yunhao Dong, Jiaxian Chen, Chenlin Ma, Amelie Chi Zhou, Rui Mao 0001
ICDE1
2024 LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV Table
abstract
Log-structured merge-tree (LSM-tree) is a storage architecture widely used in key-value (KV) stores. To enhance the read efficiency of LSM-tree, recent works utilize the learned index to learn the mapping between keys and locations. However, in existing learned-index-aided KV stores, inefficient design of the learned index and disk access significantly impact the read performance. How to design a learned KV store to improve index efficiency and minimize disk access remains a critical problem. This paper presents LeaderKV, a read-optimized LSM-tree-based KV store. LeaderKV employs decoupled KV tables (DK-Table) and efficient learned indexes for data retrieval. DKTables are storage files in Leader Kvbecause they avoid reading irrelevant data in collaboration with learned indexes during queries. A learned index called Leader is proposed to accelerate data retrieval within DKTable. Leader is composed of precise models and approximate models. A redirect mechanism is designed to reduce the cost of mispredictions in Leader. We integrate DKTable and Leader into LeaderKV and demonstrate its effectiveness using a variety of datasets and workloads. Experimental results show that LeaderKV significantly improves the read performance compared to representative schemes.
Yi Wang 0003, Jianan Yuan, Shangyu Wu, Jiaxian Chen, Chenlin Ma, Jianbin Qin
ICDE1
2024 Tackling Cold Start in Serverless Computing with Multi-Level Container Reuse
abstract
In Serverless Computing, function cold-start is a major issue that causes delay of the system. Various solutions have been proposed to address function cold-start issue, among which keeping containers alive after function completion is an easy and commonly adopted way in real serverless clouds. However, when reusing warm containers for function warm starts, existing systems only match functions to containers with the same configurations. This greatly limits the warm resource utilization. Our analysis of real-world applications reveals that many serverless applications share the same operating system and language frameworks. Thus, we propose multi-level container reuse that tries to reduce the startup latency of functions using "similar" containers to greatly improve warm resource utilization. Due to the complexity of selecting the best container reuse solutions, we designed a Deep Reinforcement Learning (DRL) based scheduler to efficiently and effectively address the problem. Moreover, we released a new serverless benchmark named FStartBench that contains detailed package information for comparing the effectiveness of different function cold-start methods. Experiments based on FStartBench show that, given a warm resource pool with fixed size, our DRL-based scheduler can achieve up to 53% reduction on the average function startup latency compared to state-of-the-art solutions.
Amelie Chi Zhou, Rongzheng Huang, Zhoubin Ke, Yusen Li, Yi Wang 0003, Rui Mao 0001
IPDPS5
2024 Low-Latency Video Conferencing via Optimized Packet Routing and Reordering
abstract
In the face of rising global demand for video meetings, managing traffic across geographically distributed (geo-distributed) data centers presents a significant challenge due to the dynamic and limited nature of inter-DC network performance. Facing these issues, this paper introduces two novel techniques, VCRoute and WMJitter, to optimize the performance of geo-distributed video conferencing systems. VCRoute is a routing method designed for audio data packets of video conferences. It treats the routing problem as a Multi-Armed Bandit issue, and utilizes a tailored Thompson Sampling algorithm for resolution. Unlike traditional approaches, VCRoute considers transmitting latency and its variance simultaneously by using Thompson Sampling algorithm, which leads to effective end-to-end latency optimization. In conjunction with VCRoute, we present WMJitter, a watermark-based mechanism for managing network jitter, which can further reduce the end-to-end delay and keep an improved balance between latency and loss rate. Evaluations based on real geo-distributed network performance demonstrate the effectiveness and scalability of VCRoute and WMJitter, offering robust solutions for optimizing video conferencing systems in geo-distributed settings.
Sitian Chen, Amelie Chi Zhou, Shuhao Zhang 0001, Yi Wang 0003, Rui Mao 0001
IWQoS5
2024 NICE: A Nonintrusive In-Storage-Computing Framework for Embedded Applications
abstract
Embedded machine learning applications face challenges related to massive data movement and high computational intensity, exacerbated by the limited performance of mobile devices. Computational storage devices (CSDs) pose huge potential for accelerating both data-intensive and computation-intensive embedded machine learning tasks by effectively reducing data movement and leveraging built-in accelerators. However, existing in-storage-computing (ISC) frameworks either require invasive customization of existing host driver layers or necessitate complex device firmware modifications, hindering the widespread deployment of CSDs. In addition, the lack of file semantics and the constrained internal resources within CSD implicitly compromise system performance and impact normal read/write performance. In this article, we aim to provide a nonintrusive in-storage-computing framework for embedded applications, named NICE. This framework includes an easy-to-use ISC programming interface that bypasses the kernel stack and requires no modification to the host NVMe driver, which is achieved through a novel hyper-addressing-based programming library and a file-aware page data layout within the CSD. In addition, we incorporate a lightweight kernel with coroutine-based command scheduling and several FPGA-based accelerators within the storage device firmware to enhance the performance of embedded machine learning applications while ensuring that the normal I/O performance remains unaffected. NICE is implemented on real CSD hardware integrated with ARM and FPGA. Experimental results demonstrate that our NICE framework can achieve an average latency performance improvement of$43.5\times $($9.32\times $) compared to CPU-(GPU-) based embedded machine learning solutions using the state-of-the-art NVIDIA Jetson NX platform, with$27.5\times $($4.3\times $) higher energy efficiency. NICE also has$34.2\times $less software and I/O performance overheads than state-of-the-art ISC frameworks.
Tianyu Wang 0009, Yongbiao Zhu, Shaoqi Li, Jin Xue, Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 A Semantic-Integrated LSM-Tree-Based Key-Value Storage Engine for Blockchain Systems
abstract
Blockchain systems play an important role in distributed ledgers, database systems, etc. As more and more blocks are mined, the storage burden of blockchain system is significantly increased. The current blockchain system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only aggravates the write amplification effect of the storage engine, but also increases the redundancy of data query steps, resulting in the performance bottleneck of blockchain system. In this paper, we propose a semantic-integrated LSM-tree based Key-Value storage engine for blockchain systems, called Block-LSM, which significantly improves the data synchronization and data query efficiency of blockchain system. Specifically, we first design a shared prefix scheme to transform blockchain data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of blockchain data, and implement memory buffer space management strategy to further improve memory efficiency. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. We also reduce step redundancy in transaction queries by modifying the body data storage format. Finally, we implement Block-LSM in a real blockchain environment and conduct a series of comparative experiments with the typical blockchain system Ethereum. The evaluation results show that Block-LSM significantly reduces up to 7.56× storage write amplification and increases throughput by 8.64× compared with the original Ethereum design. In terms of data lookups (i.e. transaction and account lookup), Block-LSM improves the throughput by 50% compared to the original Ethereum design.
Yuhao Zhang 0006, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao, Bingzhe Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2023 Lift: Exploiting Hybrid Stacked Memory for Energy-Efficient Processing of Graph Convolutional Networks
abstract
Graph Convolutional Networks (GCNs) are powerful learning approaches for graph-structured data. GCNs are both computing- and memory-intensive. The emerging 3D-stacked computation-in-memory (CIM) architecture provides a promising solution to process GCNs efficiently. The CIM architecture can provide near-data computing, thereby reducing data movement between computing logic and memory. However, previous works do not fully exploit the CIM architecture in both dataflow and mapping, leading to significant energy consumption.This paper presents Lift, an energy-efficient GCN accelerator based on 3D CIM architecture using software and hardware co-design. At the hardware level, Lift introduces a hybrid architecture to process vertices with different characteristics. Lift adopts near-bank processing units with a push-based dataflow to process vertices with strong re-usability. A dedicated unit is introduced to reduce massive data movement caused by high-degree vertices. At the software level, Lift adopts a hybrid mapping to further exploit data locality and fully utilize the hybrid computing resources. The experimental results show that the proposed scheme can significantly reduce data movement and energy consumption compared with representative schemes.
Jiaxian Chen, Zhaoyu Zhong, Kaoyi Sun, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
DAC6
2023 Meta-Block: Exploiting Cross-Layer and Direct Storage Access for Decentralized Blockchain Storage Systems
abstract
Decentralized storage systems such as blockchain storage applications adopt the distributed storage technology and use distributed storage nodes to store the persistent data. For each off-chain storage node, key-value (KV) stores are normally used to manage data. As the most common data structure for KV store, Log Structured Merge Tree (LSM-Tree) eliminates random write operations and keeps acceptable read performance. Although LSM-Tree-based decentralized storage system can provide a secure and reliable storage platform, the unique feature of blockchain applications is not fully exploited. In blockchain storage applications, the generation of keys for KV stores is based on the encrypted data, and the key determines the allocation of data. Since the granularity for a read/write request at the level of blockchain storage platform is much smaller than that at the level of LSM-Tree or flash memory, a physical block in a solid-state drive (SSD) could be filled with data from different system users. This mixture of workloads will lead to the inefficient usage of physical spaces in the SSD and cause extra compaction operations for LSM-Tree. This paper presentsMeta-Block, a cross-layer and efficient storage management strategy for decentralized blockchain storage applications. Meta-Block utilizes rich functionalities provided by the system infrastructure of open-channel SSD to provide direct storage accesses for the off-chain storage node. The objective is to capture the features of blockchain storage applications and reduce unnecessary read and write operations across different storage layers. As a cross-layer design, Meta-Block redesigns the organization of LSM-Tree, which can effectively reduce the write amplification. We also design a data prefetching strategy to speed up the indexing and enable direct storage access. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that Meta-Block can effectively reduce the write amplification and extend the lifetime of SSDs in comparison with representative schemes.
Yi Wang 0003, Jing Liao 0008, Jing Yang 0018, Zhengda Li, Chenlin Ma, Rui Mao 0001
IEEE Trans. Computers1
2023 Tidal-Tree-Mem: Toward Read-Intensive Key-Value Stores With Tidal Structure Based on LSM-Tree
abstract
The log-structured merge-tree (LSM-tree)-based key-value store has been widely adopted by many large-scale data storage applications for its excellent write performance. However, such write performance gains mainly come from scarifying read performance due to the leveled and log-structured intrinsic characteristics of the LSM-tree. Therefore, the critical challenge of the existing LSM-tree is how to improve the read efficiency by reducing read amplification. This article for the first time proposes Tidal-tree-Mem, a novel data structure where data flow inside the LSM-tree-like Tidal waves. First, a floating strategy is proposed to allow frequently accessed files at the bottom of the LSM-tree to move to higher positions, reducing read amplification. Second, a stretching strategy is proposed to vary the shape of the LSM-tree to adapt to workloads with different characteristics. To evaluate the performance of Tidal-tree-Mem, we conduct a series of experiments using standard benchmarks from YCSB. The experimental results show that Tidal-tree-Mem can effectively reduce read amplification and the overall latency by over 71.94% and 47.34%, respectively, compared with representative schemes.
Chenlin Ma, Shangyu Wu, Yi Wang 0003, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 Work-in-Progress: Lark: A Learned Secondary Index Toward LSM-tree for Resource-Constrained Embedded Storage Systems
abstract
LSM-tree-based key-value stores are popular in embedded storage systems. With the growing demands of data analysis, the secondary index is created to support non-primary-key lookups. However, the lookup efficiency and space consumption of secondary index remain for further optimization. Inspired by the learned index, this paper presents Lark, a learned secondary index toward LSM-tree for resource-constrained embedded storage systems. Lark employs machine learning to speed up the non-primary-key queries and compress secondary indexes. Our preliminary evaluations show that, in comparison with traditional secondary index schemes, Lark achieves better lookup performance with less space consumption.
Jianan Yuan, Shangyu Wu, Yiquan Lin, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
CODES+ISSS8
2022 MU-RMW: Minimizing Unnecessary RMW Operations in the Embedded Flash with SMR Disk
abstract
Emerging Shingled Magnetic Recording (SMR) Disk can improve the storage capacity significantly by overlapping multiple tracks with the shingled direction. However, the shingled-like structure leads to severe write amplification caused by RMW operations inner SMR disks. As the mainstream solid-state storage technology, NAND flash has the advantages of tiny size, cost-effective, high performance, making it suitable and promising to be incorporated into SMR disks to boost the system performance. In this hybrid embedded storage system (i.e., the Embedded Flash with SMR disk (EF-SMR) system), we observe that physical flash blocks can contain a mixture of data associated with different SMR data bands; when garbage collecting such flash blocks, multiple RMW operations are triggered to rewrite the involved SMR bands and the performance is further exacerbated. Therefore, in this paper, we for the first time present MU-RMW to guarantee data from different SMR bands will not be mixed up within the flash blocks with an aim at minimizing unnecessary RMW operations. The effectiveness of MU-RMW was evaluated with realistic and intensive I/O workloads and the results are encouraging.
Chenlin Ma, Zhuokai Zhou, Yingping Wang, Yi Wang 0003, Rui Mao 0001
DATE4
2022 GCIM: Toward Efficient Processing of Graph Convolutional Networks in 3D-Stacked Memory
abstract
Graph convolutional networks (GCNs) have become a powerful deep learning approach for graph-structured data. Different from traditional neural networks such as convolutional neural networks, GCNs handle irregular input graph data, and GCNs are both computation-bound and memory-bound. How to efficiently utilize the underlying computation and memory resource becomes a critical issue. The emerging 3D-stacked computation-in-memory (CIM) architecture can reduce the data movement between computing logic and memory, thereby presenting a promising solution for the processing of GCNs. An unsolved key challenge is how to allocate GCNs to take advantage of fast near-data processing of the 3D-stacked CIM architecture. This article presents GCIM, a software–hardware co-design approach to exploit the efficient processing of GCNs on the CIM architecture. At the level of hardware design, GCIM integrates lightweight computing units near memory banks to fully exploit bank-level bandwidth and parallelism. At the level of software design, a locality-aware data mapping algorithm is proposed to partition the input graph and achieve workload balancing. GCIM is evaluated through a set of representative GCN models and standard graph datasets. The experimental results show that GCIM can significantly reduce the processing latency and data movement overhead compared with representative schemes.
Jiaxian Chen, Yiquan Lin, Kaoyi Sun, Jiexin Chen, Chenlin Ma, Rui Mao 0001, Yi Wang 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2022 Rebirth-FTL: Lifetime Optimization via Approximate Storage for NAND Flash Memory
abstract
The lifetime of NAND flash cells significantly degrades with feature-size reductions and multilevel cell technology. On the other hand, we have more and more approximate data, such as images and videos that are more error tolerant than regular data like text. In this article, we propose Rebirth-FTL, which reuses faulty blocks that contain uncorrectable errors to store approximate data for lifetime optimization. Rebirth-FTL effectively manages two spaces, namely, the approximate space and the normal space, with an efficient address translator, a coordinated garbage collection, and a differential wear leveler. In addition, we develop an migration times restriction (MTR) policy to restrict the movement of the approximate data in the approximate space. We also develop a scheme to pass approximate information from userland to kernel space in Linux. Finally, a lifetime model is presented for lifetime analysis. Our experimental results show that Rebirth-FTL can extend the lifetime by 41.63% on average.
Chenlin Ma, Zhuokai Zhou, Zhaoyan Shen, Yi Wang 0003, Renhai Chen, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2022 MAID-Q: Minimizing Tail Latency in Embedded Flash With SMR Disk via -Learning Model
abstract
As the mainstream solid-state storage technology, NAND flash has the advantages of tiny size, cost-effective, and high performance, which make it a promising candidate to be embedded into the shingled magnetic recording (SMR) disk to build a faster, denser, and cheaper storage system. However, such an embedded flash with SMR (EF-SMR) disk system suffers from lengthy tail-latency due to “reclamation issues” in both the NAND flash and the SMR disk. Our preliminary observations reveal that tremendous idle time intervals exist in real-world scenarios, and few prior works have focused on addressing the tail-latency issue in the EF-SMR disk. In this article, we propose a novel method termed MAID-Q to fully exploit the idle time intervals to minimize the lengthy tail-latency of the EF-SMR disk based on a lightweight reinforcement learning model (i.e., the$Q$-learning model). In addition, fine-grained block-level space management and a parallel reclamation strategy are proposed to improve the reclamation efficiency and hide the reclamation overheads. The effectiveness of our proposed design was evaluated with realistic I/O traces, and the results show that the proposed design can remedy the tail-latency by 88.31% and improve the overall performance by 79.03%.
Chenlin Ma, Zhuokai Zhou, Yingping Wang, Yi Wang 0003, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2022 PVSensing: A Process-Variation-Aware Space Allocation Strategy for 3D NAND Flash Memory
abstract
Three-dimensional (3D) flash memory is an emerging memory technology that enables a number of improvements to conventional planar NAND flash memory, including larger capacity, less program disturb, and lower access latency. Despite these advantages, 3D flash memory brings a number of new challenges. First, in 3D flash memory, NAND strings punch through multiple stacked layers to form the 3D infrastructure. Current etching process is unable to manufacture perfect channels with identical feature size. Second, with more stacked layers, the cell current in 3D flash memory is only 20% compared to planar flash memory, making it difficult to give a reliable sensing margin. These issues are affected by process variation, and they pose threats to the integrity of data stored in 3D flash memory. This article presentPVSensing, a process-variation-aware space allocation strategy for open-channel SSD with 3D charge-trap flash memory. PVSensing is a novel hardware and file system interface that can transparently allocate physical space in the presence of process variation. PVSensing utilizes the rich functionalities provided by the system infrastructure of open-channel SSD to reduce the uncorrectable bit errors. Three reliability enhancement strategies (i.e., the adaptive creation of fault cubes, the physical block mining, and the live migration of write requests) are proposed. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that PVSensing can effectively reduce uncorrectable bit errors, and improve the reliability of critical data with negligible extra erase operations in comparison with representative schemes.
Yi Wang 0003, Jiangfan Huang, Rui Mao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 FarSpot: Optimizing Monetary Cost for HPC Applications in the Cloud Spot Market
abstract
Recently, we have witnessed many HPC applications developed and hosted in the cloud, which can benefit from the elastic and diversified resources on the cloud, while on the other hand confronting high costs for executing the long-running HPC applications. Although public clouds such as Amazon EC2 offer spot instances with dynamic and usually low prices compared to on-demand ones, the spot prices can vary significantly and sometimes can even be more expensive than on-demand prices of the same type. Previous work on reducing the monetary cost for HPC applications using spot instances focused on designing fault tolerance techniques or selecting appropriate instance types/bid prices to make good usage of the low spot prices. However, with the recent update of spot pricing model on Amazon EC2, these work may become either inefficient or invalid. In this paper, we present FarSpot which is an optimization framework for HPC applications in the latest cloud spot market with the goal of minimizing application cost while ensuring performance constraints. FarSpot provides accurate long-term price prediction for a wide range of spot instance types using ensemble-based learning method. It further incorporates a cost-aware deadline assignment algorithm to distribute application deadline to each task according to spot price changes. With the assigned subdeadline of each task, FarSpot dynamically migrates tasks among spot instances to reduce execution cost. Evaluation results using real HPC benchmark show that 1) the prediction error of FarSpot is very low (below 3%), 2) FarSpot reduced the monetary cost by 32% on average compared to state-of-the-art algorithms, and 3) FarSpot satisfies the user-specified deadline constraints at all time.
Amelie Chi Zhou, Jianming Lao, Zhoubin Ke, Yi Wang 0003, Rui Mao 0001
IEEE Trans. Parallel Distributed Syst.4
2021 LolliRAM: A Cross-Layer Design to Exploit Data Locality in Oblivious RAM
abstract
Oblivious RAM (ORAM) conceals memory access pattern by translating a single read/write operation into the accesses to a set of randomized locations. The obliviousness is achieved by adding redundancy to the memory system, which comes at the expense of increased performance overhead. In memory systems, locality has always been a critical factor, as accessing the data with temporal or spatial locality would result in a performance gain. Although the two design considerations of obliviousness and locality may seem contradictory at first glance, combining them in a unified design can potentially hide long memory access latency in ORAM without sacrificing provable data security.This paper presents LolliRAM, a cross-layer design to exploit data locality in Oblivious RAM. LolliRAM optimizes ORAM system through two different layers: (i) the data structure in ORAM, and (ii) the fast and secure cache in ORAM controller. The reuse of redundant memory accesses at the layer of data structure in ORAM can effectively reduce memory footprints. Both temporal and spatial locality can be exploited through the elastic grouping of blocks and the optimization at the layer of ORAM controller. We conduct a set of experiments using realistic workloads that are generated from standard benchmarks. Experimental results show that LolliRAM can reduce access latency by 71.57% on average (up to 81.50%) with negligible space overhead in comparison with representative schemes.
Yi Wang 0003, Weixuan 'Vincent' Chen, Xianhua Wang, Rui Mao 0001
DAC1
2021 Block-LSM: An Ether-aware Block-ordered LSM-tree based Key-Value Storage Engine
abstract
Ethereum as one of the largest blockchain systems plays an important role in the distributed ledger, database systems, etc. As more and more blocks are mined, the storage burden of Ethereum is significantly increased. The current Ethereum system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only exacerbates the write amplification effect of the storage engine but also hurts the performance of Ethereum. In this paper, we proposed a new Ethereum-aware storage model called Block-LSM, which significantly improves the data synchronization of the Ethereum system. Specifically, we first design a shared prefix scheme to transform Ethereum data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of Ethereum data. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. Finally, we implement Block-LSM in the real Ethereum environment and conduct a series of experiments. The evaluation results show that Block-LSM significantly reduces up to 3.7× storage write amplification and increases throughput by 3× compared with the original Ethereum design.
Bingzhe Li, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao
ICCD6
2021 Towards efficient allocation of graph convolutional networks on hybrid computation-in-memory architecture
Jiaxian Chen, Guanquan Lin, Jiexin Chen, Yi Wang 0003
Sci. China Inf. Sci.4
2021 Tiler: An Autonomous Region-Based Scheme for SMR Storage
abstract
Shingled Magnetic Recording (SMR) Disks are adopted as a high-density, non-volatile media that significantly precedes conventional disks in both the storage capacity and cost. However, inefficient read-modify-writes (RMWs) greatly challenge the management of SMR disks. This article for the first time presents an approach called Tiler to manage SMR disks by dividing the physical space into small autonomous regions (ARs). Each AR can manage its space allocation, address mapping, and cleaning independently. By managing these ARs in a log-structured way, RMWs can be avoided; besides, ARs can also help update data when the adjacent tracks contain no valid data. Tiler is capable of partitioning a large-scale cleaning into self-contained-small-scale cleaning and thus, the data that need to be relocated are limited inside independent ARs, which further minimizes the performance overhead. Our experimental results show that Tiler can shorten the overall system response time by 50.21 percent and reduce the cleaning time by 90.24 percent on average.
Chenlin Ma, Zhaoyan Shen, Yi Wang 0003, Renhai Chen, Zili Shao
IEEE Trans. Computers4
2020 Towards Read-Intensive Key-Value Stores with Tidal Structure Based on LSM-Tree
abstract
Key-value store has played a critical role in many large-scale data storage applications. The log-structured merge-tree (LSM-tree) based key-value store achieves excellent performance on write-intensive workloads which is mainly benefited from the mechanism of converting a batch of random writes into sequential writes. However, LSM-tree doesn't improve a lot in read-intensive workloads which takes a higher latency. The main reason lies in the hierarchical search mechanism in LSM-tree structure. The key challenge is how to propose new strategies based on the existing LSM-tree structure to improve read efficiency and reduce read amplifications.This paper proposes Tidal-tree, a novel data structure where data flows inside LSM-tree like Tidal waves. Tidal-tree targets at improving read efficiency in read-intensive workloads. Tidal-tree allows frequently accessed files at the bottom of LSM-tree to move to higher positions, thereby reducing read latency. Tidal-tree also makes LSM-tree into a variable shape to cater for different characteristic workloads. To evaluate the performance of Tidal-tree, we conduct a series of experiments using standard benchmarks from YCSB. The experimental results show that Tidal-tree can significantly improve read efficiency and reduce read amplifications compared with representative schemes.
Yi Wang 0003, Shangyu Wu, Rui Mao 0001
ASP-DAC1
2020 KFR: Optimal Cache Management with K-Framed Reclamation for Drive-Managed SMR Disks
abstract
Shingled Magnetic Recording (SMR) disks have been proposed as a promising solution to satisfy the increasing capacity need in the big data era. Drive-Managed SMR (DM-SMR) disk which acts as a traditional block device is favored for providing high compatibility. However, DM-SMR disks suffer from high performance recovery time (PRT) due to the "SMR space reclamation" issue. This paper proposes an optimal cache management named K-Framed Reclamation (KFR) to minimize PRT within the DM-SMR disk. The effectiveness of our proposed design was evaluated with realistic and intensive I/O workloads and the results are encouraging.
Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Zili Shao
DAC2
2020 Formalization of continuous Fourier transform in verifying applications for dependable cyber-physical systems
Jie Zhang 0074, Zhi-Ping Shi 0002, Yi Wang 0003, Yongdong Li
J. Syst. Archit.4
2020 Temperature-Aware Persistent Data Management for LSM-Tree on 3-D NAND Flash Memory
abstract
Key-value (KV) store has been widely deployed in both embedded systems and enterprise systems. Most KV stores today use log structured merge tree (LSM-Tree), as LSM-Tree can eliminate random write operations to the secondary storage and maintain acceptable read performance. LSM-Tree is originally designed for the secondary storage device with hard disk drives. As the emerging storage media, three-dimensional (3-D) flash memory has become the mainstream technology to replace hard disk drives. Different from hard disk drives and the conventional planar flash memory, 3-D flash memory is vulnerable to temperature. High temperature will introduce both charge loss and retention degradation. Since LSM-Tree transfers random write operations to the sequential ones, the access to consecutive physical address in flash memory will cause the temperature issue. This will affect the integrity of data stored in 3-D flash memory. This article presents TLSM, a temperature-aware persistent data management scheme for LSM-Tree-based KV store on 3-D NAND flash memory. TLSM offers both application-level LSM-Tree optimization and firmware-level address management to allocate persistent data to 3-D flash. At the application-level, TLSM presents a novel temperature-aware LSM data structure to reduce the amount of data issued from LSM-Tree to 3-D flash memory. At the firmware-level, TLSM reallocates the data to physical blocks with relatively low temperature. This cross-layer optimization can effectively handle the temperature issue to ensure the data integrity of LSM-Tree in 3-D flash memory. We demonstrate the viability of the proposed scheme using a set of standard benchmarks. Our extensive evaluations show that, TLSM can significantly enhance the data integrity and reduce write amplifications compared to representative schemes.
Yi Wang 0003, Jiali Tan, Rui Mao 0001, Tao Li 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 MNFTL: An Efficient Flash Translation Layer for MLC NAND Flash Memory
abstract
The write constraints of Multi-Level Cell (MLC) NAND flash memory make most of the existing flash translation layer (FTL) schemes inefficient or inapplicable. In this article, we solve several fundamental problems in the design of MLC flash translation layer. The objective is to reduce the garbage collection overhead to reduce the average system response time. We make the key observation that the valid pages copy is the essential garbage collection overhead. Based on this observation, we propose two approaches, namely, concentrated mapping and postponed reclamation, to effectively reduce the valid pages copy. Besides, we propose a progressive garbage collection that can well utilize the system idle time to reclaim more spaces. We conduct a series of experiments on an embedded developing board with a set of benchmarks. The experimental results show that our scheme can achieve a significant reduction in the average system response time compared with the previous work.
Chenlin Ma, Yi Wang 0003, Zhaoyan Shen, Renhai Chen, Zili Shao
ACM Trans. Design Autom. Electr. Syst.2
2019 PATCH: Process-Variation-Resilient Space Allocation for Open-Channel SSD with 3D Flash
abstract
Advanced three-dimensional (3D) flash memory adopts charge-trap technology that can effectively improve the hit density and reduce the coupling effect. Despite these advantages, 3D charge-trap flash brings a number of new challenges. First, current etching process is unable to manufacture perfect channels with identical feature size. Second, the cell current in 3D charge-trap flash is only 20% compared to planar flash memory, making it difficult to give a reliable sensing margin. These issues are affected by process variation, and they pose threats to the integrity of data stored in 3D charge-trap flash. This paper presents PATCH, a process-variation-resilient space allocation scheme for open-channel SSD with 3D charge-trap flash memory. PATCH is a novel hardware and file system interface that can transparently allocate physical space in the presence of process variation. PATCH utilizes the rich functionalities provided by the system infrastructure of open-channel SSD to reduce the uncorrectable bit errors. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that PATCH can effectively enhance the reliability with negligible extra erase operations in comparison with representative schemes.
Yi Wang 0003, Amelie Chi Zhou, Rui Mao 0001, Tao Li 0006
DATE2
2019 Towards Cross-Platform Inference on Edge Devices with Emerging Neuromorphic Architecture
abstract
Deep convolutional neural networks have become the mainstream solution for many artificial intelligence applications. However, they are still rarely deployed on mobile or edge devices due to the cost of a substantial amount of data movement among limited resources. The emerging processing-inmemory neuromorphic architecture offers a promising direction to accelerate the inference process. The key issue becomes how to effectively allocate the processing of inference between computing and storage resources on an edge device.This paper presents Mobile-I, a resource allocation scheme to accelerate the Inference process on Mobile or edge devices. Mobile-I targets at the emerging 3D neuromorphic architecture to reduce the processing latency among computing resources and fully utilize the limited on-chip storage resources. We formulate the target problem as a resource allocation problem and use a software-based solution to offer the cross-platform deployment across multiple mobile or edge devices. We conduct a set of experiments using realistic workloads that are generated from Intel Movidius neural compute stick. Experimental results show that Mobile-I can effectively reduce the processing latency and improve the utilization of computing resources with negligible overhead in comparison with representative schemes.
Shangyu Wu, Yi Wang 0003, Amelie Chi Zhou, Rui Mao 0001, Zili Shao, Tao Li 0006
DATE2
2019 A Thermal-Aware Physical Space Reallocation for Open-Channel SSD With 3-D Flash Memory
abstract
3-D flash memory faces a number of challenges, including thermal issues and process variation. The high temperature will cause charge loss and lead to the fluctuation of threshold voltage. To address the thermal issue of 3-D flash memory, this paper presents ThermAlloc, a novel thermal-aware physical space allocation strategy for open-channel solid-state drive with 3-D charge trapping flash memory. ThermAlloc permutes the allocation of physical blocks. Consecutively accessed logical blocks are distributed to different physical locations in order to prevent the accumulation of hotspots. The objective is to postpone garbage collection operations and keep the distribution of block temperature well under control. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that ThermAlloc can reduce the peak temperature by 26.94% with 3.15% extra worst-case response time in comparison with the baseline scheme.
Yi Wang 0003, Mingxu Zhang, Tao Li 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 DCR: Deterministic Crash Recovery for NAND Flash Storage Systems
abstract
NAND flash memory has been widely adopted as the storage medium. As power failure may occur at any time and result in data loss, crash recovery becomes vitally important in NAND flash memory storage systems. Since flash translation layers (FTLs) are used to manage flash memory, the crash recovery problem in NAND flash is how to efficiently and effectively recover FTL metadata with consistency after system crash. In this paper, we present deterministic crash recovery (DCR) that adopts a deterministic approach for crash recovery in NAND flash storage systems. The basic idea is to exploit the determinism of FTLs and reproduce events that happened between the last checkpoint and the crash point during crash recovery. Different from existing approaches that need to scan the whole flash memory chip, DCR can recover the system more efficiently by only checking a limited number of blocks based on deterministic FTL operations. We have implemented DCR in an FTL and compared it with the representative version-based and power loss recovery schemes based on an ARM-based embedded system. Experimental results show that DCR can greatly reduce the recovery time and guarantee the consistency of FTL metadata after recovery.
Renhai Chen, Yi Wang 0003, Zhaoyan Shen, Duo Liu 0002, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 Alleviating Hot Data Write Back Effect for Shingled Magnetic Recording Storage Systems
abstract
Shingled magnetic recording (SMR) is a novel technology that can significantly increase the disk capacity and reduce the cost. However, this new technology leads to write constraint that prohibits random writes to SMR disks. In order to solve this, current solutions utilize a persistent cache (PC) to change random writes to sequential ones. By doing this, nevertheless, the hot-data write-back effect will be triggered which inevitably incurs extra read-modify-write (RMW) operations. This paper presents a cache management scheme calleddual-bufferto manage the overall SMR space, by which the PC is partitioned into the persistent buffer and the filter buffer. It leaves hot data in the filter buffer and moves cold data back to the SMR disk so that the hot data write-back effect can be alleviated. Since the PC may contain a large amount of valid data,dual-bufferalso presents a prediction-based dynamic configuration strategy so hot data can be cached as much as possible. We conducted a series of experiments with synthetic traces. Experimental results show that ourdual-bufferscheme can shorten the average response time by 51.66% on average and reduce the total number of RMW operations by 98.76% on average compared to the previous work.
Chenlin Ma, Zhaoyan Shen, Yi Wang 0003, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 A Temperature-Aware Reliability Enhancement Strategy for 3-D Charge-Trap Flash Memory
abstract
Compared to the conventional planar flash memory, advanced 3-D flash memory adopts charge-trap technology that can significantly enhance cell density and storage capacity. Despite these advantages, 3-D charge-trap flash memory brings several new challenges. First, charge-trap flash is sensitive to temperature. Recent studies demonstrate that, the high temperature will incur both charge loss and retention degradation. This issue does not happen in 2-D flash memory which adopts floating gate technology. Second, current 3-D charge-trap flash integrates the extra large capacity physical block, and each block contains over 1024 physical pages. The large-capacity block infrastructure will cause extra garbage collection overhead, which makes the thermal issue more complicated. This paper presents TempCure, a temperature-aware reliability enhancement strategy for 3-D charge-trap flash memory. TempCure is a novel hardware and file system interface that can transparently allocate physical space based on the temperature status. TempCure adopts two reliability enhancement strategies, temperature mining and block allotment, to prevent the generation of hotspots and enhance the data integrity of 3-D flash memory. We conduct a set of experiments using standard benchmarks. Experimental results show that TempCure can effectively reduce the peak temperature and block erase counts with negligible timing overhead in comparison with representative schemes.
Yi Wang 0003, Jiangfan Huang, Jing Yang 0018, Tao Li 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 Exploiting Parallelism for CNN Applications on 3D Stacked Processing-In-Memory Architecture
abstract
Deep convolutional neural networks (CNNs) are widely adopted in intelligent systems with unprecedented accuracy but at the cost of a substantial amount of data movement. Although the emerging processing-in-memory (PIM) architecture seeks to minimize data movement by placing memory near processing elements, memory is still the major bottleneck in the entire system. The selection of hyper-parameters in the training of CNN applications requires over hundreds of kilobytes cache capacity for concurrent processing of convolutions. How to jointly explore the computation capability of the PIM architecture and the highly parallel property of neural networks remains a critical issue. This paper presents Para-Net, that exploits Parallelism for deterministic convolutional neural Networks on the PIM architecture. Para- Net achieves data-level parallelism for convolutions by fully utilizing the on-chip processing engine (PE) in PIM. The objective is to capture the characteristics of neural networks and present a hardware-independent design to jointly optimize the scheduling of both intermediate results and computation tasks. We formulate this data allocation problem as a dynamic programming model and obtain an optimal solution. To demonstrate the viability of the proposed Para-Net, we conduct a set of experiments using a variety of realistic CNN applications. The graph abstractions are obtained from deep learning framework Caffe. Experimental results show that Para-Net can significantly reduce processing time and improve cache efficiency compared to representative schemes.
Yi Wang 0003, Weixuan 'Vincent' Chen, Jing Yang 0018, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2018 A Reliable Video Storage Architecture in Hybrid SLC/MLC Nand Flash
abstract
In this paper, we propose a reliable video storage architecture in hybrid SLC/MLC storage systems. In this architecture, the video stream is reconstructed as the key cluster and the non-key cluster according to the importance of video restoration. The key cluster is stored in SLC blocks to ensure the reliability, while the non-key cluster is stored in MLC blocks to utilize its capacity. The experimental results show that the proposed scheme can significantly improve the storage reliability of the video files and prolong the lifetime of the storage system 18x.
Yimei Kang, Zili Shao, Renhai Chen, Yi Wang 0003
ICASSP5
2018 A reliability enhanced video storage architecture in hybrid SLC/MLC NAND flash memory
Yimei Kang, Zili Shao, Renhai Chen, Yi Wang 0003
J. Syst. Archit.5
2018 Towards Memory-Efficient Allocation of CNNs on Processing-in-Memory Architecture
abstract
Convolutional neural networks (CNNs) have been successfully applied in artificial intelligent systems to perform sensory processing, sequence learning, and image processing. In contrast to conventional computing-centric applications, CNNs are known to be both computationally and memory intensive. The computational and memory resources of CNN applications are mixed together in the network weights. This incurs a significant amount of data movement, especially for high-dimensional convolutions. The emerging Processing-in-Memory (PIM) alleviates this memory bottleneck by integrating both processing elements and memory into a 3D-stacked architecture. Although this architecture can offer fast near-data processing to reduce data movement, memory is still a limiting factor of the entire system. We observe that an unsolved key challenge is how to efficiently allocate convolutions to 3D-stacked PIM to combine the advantages of both neural and computational processing. This paper presents MemoNet, a memory-efficient data allocation strategy for convolutional neural networks on 3D PIM architecture. MemoNet offers fine-grained parallelism that can fully exploit the computational power of PIM architecture. The objective is to capture the characteristics of neural network applications and perfectly match the underlining hardware resources provided by PIM, resulting in a hardware-independent design to transparently allocate data. We formulate the target problem as a dynamic programming model and present an optimal solution. To demonstrate the viability of the proposed MemoNet, we conduct a set of experiments using a variety of realistic convolutional neural network applications. The extensive evaluations show that, MemoNet can significantly improve the performance and the cache utilization compared to representative schemes.
Yi Wang 0003, Weixuan 'Vincent' Chen, Jing Yang 0018, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2017 Temperature-aware data allocation strategy for 3D charge-trap flash memory
abstract
Three-dimensional (3D) flash memory is emerging as an attractive solution to overcome the scaling bottleneck in sub-20 nanometer design. Compared to the conventional planar flash memory, current 3D flash memory adopts charge-trap technology that can significantly enhance cell density and storage capacity. Despite these advantages, novel material and fabricate process in charge-trap flash memory bring new challenges. Recent studies demonstrate that charge-trap flash is sensitive to thermal yield. This issue does not happen in two-dimensional flash memory which adopts floating gate technology. For 3D flash memory with charge-trap technology, the high temperature will incur both charge loss and retention degradation. The large capacity data block in 3D flash memory also causes extra garbage collection overhead, which makes the temperature issue even worse. This paper presents TempLoad, a temperature-aware data allocation strategy for 3D charge-trap flash memory. TempLoad is a novel hardware and file system interface that can transparently allocate physical space based on the temperature status. TempLoad adopts several address mapping strategies to fully utilize the storage capacity and reduce the garbage collection overhead. The objective is to prevent the generation of hotspots and enhance the data integrity of 3D flash memory. Experimental results show that the proposed technique can reduce the peak temperature by 28.49% and reduce uncorrectable page errors by 83.71% with negligible timing overhead in comparison with previous work.
Yi Wang 0003, Mingxu Zhang, Jing Yang 0018
ASP-DAC1
2017 Exploiting Parallelism for Convolutional Connections in Processing-In-Memory Architecture
abstract
Deep convolutional neural networks (CNNs) are widely adopted in intelligent systems with unprecedented accuracy but at the cost of a substantial amount of data movement. Although recent development in processing-in-memory (PIM) architecture seeks to minimize data movement by computing the data at the dedicated nonvolatile device, how to jointly explore the computation capability of PIM and utilize the highly parallel property of neural network remains a critical issue.
Yi Wang 0003, Mingxu Zhang, Jing Yang 0018
DAC1
2017 Online Robust Image Alignment via Subspace Learning from Gradient Orientations
abstract
Robust and efficient image alignment remains a challenging task, due to the massiveness of images, great illumination variations between images, partial occlusion and corruption. To address these challenges, we propose an online image alignment method via subspace learning from image gradient orientations (IGO). The proposed method integrates the subspace learning, transformed IGO reconstruction and image alignment into a unified online framework, which is robust for aligning images with severe intensity distortions. Our method is motivated by principal component analysis (PCA) from gradient orientations provides more reliable low-dimensional subspace than that from pixel intensities. Instead of processing in the intensity domain like conventional methods, we seek alignment in the IGO domain such that the aligned IGO of the newly arrived image can be decomposed as the sum of a sparse error and a linear composition of the IGO-PCA basis learned from previously well-aligned ones. The optimization problem is accomplished by an iterative linearization that minimizes the L1-norm of the sparse error. Furthermore, the IGO-PCA basis is adaptively updated based on incremental thin singular value decomposition (SVD) which takes the shift of IGO mean into consideration. The efficacy of the proposed method is validated on extensive challenging datasets through image alignment and face recognition. Experimental results demonstrate that our algorithm provides more illumination- and occlusion-robust image alignment than state-of-the-art methods do.
Qingqing Zheng, Yi Wang 0003, Pheng-Ann Heng
ICCV2
2017 Towards memory-efficient processing-in-memory architecture for convolutional neural networks
abstract
Convolutional neural networks (CNNs) are widely adopted in artificial intelligent systems. In contrast to conventional computing centric applications, the computational and memory resources of CNN applications are mixed together in the network weights. This incurs a significant amount of data movement, especially for highdimensional convolutions. Although recent embedded 3D-stacked Processing-in-Memory (PIM) architecture alleviates this memory bottleneck to provide fast near-data processing, memory is still a limiting factor of the entire system. An unsolved key challenge is how to efficiently allocate convolutions to 3D-stacked PIM to combine the advantages of both neural and computational processing.
Yi Wang 0003, Mingxu Zhang, Jing Yang 0018
LCTES1
2017 Fine grained, direct access file system support for storage class memory
Yi Wang 0003, Tianzheng Wang 0001, Duo Liu 0002, Zili Shao, Jingling Xue
J. Syst. Archit.1
2017 Heating Dispersal for Self-Healing NAND Flash Memory
abstract
Substantially reduced lifetimes are becoming a critical issue in NAND flash memory with the advent of multi-level cell and triple-level cell flash memory. Researchers discovered that heating can cause worn-out NAND flash cells to become reusable and greatly extend the lifetime of flash memory cells. However, the heating process consumes a substantial amount of power, and some fundamental changes are required for existing NAND flash management techniques. In particular, all existing wear-leveling techniques are based on the principle of evenly distributing writes and erases. For self-healing NAND flash, this may cause NAND flash cells to be worn out in a short period of time. Moreover, frequently healing these cells may drain the energy quickly in battery-driven mobile devices, which is defined as the concentrated heating problem. In this paper, we propose a novel wear-leveling scheme called DHeating (Dispersed Heating) to address the problem. In DHeating, rather than evenly distributing writes and erases over a time period, write and erase operations are scheduled on a small number of flash memory cells at a time, so that these cells can be worn out and healed much earlier than other cells. In this way, we can avoid quick energy depletion caused by concentrated heating. In addition, the heating process takes several seconds and has become the new performance bottleneck. In order to address this issue, we propose a lazy heating repair scheme. The lazy heating repair scheme can ease the long time delays caused by the heating via delaying the heating operation and using the system idle time to repair. Furthermore, the flash memory's reliability becomes worse with the flash memory cells reaching the excepted worn-out time. We propose an early heating strategy to solve the reliability problem. With the extended lifetime provided by self-healing, we can trade some lifetimes for reliability. The idea is to start the healing process earlier than the expected worn-out time. We evaluate our scheme based on an embedded platform. The experimental results show that the proposed scheme can effectively prolong the consecutive heating time interval, alleviate the long time delays caused by the heating, and enhance the reliability for self-healing flash memory.
Renhai Chen, Yi Wang 0003, Duo Liu 0002, Zili Shao, Song Jiang 0001
IEEE Trans. Computers2
2017 A Block-Level Log-Block Management Scheme for MLC NAND Flash Memory Storage Systems
abstract
NAND flash memory is the major storage media for both mobile storage cards and enterprise Solid-State Drives (SSDs). Log-block-based Flash Translation Layer (FTL) schemes have been widely used to manage NAND flash memory storage systems in industry. In log-block-based FTLs, a few physical blocks called log blocks are used to hold all page updates from a large amount of data blocks. Frequent page updates in log blocks introduce big overhead so log blocks become the system bottleneck. To address this problem, this paper presents BLog, a block-level log-block management scheme for MLC NAND flash memory storage system. In BLog, with block-level management, the update pages of a data block can be collected together and put into the same log block as much as possible; therefore, we can effectively reduce the associativities of log blocks so as to reduce the garbage collection overhead. We also propose a novel partial merge operation strategy called reduced-order merge by which we can effectively postpone the garbage collection of log blocks so as to maximally utilize valid pages and reduce unnecessary erase operations in log blocks. Based on BLog, we design an FTL called BLogFTL for Multi-Level Cell (MLC) NAND flash. We conduct a set of experiments on a real hardware platform. Both representative FTL schemes and the proposed BLogFTL have been implemented in the hardware evaluation board. The experimental results show that our scheme can effectively reduce the garbage collection operations and reduce the system response time compared to the previous log-block-based FTLs for MLC NAND flash.
Chenlin Ma, Renhai Chen, Yi Wang 0003, Zili Shao
IEEE Trans. Computers5
2017 vFlash: Virtualized Flash for Optimizing the I/O Performance in Mobile Devices
abstract
I/O is becoming one of major performance bottlenecks in NAND-flash-based mobile devices. Novel nonvolatile memories (NVMs), such as phase change memory and spin-transfer torque random access memory, can provide fast read/write operations. In this paper, we propose a unified NVM/flash architecture to improve the I/O performance. A transparent scheme, virtualized flash (vFlash), is also proposed to manage the unified architecture. Within vFlash, interapp and intra-app techniques are proposed to optimize the application performance by exploiting the historical locality and I/O access patterns of applications. Since vFlash is on the bottom of the I/O stack, the application features will be lost. Therefore, we also propose a cross-layer technique to transfer the application information from the application layer to the vFlash layer. The proposed scheme is evaluated based on an Android platform, and the experimental results show that the proposed scheme can effectively improve the I/O performance of mobile devices.
Renhai Chen, Yi Wang 0003, Jingtong Hu, Duo Liu 0002, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 P-Alloc: Process-Variation Tolerant Reliability Management for 3D Charge-Trapping Flash Memory
abstract
Three-dimensional (3D) flash memory is an emerging memory technology that enables a number of improvements to conventional planar NAND flash memory, including larger capacity, less program disturbance, and lower access latency. In contrast to conventional planar flash memory, 3D flash memory adopts charge-trapping mechanism. NAND strings punch through multiple stacked layers to form the three-dimensional infrastructure. However, the etching processes for NAND strings are unable to produce perfectly vertical features, especially on the scale of 20 nanometers or less. The process variation will cause uneven distribution of electrons, which poses a threat to the integrity of data stored in flash. This paper present P-Alloc , a process-variation tolerant reliability management strategy for 3D charge-trapping flash memory. P-Alloc offers both hardware and software support to allocate data to the 3D flash in the presence of process variation. P-Alloc predicts the state of a physical page, i.e., the basic unit for each write or read operation in flash memory, and tries to assign critical data to more reliable pages. A hardware-based voltage threshold compensation scheme is also proposed to further reduce the faults. We demonstrate the viability of the proposed scheme using a variety of realistic workloads. Our extensive evaluations show that, P-Alloc significantly enhances the reliability and reduces the access latency compared to the baseline scheme.
Yi Wang 0003, Lisha Dong, Rui Mao 0001
ACM Trans. Embed. Comput. Syst.1
2017 Durable Address Translation in PCM-Based Flash Storage Systems
abstract
Phase change memory (PCM) is a promising DRAM alternative because of its non-volatility, high density, low standby power and close-to-DRAM performance. These features make PCM an attractive solution to optimize the management of NAND flash memory in embedded systems. However, PCM's limited write endurance hinders its application in embedded systems. Therefore, how to manage flash memory with PCM-particularly guarantee PCM a reasonable lifetime-becomes a challenging issue. In this paper, we propose to partially replace DRAM using PCM to optimize the management of flash memory metadata for better system reliability in the presence of power failure and system crash. To prolong PCM's lifetime, we present a write-activity-aware PCM-assisted flash memory management scheme, called PCM-FTL. By differentiating sequential and random I/O behaviors, a novel two-level mapping mechanism and a customized wear-leveling scheme are developed to reduce writes to PCM and extend its lifetime. We evaluate PCM-FTL with a variety of general-purpose and mobile I/O workloads. Experimental results show that PCM-FTL can significantly reduce write activities and achieve an even distribution of writes in PCM with very low overhead.
Duo Liu 0002, Kan Zhong, Tianzheng Wang 0001, Yi Wang 0003, Zili Shao, Edwin H.-M. Sha, Jingling Xue
IEEE Trans. Parallel Distributed Syst.4
2016 A Thermal-Aware Physical Space Allocation Strategy for 3D Flash Memory Storage Systems
abstract
Three-dimensional (3D) flash memory stacks layers of data storage cells vertically to overcome the scaling limits in conventional planar NAND flash memory. Current 3D flash memory faces new challenges including thermal issues and complex manufacturing process. This paper presents TheraPhy, a novel thermal-aware physical space allocation strategy for three-dimensional flash memory storage systems. TheraPhy permutes the allocation of physical blocks. Consecutively accessed logical blocks are distributed to different physical locations in order to prevent the accumulation of hotspots. TheraPhy requires no changes to the file system, on-chip memory hierarchy, or hardware implementation of 3D flash memory. Based on TheraPhy, we present an address mapping strategy that is capable of determining the allocation of physical blocks based on their thermal status. We demonstrate the viability of the proposed technique using a set of extensive experiments. Experimental results show that TheraPhy can reduce the peak temperature by 15.39% with less than 1% extra erase overhead in comparison with the baseline scheme.
Yi Wang 0003, Mingxu Zhang, Lisha Dong
ISLPED1
2016 An Endurance-Aware Metadata Allocation Strategy for MLC NAND Flash Memory Storage Systems
abstract
This paper presents a reliability-aware metadata allocation strategy called scatter-single-level cell (SLC) for multiple-level cell (MLC) NAND flash memory storage systems. In scatter-SLC, metadata is kept in least significant bit (LSB) pages and corresponding most significant bit (MSB) pages are bypassed. Without partitioning SLC and MLC blocks, scatter-SLC can eliminate the unbalanced lifetime between SLC and MLC blocks while achieving the similar error rate as the method to store metadata in SLC blocks. We implemented scatter-SLC on a real-hardware platform. The experiment results show that scatter-SLC can reduce uncorrectable page errors by 93.54% while incurring less than 1% time overhead on average compared with the previous work.
Min Huang 0002, Zhaoqing Liu, Liyan Qiao, Yi Wang 0003, Zili Shao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2016 Energy-Efficient In-Memory Paging for Smartphones
abstract
Smartphones are becoming increasingly energy-hungry to support feature-rich applications, posing a lot of pressure on battery lifetime and making energy consumption a non-negligible issue. In particular, dynamic random access memory (DRAM)-based main memory subsystem is a major contributor to the energy consumption of mobile devices. In this paper, we propose direct read (DR). Swap, an energy-efficient in-memory paging design to reduce energy consumption in smartphones. In DR. Swap, we adopt emerging energy-efficient nonvolatile memory (NVM) and use it as the swap area. Utilizing NVMs byte-addressability, we propose DR which guarantees zero memory copy for read-only requests when accessing a page in swap area. To better understand the energy consumption of swapping, we build an energy model to analyze the energy consumption of different paging architectures. We evaluate DR. Swap based on the Google Nexus 5 smartphone, experimental results show that our technique can reduce more than 50% energy consumption compared to DRAM backed swapping.
Kan Zhong, Duo Liu 0002, Liang Liang 0002, Linbo Long, Yi Wang 0003, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2016 Image-Content-Aware I/O Optimization for Mobile Virtualization
Renhai Chen, Yi Wang 0003, Jingtong Hu, Duo Liu 0002, Zili Shao
ACM Trans. Embed. Comput. Syst.2
2016 An Adaptive Demand-Based Caching Mechanism for NAND Flash Memory Storage Systems
abstract
During past decades, the capacity of NAND flash memory has been increasing dramatically, leading to the use of nonvolatile flash in the system’s memory hierarchy. The increasing capacity of NAND flash memory introduces a large RAM footprint to store the logical to physical address mapping. The demand-based approach can effectively reduce and well control the RAM footprint. However, extra address translation overhead is also introduced which may degrade the system performance. In this article, we present CDFTL, an adaptive Caching mechanism for Demand-based Flash Translation Layer, for NAND flash memory storage systems. CDFTL adopts both the fine-grained entry-based caching mechanism to exploit temporal locality and the coarse-grained translation-page-based caching mechanism to exploit spatial locality of workloads. By selectively caching the on-demand address mappings and adaptively changing the space configurations of two granularities, CDFTL can effectively utilize the RAM space and improve the cache hit ratio. We evaluate CDFTL under a real hardware embedded platform using a variety of I/O traces. Experimental results show that our technique can achieve an 11.13% reduction in average system response time and a 35.21% reduction in translation block erase counts compared with the previous work.
Yi Wang 0003, Zhiwei Qin 0004, Renhai Chen, Zili Shao, Laurence T. Yang
ACM Trans. Design Autom. Electr. Syst.1
2015 Unified non-volatile memory and NAND flash memory architecture in smartphones
abstract
I/O is becoming one of major performance bottlenecks in NAND-flash-based smartphones. Novel NVMs (nonvolatile memories), such as PCM (Phase Change Memory) and STT-RAM (Spin-Transfer Torque Random Access Memory), can provide fast read/write operations. In this paper, we propose an unified NVM/flash architecture to improve the I/O performance. A transparent scheme, vFlash (Virtualized Flash), is also proposed to manage the unified architecture. Within vFlash, inter-app technique is proposed to optimize the application performance by exploiting the historic locality of applications. Since vFlash is on the bottom of the I/O stack, the application features will be lost. Therefore, we also propose a cross-layer technique to transfer the application information from the application layer to the vFlash layer. The proposed scheme is evaluated based on a real Android platform, and the experimental results show that the read and write performance for the proposed scheme is 2.45 times and 3.37 times better than that of the stock Android 4.2 system, respectively.
Renhai Chen, Yi Wang 0003, Jingtong Hu, Duo Liu 0002, Zili Shao
ASP-DAC2
2015 A Garbage Collection Aware Stripping method for Solid-State Drives
abstract
2015 20th Asia and South Pacific Design Automation Conference, ASP-DAC 2015, Japan, 19-22 January 2015
Min Huang 0002, Yi Wang 0003, Zhaoqing Liu, Liyan Qiao, Zili Shao
ASP-DAC2
2015 Optimizing deterministic garbage collection in NAND flash storage systems
abstract
NAND flash has been widely adopted as storage devices in real-time embedded systems. However, garbage collection is needed to reclaim space and introduces a lot of time overhead. As the worst system latency is determined by the worst-case execution time of garbage collection in NAND flash, it is important to optimize garbage collection so as to give a deterministic worst system latency. On the other hand, since the garbage collection does not happen very often, optimizing garbage collection should not bring too much overhead to the average system latency. This paper presents for the first time a worst-case and average-case joint optimization scheme for garbage collection in NAND flash. With our scheme, garbage collection can be postponed to the latest stage so improves the average system latency. By combining partial garbage collection and over-provisioning, our scheme can guarantee that one free block is enough to hold all pages from both write requests and valid-page copies. The experiments have been conducted on a real embedded platform and the results show that our technique can improve both worstcase and average-case system latency compared with the previous works.
Xuandong Li, Linzhang Wang, Tian Zhang 0001, Yi Wang 0003, Zili Shao
RTAS5
2015 On-Demand Block-Level Address Mapping in Large-Scale NAND Flash Storage Systems
abstract
The density of flash memory chips has doubled every two years in the past decade and the trend is expected to continue. The increasing capacity of NAND flash memory leads to large RAM footprint on address mapping management. This paper proposes a novel Demand-based block-level Address mapping scheme with a two-level Caching mechanism (DAC) for large-scale NAND flash storage systems. The objective is to reduce RAM footprint without excessively compromising system response time. In our technique, the block-level address mapping table is stored in fixed pages (called the translation pages) in the flash memory. Considering temporal locality that workloads exhibit, we maintain one cache in RAM to store the on-demand address mapping entries. Meanwhile, by exploring both spatial locality and access frequency of workloads with another two caches, the second-level cache is designed to cache selected translation pages. In such a way, both the most-frequently-accessed and sequentially accessed address mapping entries can be stored in the cache so the cache hit ratio can be increased and the system response time can be improved. To the best of our knowledge, this is the first work to reduce the RAM cost by employing the demand-based approach on block-level address mapping schemes. The experiments have been conducted on a real embedded platform. The experimental results show that our technique can effectively reduce the RAM footprint while maintaining similar average system response time compared with previous work.
Renhai Chen, Zhiwei Qin 0004, Yi Wang 0003, Duo Liu 0002, Zili Shao
IEEE Trans. Computers3
2015 Temperature-Aware Data Allocation for Embedded Systems with Cache and Scratchpad Memory
abstract
The hybrid memory architecture that contains both on-chip cache and scratchpad memory (SPM) has been widely used in embedded systems. In this article, we explore this hybrid memory architecture by jointly optimizing time performance and temperature for embedded systems with loops. Our basic idea is to adaptively adjust the workload distribution between cache and SPM based on the current temperature. For a problem in which the workload can be estimated a priori, we present a nonlinear programming formulation to optimally minimize the total execution time of a loop under the constraints of SPM size and temperature. To solve a problem in which the workload is not known a priori, we propose a temperature-aware adaptive loop scheduling algorithm called TALS to dynamically allocate data to cache and SPM at runtime. The experimental results show that our algorithms can effectively achieve both performance and temperature optimization for embedded systems with cache and SPM.
Zhiping Jia, Yi Wang 0003, Meng Wang 0005, Zili Shao
ACM Trans. Embed. Comput. Syst.3
2015 Towards Write-Activity-Aware Page Table Management for Non-volatile Main Memories
abstract
Non-volatile memories such as phase change memory (PCM) and memristor are being actively studied as an alternative to DRAM-based main memory in embedded systems because of their properties, which include low power consumption and high density. Though PCM is one of the most promising candidates with commercial products available, its adoption has been greatly compromised by limited write endurance. As main memory is one of the most heavily accessed components, it is critical to prolong the lifetime of PCM. In this article, we present w rite- a ctivity-aware p age t able m anagement (WAPTM), a simple yet effective page table management scheme for reducing unnecessary writes, by redesigning system software and exploiting write-activity-aware features provided by the hardware. We implemented WAPTM in Google Android based on the ARM architecture and evaluated it with real Android applications. Experimental results show that WAPTM can significantly reduce writes in page tables, proving the feasibility and potential of prolonging the lifetime of PCM-based main memory through reducing writes at the OS level.
Tianzheng Wang 0001, Duo Liu 0002, Yi Wang 0003, Zili Shao
ACM Trans. Embed. Comput. Syst.3
2015 Lazy-RTGC: A Real-Time Lazy Garbage Collection Mechanism with Jointly Optimizing Average and Worst Performance for NAND Flash Memory Storage Systems
abstract
Due to many attractive and unique properties, NAND flash memory has been widely adopted in mission-critical hard real-time systems and some soft real-time systems. However, the nondeterministic garbage collection operation in NAND flash memory makes it difficult to predict the system response time of each data request. This article presents Lazy-RTGC , a real-time lazy garbage collection mechanism for NAND flash memory storage systems. Lazy-RTGC adopts two design optimization techniques: on-demand page-level address mappings, and partial garbage collection. On-demand page-level address mappings can achieve high performance of address translation and can effectively manage the flash space with the minimum RAM cost. On the other hand, partial garbage collection can provide the guaranteed system response time. By adopting these techniques, Lazy-RTGC jointly optimizes both the average and the worst system response time, and provides a lower bound of reclaimed free space. Lazy-RTGC is implemented in FlashSim and compared with representative real-time NAND flash memory management schemes. Experimental results show that our technique can significantly improve both the average and worst system performance with very low extra flash-space requirements.
Xuandong Li, Linzhang Wang, Tian Zhang 0001, Yi Wang 0003, Zili Shao
ACM Trans. Design Autom. Electr. Syst.5
2014 Deterministic Crash Recovery for NAND Flash Based Storage Systems
abstract
NAND flash memory has long been the dominant storage medium in mobile devices. However, power failure may occur at any time and result in loss of important data. Crash recovery therefore becomes vitally important in NAND flash memory storage systems. As flash translation layer (FTL) directly manages flash memory using various metadata, the problem of FTL crash recovery in NAND flash is how to efficiently and effectively maintain and recover the consistency of FTL metadata after system crash.
Yi Wang 0003, Tianzheng Wang 0001, Renhai Chen, Duo Liu 0002, Zili Shao
DAC2
2014 Analysis of worst-case backlog bounds for Networks-on-Chip
Yi Wang 0003, Zili Shao, Wenhua Dou, Qiang Dou
J. Syst. Archit.3
2014 Loop scheduling with memory access reduction subject to register constraints for DSP applications
abstract
SUMMARY Memory accesses introduce big‐time overhead and power consumption because of the performance gap between processors and main memory. This paper describes and evaluates a technique, loop scheduling with memory access reduction (LSMAR), that replaces hidden redundant load operations with register operations in loop kernels and performs partial scheduling for newly generated register operations subject to register constraints. By exploiting data dependence of memory access operations, the LSMAR technique can effectively reduce the number of memory accesses of loop kernels, thereby improving timing performance. The technique has been implemented into the Trimaran compiler and evaluated using a set of benchmarks from DSPstone and MiBench on the cycle‐accurate simulator of the Trimaran infrastructure. The experimental results show that when the LSMAR technique is applied, the number of memory accesses can be reduced by 18.47% on average over the benchmarks when it is not applied. The measurements also indicate that the optimizations only lead to an average 1.41% increase in code size. With such small code size expansion, the technique is more suitable for embedded systems compared with prior work.Copyright © 2013 John Wiley & Sons, Ltd.
Yi Wang 0003, Zhiping Jia, Renhai Chen, Meng Wang 0005, Duo Liu 0002, Zili Shao
Softw. Pract. Exp.1
2014 Application-Specific Wear Leveling for Extending Lifetime of Phase Change Memory in Embedded Systems
abstract
Phase change memory (PCM) has been proposed to replace NOR flash and DRAM in embedded systems because of its attractive features. However, the endurance of PCM greatly limits its adoption in embedded systems. As most embedded systems are application-oriented, we can tackle the endurance problem of PCM by exploring application-specific features such as fixed access patterns and update frequencies. In this paper, we propose an application-specific wear leveling technique, called Curling-PCM, to evenly distribute write activities across the whole PCM chip to improve the endurance of PCM in embedded systems. The basic idea is to exploit application-specific features in embedded systems and periodically move the hot region across the whole PCM chip. To reduce the overhead of moving the hot region and improve the performance of PCM-based embedded systems, a fine-grained partial wear leveling policy is proposed for Curling-PCM, by which only part of the hot region is moved during each request handling period. Experimental results show that Curling-PCM can effectively evenly distribute write traffic for a prime application of PCM in embedded systems. We expect this paper can serve as a first step toward the full exploration of application-specific features in PCM-based embedded systems.
Duo Liu 0002, Tianzheng Wang 0001, Yi Wang 0003, Zili Shao, Qingfeng Zhuge, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 A Reliability-Aware Address Mapping Strategy for NAND Flash Memory Storage Systems
abstract
The increasing density of NAND flash memory leads to a dramatic increase in the bit error rate of flash, which greatly reduces the ability of error correcting codes (ECC) to handle multibit errors. NAND flash memory is normally used to store the file system metadata and page mapping information. Thus, a broken physical page containing metadata may cause an unintended and severe change in functionality of the entire flash. This paper presents Meta-Cure, a novel hardware and file system interface that transparently protects metadata in the presence of multibit faults. Meta-Cure exploits built-in ECC and replication in order to protect pages containing critical data, such as file system metadata. Redundant pairs are formed at run time and distributed to different physical pages to protect against failures. Meta-Cure requires no changes to the file system, on-chip hierarchy, or hardware implementation of flash memory chip. We evaluate Meta-Cure under a real-embedded platform using a variety of I/O traces. The evaluation platform adopts dual ARM Cortex A9 processor cores with 64 Gb NAND flash memory. We have evaluated the effectiveness of Meta-Cure on the new technology file system file system. Experimental results show that the proposed technique can reduce uncorrectable page errors by 70.38% with less than 7.86% time overhead in comparison with conventional error correction techniques.
Yi Wang 0003, Min Huang 0002, Zili Shao, Henry C. B. Chan, Luis Angel D. Bathen, Nikil Dutt
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 Memory-Aware Task Scheduling with Communication Overhead Minimization for Streaming Applications on Bus-Based Multiprocessor System-on-Chips
abstract
Inter-core communication introduces overheads in task schedules on Multiprocessor System-on-Chips (MPSoCs). Inter-core communication overhead not only negatively impacts the timing performance but also significantly degrades the memory usage for streaming applications running on MPSoC architectures. By minimizing inter-core communication overhead, a shorter period can be applied and system performance (e.g., throughput, memory usage) can be improved. In this paper, we focus on solving the problem of minimizing inter-core communication overhead for streaming applications on bus-based MPSoCs. The objective is to minimize inter-core communication overhead while minimizing the overall memory usage. To solve the problem, we first let tasks with intra-period data dependencies transform to inter-period data dependencies so as to overlap the execution of computation and inter-core communication tasks. By doing this, inter-core communication overhead can be effectively removed. To minimize the overall memory usage, we then perform schedulability analysis and obtain the bounds of the times needed to reschedule each task. Based on the schedulability analysis, we formulate the scheduling problem as an integer linear programming (ILP) model and obtain an optimal schedule. In addition, we propose a heuristic approach to efficiently obtain a near-optimal solution. We conduct experiments on a set of benchmarks from both real-life streaming applications and synthetic task graphs. The experimental results show that the proposed approach can significantly reduce the schedule length and improve the memory usage compared with the previous work.
Yi Wang 0003, Zili Shao, Henry C. B. Chan, Duo Liu 0002
IEEE Trans. Parallel Distributed Syst.1
2014 A Reliability Enhanced Address Mapping Strategy for Three-Dimensional (3-D) NAND Flash Memory
abstract
The linear scaling down of NAND flash memory is approaching its physical, electrical, and reliability limitations. To maintain the current trend of increasing bit density and reducing bit per cost, 3-D flash memory is emerging as a viable solution to fulfill the ever-increasing demands of storage capacity. In 3-D NAND flash memory, multiple layers are stacked to provide ultrahigh density storage devices. However, the physical architecture of 3-D flash memory leads to a higher probability of disturbance to adjacent physical pages and greatly increases bit error rates. This paper presents a novel physical-location-aware address mapping strategy for 3-D NAND flash memory. It permutes the physical mapping of pages and maximizes the distance between the consecutively logical pages, which can significantly reduce the disturbance to adjacent physical pages and effectively enhance the reliability. The proposed mapping strategy is applied to a representative flash storage system. Experimental results show that the proposed scheme can reduce uncorrectable page errors by 70.16% with less than 10.01% space overhead in comparison with the baseline scheme.
Yi Wang 0003, Zili Shao, Henry C. B. Chan, Luis Angel D. Bathen, Nikil Dutt
IEEE Trans. Very Large Scale Integr. Syst.1
2013 Curling-PCM: Application-specific wear leveling for phase change memory based embedded systems
abstract
Phase change memory (PCM) has been used as NOR flash replacement in embedded systems with its attractive features. However, the endurance of PCM keeps drifting down and greatly limits its adoption in embedded systems. As most embedded systems are application-oriented, we can better utilize PCM by exploring application-specific features such as fixed access patterns and update frequencies to prolong the lifetime of PCM. In this paper, we propose an application-specific wear leveling technique, called Curling-PCM, to evenly distribute write activities across the PCM chip in order to improve the endurance of PCM. The basic idea is to exploit application-specific features in embedded systems and periodically move the hot region across the whole PCM chip. To further reduce the overhead of moving the hot region and improve the performance of PCM-based embedded systems, a fine-grained partial wear leveling policy is proposed in Curling-PCM, by which only part of the hot region is moved during each request handling period. The experimental results show that Curling-PCM can effectively evenly distribute write traffic in PCM chips compared with previous work. We expect this work can serve as a first step towards the full exploration of application-specific features in PCM-based embedded systems.
Duo Liu 0002, Tianzheng Wang 0001, Yi Wang 0003, Zili Shao, Qingfeng Zhuge, Edwin H.-M. Sha
ASP-DAC3
2013 Optimizing translation information management in NAND flash memory storage systems
abstract
Address mapping is one of the major functions in managing NAND flash. With the capacity increase of NAND flash, it becomes vitally important to reduce the RAM print of the address mapping table while not introducing big performance overhead. Demand-based address mapping is an effective approach to solve this problem, in which the address mapping table is stored in NAND flash (called translation pages), and mapping items are cached on-demand in RAM. Therefore, it is critical to manage translation pages in demand-based address mapping. This paper solves two most important problems in translation page management. First, to reduce frequent translation page updates caused by data requests, we propose a page-level caching mechanism to exploit the fundamental property of NAND flash where the basic read/write unit is one page. Second, to reduce the garbage collection overhead from translation pages, we propose a multiple write pointers strategy to group data pages corresponding to the same translation page into one data block, by which, when the data block is reclaimed via the garbage collection, we only need to update one translation page. We evaluate our scheme using a set of benchmarks from both real-world and synthetic traces. Experimental results show that our techniques can achieve significant reduction in the extra translation operations and improve the system response time.
Xuandong Li, Linzhang Wang, Tian Zhang 0001, Yi Wang 0003, Zili Shao
ASP-DAC5
2013 BLog: block-level log-block management for NAND flash memorystorage systems
abstract
Log-block-based FTL (Flash Translation Layer) schemes have been widely used to manage NAND flash memory storage systems in industry. In log-block-based FTLs, a few physical blocks called log blocks are used to hold all page updates from a large amount of data blocks. Frequent page updates in log blocks introduce big overhead so log blocks become the system bottleneck.
Yi Wang 0003, Renhai Chen, Zili Shao
LCTES3
2013 FTL2: a hybrid flash translation layer with logging for write reduction in flash memory
abstract
NAND flash memory has been widely used to build embedded devices such as smartphones and solid state drives (SSD) because of its high performance, low power consumption, great shock resistance and small form factor. However, its lifetime and performance are greatly constrained by partial page updates, which will lead to early depletion of free pages and frequent garbage collections. On the one hand, partial page updates are prevalent as a large portion of I/O does not modify file contents drastically. On the other hand, general-purpose cache usually does not specifically consider and eliminate duplicated contents, despite its popularity.
Tianzheng Wang 0001, Duo Liu 0002, Yi Wang 0003, Zili Shao
LCTES3
2013 SolarTune: Real-time scheduling with load tuning for solar energy powered multicore systems
abstract
In this paper, we target at the direct-coupled solar energy powered multicore architectures that provide direct power supply between photovoltaic (PV) generation and the load without the adoption of battery. We present Solar-Tune, a real-time scheduling technique with load tuning for sporadic tasks on solar energy powered multicore systems. The objective is to fully utilize the available solar energy while meeting the deadlines of tasks. To solve the problem, we first perform analysis and formulate the scheduling problem in each duration as an integer linear programming (ILP) model to obtain an optimal schedule. Then we present a heuristic algorithm to dynamically refine the task scheduling based on the predictions of the availability of solar energy. We conduct experiments using real-world meteorological data across different geographic sites. The experimental results show that SolarTune can significantly improve the solar energy utilization ratio and reduce the number of deadline misses compared to the conventional task scheduler.
Yi Wang 0003, Renhai Chen, Zili Shao, Tao Li 0006
RTCSA1
2013 Optimally Removing Intercore Communication Overhead for Streaming Applications on MPSoCs
abstract
This paper aims to totally remove intercore communication overhead with joint computation and communication task scheduling for streaming applications on Multiprocessor System-on-Chips (MPSoCs). Our basic idea is to let some computation and communication tasks be executed in earlier periods (the added periods are called the prologue) such that intercore data transfer can be finished before the execution of the tasks that need the data to start. In particular, we solve the following problem: how to do rescheduling in such a way that the schedule length can be minimized with the minimum prologue length (the number of periods in the prologue) while the intercore communication overhead can be totally removed? To solve this problem, we first perform schedulability analysis and obtain the upper bound of the times needed to reschedule each computation task. Then we formulate the problem as an Integer Linear Programming (ILP) formulation and obtain an optimal solution. We evaluate our technique with a set of benchmarks from both real-life streaming applications and synthetic task graphs. The experimental results show that our technique can achieve significant reductions in schedule length and energy consumption compared with the previous work.
Yi Wang 0003, Duo Liu 0002, Zhiwei Qin 0004, Zili Shao
IEEE Trans. Computers1
2012 Meta-Cure: a reliability enhancement strategy for metadata in NAND flash memory storage systems
abstract
The increasing density of NAND flash memory leads to a dramatic increase in the bit error rate of flash, which greatly reduces the ability of error correcting codes (ECC) to handle multi-bit errors. To ensure the functionality and reliability of flash memory, the pages containing address mapping information and other metadata should be carefully stored in flash memory. This paper presents Meta-Cure, a novel hardware and file system interface that transparently protects metadata in the presence of multi-bit faults. Meta-Cure exploits built-in ECC and replication in order to protect pages containing critical data. Redundant pairs are formed at run time and distributed to different physical pages to protect against failures. Meta-Cure requires no changes to the file system, on-chip hierarchy, or hardware implementation of flash memory chip. Experimental results show that the proposed technique can reduce uncorrectable page errors by 92% with less than 1% space overhead in comparison with conventional error correction techniques.
Yi Wang 0003, Luis Angel D. Bathen, Nikil Dutt, Zili Shao
DAC1
2012 A block-level flash memory management scheme for reducing write activities in PCM-based embedded systems
abstract
This paper targets at an embedded system with phase change memory (PCM) and NAND flash memory. Although PCM is a promising main memory alternative and is recently introduced to embedded system designs, its endurance keeps drifting down and greatly limits the lifetime of the whole system. Therefore, this paper presents a block-level flash memory management scheme, WAB-FTL, to effectively manage NAND flash memory while reducing write activities of the PCM-based embedded systems. The basic idea is to preserve each bit in flash mapping table hosted by PCM from being inverted frequently during the process of mapping table update. To achieve this, a new merge strategy is adopted in WAB-FTL to delay the mapping table update, and a tiny mapping buffer is used for caching frequently updated mapping records. Experimental results based on Android traces show that WAB-FTL can effectively reduce write activities when compared with the baseline scheme.
Duo Liu 0002, Tianzheng Wang 0001, Yi Wang 0003, Zhiwei Qin 0004, Zili Shao
DATE3
2012 3D-FlashMap: A physical-location-aware block mapping strategy for 3D NAND flash memory
abstract
Three-dimensional (3D) flash memory is emerging to fulfil the ever-increasing demands of storage capacity. In 3D NAND flash memory, multiple layers are stacked to increase bit density and reduce bit cost of flash memory. However, the physical architecture of 3D flash memory leads to a higher probability of disturbance to adjacent physical pages and greatly increases bit error rates. This paper presents 3D-FlashMap, a novel physical-location-aware block mapping strategy for three-dimensional NAND flash memory. 3D-FlashMap permutes the physical mapping of blocks and maximizes the distance between consecutively logical blocks, which can significantly reduce the disturbance to adjacent physical pages and effectively enhance the reliability. We apply 3D-FlashMap to a representative flash storage system. Experimental results show that the proposed scheme can reduce uncorrectable page errors by 85% with less than 2% space overhead in comparison with the baseline scheme.
Yi Wang 0003, Luis Angel D. Bathen, Zili Shao, Nikil Dutt
DATE1
2012 Real-Time Flash Translation Layer for NAND Flash Memory Storage Systems
abstract
Due to the variable garbage collection latency, NAND flash memory storage systems may suffer long system response time, especially when the flash memory is close to be full. Most of existing flash translation layer (FTL) schemes focus on improving the average response time but ignore to provide a desirable worst case response time upper bound. This paper proposes a Real-time Flash Translation Layer (RFTL) scheme to hide the long garbage collection latency while satisfying a worst case response time upper bound that achieves an ideal case. We achieve this by using a distributed partial garbage collection policy that enables RFTL to reclaim the space and to serve the write requests simultaneously. A new block-level address mapping approach is designed to guarantee enough free space to serve the write request arriving at any time period. Experimental results show that our scheme improves both the worst case system response time and the average system response time compared with previous work.
Zhiwei Qin 0004, Yi Wang 0003, Duo Liu 0002, Zili Shao
IEEE Real-Time and Embedded Technology and Applications Symposium2
2012 Staying-alive path planning with energy optimization for mobile robots
Hongxing Wei, Bin Wang 0065, Yi Wang 0003, Zili Shao, Keith C. C. Chan
Expert Syst. Appl.3
2012 Optimally Maximizing Iteration-Level Loop Parallelism
abstract
Loops are the main source of parallelism in many applications. This paper solves the open problem of extracting the maximal number of iterations from a loop to run parallel on chip multiprocessors. Our algorithm solves it optimally by migrating the weights of parallelism-inhibiting dependences on dependence cycles in two phases. First, we model dependence migration with retiming and formulate this classic loop parallelization into a graph optimization problem, i.e., one of finding retiming values for its nodes so that the minimum nonzero edge weight in the graph is maximized. We present our algorithm in three stages with each being built incrementally on the preceding one. Second, the optimal code for a loop is generated from the retimed graph of the loop found in the first phase. We demonstrate the effectiveness of our optimal algorithm by comparing with a number of representative nonoptimal algorithms using a set of benchmarks frequently used in prior work and a set of graphs generated by TGFF.
Duo Liu 0002, Yi Wang 0003, Zili Shao, Minyi Guo, Jingling Xue
IEEE Trans. Parallel Distributed Syst.2
2012 A Space Reuse Strategy for Flash Translation Layers in SLC NAND Flash Memory Storage Systems
abstract
This paper presents a space reuse strategy for flash translation layers in SLC nand flash storage systems. The basic idea is to prevent a block with many free pages from being erased in a merge operation. The preserved blocks are further reused as replacement blocks. In such a way, the space utilization and the number of erase counts of each block in a nand flash are enhanced. By employing the reuse strategy, we propose a reuse-aware flash translation layer (FTL) called reuse-aware NFTL (RNFTL) to improve the endurance and space utilization of single level cell (SLC) nand flash. We provide the performance analysis of RNFTL for frequent update operations and sequential write operations, and theoretically compare RNFTL with representative FTL schemes. We also discuss the opportunity to apply the reuse strategy in log-block-based FTL schemes. To the best of our knowledge, this is the first work to employ a space reuse strategy in FTLs to improve the space utilization and endurance of nand flash. The experiments have been conducted on a set of traces collected from real workload in daily life. The results show that the space reuse strategy can effectively improve space utilization, block lifetime and wear-leveling compared with the previous work.
Duo Liu 0002, Yi Wang 0003, Zhiwei Qin 0004, Zili Shao
IEEE Trans. Very Large Scale Integr. Syst.2
2011 MNFTL: an efficient flash translation layer for MLC NAND flash memory storage systems
abstract
The new write constraints of multi-level cell (MLC) NAND flash memory make most of the existing flash translation layer (FTL) schemes inefficient or inapplicable. In this paper, we solve several fundamental problems in the design of MLC flash translation layer. The objective is to reduce the garbage collection overhead so as to reduce the average system response time. We make the key observation that the valid page copy is the essential garbage collection overhead. Based on this observation, we propose two approaches, namely, concentrated mapping and postponed reclamation, to effective reduce the valid page copies. We conduct experiments on a set of benchmarks from both the real world and synthetic traces. The experimental results show that our scheme can achieve a significant reduction in the average system response time compared with the previous work.
Zhiwei Qin 0004, Yi Wang 0003, Duo Liu 0002, Zili Shao
DAC2
2011 An endurance-enhanced Flash Translation Layer via reuse for NAND flash memory storage systems
abstract
NAND flash memory is widely used in embedded systems due to its non-volatility, shock resistance and high cell density. In recent years, various Flash Translation Layer (FTL) schemes (especially hybrid-level FTL schemes) have been proposed. Although these FTL schemes provide good solutions in terms of endurance and wear-leveling, none of them have considered to reuse free pages in both data blocks and log blocks during a merge operation. By reusing these free pages, less free blocks are needed and the endurance of NAND flash memory is enhanced. We evaluate our reuse strategy using a variety of application specific I/O traces from Windows systems. Experimental results show that the proposed scheme can effectively reduce the erase counts and enhance the endurance of flash memory.
Yi Wang 0003, Duo Liu 0002, Zhiwei Qin 0004, Zili Shao
DATE1
2011 A Two-Level Caching Mechanism for Demand-Based Page-Level Address Mapping in NAND Flash Memory Storage Systems
abstract
The increasing capacity of NAND flash memory leads to large RAM footprint on address mapping in the Flash Translation Layer (FTL) design. The demand-based approach can reduce the RAM footprint, but extra address translation overhead is also introduced which may degrade the system performance. This paper proposes a two-level caching mechanism to selectively cache the on-demand page-level address mappings by jointly exploiting the temporal locality and the spatial locality of workloads. The objective is to improve the cache hit ratio so as to shorten the system response time and reduce the block erase counts for NAND flash memory storage systems. By exploring the optimized temporal-spatial cache configurations, our technique can well capture the reference locality in workloads so that the hit ratio can be improved. Experimental results show that our technique can achieve a 31.51% improvement in hit ratio, which leads to a 31.11% reduction in average system response time and a 50.83% reduction in block erase counts compared with the previous work.
Zhiwei Qin 0004, Yi Wang 0003, Duo Liu 0002, Zili Shao
IEEE Real-Time and Embedded Technology and Applications Symposium2
2011 PCM-FTL: A Write-Activity-Aware NAND Flash Memory Management Scheme for PCM-Based Embedded Systems
abstract
Due to its properties of high density, in-place update, and low standby power, phase change memory (PCM) becomes a promising main memory alternative in embedded systems. On the other hand, NAND flash memory is widely used as a secondary storage and has been integrated into PCM-based embedded systems. Since both NAND flash memory and PCM have limited lifetime, how to effectively manage NAND flash memory in PCM-based embedded systems, while considering the endurance issue is very important. In this paper, we present for the first time a write-activity-aware NAND flash memory management scheme, called PCM-FTL, to effectively manage NAND flash memory and enhance the endurance of PCM-based embedded systems. The basic idea is to preserve each bit in flash mapping table, which is stored in PCM, from being inverted frequently, i.e., we focus on minimizing the number of bit flips in a PCM cell when updating the flash mapping table. PCM-FTL employs a two-level mapping mechanism, which not only focuses on minimizing the write activities of PCM but also considers the access behavior of I/O requests. We evaluate PCM-FTL using a variety of realistic I/O traces. Experimental results show that the proposed technique can achieve an average reduction of 93.10% and a maximum reduction of 98.98% in the maximum number of bit flips for a PCM-based embedded system with 1GB NAND flash memory. We hope this work can serve as a first step towards the design of write-activity-aware FTL for the PCM-based embedded systems via simple and feasible modifications.
Duo Liu 0002, Tianzheng Wang 0001, Yi Wang 0003, Zhiwei Qin 0004, Zili Shao
RTSS3
2011 On Improving Real-Time Interrupt Latencies of Hybrid Operating Systems with Two-Level Hardware Interrupts
abstract
In this paper, we propose to implement hybrid operating systems based on two-level hardware interrupts. We analyze and model the worst-case real-time interrupt latency for RTAI and identify the key component for its optimization. Then, we propose our methodology to implement hybrid operating systems with two-level hardware interrupts by combining the real-time kernel and the time sharing OS (Operating System) kernel. Based on the methodology, we discuss the important issues for the implementation. Finally, we implement a hybrid system called RTLinux-THIN (Real-Time LINUX with Two-level Hardware INterrupts) on the ARM architecture by combining ARM Linux kernel 2.6.9 and μC/OS-II. We conduct experiments on a set of real application programs including mplayer, Bonnie, and iperf, and compare the interrupt latency and interrupt task distributions for RTLinux-THIN (with and without cache locking), RTAI, Linux, and Linux with RT patch on a hardware platform based on Intel PXA270 processor. The results show that our scheme not only provides an easy method for implementing hybrid systems but also achieves the performance improvement for both the time sharing and real-time subsystems.
Duo Liu 0002, Yi Wang 0003, Meng Wang 0005, Zili Shao
IEEE Trans. Computers3
2011 Overhead-aware energy optimization for real-time streaming applications on multiprocessor System-on-Chip
abstract
In this article, we focus on solving the energy optimization problem for real-time streaming applications on multiprocessor System-on-Chip by combining task-level coarse-grained software pipelining with DVS (Dynamic Voltage Scaling) and DPM (Dynamic Power Management) considering transition overhead, inter-core communication and discrete voltage levels. We propose a two-phase approach to solve the problem. In the first phase, we propose a coarse-grained task parallelization algorithm called RDAG to transform a periodic dependent task graph into a set of independent tasks by exploiting the periodic feature of streaming applications. In the second phase, we propose a scheduling algorithm, GeneS, to optimize energy consumption. GeneS is a genetic algorithm that can search and find the best schedule within the solution space generated by gene evolution. We conduct experiments with a set of benchmarks from E3S and TGFF. The experimental results show that our approach can achieve a 24.4% reduction in energy consumption on average compared with the previous work.
Yi Wang 0003, Hui Liu 0006, Duo Liu 0002, Zhiwei Qin 0004, Zili Shao, Edwin H.-M. Sha
ACM Trans. Design Autom. Electr. Syst.1
2010 RNFTL: a reuse-aware NAND flash translation layer for flash memory
abstract
In this paper, we propose a hybrid-level flash translation layer (FTL) called RNFTL (Reuse-Aware NFTL) to improve the endurance and space utilization of NAND flash memory. Our basic idea is to prevent a primary block with many free pages from being erased in a merge operation. The preserved primary blocks are further reused as replacement blocks. In such a way, the space utilization and the number of erase counts for each block in NAND flash can be enhanced. To the best of our knowledge, this is the first work to employ a reuse-aware strategy in FTL for improving the space utilization and endurance of NAND flash. We conduct experiments on a set of traces that collected from real workload in daily life. The experimental results show that our technique has significant improvement on space utilization, block lifetime and wear-leveling compared with the previous work.
Yi Wang 0003, Duo Liu 0002, Meng Wang 0005, Zhiwei Qin 0004, Zili Shao
LCTES1
2010 Optimal Task Scheduling by Removing Inter-Core Communication Overhead for Streaming Applications on MPSoC
abstract
In this paper, we jointly optimize computation and communication task scheduling for streaming applications on MPSoC. The objective is to minimize schedule length by totally removing inter-core communication overhead. By minimizing schedule length, the system performance can be improved by adopting a smaller period or exploring the slacks generated for energy reduction with DVS. To guarantee the schedulability of communication tasks, we perform the schedulability analysis, and theoretically obtain the upper bound of the times needed to reschedule each computation task. Based on the analysis, we formulate the scheduling problem as an ILP (Integer Linear Programming) formulation and obtain an optimal solution. We evaluate our technique with a set of benchmarks from both real-life streaming applications and synthetic task graphs. The simulation results show that our technique can achieve a 27.72% reduction in schedule length and a 13.38% reduction in energy consumption on average compared with the previous work.
Yi Wang 0003, Duo Liu 0002, Meng Wang 0005, Zhiwei Qin 0004, Zili Shao
IEEE Real-Time and Embedded Technology and Applications Symposium1
2010 Memory-Aware Optimal Scheduling with Communication Overhead Minimization for Streaming Applications on Chip Multiprocessors
abstract
In this paper, we focus on solving the problem of removing inter-core communication overhead for streaming applications on chip multiprocessors. The objective is to totally remove inter-core communication overhead while minimizing the overall memory usage. By totally removing inter-core communication overhead, a shorter period can be applied and system throughput can be improved. Our basic idea is to let tasks with intra-period data dependencies transform to inter-period data dependencies so as to overlap the execution of computation and inter-core communication tasks. To solve the problem, we first perform analysis and obtain the bounds of the times needed to reschedule each task. Then we formulate the scheduling problem as an integer linear programming (ILP) model and obtain an optimal schedule. We perform simulations on a set of benchmarks from both real-life streaming applications and synthetic task graphs. The simulation results show that the proposed approach can achieve significant reduction in schedule length and improve the memory usage compared with the previous work.
Yi Wang 0003, Duo Liu 0002, Zhiwei Qin 0004, Zili Shao
RTSS1
2010 Compiler-assisted leakage-aware loop scheduling for embedded VLIW DSP processors
Meng Wang 0005, Yi Wang 0003, Duo Liu 0002, Zhiwei Qin 0004, Zili Shao
J. Syst. Softw.2
2009 Improving the Reliability of Embedded Systems with Cache and SPM
abstract
In this paper, we develop a compiler-assisted thermal-aware data allocation algorithm to improve the reliability of embedded systems with cache and SPM (scratch-pad memory). Our basic idea is to distribute the workload evenly between the cache and SPM in order to alleviate the temperature hot spots in the on-chip memory system. In the algorithm, considering the size of SPM, we first divide the loop iterations into two parts, and put the accessed data of the first part into SPM. Then we perform code transformation based on the partitioning of iterations. By alternatively using the data cache and SPM, the peak temperature is reduced. We implement our technique and simulate them using the Trimaran infrastructure with power models for cache and SPM, and the thermal simulator, HotSpot, on a set of benchmarks from DSPstone and MiBench. The experimental results show that our technique can significantly improve the reliability of the on-chip memory system.
Meng Wang 0005, Yi Wang 0003, Duo Liu 0002, Zili Shao
MASS2