Jen-Wei Hsieh

dblp:63/4379 · DBLP profile ↗
← Back
40ranked-venue papers
15as first author
8since 2021 · last 2025
0000-0002-1803-6947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 13 first-author · 7 since 2021Software engineering, systems software and programming languages · 4Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 AdaGray: An Energy-Efficient Adaptive Gray-Code Strategy for QLC Flash-Memory Storage Systems
abstract
In recent years, solid-state drives (SSDs) are gradually replacing traditional hard disk drives (HDDs) as the primary storage devices. It offers advantages such as shock resistance, higher speed, and a more compact size. As storage demands escalate, the concept of multi-level cells (MLC, TLC, QLC, etc.) has begun to emerge. While this approach increases capacity, it also introduces significant challenges: greater bits per cell lead to faster wear-out, longer read/write latency, and elevated energy consumption. To address these issues, the integration of various Gray codes into NAND flash memory encoding has shown promise. In this paper, we propose an Adaptive Gray code strategy (AdaGray), which leverages two distinct Gray code encoding schemes and dynamically allocates data into suitable coding blocks based on their characteristics. Experimental results demonstrate that AdaGray achieves a 24.31% reduction in write latency and lowers the erase count by up to 35.77%. More importantly, it significantly reduces the energy overhead from garbage collection by up to 45%, resulting in as much as 6982.8 mJ of energy savings.
Han-Yu Liao, Jen-Wei Hsieh, Yi-Shen Chen, Chang-Lin Tsai, Yuan-Hao Chang 0001
ISLPED2
2024 CellRejuvo: Rescuing the Aging of 3D NAND Flash Cells with Dense-Sparse Cell Reprogramming
abstract
3D NAND flash memory is one of the most important storage technologies in modern computer systems because of its non-volatile nature and excellent data access performance. However, it suffers from aging and reliability issues due to its inherent property. In contrast to the previous research that tried to recover the data with additional encoding techniques, we propose a novel reprogramming technique, called CellRejuvo, to improve the reliability of NAND flash cells. To the best of our knowledge, CellRejuvo is the pioneer for data recovery technique that cleverly leverages reprogramming to alleviate cell aging, extending the lifespan of solid-state drives. We implement CellRejuvo on a real 3D NAND flash-based SSD and evaluate its capability on various realistic workloads. The extensive experimental results show that CellRejuvo successfully reduces the error rate of SSD by an average of 38.28% under various retention times.
Han-Yu Liao, Yi-Shen Chen, Jen-Wei Hsieh, Yuan-Hao Chang 0001
ICCAD3
2024 CDS: Coupled Data Storage to Enhance Read Performance of 3D TLC NAND Flash Memory
abstract
Due to the strong demand of massive storage capacity, the density of flash memory has been improved in terms of technology node scaling, multi-bit per cell technique, and 3D stacking. However, these techniques also degrade read performance and reliability. The long read latency comes from increased data sensing time and time-consuming ECC decoding time. Storing multiple bits per cell results in more read reference voltages and increased latency of identifying appropriate threshold voltages. To deal with error correction, LDPC is widely used in flash memory to provide stronger ECC capability. However, LDPC incurs a long decoding latency when bit errors are numerous. In this work, we propose coupled data storage (CDS) to improve the read performance of 3D NAND flash-memory storage devices. CDS supports two modes to improve read latency: The high read-speed mode is designed to improve data sensing time with reduced voltage states, while the data correction mode is designed to mitigate bit errors and LDPC overhead. Experiment results showed that CDS could reduce 50$\sim$66.6% read latency and 25.7$\sim$27.5% write latency under the high read-speed mode. For the data correction mode, RBER could be decreased by 37$\sim$52% and the lifetime could be prolonged to 1.6 to 3 times.
Wan-Ling Wu, Jen-Wei Hsieh, Hao-Yu Ku
IEEE Trans. Computers2
2022 EMT: Elegantly Measured Tanner for Key-Value Store on SSD
abstract
With the emergence of big data era, NoSQL key-value database is considered as a promising candidate for replacing relational database management system (RDBMS). As cost per GB of flash memory is getting closer to HDD, high performance solid state drive (SSD) is regarded as the best substitute of HDD. However, there are still some issues that need to be taken into consideration.et al.have reported that applying key-value store to SSD with conventional FTL would incur internal fragmentation and further cause the degradation of device lifespan. Although they proposed KVFTL to deal with the above issues, their work gave rise to the read amplification problem. In this article, we investigate the root cause of the read amplification problem and propose elegantly measured tanner FTL (EMT-FTL) to mitigate internal fragmentation and read amplification with acceptable memory overhead. Different from KVFTL that slices a variable-sized value into multiple partitions (up to 20) of different sizes and manages them as a linked chain, EMT-FTL slices a value into 16 KiB full partitions with its remaining bytes treated as a fragment partition and appended to the fragment buffer. For a 64 GiB SSD with 16 KiB pages, the overall memory usage of EMT-FTL is only 8.81% of KVFTL. The experiments showed that EMT-FTL achieved almost the optimal space utilization as KVFTL did under most of traces and averagely improved thegetperformance by 78.57% (compared with KVFTL) under the traces with request sizes ranging from 1 byte to 16 KiB.
Tai Chang, Jen-Wei Hsieh, Tai-Chieh Chang, Liang-Wei Lai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Alternative Encoding: A Two-Step Transition Reduction Scheme for MLC STT-RAM Cache
abstract
Although multiple-level-cell (MLC) STT-RAM increases data density, it suffers from the two-step transition (TT) issue. It is because hard domain and soft domain of an MLC STT-RAM cell cannot be flipped to the opposite magnetization direction at the same time. Thus, the soft domain has to be flipped twice to the opposite magnetization direction of the hard domain. The TT problem hurts the lifetime of MLC STT-RAM due to additional flips on soft domains. To mitigate the TT problem of MLC STT-RAM, we propose an alternative encoding scheme (AES) to reduce the occurrence of TTs. AES utilizes the encoding method to eliminate most TTs and distributes unavoidable TTs among cells evenly to improve the lifetime of MLC STT-RAM cache. The experimental results showed that the proposed scheme achieved a great lifetime improvement than the related work.
Jen-Wei Hsieh, Yueh-Ting Hou, Tai-Chieh Chang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Differential Evolution Algorithm With Asymmetric Coding for Solving the Reliability Problem of 3D-TLC CT Flash-Memory Storage Systems
abstract
In recent years, NAND flash memory has been widely used in mobile devices, laptops, desktops, and data center storage systems due to its low-power consumption, high performance, high density, lightweight, shock resistance, and high-reliability natures. However, as the stacked layers and the storage density increase, flash memory also suffers from a higher raw bit error rate (RBER) and shorter lifetime. Observing that reliable cell states suffer from less data retention errors and program disturbance, we propose a differential evolution coding scheme to increase the probability of storing data in more reliable cell states, thereby reducing the RBER. We conducted the experiments over a development platform of SSD storage device with 3D-TLC charge trap (CT) NAND flash memory. The experimental results showed that the proposed differential evolution algorithm with asymmetric coding scheme could averagely reduce RBER by 48.88%, 65.45%, 52.61%, 61.99%, 80.19%, and 33.18% compared with baseline, asymmetric coding algorithm, asymmetric coding scheme with stripe-pattern elimination algorithm, UAC-$n$LC, word-line batch score modulation programming, and SCB schemes.
David Kuang-Hui Yu, Jen-Wei Hsieh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Read/Write Disturbance-Aware Design for MLC STT-RAM-based Cache
abstract
Spin-transfer torque RAM (STT-RAM) has been considered as a promising candidate for the next generation on-chip last-level cache (LLC) due to its high cell density, non-volatility, and near-zero standby power. To further improve cell density, multi-level cell (MLC) STT-RAM has been proposed and widely adopted. However, applying MLC STT-RAM to LLC might suffer from both write disturbance (WD) and read disturbance (RD). WD that needs two-step write operations to write data in MLC STT-RAM cell incurs extra energy consumption and latency overhead. RD means that reading data from a cell will also disturb the original data. In this paper, we propose a read/write disturbance-aware (RWDA) design for MLC STT-RAM-based cache to reduce the overhead caused by the WD and RD. We delay restore operations to mitigate the adverse impacts of disturbances. Instead of the typical LRU replacement policy, we propose a priority-based victim selection policy to meet the very distinct characteristics of MLC STT-RAM. Since accessing soft bits is much more beneficial than accessing hard bits in terms of access latency and energy consumption, we adopt a swapping mechanism to exchange frequently accessed data from hard bits to soft bits. The experimental results showed that the proposed design could averagely achieve 26.6% energy-consumption reduction and 29.5% IPC of system-performance improvement, compared with the conventional design of MLC STT-RAM cache.
Yao-Hung Huang, Jen-Wei Hsieh
RTCSA2
2021 TSE: Two-Step Elimination for MLC STT-RAM Last-Level Cache
abstract
Spin-transfer torque RAM (STT-RAM) is an emerging non-volatile memory that has been recognized as the potential candidate to replace SRAM. Compared with SRAM, STT-RAM has advantages of non-volatility, zero leakage power, and higher density. To further improve data density, multi-level cell (MLC) STT-RAM that can store two bits per cell has been proposed. However, writing hard bit of a cell would write its soft bit to the same value as well, which complicates the write operation of MLC STT-RAM. Although two-step transition (TT) is usually adopted to ensure the data correctness during a write operation, it incurs overhead of additional energy consumption and performance degradation. In this article, we propose the two-step elimination (TSE) scheme to eliminate TTs while ensure data integrity. By flipping hard bits of the cells that suffered from TTs, the TSE scheme could reduce TTs to soft transitions (STs) or zero transitions (ZTs), which incur much less overhead than TTs. To keep track of the flipped cells effectively, 6-bit TSE tag is introduced. We exploit tag reversing and advanced mode of the TSE scheme to further improve the performance. The experimental results showed that our scheme could reduce 61 percent TTs and achieve significant lifetime improvement, compared with conventional MLC STT-RAM (CMLC) scheme.
Jen-Wei Hsieh, Yi-Yu Liu, Hung-Tse Lee, Tai Chang
IEEE Trans. Computers1
2020 A Management Scheme of Multi-Level Retention-Time Queues for Improving the Endurance of Flash-Memory Storage Devices
abstract
As flash memory technology has been scaled down to 1x nm and more bits can be stored in a cell, the storage density of flash memory has been significantly improved. However, these technical trends also severely hurt the programming speed and endurance of flash memory. The internal data retention time is the duration for which a flash cell can correctly hold data. By relaxing internal data retention time, both the page programming speed and the block endurance could be improved. However, the retention time of flash memory typically requires to last for several years according to the industrial standard. Thus a refreshment scheme is required to deal with the decreasing of retention time. In this article, we propose multi-level retention-time queues with a management scheme to meet the retention-time requirement for a reliable storage system. Observing that many data are overwritten in hours or days in real workloads, multiple retention-time queues could effectively separate data with different update frequencies. There are three challenge issues for a proper design: (1) Since access pattern might change from time to time, a technical issue is how to promote/demote data so that data could be maintained in the proper retention-time queue to minimize the refreshment overhead. (2) Another technical issue is how to refresh each retention-time queue in time to guarantee data integrity. (3) Since blocks resided in different retention-time queue would suffer from different level of wearing, the third technical issue is how to estimate wearing status of flash-memory blocks in an effective and efficient manner to achieve wear leveling. In our scheme, data allocator, multi-level refresh module, garbage collector, and wear leveler are introduced to deal with these technical issues. Based on our experimental results, not only endurance and performance but also energy consumption of the flash-memory storage system could be significantly improved by our scheme.
David Kuang-Hui Yu, Jen-Wei Hsieh
IEEE Trans. Computers2
2019 Revive Bad Flash-Memory Pages by HLC Scheme
abstract
In recent years, flash memory has been widely used in embedded systems, portable devices, and high-performance storage products due to its nonvolatility, shock resistance, low power consumption, and high performance natures. To reduce the product cost, multi-level-cell (MLC) flash memory has been proposed; compared with the traditional single-level-cell (SLC) flash memory that only stores one bit of data per cell, each MLC cell can store two or more bits of data. Thus MLC can achieve a larger capacity and reduce the cost per unit. However, MLC also suffers from the degradation in both performance and reliability. In this paper, we try to enhance the reliability and reduce the product cost of flash-memory-based solid-state drive (SSD) from a totally different perspective. We propose the half-level-cell (HLC) scheme to manage and reuse the worn-out space in SSD; through our management scheme, the system can treat two bad pages as a normal page without sacrificing performance and reliability. The proposed scheme is purely on software/firmware-level, thus there is no need to change the hardware. The experiment results show that the lifetime of SSD with our proposed HLC scheme can be extended to 50.56% under the Windows workload and up to 65.45% under the multimedia workload. When we apply the HLC scheme to flash-memory cache of hybrid storage systems, the response time can be improved up to 20.57%.
Han-Yi Lin, Jen-Wei Hsieh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Introduction to the Special Issue on Real-Time aspects in Cyber-Physical Systems
abstract
No abstract available.
Luís Almeida 0001, Björn Andersson, Jen-Wei Hsieh, Li-Pin Chang, Xiaobo Sharon Hu
ACM Trans. Cyber Phys. Syst.3
2018 Retention-Time Relaxation Scheme for MLC Flash-Memory Storage Systems
abstract
As flash memory technology has been scaled to 1xnm and more bits can be stored in a cell, the storage density of flash memory has been significantly improved. However, these technical trends also severely hurt the programming speed and endurance of flash memory. The internal data retention time is the duration for which a flash cell can correctly hold data. By relaxing internal data retention time, both the page programming speed and the block endurance could be improved. However, the retention time of flash memory typically requires to last for several years according to the industrial standard. Thus a refreshment scheme is required to deal with the decreasing of retention time. In this paper, we present a retention-time relaxation scheme to meet the retention-time requirement for a reliable storage system. Observing that many data are overwritten in hours or days in real workloads, multiple retention-time queues are maintained to store data with different update frequency separately. Since access pattern might change from time to time, one technical issue is how to promote/demote data so that data could be maintained in the proper retention-time queue to minimize the refreshment overhead. Another technical issue is how to efficiently and effectively refresh each retention-time queue in time to guarantee data integrity. In our scheme, data allocator and multi-level refresh module are proposed to deal with these technical issues. Based on the experimental results, not only endurance and performance but also energy consumption of the flash-memory storage system could be significantly improved by our scheme.
David Kuang-Hui Yu, Jen-Wei Hsieh
RTCSA2
2015 HLC: software-based half-level-cell flash memory
Han-Yi Lin, Jen-Wei Hsieh
DATE2
2015 Adaptive ECC Scheme for Hybrid SSD's
abstract
In recent years, multi-level cell flash memory (MLC) has been widely adopted in solid state drives (SSD's) as the major storage medium due to its lower cost and higher density, compared with single-level cell flash memory (SLC). However, MLC has reliability concerns since it has lower endurance and higher disturb failure rate. Researchers thus proposed SLC/MLC hybrid SSD to exploit the advantages of SLC and MLC by separating frequently updated data in SLC and seldom modified data in MLC. In this paper, we propose an adaptive error correction code (ECC) scheme with four ECC levels to enhance the reliability of SSD's. Different from past researches, SLC is dedicated for the management of ECC, not for the user data. Since ECC is maintained in the data area of SLC (2 KB), rather than the spare area of MLC (128 Bytes), ECC capability is no longer confined by the limited size of the spare area. With adaptive management, ECC capability for a page would be upgraded whenever the current ECC cannot guarantee its reliability. A quantitative analysis is conducted to explore the impacts of different settings. The experiment results show that the lifetime of SSD can be extended by 318 percent for the trace of OLTP applications with our adaptive ECC scheme.
Jen-Wei Hsieh, Chung-Wei Chen, Han-Yi Lin
IEEE Trans. Computers1
2015 DCCS: Double Circular Caching Scheme for DRAM/PRAM Hybrid Cache
abstract
DRAM is widely adopted as a cache for secondary storage due to its small access latency. Compared with DRAM, PRAM draws a lot of attention recently, since it provides higher density and has no need to refresh the capacitor charge periodically. The non-volatile nature of PRAM can even reduce compulsory miss, which cannot be avoided by DRAM cache. However, PRAM cache cannot replace DRAM cache due to its endurance issue. Thus DRAM/PRAM hybrid cache becomes a good alternative for traditional DRAM cache. Least recently used (LRU) replacement algorithm and CLOCK-Pro algorithm work well for traditional DRAM cache. But these algorithms shall not be directly applied to DRAM/PRAM hybrid cache since the characteristics of PRAM are not considered. This paper proposed a double circular caching scheme (DCCS) to manage DRAM/PRAM hybrid cache. In our scheme, cached data migrate between DRAM cache and PRAM cache adaptively to achieve good hit ratio while frequent writes to PRAM cache are avoided for endurance concern. The experimental results showed that our scheme can reduce up to 87.10 percent PRAM write accesses for readintensive access pattern and up to 44.90 percent energy consumption for write-intensive access pattern, compared with other caching schemes.
Jen-Wei Hsieh, Yuan-Hung Kuan
IEEE Trans. Computers1
2015 Block-Based Multi-Version B+-Tree for Flash-Based Embedded Database Systems
abstract
In this paper, we propose a novel multi-version B$^+$-tree index structure, called block-based multi-version B$^+$-tree ( BbMVBT), for indexing multi-versions of data items in an embedded multi-version database (EMVDB ) on flash memory. An EMVDB needs to support streams of update transactions and version-range queries to access different versions of data items maintained in the database. In BbMVBT, the index is divided into two levels. At the higher level, a multi-version index is maintained for keeping successive versions of each data item. These versions are allocated consecutively in a version block. At the lower level, a version array is used to search for a specific data version within a version block. With the reduced index structure of BbMVBT, the overhead for managing the index in processing update operations can be greatly reduced. At the same time, BbMVBT can also greatly reduce the number of accesses to the index in processing version-range queries. To ensure sufficient free blocks for creating version blocks for efficient execution of BbMVBT, in this paper, we also discuss how to perform garbage collection using the purging-range queries for reclaiming “old” versions of data items and their associated entries in the index nodes. Analysis of the performance of BbMVBT is presented and verified with performance studies using both synthetic and real workloads. The performance results illustrate that BbMVBT can significantly improve the read and write performance to the multi-version index as compared with MVBT even though the sizes of the version blocks are not large.
Kam-yiu Lam, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang
IEEE Trans. Computers4
2014 Garbage collection for multi-version index on flash memory
abstract
In this paper, we study the important performance issues in using the purging-range query to reclaim old data versions to be free blocks in a flash-based multi-version database. To reduce the overheads for using the purging-range query in garbage collection, the physical block labeling (PBL) scheme is proposed to provide a better estimation on the purging version number to be used for purging old data versions. With the use of the frequency-based placement (FBP) scheme to place data versions in a block, the efficiency in garbage collection can be further enhanced by increasing the deadspans of data versions and reducing reallocation cost especially when the spaces of the flash memory for the databases are limited.
Kam-yiu Lam, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang, Chung Keung Poon, Chun Jiang Zhu
DATE4
2014 Garbage collection of multi-version indexed data on flash memory
Kam-yiu Lam, Chun Jiang Zhu, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang, Chung Keung Poon
J. Syst. Archit.4
2014 Multi-Channel Architecture-Based FTL for Reliable and High-Performance SSD
abstract
Several excellent researches have been proposed to improve the performance of solid-state drives (SSDs) by exploiting I/O parallelism of multi-channel architecture. However, these researches do not fully explore the internal parallelism and do not take wear leveling into consideration. In this paper, I/O performance is further improved by interleaving requests in channel level and striping sub-requests in plane level. A wear-leveling-aware distributed garbage collector is proposed to improve SSD lifetime and reclamation efficiency. To balance the utilization of user space among all channels, data migration is performed implicitly during channel selection and explicitly during garbage collection. To the best of our knowledge, this is the first paper on the design of distributed garbage collector for multi-channel flash-memory storage system. The experimental results showed that the proposed scheme can achieve good wear leveling and improve the overall performance by 34% for the Windows workload, 56.5% for the Linux workload, 88.4% for the multimedia workload, and 9.3% for the on-line transaction processing (OLTP) workload under the two-die-two-plane architecture, compared with the related work.
Jen-Wei Hsieh, Han-Yi Lin, Dong-Lin Yang
IEEE Trans. Computers1
2013 VAST: Virtually Associative Sector Translation for MLC Storage Systems
abstract
In recent years, multilevel cell Flash memory (MLC), which stores two or more bits per cell, has gradually replaced single-level cell flash memory due to its lower cost and higher density. However, MLC also brings new constraints, i.e., no partial programming and sequential page writes within a block, to the management. This paper proposes a virtual log-block-based hybrid-mapping scheme, referred to as virtually associative sector translation (VAST), for MLC storage systems. Unlike traditional hybrid-mapping schemes, VAST is a combination of block-level and segment-level mappings and manages log blocks in a flexible manner. The goals of our research are to avoid timeout by decreasing dummy-page writes, to get a better response time by decreasing live-page copies, and to prolong the life span of flash memory by decreasing total block erasures. Our trace-driven simulation shows that VAST could reduce up to 90% of dummy-page writes, 22%~52% of live-page copies, and 55%~83% of block erasures, compared to well-known hybrid-mapping schemes.
Jen-Wei Hsieh, Yu-Cheng Zheng, Yong-Sheng Peng, Po-Hung Yeh
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2013 Implementation strategy for downgraded flash-memory storage devices
abstract
In recent years, low-cost flash-memory devices have contributed greatly to the rapid growth of the flash memory market. Given that the most of the cost of such devices is the cost of the flash-memory chips, many vendors are managing the cost of such devices by using flash-memory chips of low quality, and they will continue to do so in the near future. Recognizing strong market demand, this work presents a set-based mapping strategy with an effective implementation and low hardware resource requirements for making downgraded flash-memory chips useable in products. A configurable management design for managing chips of various qualities with improved lifetime is presented. The effectiveness of the proposed strategy is evaluated by performing a series of experiments and analyzed with reference to popular implementations in industry.
Jen-Wei Hsieh, Yuan-Hao Chang 0001, Yuan-Sheng Chu
ACM Trans. Embed. Comput. Syst.1
2012 Double Circular Caching Scheme for DRAM/PRAM Hybrid Cache
abstract
DRAM is widely adopted as a cache for secondary storage due to its small access latency. Compared with DRAM, PRAM draws a lot of attention recently, since it provides higher density and has no need to refresh the capacitor charge periodically. The non-volatile nature of PRAM can even reduce compulsory miss, which cannot be avoided by DRAM cache. However, PRAM cache cannot replace DRAM cache due to its endurance issue. Thus DRAM/PRAM hybrid cache becomes a good alternative for traditional DRAM cache. Least recently used (LRU) replacement algorithm and CLOCK-Pro algorithm work well for traditional DRAM cache. But these algorithms shall not be directly applied to DRAM/PRAM hybrid cache since the characteristics of PRAM are not considered. This paper proposed a double circular caching scheme to manage DRAM/PRAM hybrid cache. In our scheme, cached data migrate between DRAM cache and PRAM cache adaptively to achieve good hit ratio while frequent writes to PRAM cache are avoided for endurance concern.
Jen-Wei Hsieh, Yuan-Hung Kuan
RTCSA1
2012 MFTL: A Design and Implementation for MLC Flash Memory Storage Systems
abstract
NAND flash memory has gained its popularity in a variety of applications as a storage medium due to its low power consumption, nonvolatility, high performance, physical stability, and portability. In particular, Multi-Level Cell (MLC) flash memory, which provides a lower cost and higher density solution, has occupied the largest part of NAND flash-memory market share. However, MLC flash memory also introduces new challenges: (1) Pages in a block must be written sequentially. (2) Information to indicate a page being obsoleted cannot be recorded in its spare area due to the limitation on the number of partial programming. Since most of applications access NAND flash memory under FAT file system, this article designs an MLC Flash Translation Layer (MFTL) for flash-memory storage systems which takes constraints of MLC flash memory and access behaviors of FAT file system into consideration. A series of trace-driven simulations was conducted to evaluate the performance of the proposed scheme. Although MFTL is designed for MLC flash memory and FAT file system, it is applicable to SLC flash memory and other file systems as well. Our experiment results show that the proposed MFTL could achieve a good performance for various access patterns even on SLC flash memory.
Jen-Wei Hsieh, Chung-Hsien Wu 0001, Ge-Ming Chiu
ACM Trans. Storage1
2011 An enhanced leakage-aware scheduler for dynamically reconfigurable FPGAs
abstract
The FPGAs (Field-Programmable Gate Array) are popular in hardware designs and even hardware/software co-designs. Due to the advance of manufacturing technologies, leakage power has become an important issue in the design of modern FPGAs. In particular, the partially dynamical reconfigurable FPGAs allow the latency between FPGA reconfiguration and task execution for the performance consideration. However, this latency introduces unnecessary leakage power called leakage waste. In this work, we propose a leakage-aware scheduling algorithm to minimize the leakage waste without increasing the schedule length of tasks. In this algorithm, a priority dispatcher with a split-aware placement is proposed to reduce the scheduling complexity with considering the hardware constraints of FPGAs. A series of experiments based on synthetic designs demonstrates that the proposed algorithm could effectively reduce leakage waste with limited sacrifices on the task schedulability.
Jen-Wei Hsieh, Yuan-Hao Chang 0001, Wei-Li Lee
ASP-DAC1
2011 Detecting Solid-State Disk Geometry for Write Pattern Optimization
abstract
Solid-state disks use flash memory as their storage medium, and adopt a firmware layer that makes data mapping and wear leveling transparent to the hosts. Even though solid-state disks emulate a collection of logical sectors, the I/O delays of accessing all these logical sectors are not uniform because the management of flash memory is subject to many physical constraints of flash memory. This work proposes a collection of black-box tests can detect the geometry inside of a solid-state disk. The host system software can arrange data in the logical disk space according to the detected geometry information to match the host write pattern with the device characteristic for reducing the flash management overhead in solid-state disks.
Chun-Chieh Kuo, Jen-Wei Hsieh, Li-Pin Chang
RTCSA (2)2
2010 Design and Implementation for Multi-level Cell Flash Memory Storage Systems
abstract
NAND flash memory has gained its popularity in a variety of applications as a storage medium due to its low power consumption, non-volatility, high performance, physical stability, and portability. In particular, Multi-Level Cell (MLC) flash memory, which provides a lower cost and higher density solution, has occupied the largest part of NAND flash-memory market share. However, MLC flash memory also introduces new challenges: (1) Pages in a block must be written sequentially. (2) Information to indicate a page being obsoleted cannot be recorded in its spare area. This paper designs an MLC Flash Translation Layer (MFTL) for flash-memory storage systems which takes new constraints of MLC flash memory and access behaviors of file system into consideration. A series of trace-driven simulations is conducted to evaluate the performance of the proposed scheme. Our experiment results show that the proposed MFTL outperforms other related works in terms of the number of extra page writes, the number of total block erasures, and the memory requirement for the management.
Jen-Wei Hsieh, Chung-Hsien Wu 0001, Ge-Ming Chiu
RTCSA1
2010 Improving Flash Wear-Leveling by Proactively Moving Static Data
abstract
Motivated by the strong demand for flash memory with enhanced reliability, this work attempts to achieve improved flash-memory endurance without substantially increasing overhead and without excessively modifying popular implementation designs such as the flash translation layer protocol (FTL), NAND flash translation layer protocol (NFTL), and block-level flash translation layer protocol (BL). A wear-leveling mechanism for moving data that are not updated is proposed to distribute wear-leveling actions over the entire physical address space, so that static or rarely updated data can be proactively moved and memory-space requirements can be minimized. The properties of the mechanism are then explored with various implementation considerations. A series of experiments based on a realistic trace demonstrates the significantly improved endurance of FTL, NFTL, and BL with limited system overhead.
Yuan-Hao Chang 0001, Jen-Wei Hsieh, Tei-Wei Kuo
IEEE Trans. Computers2
2010 A strategy to emulate NOR flash with NAND flash
abstract
This work is motivated by a strong market demand for the replacement of NOR flash memory with NAND flash memory to cut down the cost of many embedded-system designs, such as mobile phones. Different from LRU-related caching or buffering studies, we are interested in prediction-based prefetching based on given execution traces of application executions. An implementation strategy is proposed for the storage of the prefetching information with limited SRAM and run-time overheads. An efficient prediction procedure is presented based on information extracted from application executions to reduce the performance gap between NAND flash memory and NOR flash memory in reads. With the behavior of a target application extracted from a set of collected traces, we show that data access to NOR flash memory can respond effectively over the proposed implementation.
Yuan-Hao Chang 0001, Jian-Hong Lin, Jen-Wei Hsieh, Tei-Wei Kuo
ACM Trans. Storage3
2009 A set-based mapping strategy for flash-memory reliability enhancement
abstract
With wide applicability of flash memory in various application domains, reliability has become a very critical issue. This research is motivated by the needs to resolve the lifetime problem of flash memory and a strong demand in turning thrown-away flash-memory chips into downgraded products. We proposes a set-based mapping strategy with an effective implementation and low resource requirements, e.g., SRAM. A configurable management design and wear-leveling issue are considered. The behavior of the proposed method is also analyzed with respect to popular implementations in the industry.We show that the endurance of flash memory can be significantly improved by a series of experiments over a realistic trace. Our experiments show that the read performance is even largely improved.
Yuan-Sheng Chu, Jen-Wei Hsieh, Yuan-Hao Chang 0001, Tei-Wei Kuo
DATE2
2008 The Behavior Analysis of Flash-Memory Storage Systems
abstract
Performance and reliability are two major design concerns of flash-memory storage systems, especially for low-cost products. Although various excellent flash- memory management schemes are proposed, there is little work done on how to evaluate the designs or implementations of flash-memory storage systems. Many of the existing evaluation workloads for flash-memory storage systems still rely on those based on hard disks. This work aims at the needs of behavior analysis of flash-memory storage systems and their evaluations. In particular, a set of evaluation metrics and their corresponding access patterns are proposed. The behaviors of flash memory are also analyzed in terms of performance and reliability issues.
Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Jen-Wei Hsieh, Miller Lin
ISORC4
2008 Configurable Flash-Memory Management: Performance versus Overheads
abstract
Flash memory is widely adopted in various consumer products for information storage, especially for embedded systems. With strong demands on product designs for overhead control and performance requirements, vendors must have an effective design for the mapping of logical block addresses (LBA's) and physical addresses of data over flash memory. This paper targets such an essential issue by proposing a configurable mapping method that could trade the main-memory overhead with the system performance under the best needs of vendors. A series of experiments is conducted to provide insights on different configurations and the proposed method, compared to existing implementations.
Jen-Wei Hsieh, Yi-Lin Tsai, Tei-Wei Kuo, Tzao-Lin Lee
IEEE Trans. Computers1
2007 Endurance Enhancement of Flash-Memory Storage, Systems: An Efficient Static Wear Leveling Design
abstract
This work is motivated by the strong demand of reliability enhancement over flash memory. Our objective is to improve the endurance of flash memory with limited overhead and without many modifications to popular implementation designs, such as Flash Translation Layer protocol (FTL) and NAND Flash Translation Layer protocol (NFTL). A static wear leveling mechanism is proposed with limited memory-space requirements and an efficient implementation. The propreties of the mechanism are then explored with various implementation considerations. Through a series of experiments based on a realistic trace, we show that the endurance of FTL and NFTL could be significantly improved with limited system overheads.
Yuan-Hao Chang 0001, Jen-Wei Hsieh, Tei-Wei Kuo
DAC2
2007 Energy-efficient and performance-enhanced disks using flash-memory cache
abstract
This work explores the unique characteristics of flash memory in serving as a cache layer for disks. The experiments show that the proposed management scheme could save up to 20% energy consumption while reduce the read response time by the two third and the write response time by the five sixth of their counterparts. The estimated lifetime of the flash-memory cache is significantly improved as well.
Jen-Wei Hsieh, Tei-Wei Kuo, Po-Liang Wu, Yu-Chung Huang
ISLPED1
2007 A NOR Emulation Strategy over NAND Flash Memory
abstract
This work is motivated by a strong market demand in the replacement of NOR flash memory with NAND flash memory to cut down the cost in many embedded-system designs, such as mobile phones. Different from LRU-related caching or buffering studies, we are interested in prediction-based prefetching based on given execution traces of application executions. An implementation strategy is proposed in the storage of the prefetching information with limited SRAM and run-time overheads. An efficient prediction procedure is presented based on information extracted from application executions to reduce the performance gap between NAND flash memory and NOR flash memory in reads. With the behavior of a target application extracted from a set of collected traces, we show that data access to NOR flash memory can be responded effectively over the proposed implementation.
Jian-Hong Lin, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Tei-Wei Kuo, Cheng-Chih Yang
RTCSA3
2006 Configurability of performance and overheads in flash management
abstract
Flash memory has been widely considered as a good alternative for storage system implementations because it offers superior vibration tolerance and power efficiency, compared to hard-disks. Because of its unique characteristics, direct applications of disk management methods over flash memory might result in performance degradation and even the reducing of the lifetime. The management issues become even more challenging, especially when the capacity of flash memory increases significantly in the past few years. In this paper, we summarize our work on several important issues in flash memory management, where system performance and management overheads are considered. The capability of the proposed methodology was evaluated by a series of experiments to provide more insights in system designs
Tei-Wei Kuo, Jen-Wei Hsieh, Li-Pin Chang, Yuan-Hao Chang 0001
ASP-DAC2
2006 A faster exact schedulability analysis for fixed-priority scheduling
Wan-Chen Lu, Jen-Wei Hsieh, Wei-Kuan Shih, Tei-Wei Kuo
J. Syst. Softw.2
2006 Efficient identification of hot data for flash memory storage systems
abstract
Hot data identification for flash memory storage systems not only imposes great impacts on flash memory garbage collection but also strongly affects the performance of flash memory access and its lifetime (due to wear-levelling). This research proposes a highly efficient method for on-line hot data identification with limited space requirements. Different from past work, multiple independent hash functions are adopted to reduce the chance of false identification of hot data and to provide predictable and excellent performance for hot data identification. This research not only offers an efficient implementation for the proposed framework, but also presents an analytic study on the chance of false hot data identification. A series of experiments was conducted to verify the performance of the proposed method, and very encouraging results are presented.
Jen-Wei Hsieh, Tei-Wei Kuo, Li-Pin Chang
ACM Trans. Storage1
2003 Using FPGA to implement a n-channel arbitrary waveform generator with various add-on functions
abstract
This paper is a prototype design example to demonstrate the method to implement a PC/FPGA/spl I.bar/based n-channel arbitrary waveform generator (AWG) with add-on functions. Any arbitrary waveforms described by mathematical equations or piecewise-linear functions from the PC platform can be generated by FPGA/DDS technique. With USB interface, we can provide flexible n-channel waveform outputs. Add-on functions can be added by implementing specified control codes. Adopting 50 MHz clock frequency and 32-b phase accumulator word length, we have a 0.01164 Hz frequency resolution.
Jen-Wei Hsieh, Guo-Ruey Tsai, Min-Chuan Lin
FPT1
2003 The Design and Implementation of A Real-Time Data Dispatching System
abstract
This paper describes the design of a real-time data dispatching system (RTDDS), which is motivated by the needs of a performance guarantee on real-time data services by front-end clients. RTDDS consists of homogeneous machines with a quality of service performance guarantee. We address the resource allocation problems and consider Bin-Packing algorithms to assign requests to machines of RTDDS. Processes scheduling over each machine is done autonomously by SRP-based algorithms. The goal is to maximize the number of concurrent clients and to meet the individual quality of service requirements of clients at the same time.
Nei-Chiung Perng, Neung-Tsung Tsai, Jen-Wei Hsieh, Tei-Wei Kuo
ISORC3
2003 Resource Reservation and Enforcement for Framebuffer-Based Devices
Chung-You Wei, Jen-Wei Hsieh, Tei-Wei Kuo, I-Hsiang Lee, Yian-Nien Wu, Mei-Chin Tsa
RTCSA2