Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yu-Ming Chang

dblp:73/5531 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
2since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 5Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Storage systems · 80% Memory systems · 18% Hardware reliability and fault tolerance · 2%
Artificial intelligence
2 papers
Optimization for machine learning · 82% Learning theory · 18%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
flash and SSD
1.542023
Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash · IEEE Trans. Computers 2023
Achieving defect-free multilevel 3D flash memories with one-shot program design · DAC 2018
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › data reduction
data deduplication
0.912025
PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System · DAC 2025
Memory systems › processing-in-memory
near-data processing
0.912025
PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System · DAC 2025
Memory systems
processing-in-memory
0.912025
PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System · DAC 2025
Storage systems › flash and SSD
flash memory
0.942018
An SLC-Like Programming Scheme for MLC Flash Memory · ACM Trans. Storage 2018
Disturbance Relaxation for 3D Flash Memory · IEEE Trans. Computers 2016
Achieving SLC performance with MLC flash memory · DAC 2015
Storage systems › flash and SSD
data sanitization
0.712023
Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash · IEEE Trans. Computers 2023
Storage systems › flash and SSD › flash memory
NAND flash
0.712023
Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash · IEEE Trans. Computers 2023
Storage systems › flash and SSD › flash memory
threshold voltage distribution
0.712023
Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash · IEEE Trans. Computers 2023
Storage systems
storage reliability
0.432018
On Trading Wear-leveling with Heal-leveling · DAC 2014
Achieving defect-free multilevel 3D flash memories with one-shot program design · DAC 2018
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › flash and SSD
flash memory reliability
0.322018
Disturbance Relaxation for 3D Flash Memory · IEEE Trans. Computers 2016
An SLC-Like Programming Scheme for MLC Flash Memory · ACM Trans. Storage 2018
Storage systems › flash and SSD › flash memory › NAND flash
3D NAND flash
0.312017
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › flash and SSD
flash memory management
0.312017
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › flash and SSD › flash memory management
garbage collection
0.312017
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › storage reliability › durability
flash lifetime
0.322017
On Trading Wear-leveling with Heal-leveling · DAC 2014
VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory · DAC 2017
Storage systems › flash and SSD › flash memory
3d flash memory
0.212016
Disturbance Relaxation for 3D Flash Memory · IEEE Trans. Computers 2016
Storage systems › flash and SSD › flash memory management
flash translation layer
0.212015
Achieving SLC performance with MLC flash memory · DAC 2015
Storage systems › flash and SSD › flash memory
multi-level cell flash
0.212015
Achieving SLC performance with MLC flash memory · DAC 2015
Storage systems › flash and SSD › flash memory management
wear leveling
0.212015
Achieving SLC performance with MLC flash memory · DAC 2015
Machine learning › Optimization for machine learning › adaptive optimization
adaptive learning rate
0.222009
Periodic Step Size Adaptation for Single Pass On-line Learning · NIPS 2009
Training Conditional Random Fields by Periodic Step Size Adaptation for Large-Scale Text Mining · ICDM 2007
Machine learning › Optimization for machine learning
stochastic gradient descent
0.222009
Periodic Step Size Adaptation for Single Pass On-line Learning · NIPS 2009
Training Conditional Random Fields by Periodic Step Size Adaptation for Large-Scale Text Mining · ICDM 2007
Machine learning › Learning theory
online learning
0.112009
Periodic Step Size Adaptation for Single Pass On-line Learning · NIPS 2009
Machine learning › Optimization for machine learning
second-order optimization
0.112009
Periodic Step Size Adaptation for Single Pass On-line Learning · NIPS 2009

Methods — techniques the papers use, named apart from their topics

parallelization · 0.9DPU-friendly chunking · 0.9threshold voltage distribution merging · 0.7overwriting-based sanitization · 0.7threshold-voltage control · 0.5prophetic programming · 0.3one-shot program design · 0.3classification programming · 0.3SLC-like programming scheme · 0.3erase-free scheme · 0.3jacobian approximation · 0.1hessian approximation · 0.1stochastic gradient descent · 0.1online learning · 0.1conditional random field · 0.1
YearPublicationVenuePosition
2025 PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System
abstract
Data deduplication enhances storage efficiency through non-destructive compression but is often hindered by the chunking process, which requires scanning the entire dataset. While traditional methods leveraging conventional architectures and hardware accelerators (e.g., GPUs and FPGAs) have been developed to address this issue, they continue to face challenges related to excessive data movement and associated performance degradation. These limitations stem from the von Neumann architecture, where computation and storage are separated in a processor-centric design, necessitating multiple memory hierarchy traversals and causing inefficiencies. To overcome these challenges, we explore UPMEM’s DPU, a processing-in-memory (PIM) technology that reduces data movement by performing computations directly within memory. However, designing a deduplication system for DPUs presents unique obstacles, including restricted inter-DPU data sharing, the absence of native multiplication support, and significant DPU-CPU communication overhead. In response, we propose PIMDup, a DPU-optimized deduplication system that addresses these constraints through efficient parallelization, DPU-friendly chunking techniques, and reduced data transfer volumes. Experimental results demonstrate that PIMDup improves chunking performance without compromising deduplication accuracy, achieving a $1.67 \times$ speedup over CPU-based systems while maintaining 100% result consistency.
Chun-Le Yeh, Liang-Chi Chen, Chien-Chung Ho, Yu-Ming Chang, Da-Wei Chang
DAC4
2023 Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash
abstract
With the rising security concern on modern storage systems, the concept of data sanitization has been widely investigated recently. Among the existing works targeting data sanitization, an overwriting-based approach, namely one-shot sanitization, is one of the most efficient sanitization approaches. Nonetheless, we find that the one-shot sanitization approach would fail to achieve precise data sanitization for 3D TLC NAND flash, because of incurring undesired data errors. That is, how to simultaneously realize precise sanitization and high security with decent latency and reliability on emerging storage devices remains unsolved. This work proposes an elaborate sanitization design that skillfully manipulates the threshold voltage ($V_{t}$) distribution of sanitized pages. Not only does the proposed design achieve precise sanitization and high security, but it also enhances read performance and data reliability. Specifically, this work elaborately sanitizes data by merging specific$V_{t}$distributions of the target physical page on 3D TLC NAND flash. Besides, the proposed approach further takes lateral charge migration into consideration to improve data reliability. We conduct a series of experiments to evaluate our proposed approach on real 3D TLC NAND flash. The experiment results demonstrate the proposed approach can achieve elaborate data sanitization under various scenarios and improve read performance by 29%.
Wei-Chen Wang 0002, Chien-Chung Ho, Yung-Chun Li, Liang-Chi Chen, Yu-Ming Chang
IEEE Trans. Computers5
2019 Toward Instantaneous Sanitization through Disturbance-induced Errors and Recycling Programming over 3D Flash Memory
abstract
As data security has become one of the most crucial issues in modern storage system/application designs, the data sanitization techniques are regarded as the promising solution on 3D NAND flash-memory-based devices. Many excellent works had been proposed to exploit the in-place reprogramming, erasure and encryption techniques to achieve and implement the sanitization functionalities. However, existing sanitization approaches could lead to performance, disturbance overheads or even deciphered issues. Different from existing works, this work aims at exploring an instantaneous data sanitization scheme by taking advantage of programming disturbance properties. Our proposed design can not only achieve the instantaneous data sanitization by exploiting programming disturbance and error correction code properly, but also enhance the performance with the recycling programming design. The feasibility and capability of our proposed design are evaluated by a series of experiments on 3D NAND flash memory chips, for which we have very encouraging results. The experiment results show that the proposed design could achieve the instantaneous data sanitization with low overhead; besides, it improves the average response time and reduces the number of block erase count by up to 86.8% and 88.8%, respectively.
Wei-Chen Wang 0002, Ping-Hsien Lin, Yung-Chun Li, Chien-Chung Ho, Yu-Ming Chang, Yuan-Hao Chang 0001
ICCAD5
2019 Achieving Lossless Accuracy with Lossy Programming for Efficient Neural-Network Training on NVM-Based Systems
abstract
Neural networks over conventional computing platforms are heavily restricted by the data volume and performance concerns. While non-volatile memory offers potential solutions to data volume issues, challenges must be faced over performance issues, especially with asymmetric read and write performance. Beside that, critical concerns over endurance must also be resolved before non-volatile memory could be used in reality for neural networks. This work addresses the performance and endurance concerns altogether by proposing a data-aware programming scheme. We propose to consider neural network training jointly with respect to the data-flow and data-content points of view. In particular, methodologies with approximate results over Dual-SET operations were presented. Encouraging results were observed through a series of experiments, where great efficiency and lifetime enhancement is seen without sacrificing the result accuracy.
Wei-Chen Wang 0002, Yuan-Hao Chang 0001, Tei-Wei Kuo, Chien-Chung Ho, Yu-Ming Chang, Hung-Sheng Chang
ACM Trans. Embed. Comput. Syst.5
2019 mwJFS: A Multiwrite-Mode Journaling File System for MLC NVRAM Storages
abstract
At present, nonvolatile random access memory (NVRAM) is widely considered as a promising candidate for the next-generation storage medium due to its appealing characteristics, including short read/write latency, byte addressability, and low idle energy consumption. In addition, to provide a higher bit density, multilevel-cell (MLC) NVRAM has also been proposed. Nevertheless, when compared with conventional single-level-cell (SLC) NVRAM, MLC NVRAM has longer write latency and higher energy consumption. Hence, the performance of MLC NVRAM-based storage systems could be degraded due to the lengthened write latency. The performance degradation is further magnified by existing journaling file systems (JFS) on MLC NVRAM-based storage devices due to the JFS's fail-safe policy of writing the same data twice. Such observations motivate us to propose multiwrite-mode JFSs (mwJFSs) to alleviate the drawbacks of MLC NVRAM and boost the performance of MLC NVRAM-based JFS. The proposed mwJFS differentiates the data retention requirement of journaled data and applies different write modes to enhance the access performance with lower energy consumption. A series of experiments was conducted to demonstrate the capability of mwJFS on MLC NVRAM-based storage systems.
Shuo-Han Chen, Yuan-Hao Chang 0001, Yu-Ming Chang, Wei-Kuan Shih
IEEE Trans. Very Large Scale Integr. Syst.3
2018 Achieving defect-free multilevel 3D flash memories with one-shot program design
abstract
To store the desired data on MLC and TLC flash memories, the conventional programming strategies need to divide a fixed range of threshold voltage (Vt) window into several parts. The narrowly partitioned Vt window in turn limits the design of programming strategy and becomes the main reason to cause flash-memory defects, i.e., the longer read/write latency and worse data reliability. This motivates this work to explore the innovative programming design for solving the flash-memory defects. Thus, to achieve the defect-free 3D NAND flash memory, this paper presents and realizes a one-shot program design to significantly eliminate the negative impacts caused by conventional programming strategies. The proposed one-shot program design includes two strategies, i.e., prophetic and classification programming, for MLC flash memories, and the idea is extended to TLC flash memories. The measurement results show that it can accelerate programming speed by 31x and reduce RBER by 1000x for the MLC flash memory, and it can broaden the available window of threshold voltage up to 5.1x for the TLC flash memory.
Chien-Chung Ho, Yung-Chun Li, Yuan-Hao Chang 0001, Yu-Ming Chang
DAC4
2018 Achieving fast sanitization with zero live data copy for MLC flash memory
abstract
As data security has become the major concern in modern storage systems with low-cost multi-level-cell (MLC) flash memories, it is not trivial to realize data sanitization in such a system. Even though some existing works employ the encryption or the built-in erase to achieve this requirement, they still suffer the risk of being deciphered or the issue of performance degradation. In contrast to the existing work, a fast sanitization scheme is proposed to provide the highest degree of security for data sanitization; that is, every old version of data could be immediately sanitized with zero live-data-copy overhead once the new version of data is created/written. In particular, this scheme further considers the reliability issue of MLC flash memories; the proposed scheme includes a one-shot sanitization design to minimize the disturbance during data sanitization. The feasibility and the capability of the proposed scheme were evaluated through extensive experiments based on real flash chips. The results demonstrate that this scheme can achieve the data sanitization with zero live-data-copy, where performance overhead is less than 1%.
Ping-Hsien Lin, Yu-Ming Chang, Yung-Chun Li, Wei-Chen Wang 0002, Chien-Chung Ho, Yuan-Hao Chang 0001
ICCAD2
2018 The Impacts of an Academic English Competitive Mahjong Game on Learners' Motivation
Yu-Ming Chang, Sherry Y. Chen
ICCE1
2018 Enhancing the Energy Efficiency of Journaling File System via Exploiting Multi-Write Modes on MLC NVRAM
abstract
Non-volatile random-access memory (NVRAM) is regarded as a great alternative storage medium owing to its attractive features, including low idle energy consumption, byte addressability, and short read/write latency. In addition, multi-level-cell (MLC) NVRAM has also been proposed to provide higher bit density. However, MLC NVRAM has lower energy efficiency and longer write latency when compared with single-level-cell (SLC) NVRAM. These drawbacks could lead to higher energy consumption of MLC NVRAM-based storage systems. The energy consumption is magnified by existing journaling file systems (JFS) on MLC NVRAM-based storage devices due to the JFS's fail-safe policy of writing the same data twice. Such observations motivate us to propose a multi-write-mode journaling file systems (mwJFS) to alleviate the drawbacks of MLC NVRAM and lower the energy consumption of MLC NVRAM-based JFS. The proposed mwJFS differentiates the data retention requirement of journaled data and applies different write modes to enhance the energy efficiency with better access performance. A series of experiments was conducted to demonstrate the capability of mwJFS on a MLC NVRAM-based storage system.
Shuo-Han Chen, Yuan-Hao Chang 0001, Tseng-Yi Chen, Yu-Ming Chang, Pei-Wen Hsiao, Hsin-Wen Wei, Wei-Kuan Shih
ISLPED4
2018 Enhancing Flash Memory Reliability by Jointly Considering Write-back Pattern and Block Endurance
abstract
Owing to high cell density caused by the advanced manufacturing process, the reliability of flash drives turns out to be rather challenging in flash system designs. To enhance the reliability of flash drives, error-correcting code (ECC) has been widely utilized in flash drives to correct error bits during programming/reading data to/from flash drives. Although ECC can effectively enhance the reliability of flash drives by correcting error bits, the capability of ECC would degrade while the program/erase (P/E) cycles of flash blocks is increased. Finally, ECC could not correct a flash page, because a flash page contains too many error bits. As a result, reducing error bits is an effective solution to further improve the reliability of flash drives when a specific ECC is adopted in the flash drive. This work focuses on how to reduce the probability of producing error bits in a flash page. Thus, we propose a pattern-aware write strategy for flash reliability enhancement. The proposed write strategy considers both the P/E cycle of blocks and the pattern of written data while a flash block is allocated to store the written data. Since the proposed write strategy allocates young blocks (respectively, old blocks) for hot data (respectively, cold data) and flips the bit pattern of the written data to the appropriate bit pattern, the proposed strategy can effectively improve the reliability of flash drives. The experimental results show that the proposed strategy can reduce the number of error pages by up to 50%, compared with the well-known DFTL solution. Moreover, the proposed strategy is orthogonal with all ECC mechanisms so that the reliability of the flash drives with ECC mechanisms can be further improved by the proposed strategy.
Tseng-Yi Chen, Yuan-Hao Chang 0001, Yuan-Hung Kuan, Ming-Chang Yang, Yu-Ming Chang, Pi-Cheng Hsiu
ACM Trans. Design Autom. Electr. Syst.5
2018 An SLC-Like Programming Scheme for MLC Flash Memory
abstract
Although the multilevel cell (MLC) technique is widely adopted by flash-memory vendors to boost the chip density and lower the cost, it results in serious performance and reliability problems. Different from past work, a new cell programming method is proposed to not only significantly improve chip performance but also reduce the potential bit error rate. In particular, a single-level cell (SLC)-like programming scheme is proposed to better explore the threshold-voltage relationship to denote different MLC bit information, which in turn drastically provides a larger window of threshold voltage similar to that found in SLC chips. It could result in less programming iterations and simultaneously a much less reliability problem in programming flash-memory cells. In the experiments, the new programming scheme could accelerate the programming speed up to 742% and even reduce the bit error rate up to 471% for MLC pages.
Chien-Chung Ho, Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo
ACM Trans. Storage2
2017 VirtualGC: Enabling Erase-free Garbage Collection to Upgrade the Performance of Rewritable SLC NAND Flash Memory
abstract
Since 3D NAND flash memory could provide more reliable storage than a 2D planar flash memory by relaxing the design rule of a memory cell, a kind of brand new programming technique, namely erase-free scheme, has been proposed to further enhance the endurance of a 3D SLC NAND flash memory. The erase-free scheme brings tons of benefits to flash memory performance and endurance. For example, the erase-free scheme could reclaim invalid (page) space without physically erasing a flash block. However, current flash management designs could not fully exploit the benefits of the erase-free scheme. With the considerations of the features of the erase-free scheme, this paper is the first work to propose a novel flash management design, namely VirtualGC strategy, to deal with the erase-free garbage collection process. By taking the advantages of the erase-free scheme, the proposed strategy reduces the overhead of copying live pages so as to increase flash memory performance. The results show that the proposed strategy significantly improves the performance of rewritable 3D flash memory drives.
Tseng-Yi Chen, Yuan-Hao Chang 0001, Yuan-Hung Kuan, Yu-Ming Chang
DAC4
2016 Disturbance Relaxation for 3D Flash Memory
abstract
Even though 3D flash memory presents a grand opportunity to huge-capacity non-volatile memory, it suffers from serious program disturbance problems. In contrast to the past efforts in error correction codes and the work in trading the space utilization for reliability, we propose a disturbance-relaxation scheme that can alleviate the negative effects caused by program disturbance inside a physical block. This scheme does not introduce any extra overheads on encoding or storing of extra redundant data. In particular, a methodology is proposed to reduce the data error rate by distributing unavoidable disturbance errors to the flash-memory space of invalid data, with the considerations of the physical organization of 3D flash memory. A series of experiments was conducted based on real multi-layer 3D flash chips, and it showed that the proposed scheme could significantly enhance the reliability of 3D flash memory.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Yung-Chun Li, Hsiang-Pang Li
IEEE Trans. Computers1
2016 Improving PCM Endurance with a Constant-Cost Wear Leveling Design
abstract
Improving PCM endurance is a fundamental issue when it is considered as an alternative to replace DRAM as main memory. Memory-based wear leveling (WL) is an effective way to improve PCM endurance, but its major challenge is how to efficiently determine the appropriate memory pages for allocation or swapping. In this article, we present a constant-cost WL design that is compatible with existing memory management. Two implementations, namely bucket-based and array-based WL, with constant-time (or nearly zero) search cost are proposed to be integrated into the OS layer and the hardware layer, respectively, as well as to trade between time and space complexity. The results of experiments conducted based on an implementation in Android, as well as simulations with popular benchmarks, to evaluate the effectiveness of the proposed design are very encouraging.
Yu-Ming Chang, Pi-Cheng Hsiu, Yuan-Hao Chang 0001, Chi-Hao Chen, Tei-Wei Kuo, Cheng-Yuan Michael Wang
ACM Trans. Design Autom. Electr. Syst.1
2015 Achieving SLC performance with MLC flash memory
abstract
Although the Multi-Level-Cell technique is widely adopted by flash-memory vendors to boost the chip density and to lower the cost, it results in serious performance and reliability problems. Different from the past work, a new cell programming method is proposed to not only significantly improve the chip performance but also reduce the potential bit error rate. In particular, a Single-Level-Cell-like programming style is proposed to better explore the threshold-voltage relationship to denote different Multi-Level-Cell bit information, which in turn drastically provides a larger window of threshold voltage similar to that found in Single-Level-Cell chips. It could result in less programming iterations and simultaneously a much less reliability problem in programming flash-memory cells. In the experiments, the new programming style could accelerate the programming speed up to 742% and even reduce the bit error rate up to 471% for Multi-Level-Cell pages.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Yung-Chun Li, Hsiang-Pang Li
DAC1
2015 On Relaxing Page Program Disturbance over 3D MLC Flash Memory
abstract
With the rapidly-increasing capacity demand over flash memory, 3D NAND flash memory has drawn tremendous attention as a promising solution to further reduce the bit cost and to increase the bit density. However, such advanced 3D devices will suffer more intensive program disturbance, compared to 2D NAND flash memory. Especially when multi-level-cell (MLC) technology is adopted, the deteriorated disturbance due to the program operations of intra and inter pages will become even more critical for reliability. In contrast to the past efforts that try to resolve the reliability issue with error correction codes or hardware designs, this work seeks for the redesign of the program operation. A disturb-aware programming scheme is proposed to not only relax the disturbance induced by slow cells as much as possible but also reduce the possibility in requiring a high voltage to program the slow cells. A series of experiments was conducted based on real 3D MLC flash chips, and the results demonstrate that the proposed scheme is extremely effective on reducing the disturbance as well as the bit error rate.
Yu-Ming Chang, Yung-Chun Li, Yuan-Hao Chang 0001, Tei-Wei Kuo, Chih-Chang Hsieh, Hsiang-Pang Li
ICCAD1
2015 Energy stealing - an exploration into unperceived activities on mobile systems
abstract
Understanding the implications in smartphone usage and the power breakdown among hardware components has led to various energy-efficient designs for mobile systems. While energy consumption has been extensively explored, one critical dimension is often overlooked - unperceived activities that could steal a significant amount of energy behind users' back potentially. In this paper, we conduct the first exploration of unperceived activities in mobile systems. Specifically, we design a series of experiments to reveal, characterize, and analyze unperceived activities invoked by popular resident applications when an Android smartphone is left unused. We draw possible solutions inspired by the exploration and demonstrate that even an immediate remedy can mitigate energy dissipation to some extent.
Chi-Hsuan Lin, Yu-Ming Chang, Pi-Cheng Hsiu, Yuan-Hao Chang 0001
ISLPED2
2015 Read leveling for flash storage systems
abstract
Due to its several attractive benefits such as shock resistance, energy efficiency, and space-efficient form factor, flash memory is now applied to a wide range of electronics. Typically, since write requests are harmful to the health of flash memory, some flash-based storage devices tend to be deployed for read-intensive applications recently. However, as the technology node keeps going, read disturbance becomes a worsening problem in flash memory. Even under a pure read workload, flash memory often needs to refresh disturbed data, which brings about additionalwrite and erase operations. In this work, we propose a new design direction, read leveling, that aims at distributing read-hot data over different flash blocks. Thus, all read operations could be issued to different blocks as evenly as possible, so as to minimize the interference between read-hot data and other valid data on the same block and avoid refreshing cost. A series of experiments were conducted to prove the effectiveness of the proposed concept, and the results are very encouraging.
Chun-Yi Liu 0002, Yu-Ming Chang, Yuan-Hao Chang 0001
SYSTOR2
2014 On Trading Wear-leveling with Heal-leveling
abstract
Manufacturers are constantly seeking to increase flash memory density in order to fulfill the ever growing demand for storage capacity. However, this trend significantly reduces the reliability and endurance of flash memory chips. The lifetime degradation worsens as the number of erase cycles grows, even with wear leveling technology being adopted to extend flash memory lifetime by evenly distributing erase cycles to every flash block. To address this issue, self-healing technology is proposed to recover a flash block before the flash block is worn out, but such a technology still has its limitation when recovering flash blocks. In contrast to the existing wear leveling designs, we adopt the self-healing technology to propose a heal-leveling design that evenly distributes healing cycles to flash blocks. Ultimately, heal-leveling aims to extend the lifetime of flash memory without introducing a large amount of live-data copying overheads. We conducted a series of experiments to evaluate the capability of the proposed design. The results show that our design can significantly improve the access performance and the effective lifetime of flash memory without the unnecessary overheads caused by wear leveling technology.
Yu-Ming Chang, Yuan-Hao Chang 0001, Jian-Jia Chen, Tei-Wei Kuo, Hsiang-Pang Li, Hang-Ting Lue
DAC1
2014 Garbage collection and wear leveling for flash memory: Past and future
abstract
Recently, storage systems have observed a great leap in performance, reliability, endurance, and cost, due to the advance in non-volatile memory technologies, such as NAND flash memory. However, although delivering better performance, shock resistance, and energy efficiency than mechanical hard disks, NAND flash memory comes with unique characteristics and operational constraints, and cannot be directly used as an ideal block device. In particular, to address the notorious write-once property, garbage collection is necessary to clean the outdated data on flash memory. However, garbage collection is very time-consuming and often becomes the performance bottleneck of flash memory. Moreover, because flash memory cells endure very limited writes (as compared to mechanical hard disks) before they are worn out, the wear-leveling design is also indispensable to equalize the use of flash memory space and to prolong the flash memory lifetime. In response, this paper surveys state-of-the-art garbage collection and wear-leveling designs, so as to assist the design of flash memory management in various application scenarios. The future development trends of flash memory, such as the widespread adoption of higher-level flash memory and the emerging of three-dimensional (3D) flash memory architectures, are also discussed.
Ming-Chang Yang, Yu-Ming Chang, Che-Wei Tsao, Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
SMARTCOMP2
2013 A disturb-alleviation scheme for 3D flash memory
abstract
Even though 3D flash memory presents a grand opportunity for huge-capacity non-volatile memory, it suffers from serious program disturb problems. Different from the past efforts in error correction codes or the work in trading the space utilization with reliability, we propose a disturb-alleviation scheme that can alleviate the negative effects caused by program disturb, especially inside a block, without introducing extra overheads on encoding or storing of extra redundant data. In particular, a methodology is proposed to reduce the data error rate by distributing unavoidable disturb errors over the flash-memory space of invalid data, with the considerations of the physical organization of 3D flash memory. A series of experiments was conducted based on real multi-layer 3D flash chips, and it showed that the proposed scheme could significantly enhance the reliability of 3D flash memory.
Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Hsiang-Pang Li, Yung-Chun Li
ICCAD1
2013 A resource-driven DVFS scheme for smart handheld devices
abstract
Reducing the energy consumption of the emerging genre of smart handheld devices while simultaneously maintaining mobile applications and services is a major challenge. This work is inspired by an observation on the resource usage patterns of mobile applications. In contrast to existing DVFS scheduling algorithms and history-based prediction techniques, we propose a resource-driven DVFS scheme in which resource state machines are designed to model the resource usage patterns in an online fashion to guide DVFS. We have implemented the proposed scheme on Android smartphones and conducted experiments based on real-world applications. The results are very encouraging and demonstrate the efficacy of the proposed scheme.
Yu-Ming Chang, Pi-Cheng Hsiu, Yuan-Hao Chang 0001
ACM Trans. Embed. Comput. Syst.1
2009 Exploring GPS Data for Traffic Status Estimation
abstract
Traffic status plays an important role in navigation systems. To estimate traffic status, expensive sensors are deployed, which is not cost efficient. In view of the growth of GPS navigation services, in this paper, we propose two algorithms to estimate traffic status of road segments in our CarWeb platform. To evaluate our proposed algorithms, we implement the proposed algorithms in our CarWeb platform that is used to collect GPS data points from cars. Extensive experiments are conducted on real datasets and experimental results indicate that our algorithms can provide desirable predictions of traffic status.
Yu-Ming Chang, Ling-Yin Wei, Chun-Shuo Lin, Chen-Hen Jung, I-Hung Chen, Wen-Chih Peng
Mobile Data Management1
2009 Periodic Step Size Adaptation for Single Pass On-line Learning
abstract
It has been established that the second-order stochastic gradient descent (2SGD) method can potentially achieve generalization performance as well as empirical optimum in a single pass (i.e., epoch) through the training examples. However, 2SGD requires computing the inverse of the Hessian matrix of the loss function, which is prohibitively expensive. This paper presents Periodic Step-size Adaptation (PSA), which approximates the Jacobian matrix of the mapping function and explores a linear relation between the Jacobian and Hessian to approximate the Hessian periodically and achieve near-optimal results in experiments on a wide variety of models and tasks.
Chun-Nan Hsu, Yu-Ming Chang, Han-Shen Huang, Yuh-Jye Lee
NIPS2
2009 Global and componentwise extrapolations for accelerating training of Bayesian networks and conditional random fields
Han-Shen Huang, Bo-Hou Yang, Yu-Ming Chang, Chun-Nan Hsu
Data Min. Knowl. Discov.3
2009 Periodic step-size adaptation in second-order gradient descent for single-pass on-line structured learning
abstract
It has been established that the second-order stochastic gradient descent (SGD) method can potentially achieve generalization performance as well as empirical optimum in a single pass through the training examples. However, second-order SGD requires computing the inverse of the Hessian matrix of the loss function, which is prohibitively expensive for structured prediction problems that usually involve a very high dimensional feature space. This paper presents a new second-order SGD method, called Periodic Step-size Adaptation (PSA). PSA approximates the Jacobian matrix of the mapping function and explores a linear relation between the Jacobian and Hessian to approximate the Hessian, which is proved to be simpler and more effective than directly approximating Hessian in an on-line setting. We tested PSA on a wide variety of models and tasks, including large scale sequence labeling tasks using conditional random fields and large scale classification tasks using linear support vector machines and convolutional neural networks. Experimental results show that single-pass performance of PSA is always very close to empirical optimum.
Chun-Nan Hsu, Han-Shen Huang, Yu-Ming Chang, Yuh-Jye Lee
Mach. Learn.3
2008 Integrating high dimensional bi-directional parsing models for gene mention tagging
abstract
MOTIVATION: Tagging gene and gene product mentions in scientific text is an important initial step of literature mining. In this article, we describe in detail our gene mention tagger participated in BioCreative 2 challenge and analyze what contributes to its good performance. Our tagger is based on the conditional random fields model (CRF), the most prevailing method for the gene mention tagging task in BioCreative 2. Our tagger is interesting because it accomplished the highest F-scores among CRF-based methods and second over all. Moreover, we obtained our results by mostly applying open source packages, making it easy to duplicate our results. RESULTS: We first describe in detail how we developed our CRF-based tagger. We designed a very high dimensional feature set that includes most of information that may be relevant. We trained bi-directional CRF models with the same set of features, one applies forward parsing and the other backward, and integrated two models based on the output scores and dictionary filtering. One of the most prominent factors that contributes to the good performance of our tagger is the integration of an additional backward parsing model. However, from the definition of CRF, it appears that a CRF model is symmetric and bi-directional parsing models will produce the same results. We show that due to different feature settings, a CRF model can be asymmetric and the feature setting for our tagger in BioCreative 2 not only produces different results but also gives backward parsing models slight but constant advantage over forward parsing model. To fully explore the potential of integrating bi-directional parsing models, we applied different asymmetric feature settings to generate many bi-directional parsing models and integrate them based on the output scores. Experimental results show that this integrated model can achieve even higher F-score solely based on the training corpus for gene mention tagging. AVAILABILITY: Data sets, programs and an on-line service of our gene mention tagger can be accessed at http://aiia.iis.sinica.edu.tw/biocreative2.htm.
Chun-Nan Hsu, Yu-Ming Chang, Cheng-Ju Kuo, Yu-Shi Lin, Han-Shen Huang, I-Fang Chung
ISMB2
2007 Training Conditional Random Fields by Periodic Step Size Adaptation for Large-Scale Text Mining
abstract
For applications with consecutive incoming training examples, on-line learning has the potential to achieve a likelihood as high as off-line learning without scanning all available training examples and usually has a much smaller memory footprint. To train CRFson-line, this paper presents the Periodic Step size Adaptation (PSA) method to dynamically adjust the learning rates in stochastic gradient descent. We applied our method to three large scale text mining tasks. Experimental results show that PSA outperforms the best off-line algorithm, L-BFGS, by many hundred times, and outperforms the best on-line algorithm, SMD, by an order of magnitude in terms of the number of passes required to scan the training data set.
Han-Shen Huang, Yu-Ming Chang, Chun-Nan Hsu
ICDM2
2006 Enhancing Automatic Chinese Essay Scoring System from Figures-of-Speech
Tao-Hsing Chang, Chia-Hoang Lee, Yu-Ming Chang
PACLIC3
2003 Performance evaluation of ring-structure register file in multimedia applications
abstract
As the concurrent functional units in media processors increase continuously to meet the performance needs, the required access (i.e. read or write) ports of the centralized register file (RF) multiply rapidly and cannot be efficiently implemented. We propose a novel ring-structure RF, which is composed of register sub-blocks identical to the RF for a single functional unit. Data exchanges among functional units occur on the switch network of the ring registers. The proposed RF has been integrated into a four-way VLIW DSP processor successfully that demonstrates its effectiveness in DSP kernels. The synthesis result shows that our proposed ring-structure RF saves 91.88% silicon area of the centralized one, while reducing its access time by 77.35%.
Tay-Jyi Lin, Chin-Chi Chang, Tsung-Hsun Yang, Yu-Ming Chang, Chen-Chia Lee, Hung-Yueh Lin, Chein-Wei Jen
ICME4
1981 Design and Implementation of a Chinese Terminal Controller
Jong-Chuang Tsay, Yu-Ming Chang
Comput. Lang.2