Po-Chun Huang

dblp:43/1486 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-1076-2271ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 30 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 CAPR: Confidence-Aware Prompt Refinement in Large Language Models
Jen-Tzung Chien, Po-Chun Huang
INTERSPEECH2
2025 A Survey on Flash-Memory Storage Systems: A Host-Side Perspective
abstract
NAND flash memory has become the dominant storage media choice in a vast majority of application scenarios. Compared to mechanical hard disks, flash offers better access performance, energy efficiency, and shock resistance. However, the unique hardware peculiarities of this technology require dedicated facilities to manage the flash space and data. The implementation of flash management facilities has alternatively been realized either at the device or host computer level. Managing flash on the device side eases integration/compatibility and increases performance in certain scenarios. However, the limited computing resources inherent to devices and the lack of higher-level file system/application information make these solutions suboptimal in many situations. Managing flash on the host allows leveraging its abundant resources, and host-side knowledge such as data access patterns can be exploited to optimize flash management, at the cost of increased host-side complexity. The pros and cons of each approach also led to the appearance of hybrid, cross-layer solutions, enabling the collaboration of different layers of the storage stack. Recently, the pressure on modern storage systems requires that an increasing amount of flash management responsibilities is offloaded to the host, and the development of application-specific cross-layer solutions: In that context, it is crucial to review these developments. In this article, we make a comprehensive survey of the host-side management technologies of flash memory, application-/system-level flash-friendly designs, and emergent applications based on flash memory.
Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun-Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001
ACM Trans. Storage8
2024 PRESS: Persistence Relaxation for Efficient and Secure Data Sanitization on Zoned Namespace Storage : (Invited Paper)
abstract
Recently, secure data deletion or data sanitization has been identified as a key technology of storage devices to securely delete obsolete sensitive data that are no longer used. However, secure data deletion requires extra management efforts on flash memory storage devices, due to the deferred reclamation of flash blocks in many flash translation layer schemes. The emerging zoned namespace storage further exacerbates the design complexity of secure data deletion, due to the much larger size of a zone than that of a flash block. Concerning the very long latency to reset an entire zone, once some data have been written into a zone, it is very difficult to securely delete them from the zone. To achieve efficient and secure data deletion on zoned namespace storage, we propose persistence relaxation for efficient and secure sanitization (PRESS), which considers the working principle of zones of zoned namespace storage and allows the fine-grained control of deferred data persistence. As a result, applications can efficiently delete their recently written data or make the data persistent for long-term storage. Our proposal, PRESS, is evaluated through a series of experimental studies, where the results are quite encouraging.
Yun-Shan Hsieh, Bo-Jun Chen, Po-Chun Huang, Yuan-Hao Chang 0001
ASPDAC3
2024 FIRM-Tree: A Multidimensional Index Structure for Reprogrammable Flash Memory
abstract
For many emerging data-centric computing applications, it is a key capability to efficiently store, manage, and access multidimensional data. To achieve this, many multidimensional index data structures have been proposed. However, when existing multidimensional index data structures are maintained on modern nonvolatile memories (NVMs), such as NAND flash memory, they often face challenges in effective management of multidimensional data and handling of memory medium peculiarities, such as the write-once property and the need for block reclamation of NAND flash memory. Without appropriate management, these challenges often result in serious amplification of the read/write traffic, which degrades the performance of multidimensional data structures. Motivated by the urgent needs of efficient multidimensional index data structures on modern NVMs, we propose the FIRM-tree, a time-efficient and space-economic index data structure for multidimensional point data on NAND flash memory. Unique to the prior work, the FIRM-tree holistically utilizes RAM and flash memory space, and dedicatedly leverages the page reprogrammability of modern NAND flash memory, to enhance data access performance and flash management overheads. We then verify our proposal through analytical and experimental studies, where the results are quite encouraging.
Shin-Ting Wu, Pin-Jung Chen, Po-Chun Huang, Wei-Kuan Shih, Yuan-Hao Chang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 HF-Dedupe: Hierarchical Fingerprint Scheme for High Efficiency Data Deduplication on Flash-based Storage Systems
abstract
Even though flash memory is widely used in many applications as storage due to its high performance, demands for lower storage cost and better I/O performance are still high because of the continuous growth of data. Data deduplication has the potential to address these issues by eliminating redundant writes in I/O workloads and different strategies have been proposed to improve its efficiency. However, existing designs mainly rely on time-consuming SHA-1 fingerprint scheme or byte-by-byte comparison to identify duplicate data, and these methods cause much overhead and become a bottleneck in data deduplication. To tackle this issue, we propose the hierarchical fingerprint scheme (HF-Dedupe) to improve the efficiency of data deduplication for flash-based storage systems. By leveraging multiple levels of light-weight hashes in the fingerprint, our design only takes the minimal effort to distinguish different data in write traffic. In order to evaluate our design, a series of experiments were conducted based on trace-driven simulations. Compared with other designs, the experimental results show that HF-Dedupe further reduces the deduplication time by 34.76%-65.02 % while retaining high deduplication ratio, and therefore achieves the most improvement to overall I/O performance.
Kai-Ting Weng, Yun-Shan Hsieh, Yen-Ting Chen, Yu-Pei Liang, Yuan-Hao Chang 0001, Po-Chun Huang, Wei-Kuan Shih
ICCAD6
2023 A Novel Compact Current Driver Circuit with Temperature Feedback Control for 2D Nanophotonic Phased Arrays
abstract
This paper presents a compact driver circuit with independent pixel-level temperature regulation for thermo-optic based 2D Nanophotonic phased arrays (NPAs) in light detection and ranging (LIDAR) and Virtual Reality (VR) applications. To minimize the interconnection density, the proposed driver unit uses only a single electrical contact to its corresponding NPA pixel for both heating and temperature measurement functions. The driver was fabricated using TSMC 65 nm technology and each unit is realized in an area of$15\ \mu\mathrm{m}\times 15\ \mu\mathrm{m}$. The design enables scalable 3D heterogeneous integration between any tile-based NPA with pixel pitch below$15\ \mu\mathrm{m}$and its electrical control system. The temperature regulation performance of the proposed circuit was characterized by intentionally introducing a ±20% variation to the load resistance to simulate the temperature deviation in the NPA. The measured phase errors are suppressed by the feedback controller to a maximum of$0.07\pi$and an average of$0.02\pi$within the full$2\pi$phase shift operation range.
Po-Chun Huang, Xuetong Sun, Amitabh Varshney, Mario Dagenais, Martin Peckerar
ISCAS1
2023 WARM-tree: Making Quadtrees Write-efficient and Space-economic on Persistent Memories
abstract
Recently, the value of data has been widely recognized, which highlights the significance of data-centric computing in diversified application scenarios. In many cases, the data are multidimensional, and the management of multidimensional data often confronts greater challenges in supporting efficient data access operations and guaranteeing the space utilization. On the other hand, while many existing index data structures have been proposed for multidimensional data management, however, their designs are not fully optimized for modern nonvolatile memories, in particular the byte-addressable persistent memories. As a result, they might undergo serious access performance degradation or fail to guarantee space utilization. This observation motivates the redesigning of index data structures for multidimensional point data on modern persistent memories, such as the phase-change memory. In this work, we present the WARM-tree , a m ultidimensional t ree for r educing the w rite a mplification effect, for multidimensional point data. In our evaluation studies, as compared to the bucket PR quadtree and R*-tree, the WARM-tree can provide any worst-case space utilization guarantees in the form of \(\frac{m-1}{m}\) ( m ∈ ℤ^+) and effectively reduces the write traffic of key insertions by up to 48.10% and 85.86%, respectively, at the price of degraded average space utilization and prolonged latency of query operations. This suggests that the WARM-tree is a potential multidimensional index structure for insert-intensive workloads.
Shin-Ting Wu, Liang-Chi Chen, Po-Chun Huang, Yuan-Hao Chang 0001, Chien-Chung Ho, Wei-Kuan Shih
ACM Trans. Embed. Comput. Syst.3
2022 Enabling Efficient Random Data Insertion/Deletion on Block-Based File Systems
Yi-Han Lien, Yi-Hua Chen, Po-Chun Huang
IEEE Trans. Computers3
2021 Go Gig or Go Home: Enabling Social Sensing to Share Personal Data with Intimate Partner for the Health and Wellbeing of Long-Hour workers
abstract
Maintaining an awareness of one’s well-being and making work-related decisions to achieve work-life balance is critical for flexible long-hour workers. In this study, we propose that social sensing could address bottlenecks in worker’s awareness, interpretation of the informatics, and subsequent behavioral change. We conducted a four-week technology probe study by recruiting flexible long-hour professional drivers (Taxi and Uber drivers) and their significant others to use a social sensing prototype which collects data from the drivers and shares it with their partners as well as incorporates partners’ observations. We interviewed them before and after the probe study and found that while technological sensing was able to increase drivers’ awareness of their well-being status and intention to modify behaviors. The “social sensing” design was able to further shape such awareness or intention into action, highlighting the potential of using the sociotechnical approach in promoting work-life balance among long-hour workers.
Chuang-Wen You, Tina Chien-Wen Yuan, Nanyi Bi, Min-Wei Hung, Po-Chun Huang, Hao-Chuan Wang
CHI5
2021 Proximity Effect Correction for Fresnel Holograms on Nanophotonic Phased Arrays
abstract
Holographic displays and computer-generated holography offer a unique opportunity in improving optical resolutions and depth characteristics of near-eye displays. The thermally-modulated Nanopho-tonic Phased Array (NPA), a new type of holographic display, affords several advantages, including integrated light source and higher refresh rates, over other holographic display technologies. However, the thermal phase modulation of the NPA makes it susceptible to the thermal proximity effect where heating one pixel affects the temperature of nearby pixels. Proximity effect correction (PEC) methods have been proposed for 2D Fourier holograms in the far field but not for Fresnel holograms at user-specified depths. Here we extend an existing PEC method for the NPA to Fresnel holograms with phase-only hologram optimization and validate it through computational simulations. Our method is not only effective in correcting the proximity effect for the Fresnel holograms of 2D images at desired depths but can also leverage the fast refresh rate of the NPA to display 3D scenes with time-division multiplexing.
Xuetong Sun, Po-Chun Huang, Niloy Acharjee, Mario Dagenais, Martin Peckerar, Amitabh Varshney
VR3
2021 Making Frequent-Pattern Mining Scalable, Efficient, and Compact on Nonvolatile Memories
abstract
Frequent-pattern mining is a common means to reveal the hidden trends behind data. However, most frequent-pattern mining algorithms are designed for dynamic random-access memory (DRAM), instead of nonvolatile memories (NVMs) which are preferred by energy-limited systems. Due to the huge differences between the characteristics of NVMs and those of DRAM, existing frequent-pattern mining algorithms encounter the issues of write amplification and energy waste when they are run on NVMs. Moreover, the design complexity is exaggerated when parallel computing architecture is introduced to speedup the mining process. A scalable, time-efficient, and energy-economic solution to the frequent-pattern mining problem is thus urgently needed. Based on the well-known frequent-pattern tree (FP-tree) approach to frequent-pattern mining, this article proposes parallel EvFP-tree (PevFP-tree), a parallel frequent-pattern mining solution for NVMs. By considering the NVM characteristics, PevFP-tree accelerates the mining process and enhances the energy efficiency, as compared to a straightforward design of FP-trees on the parallel architecture. Moreover, PevFP-tree offers superior scalability in terms of the degrees of parallelism of the mining algorithm and the branching factor of its tree structure. Observing that keys are often sparsely distributed in FP-trees, we also propose a compression technique to PevFP-tree, namely, compressed PevFP-tree (CpevFP-tree), which further enhances the time and energy efficiencies of PevFP-tree. The proposed PevFP-tree and CpevFP-tree are evaluated by a series of experiments based on realistic datasets from diversified application scenarios, where CpevFP-tree achieves 88.73% of performance improvements over a straightforward design of FP-trees in the parallel architecture, and 79.47% of performance improvements over PevFP-tree, on average.
Chaoshu Yang, Po-Chun Huang, Duo Liu 0002, Yujuan Tan, Liang Liang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Probing User Perceptions of On-Skin Notification Displays
abstract
On-skin displays are emerging as a wearable form factor for the display of information; however, the perception of using such devices in public could determine whether they are eventually adopted or rejected. This study investigated the means by which on-skin notification displays are perceived by the general public. We adopted a mixed-methods approach to the analysis of results from an online survey (n = 254) and in-lab interviews (n = 36) pertaining to the novel form factor, device materiality, and envisioned use cases. The study was conducted in the US and Taiwan in order to examine cross-cultural attitudes toward device usage. The results of this structured examination provide valuable insights into the design of on-skin notification displays for everyday use across cultures.
Hsin-Liu Cindy Kao, Min-Wei Hung, Ximeng Zhang, Po-Chun Huang, Chuang-Wen You
Proc. ACM Hum. Comput. Interact.4
2020 Shift-Limited Sort: Optimizing Sorting Performance on Skyrmion Memory-Based Systems
abstract
Modern nonvolatile memories (NVMs) are widely recognized as energy-efficient replacements of classical memory/storage media, such as SRAM, DRAM, and mechanical hard disk. Among the popular NVMs, the skyrmion racetrack memory (SK-RM) is well known for its high storage density and unique supports of insert/delete operations. However, the existing algorithms designed for classical media might experience serious performance degradation when working on the SK-RM, due to the distinct characteristics of SK-RM. Thus, the existing algorithms should be redesigned to adapt to the brand-new memory model based on the SK-RM, so as to fully reveal the potentials of SK-RM. In particular, many existing algorithms tend to access the in-memory data in a random-hopping fashion, which generates many time-consuming shift operations of SK-RM. It is therefore crucial for the existing algorithms to eliminate unnecessary shift operations of SK-RM to boost the performance of the algorithms. In many modern applications, such as multimedia and data analysis, it is a common operation to process two or more arrays/vectors of data to perform certain computation tasks. In the arrays/vectors, an appropriate data placement strategy is critical for avoiding unnecessary shift operations of SK-RM. The observation thus motivates this work in proposing a recursive back-to-back data placement manner to effectively reduces the shift operations of SK-RM. To demonstrate the back-to-back data placement, we take sorting algorithms as a case study, and propose a novel shift-limited sorting algorithm for SK-RM. Analytical studies show that the shift-limited sort effectively enhances the time complexity of classical merge sort from O(dn lg n) to O(n lg n), where d is the bit distance between adjacent access ports on the nanotracks of the SK-RM. After that, the efficacy of the proposed shift-limited sort is then verified by experimental studies, where the results are encouraging.
Yun-Shan Hsieh, Po-Chun Huang, Ping-Xiang Chen, Yuan-Hao Chang 0001, Wang Kang 0001, Ming-Chang Yang, Wei-Kuan Shih
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 Downsizing Without Downgrading: Approximated Dynamic Time Warping on Nonvolatile Memories
abstract
In recent years, time-series data have emerged in a variety of application domains, such as wireless sensor networks and surveillance systems. To identify the similarity between time-series data, the Euclidean distance and its variations are common metrics that quantify the differences between time-series data. However, the Euclidean distance is limited by its inability to elastically shift with the time axis, which motivates the development of dynamic time warping (DTW) algorithms. While DTW algorithms have been proven very useful in diversified applications like speech recognition, their efficacy might be seriously affected by the resolution of the time-series data. However, high-resolution time-series data might take up a gigantic amount of main memory and storage space, which will slow down the DTW analysis procedure. This makes the upscaling of DTW analysis more challenging, especially for in-memory data analytics platforms with limited nonvolatile memory space. In this paper, we propose a strategy to downsample time-series data to significantly reduce their size without seriously affecting the precision of the results obtained by DTW algorithms (downsizing without downgrading). In other words, this paper proposes a technique to remove the unimportant details that are largely ignored by DTW algorithms. The efficacy of the proposed technique is verified by a series of experimental studies, where the results are quite encouraging.
Duo Liu 0002, Xingni Li, Po-Chun Huang, Yingjian Ling, Kan Zhong, Renping Liu 0002, Xianzhang Chen, Liang Liang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Correcting the Proximity Effect in Nanophotonic Phased Arrays
abstract
Thermally modulated Nanophotonic Phased Arrays (NPAs) can be used as phase-only holographic displays. Compared to the holographic displays based on Liquid Crystal on Silicon Spatial Light Modulators (LCoS SLMs), NPAs have the advantage of integrated light source and high refresh rate. However, the formation of the desired wavefront requires accurate modulation of the phase which is distorted by the thermal proximity effect. This problem has been largely overlooked and existing approaches to similar problems are either slow or do not provide a good result in the setting of NPAs. We propose two new algorithms based on the iterative phase retrieval algorithm and the proximal algorithm to address this challenge. We have carried out computational simulations to compare and contrast various algorithms in terms of image quality and computational efficiency. This work is going to benefit the research on NPAs and enable the use of large-scale NPAs as holographic displays.
Xuetong Sun, Po-Chun Huang, Niloy Acharjee, Mario Dagenais, Martin Peckerar, Amitabh Varshney
IEEE Trans. Vis. Comput. Graph.3
2018 VLSI design and implementation of a reconfigurable hardware-friendly Polar encoder architecture for emerging high-speed 5G system
Xin-Yu Shih, Po-Chun Huang, Hong-Ru Chou
Integr.2
2017 Scalable frequent-pattern mining on nonvolatile memories
abstract
Frequent-pattern mining is a common means to reveal the hidden trends behind data. However, most frequent-pattern mining algorithms are designed for DRAM, instead of the energy-economic nonvolatile memories (NVMs). Due to the huge differences between the characteristics of NVMs and those of DRAM, existing frequent-pattern mining algorithms suffer from serious overheads of write amplification or energy consumption as used on NVMs. The design complexity is exaggerated when parallel computing is used to speedup the mining process. This paper proposes PevFP-tree, a parallel frequent-pattern mining solution for NVMs, e.g., phase-change memory (PCM). By considering the NVM characteristics, PevFP-tree accelerates the mining process and enhance the energy efficiency. Moreover, PevFP-tree offers superior scalability in terms of the degree of parallelism of the mining algorithm and the branching factor of its tree structure. The efficacy of PevFP-tree is evaluated by experiments based on realistic datasets.
Po-Chun Huang, Duo Liu 0002, Liang Liang 0002
ASP-DAC2
2017 Durable and Energy Efficient In-Memory Frequent-Pattern Mining
abstract
It is a significant problem to efficiently identify the frequently occurring patterns in a given dataset, so as to unveil the trends hidden behind the dataset. This paper is motivated by the serious demands of a high-performance in-memory frequent-pattern mining strategy, with joint optimization over the mining performance and system durability. While the widely used frequent-pattern tree (FP-tree) serves as an efficient approach for frequent-pattern mining, its construction procedure often makes it unfriendly for nonvolatile memories (NVMs). In particular, the incremental construction of FP-tree could generate many unnecessary writes to the NVM and greatly degrade the energy efficiency, because NVM writes typically take more time and energy than reads. To overcome the drawbacks of FP-tree on NVMs, this paper proposes evergreen FP-tree (EvFP-tree), which includes a lazy counter and a minimum-bit-altered (MBA) encoding scheme to make FP-tree friendly for NVMs. The basic idea of the lazy counter is to greatly eliminate the redundant writes generated in FP-tree construction. On the other hand, the MBA encoding scheme is to complement existing wear-leveling techniques to evenly write each memory cell to extend the NVM lifetime. As verified by experiments, EvFP-tree greatly enhances the mining performance and system lifetime by 40.28% and 87.20% on average, respectively. And EvFP-tree reduces the energy consumption by 50.30% on average.
Duo Liu 0002, Po-Chun Huang, Liang Liang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 Multi-version checkpointing for flash file systems
abstract
Reliability has become a critical design issue in flash storage systems, because of the adoption of the low-cost, high-error-rate flash chips to fulfill the needs of the fast-growing storage capacity. In this paper, a multi-version checkpointing strategy is proposed to resolve the reliability issue of flash storage systems from the perspective of flash file systems. The proposed strategy can efficiently and effectively utilize checkpoints of file systems to guarantee the integrity and consistency of flash file systems after files or flash pages are corrupted. By utilizing the coexistence fact of multiple versions of the same data in flash memory, a control/recovery mechanism is presented to maintain checkpoints and to recover file systems with minimized management and recovery time overheads. A series of experiments was conducted based on realistic traces that were collected from benchmarks running over flash file systems in Linux operating systems. The results illustrate that the proposed strategy can significantly improve the reliability of flash file systems, as compared with other existing designs.
Shih-Chun Chou, Yuan-Hao Chang 0001, Yuan-Hung Kuan, Po-Chun Huang, Che-Wei Tsao
ASP-DAC4
2016 Relay-based key management to support secure deletion for resource-constrained flash-memory storage devices
abstract
The support of secure deletion on formatting a file system is to make sure that when a file system is formatted, there is no way to get any file content back again. Due to the fast-growing storage capacity, the performance of secure deletion to file systems on resource-constrained flash storage devices has become a critical issue. In contrast to the existing works that take a long time on overwriting/resetting all the file contents of a file system, we propose an efficient secure deletion scheme to securely delete all the contents of a file system without rewriting file contents. Thus, secure deletion to file systems can be efficiently achieved and can be independent of the device capacity and file systems. A series of experiments was conducted with realistic workloads to evaluate the capability of the proposed scheme. The results show that the proposed scheme achieves secure deletion with limited performance overheads in most cases.
Wei-Lin Wang, Yuan-Hao Chang 0001, Po-Chun Huang, Chia-Heng Tu, Hsin-Wen Wei, Wei-Kuan Shih
ASP-DAC3
2016 Making In-Memory Frequent Pattern Mining Durable and Energy Efficient
abstract
It is a significant problem to efficiently identifythe frequently-occurring patterns in a given dataset, so as tounveil the trends hidden behind the dataset. This work ismotivated by the serious demands of a high-performance inmemoryfrequent-pattern mining strategy, with joint optimizationover the mining performance and system durability. While thewidely-used frequent-pattern tree (FP-tree) serves as an efficientapproach for frequent-pattern mining, its construction procedureoften makes it unfriendly for nonvolatile memories (NVMs). Inparticular, the incremental construction of FP-tree could generatemany unnecessary writes to the NVM and greatly degrade theenergy efficiency, because NVM writes typically take more timeand energy than reads. To overcome the drawbacks of FP-treeon NVMs, this paper proposes evergreen FP-tree (EvFP-tree), which includes a lazy counter and a minimum-bit-altered (MBA) encoding scheme to make FP-tree friendly for NVMs. The basicidea of the lazy counter is to greatly eliminate the redundantwrites generated in FP-tree construction. On the other hand, theMBA encoding scheme is to complement existing wear-levelingtechniques to evenly write each memory cell to extend the NVMlifetime. As verified by experiments, EvFP-tree greatly enhancesthe mining performance and system lifetime by 28.01% and82.10% on average, respectively.
Po-Chun Huang, Duo Liu 0002, Liang Liang 0002
ICPP2
2016 Capacity-Independent Address Mapping for Flash Storage Devices with Explosively Growing Capacity
abstract
Address mapping for flash storage devices has been a challenging design issue for controllers because of rapidly growing device capacity. In contrast with existing mapping methods, this study proposes a capacity-independent address mapping method to decouple the required on-device RAM space from the capacity of a flash storage device. Especially, the required RAM size of the proposed method depends only on the user accessed data set, which is also referred to as the working set, while the page-level performance can be nearly achieved. In addition, a simple but practical wear-leveling design is proposed with the capability in lifetime estimation of flash storage devices. Experiments of the proposed scheme obtained encouraging results.
Ming-Chang Yang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Po-Chun Huang
IEEE Trans. Computers4
2016 Space-Efficient Index Scheme for PCM-Based Multiversion Databases in Cyber-Physical Systems
Yuan-Hung Kuan, Yuan-Hao Chang 0001, Tseng-Yi Chen, Po-Chun Huang, Kam-yiu Lam
ACM Trans. Embed. Comput. Syst.4
2015 Efficient Victim Block Selection for Flash Storage Devices
abstract
Motivated by the needs to enhance the performance of garbage collection in low-cost flash storage devices, we propose a victim block selection design to efficiently identify the blocks for erases and reclaim the space of invalid data without extensively scanning flash memory for the data status stored in the storage, so as to improve the garbage collection performance on reclaiming the space of invalid data. Moreover, this design could easily identify and reclaim the space released by file systems. Experiments based on benchmark traces show significant performance improvement of garbage collection with limited system overheads.
Che-Wei Tsao, Yuan-Hao Chang 0001, Ming-Chang Yang, Po-Chun Huang
IEEE Trans. Computers4
2015 Block-Based Multi-Version B+-Tree for Flash-Based Embedded Database Systems
abstract
In this paper, we propose a novel multi-version B$^+$-tree index structure, called block-based multi-version B$^+$-tree ( BbMVBT), for indexing multi-versions of data items in an embedded multi-version database (EMVDB ) on flash memory. An EMVDB needs to support streams of update transactions and version-range queries to access different versions of data items maintained in the database. In BbMVBT, the index is divided into two levels. At the higher level, a multi-version index is maintained for keeping successive versions of each data item. These versions are allocated consecutively in a version block. At the lower level, a version array is used to search for a specific data version within a version block. With the reduced index structure of BbMVBT, the overhead for managing the index in processing update operations can be greatly reduced. At the same time, BbMVBT can also greatly reduce the number of accesses to the index in processing version-range queries. To ensure sufficient free blocks for creating version blocks for efficient execution of BbMVBT, in this paper, we also discuss how to perform garbage collection using the purging-range queries for reclaiming “old” versions of data items and their associated entries in the index nodes. Analysis of the performance of BbMVBT is presented and verified with performance studies using both synthetic and real workloads. The performance results illustrate that BbMVBT can significantly improve the read and write performance to the multi-version index as compared with MVBT even though the sizes of the version blocks are not large.
Kam-yiu Lam, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang
IEEE Trans. Computers5
2014 Space-Efficient Multiversion Index Scheme for PCM-based Embedded Database Systems
abstract
Embedded database systems are widely adopted in various control and motoring systems, e.g., cyber-physical systems (CPSes). To support the functionality to access the historical data, a multiversion index is adopted to maintain multiple versions of data items and their index information. However, CPSes are usually battery-powered embedded systems that have limited energy, computing power, and storage space. In this work, we consider the systems with phase-change memory (PCM) as their storage due to its non-volatility and low energy consumption. In order to resolve the problem of the limited storage space and the fact that existing multiversion index designs are lack of space efficiency, we propose a space-efficient multiversion index scheme to enhance the space utilization and access performance of embedded multiversion database systems on PCM by utilizing the byte-addressability and write asymmetry of PCM. A series of experiments was conducted to evaluate the efficacy of the proposed scheme. The results show that the proposed scheme achieves very high space utilization and has good performance on serving update transactions and range queries.
Yuan-Hung Kuan, Yuan-Hao Chang 0001, Po-Chun Huang, Kam-yiu Lam
DAC3
2014 Garbage collection for multi-version index on flash memory
abstract
In this paper, we study the important performance issues in using the purging-range query to reclaim old data versions to be free blocks in a flash-based multi-version database. To reduce the overheads for using the purging-range query in garbage collection, the physical block labeling (PBL) scheme is proposed to provide a better estimation on the purging version number to be used for purging old data versions. With the use of the frequency-based placement (FBP) scheme to place data versions in a block, the efficiency in garbage collection can be further enhanced by increasing the deadspans of data versions and reducing reallocation cost especially when the spaces of the flash memory for the databases are limited.
Kam-yiu Lam, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang, Chung Keung Poon, Chun Jiang Zhu
DATE5
2014 Current-aware scheduling for flash storage devices
abstract
As NAND flash memory has become a major choice of storage media in diversified computing environments, the performance issue of flash memory has been extensively addressed in many excellent designs. Among them, an effective strategy is to adopt multiple channels and flash-memory chips to improve the performance on data accesses. However, the degree of data access parallelism cannot be increased by simply increasing the number of channels and chips in the storage device, because it is seriously limited by the maximum current constraint of the bus interface and affected by the access patterns of user data. As a consequence, to maximize the degree of access parallelism, it is of paramount significance to have a proper scheduling strategy to determine the order that read/write requests are served. In this paper, a current-aware scheduling strategy for read/write requests is proposed to maximize the read performance without violating the bus current constraint and without missing (the deadline of) written data. The proposed strategy is then evaluated through a series of experiments, in which the results are quite encouraging.
Tzu-Jung Huang, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
RTCSA3
2014 Garbage collection and wear leveling for flash memory: Past and future
abstract
Recently, storage systems have observed a great leap in performance, reliability, endurance, and cost, due to the advance in non-volatile memory technologies, such as NAND flash memory. However, although delivering better performance, shock resistance, and energy efficiency than mechanical hard disks, NAND flash memory comes with unique characteristics and operational constraints, and cannot be directly used as an ideal block device. In particular, to address the notorious write-once property, garbage collection is necessary to clean the outdated data on flash memory. However, garbage collection is very time-consuming and often becomes the performance bottleneck of flash memory. Moreover, because flash memory cells endure very limited writes (as compared to mechanical hard disks) before they are worn out, the wear-leveling design is also indispensable to equalize the use of flash memory space and to prolong the flash memory lifetime. In response, this paper surveys state-of-the-art garbage collection and wear-leveling designs, so as to assist the design of flash memory management in various application scenarios. The future development trends of flash memory, such as the widespread adoption of higher-level flash memory and the emerging of three-dimensional (3D) flash memory architectures, are also discussed.
Ming-Chang Yang, Yu-Ming Chang, Che-Wei Tsao, Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
SMARTCOMP4
2014 Garbage collection of multi-version indexed data on flash memory
Kam-yiu Lam, Chun Jiang Zhu, Yuan-Hao Chang 0001, Jen-Wei Hsieh, Po-Chun Huang, Chung Keung Poon
J. Syst. Archit.5
2014 Garbage Collection for Multiversion Index in Flash-Based Embedded Databases
abstract
Recently, flash-based embedded databases have gained their momentum in various control and monitoring systems, such as cyber-physical systems (CPSes). To support the functionality to access the historical data, a multiversion index is adopted to simultaneously maintain multiple versions of data items, as well as their index information. However, maintaining a multiversion index on flash memory incurs considerable performance overheads on garbage collection, which is to reclaim the spaces occupied by the outdated/invalid data items and their index information on flash memory. In this work, we propose an efficient garbage collection strategy to solve the garbage collection issues of flash-based multiversion databases. In particular, a version-tracking method is proposed to accelerate the performance on the process on identifying/reclaiming the space of invalid data and their indexes, and a pre-summary method is also designed to solve the cascading update problem that is caused by the write-once nature of flash memory and is worsened when more versions refer to the same data item. The capability of the proposed strategy is then verified by analytical and experimental studies.
Po-Chun Huang, Yuan-Hao Chang 0001, Kam-yiu Lam, Chien-Chin Huang
ACM Trans. Design Autom. Electr. Syst.1
2013 New ERA: new efficient reliability-aware wear leveling for endurance enhancement of flash storage devices
abstract
As the program/erase (P/E) cycles of flash memory keep decreasing, improving the lifetime/endurance of flash memory has become a fundamental issue in the design of flash devices. This work is motivated by the observation that flash blocks endured the same P/E cycles usually have different bit error rates. In contrast to the existing wear-leveling techniques that try to distribute erases to flash blocks as evenly as possible, we propose an efficient reliability-aware wear-leveling scheme to distribute block erases based on the bit error rates of blocks so as to even out the error rate among flash blocks, to maximize the number of good blocks, and thus to ultimately prolong the lifetime of flash storage devices. The experiments were conducted based on representative realistic workloads to evaluate the efficacy of the proposed scheme, for which the results are very encouraging.
Ming-Chang Yang, Yuan-Hao Chang 0001, Che-Wei Tsao, Po-Chun Huang
DAC4
2013 Reliability Enhancement of Flash-Memory Storage Systems: An Efficient Version-Based Design
abstract
In recent years, reliability has become one critical issue in the designs of flash-memory file/storage systems, due to the growing unreliability of advanced flash-memory chips. In this paper, a version-based design is proposed to effectively and efficiently maintain the consistency among page versions of a file for potential recovery needs. In particular, a two-version one for a native file system is presented with the minimal overheads in version maintenance. A recovery scheme is then presented to restore a corrupted file back to the latest consistent version. The design is later extended to maintain multiple data versions with the considerations of the write constraints of multilevel-cell flash memory. It was shown that the proposed design could significantly improve the reliability of flash memory with limited management and space overheads.
Yuan-Hao Chang 0001, Po-Chun Huang, Pei-Han Hsu, Lue-Jane Lee, Tei-Wei Kuo, David Hung-Chang Du
IEEE Trans. Computers2
2013 An index-based management scheme with adaptive caching for huge-scale low-cost embedded flash storages
abstract
Due to its remarkable access performance, shock resistance, and costs, NAND flash memory is now widely adopted in a variety of computing environments, especially in mobile devices such as smart phones, media players and electronic book readers. For the consideration of costs, low-cost embedded flash storages such as flash memory cards are often employed on such devices. Different from solid-state disks, the RAM buffer equipped on low-cost embedded flash storages are very small, for example, limited under several dozens of kilobytes, despite of the rapidly growing capacity of the storages. The significance of effectively utilizing the very limited on-device RAM buffers of embedded flash storages is therefore highlighted, and a novel design of scalable flash management schemes is needed to tackle the new access constraints of MLC NAND flash memory. In this work, a highly scalable design of the flash translation layer is presented with the considerations of the on-device RAM size, user access patterns, address-mapping-information caching and MLC access constraints. Through a series of experiments, it is verified that, with appropriate settings of cache sizes, the proposed management scheme provides comparable performance results to prior arts with much lower requirements on the on-device RAM. In other words, the proposed scheme suggests a strategy to make better use of the on-device RAM, and is suitable for embedded flash storages.
Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
ACM Trans. Design Autom. Electr. Syst.1
2012 Joint management of RAM and flash memory with access pattern considerations
abstract
The popularity of flash memory has triggered the emerging of various products with flash memory as storage medium. More advanced architectures with better hardware resources are now explored by vendors to fit different market needs. Different from the past work, this paper proposes to consider RAM as a storage medium together with flash memory to take advantage of the characteristics of both RAM and flash memory. In particular, an adaptive management strategy is proposed with the considerations of access patterns to improve both the system performance and the system endurance. The capability of the proposed approach is evaluated by a series of experiments, for which we have very encouraging results.
Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
DAC1
2012 A caching-oriented management design for the performance enhancement of solid-state drives
abstract
While solid-state drives are excellent alternatives to hard disks in mobile devices, a number of performance and reliability issues need to be addressed. In this work, we design an efficient flash management scheme for the performance improvement of low-cost MLC flash memory devices. Specifically, we design an efficient flash management scheme for multi-chipped flash memory devices with cache support, and develop a two-level address translation mechanism with an adaptive caching policy. We evaluated the approach on real workloads. The results demonstrate that it can improve the performance of multi-chipped solid-state drives through logical-to-physical mappings and concurrent accesses to flash chips.
Yuan-Hao Chang 0001, Cheng-Kang Hsieh, Po-Chun Huang, Pi-Cheng Hsiu
ACM Trans. Storage3
2011 A version-based strategy for reliability enhancement of flash file systems
abstract
In recent years, reliability has become one critical issue in the designs of flash file systems due to the growing unreliability of advanced flash-memory chips. In this paper, a version-based strategy with optimal space utilization is proposed to maintain the consistency among page versions of a file for potential recovery needs with the considerations of the write constraints of multi-level-cell flash memory. A series of experiments was conducted to show that the proposed strategy could improve the reliability of flash file systems with limited management and space overheads.
Pei-Han Hsu, Yuan-Hao Chang 0001, Po-Chun Huang, Tei-Wei Kuo, David Hung-Chang Du
DAC3
2010 An Efficient FTL Design for Multi-chipped Solid-State Drives
abstract
Although solid-state drives seem being excellent alternatives to replace hard disks in mobile devices, serious challenges arise due to performance and reliability concerns. This work targets performance enhancement designs with the considerations of low-cost MLC flash memory. In particular, an efficient flash management design is proposed to manage multi-chipped flash memory with cache support, where a two-level address translation mechanism is presented with an adaptive caching policy. The capability of the proposed approach is evaluated with a SystemC-based solid-state-drive simulator based on realistic workloads and benchmarks. It was shown that the proposed approach could significantly improve the performance of multi-chipped solid-state drives over various hardware configurations.
Yuan-Hao Chang 0001, Wei-Lun Lu, Po-Chun Huang, Lue-Jane Lee, Tei-Wei Kuo
RTCSA3
2009 Component-based software version management based on a Component-Interface Dependency Matrix
Shi-Ming Huang, Chih-Fong Tsai, Po-Chun Huang
J. Syst. Softw.3
2008 The Behavior Analysis of Flash-Memory Storage Systems
abstract
Performance and reliability are two major design concerns of flash-memory storage systems, especially for low-cost products. Although various excellent flash- memory management schemes are proposed, there is little work done on how to evaluate the designs or implementations of flash-memory storage systems. Many of the existing evaluation workloads for flash-memory storage systems still rely on those based on hard disks. This work aims at the needs of behavior analysis of flash-memory storage systems and their evaluations. In particular, a set of evaluation metrics and their corresponding access patterns are proposed. The behaviors of flash memory are also analyzed in terms of performance and reliability issues.
Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo, Jen-Wei Hsieh, Miller Lin
ISORC1