Chien-Chung Ho

dblp:137/1408 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0003-3460-8674ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LAMP: An Adaptive Near-Memory Processing System for High-Performance Long-Read Mapping
Jo-Ling Huang, Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
DATE3
2026 CRC-Based Error Detection Mechanism for Multiplicative Inverse in Redundant Basis S -Box
Ying-Sheng Huang, Chia-Chou Chuang, Chien-Chung Ho, Pei-Yin Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System
abstract
Data deduplication enhances storage efficiency through non-destructive compression but is often hindered by the chunking process, which requires scanning the entire dataset. While traditional methods leveraging conventional architectures and hardware accelerators (e.g., GPUs and FPGAs) have been developed to address this issue, they continue to face challenges related to excessive data movement and associated performance degradation. These limitations stem from the von Neumann architecture, where computation and storage are separated in a processor-centric design, necessitating multiple memory hierarchy traversals and causing inefficiencies. To overcome these challenges, we explore UPMEM’s DPU, a processing-in-memory (PIM) technology that reduces data movement by performing computations directly within memory. However, designing a deduplication system for DPUs presents unique obstacles, including restricted inter-DPU data sharing, the absence of native multiplication support, and significant DPU-CPU communication overhead. In response, we propose PIMDup, a DPU-optimized deduplication system that addresses these constraints through efficient parallelization, DPU-friendly chunking techniques, and reduced data transfer volumes. Experimental results demonstrate that PIMDup improves chunking performance without compromising deduplication accuracy, achieving a $1.67 \times$ speedup over CPU-based systems while maintaining 100% result consistency.
Chun-Le Yeh, Liang-Chi Chen, Chien-Chung Ho, Yu-Ming Chang, Da-Wei Chang
DAC3
2025 AridWalk: Efficient Graph Random Walks on a Resource-Limited Computational Storage Device
abstract
The effective utilization of graph structures relies on obtaining high-quality graph embeddings. Traditional embedding algorithms, such as DeepWalk and Node2Vec, which rely on random walk sampling, encounter significant challenges when applied to large-scale graphs due to the substantial data transfer demands between storage and memory. To address these limitations, we propose AridWalk, which enables a Computational Storage Device (CSD) to perform random walks directly at the storage level, minimizing external data transfers by only transferring essential data. To address the constraints of limited computational resources in the CSD, AridWalk is designed to maximize DRAM utilization while significantly reducing internal data movements, specifically between internal DRAM and flash memory. Experimental results demonstrate that AridWalk substantially decreases internal data movement, providing an efficient and scalable solution for conducting in-storage random walks on large graphs.
Liang-Chi Chen, Chien-Chung Ho, Tei-Wei Kuo, Yuan-Hao Chang 0001
ISLPED2
2025 Accelerating RNA-Seq Quantification on a Real Processing-in-Memory System
abstract
Recently, with the growth of the required data size for emerging applications (e.g., graph processing and machine learning), the von Neumann bottleneck has become a main problem for restricting the throughput of the applications. To address the problem, an acceleration technique called Processing in Memory (PIM) has garnered attention due to its potential to reduce off-chip data movement between the processing unit (e.g., CPU) and memory device (e.g., DRAM). In 2019, UPMEM introduced the commercially available processing-in-memory product, the DRAM Processing Unit (DPU) [8], showing a new chance for accelerating data-intensive applications. Among data-intensive applications, RNA sequence (RNA-seq) quantification is used to measure the abundance of RNA sequences, and it also plays a critical role in the field of bioinformatics. We aim to leverage UPMEM DPU to accelerate RNA-seq Quantification. However, due to the DPU usage limitations caused by DPU hardware, there are some challenges to realizing RNA-seq Quantification on the DPU system. To overcome these challenges, we propose UpPipe, which consists of the DPU-friendly transcriptome allocation, the DPU-aware pipeline management, and the WRAM prefetching scheme. The UpPipe considers the hardware limitations of DPUs, enabling efficient sequence alignment even within the resource-constrained DPUs. The experimental results demonstrate the feasibility and efficiency of our proposed design. We also provide an evaluation study on the impact of data granularity selection on pipeline management and the optimal size for the WRAM prefetching scheme.
Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
IEEE Trans. Computers2
2025 A Survey on Flash-Memory Storage Systems: A Host-Side Perspective
abstract
NAND flash memory has become the dominant storage media choice in a vast majority of application scenarios. Compared to mechanical hard disks, flash offers better access performance, energy efficiency, and shock resistance. However, the unique hardware peculiarities of this technology require dedicated facilities to manage the flash space and data. The implementation of flash management facilities has alternatively been realized either at the device or host computer level. Managing flash on the device side eases integration/compatibility and increases performance in certain scenarios. However, the limited computing resources inherent to devices and the lack of higher-level file system/application information make these solutions suboptimal in many situations. Managing flash on the host allows leveraging its abundant resources, and host-side knowledge such as data access patterns can be exploited to optimize flash management, at the cost of increased host-side complexity. The pros and cons of each approach also led to the appearance of hybrid, cross-layer solutions, enabling the collaboration of different layers of the storage stack. Recently, the pressure on modern storage systems requires that an increasing amount of flash management responsibilities is offloaded to the host, and the development of application-specific cross-layer solutions: In that context, it is crucial to review these developments. In this article, we make a comprehensive survey of the host-side management technologies of flash memory, application-/system-level flash-friendly designs, and emergent applications based on flash memory.
Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun-Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001
ACM Trans. Storage7
2023 UpPipe: A Novel Pipeline Management on In-Memory Processors for RNA-seq Quantification
abstract
RNA sequence quantification is an important analysis method to measure transcript abundances. A key overhead in RNA-seq quantification is to map a set of RNA reads to multiple reference transcripts, i.e., transcriptome. Besides, the performance of RNA-seq quantification is strictly limited by the excessive amounts of data movement between CPU and memory, i.e., memory wall problem on the conventional architecture. As the first publicly commercial processing-in-memory (PIM) system, UPMEM DPU, is proposed, the PIM gradually becomes a promising solution to overcome the memory wall problem. DPUs show great potential to accelerate data-intensive workloads by minimizing off-chip data movement between CPU and memory. Thus, this paper aims to improve the performance of RNA-seq quantification by fully exploiting the strengths of DPU. To achieve that, we propose a novel DPU-aware pipeline design "UpPipe" built on the software layer to address the hardware constraints of DPU. To the best of our knowledge, this is the first work to enable pipeline management on the DPU system. The evaluation results demonstrate the feasibility of our proposed design and provide a comprehensive study on how to utilize the limited hardware resources of DPUs efficiently.
Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
DAC2
2023 Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash
abstract
With the rising security concern on modern storage systems, the concept of data sanitization has been widely investigated recently. Among the existing works targeting data sanitization, an overwriting-based approach, namely one-shot sanitization, is one of the most efficient sanitization approaches. Nonetheless, we find that the one-shot sanitization approach would fail to achieve precise data sanitization for 3D TLC NAND flash, because of incurring undesired data errors. That is, how to simultaneously realize precise sanitization and high security with decent latency and reliability on emerging storage devices remains unsolved. This work proposes an elaborate sanitization design that skillfully manipulates the threshold voltage ($V_{t}$) distribution of sanitized pages. Not only does the proposed design achieve precise sanitization and high security, but it also enhances read performance and data reliability. Specifically, this work elaborately sanitizes data by merging specific$V_{t}$distributions of the target physical page on 3D TLC NAND flash. Besides, the proposed approach further takes lateral charge migration into consideration to improve data reliability. We conduct a series of experiments to evaluate our proposed approach on real 3D TLC NAND flash. The experiment results demonstrate the proposed approach can achieve elaborate data sanitization under various scenarios and improve read performance by 29%.
Wei-Chen Wang 0002, Chien-Chung Ho, Yung-Chun Li, Liang-Chi Chen, Yu-Ming Chang
IEEE Trans. Computers2
2023 WARM-tree: Making Quadtrees Write-efficient and Space-economic on Persistent Memories
abstract
Recently, the value of data has been widely recognized, which highlights the significance of data-centric computing in diversified application scenarios. In many cases, the data are multidimensional, and the management of multidimensional data often confronts greater challenges in supporting efficient data access operations and guaranteeing the space utilization. On the other hand, while many existing index data structures have been proposed for multidimensional data management, however, their designs are not fully optimized for modern nonvolatile memories, in particular the byte-addressable persistent memories. As a result, they might undergo serious access performance degradation or fail to guarantee space utilization. This observation motivates the redesigning of index data structures for multidimensional point data on modern persistent memories, such as the phase-change memory. In this work, we present the WARM-tree , a m ultidimensional t ree for r educing the w rite a mplification effect, for multidimensional point data. In our evaluation studies, as compared to the bucket PR quadtree and R*-tree, the WARM-tree can provide any worst-case space utilization guarantees in the form of \(\frac{m-1}{m}\) ( m ∈ ℤ^+) and effectively reduces the write traffic of key insertions by up to 48.10% and 85.86%, respectively, at the price of degraded average space utilization and prolonged latency of query operations. This suggests that the WARM-tree is a potential multidimensional index structure for insert-intensive workloads.
Shin-Ting Wu, Liang-Chi Chen, Po-Chun Huang, Yuan-Hao Chang 0001, Chien-Chung Ho, Wei-Kuan Shih
ACM Trans. Embed. Comput. Syst.5
2021 RVO: Unleashing SSD's Parallelism by Harnessing the Unused Power
abstract
Analytic video surveillance system is one of the fastest-growing cyber-physical applications worldwide. A video surveillance system must have a scalable storage backend to simultaneously ingest video frames and serve read requests for video analytics. As a result, 3D NAND flash-based storage devices, i.e., Solid-State Drives (SSD), are gradually regarded as promising candidates thanks to their rapidly growing density and parallelism. However, when deployed in power-constrained environments like the network edges, SSDs’ parallelism often cannot be fully unleashed. In edge video analytics systems, the lost parallelism can cause video frame drop and untimely analytics. To tackle this limitation, we first reveal that a flash program operation’s power usage is over-estimated in the conventional SSD design, leading to a limited degree of I/O parallelism. Based on the observation, we propose a novel command set, RVO (Read-Verify Overlap), which reclaims the unused power from the overestimation to amend the lost parallelism. To realize feasible fine-grained power management, we further accompany RVO with a generic power-aware scheduler. Through experiments, we show how video analytics systems equipped with RVO can achieve zero frame drop while ensuring compliance with industrial read latency requirements, even in write-intensive workloads.
Hasan Alhasan, Yun-Chih Chen, Chien-Chung Ho
ISLPED3
2019 Toward Instantaneous Sanitization through Disturbance-induced Errors and Recycling Programming over 3D Flash Memory
abstract
As data security has become one of the most crucial issues in modern storage system/application designs, the data sanitization techniques are regarded as the promising solution on 3D NAND flash-memory-based devices. Many excellent works had been proposed to exploit the in-place reprogramming, erasure and encryption techniques to achieve and implement the sanitization functionalities. However, existing sanitization approaches could lead to performance, disturbance overheads or even deciphered issues. Different from existing works, this work aims at exploring an instantaneous data sanitization scheme by taking advantage of programming disturbance properties. Our proposed design can not only achieve the instantaneous data sanitization by exploiting programming disturbance and error correction code properly, but also enhance the performance with the recycling programming design. The feasibility and capability of our proposed design are evaluated by a series of experiments on 3D NAND flash memory chips, for which we have very encouraging results. The experiment results show that the proposed design could achieve the instantaneous data sanitization with low overhead; besides, it improves the average response time and reduces the number of block erase count by up to 86.8% and 88.8%, respectively.
Wei-Chen Wang 0002, Ping-Hsien Lin, Yung-Chun Li, Chien-Chung Ho, Yu-Ming Chang, Yuan-Hao Chang 0001
ICCAD4
2019 Achieving Lossless Accuracy with Lossy Programming for Efficient Neural-Network Training on NVM-Based Systems
abstract
Neural networks over conventional computing platforms are heavily restricted by the data volume and performance concerns. While non-volatile memory offers potential solutions to data volume issues, challenges must be faced over performance issues, especially with asymmetric read and write performance. Beside that, critical concerns over endurance must also be resolved before non-volatile memory could be used in reality for neural networks. This work addresses the performance and endurance concerns altogether by proposing a data-aware programming scheme. We propose to consider neural network training jointly with respect to the data-flow and data-content points of view. In particular, methodologies with approximate results over Dual-SET operations were presented. Encouraging results were observed through a series of experiments, where great efficiency and lifetime enhancement is seen without sacrificing the result accuracy.
Wei-Chen Wang 0002, Yuan-Hao Chang 0001, Tei-Wei Kuo, Chien-Chung Ho, Yu-Ming Chang, Hung-Sheng Chang
ACM Trans. Embed. Comput. Syst.4
2018 Achieving defect-free multilevel 3D flash memories with one-shot program design
abstract
To store the desired data on MLC and TLC flash memories, the conventional programming strategies need to divide a fixed range of threshold voltage (Vt) window into several parts. The narrowly partitioned Vt window in turn limits the design of programming strategy and becomes the main reason to cause flash-memory defects, i.e., the longer read/write latency and worse data reliability. This motivates this work to explore the innovative programming design for solving the flash-memory defects. Thus, to achieve the defect-free 3D NAND flash memory, this paper presents and realizes a one-shot program design to significantly eliminate the negative impacts caused by conventional programming strategies. The proposed one-shot program design includes two strategies, i.e., prophetic and classification programming, for MLC flash memories, and the idea is extended to TLC flash memories. The measurement results show that it can accelerate programming speed by 31x and reduce RBER by 1000x for the MLC flash memory, and it can broaden the available window of threshold voltage up to 5.1x for the TLC flash memory.
Chien-Chung Ho, Yung-Chun Li, Yuan-Hao Chang 0001, Yu-Ming Chang
DAC1
2018 Achieving fast sanitization with zero live data copy for MLC flash memory
abstract
As data security has become the major concern in modern storage systems with low-cost multi-level-cell (MLC) flash memories, it is not trivial to realize data sanitization in such a system. Even though some existing works employ the encryption or the built-in erase to achieve this requirement, they still suffer the risk of being deciphered or the issue of performance degradation. In contrast to the existing work, a fast sanitization scheme is proposed to provide the highest degree of security for data sanitization; that is, every old version of data could be immediately sanitized with zero live-data-copy overhead once the new version of data is created/written. In particular, this scheme further considers the reliability issue of MLC flash memories; the proposed scheme includes a one-shot sanitization design to minimize the disturbance during data sanitization. The feasibility and the capability of the proposed scheme were evaluated through extensive experiments based on real flash chips. The results demonstrate that this scheme can achieve the data sanitization with zero live-data-copy, where performance overhead is less than 1%.
Ping-Hsien Lin, Yu-Ming Chang, Yung-Chun Li, Wei-Chen Wang 0002, Chien-Chung Ho, Yuan-Hao Chang 0001
ICCAD5
2018 Scrubbing-Aware Secure Deletion for 3-D NAND Flash
abstract
Due to the increasing security concerns, the conventional deletion operations in NAND flash memory can no longer afford the requirement of secure deletion. Although existing works exploit secure deletion and scrubbing operations to achieve the security requirement, they also result in performance and disturbance problems. The predicament becomes more severe as the growing of page numbers caused by the aggressive use of 3-D NAND flash-memory chips which stack flash cells into multiple layers in a chip. Different from existing works, this paper aims at exploring a scrubbing-aware secure deletion design so as to improve the efficiency of secure deletion by exploiting properties of disturbance. The proposed design could minimize secure deletion/scrubbing overheads by organizing sensitive data to create the scrubbing-friendly patterns, and further choose a proper operation by the proposed evaluation equations for each secure deletion command. The capability of our proposed design is evaluated by a series of experiments, for which we have very encouraging results. In a 128 Gbits 3-D NAND flash-memory device, the simulation results show that the proposed design could achieve 82% average response time reduction of each secure deletion command.
Wei-Chen Wang 0002, Chien-Chung Ho, Yuan-Hao Chang 0001, Tei-Wei Kuo, Ping-Hsien Lin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 An SLC-Like Programming Scheme for MLC Flash Memory
abstract
Although the multilevel cell (MLC) technique is widely adopted by flash-memory vendors to boost the chip density and lower the cost, it results in serious performance and reliability problems. Different from past work, a new cell programming method is proposed to not only significantly improve chip performance but also reduce the potential bit error rate. In particular, a single-level cell (SLC)-like programming scheme is proposed to better explore the threshold-voltage relationship to denote different MLC bit information, which in turn drastically provides a larger window of threshold voltage similar to that found in SLC chips. It could result in less programming iterations and simultaneously a much less reliability problem in programming flash-memory cells. In the experiments, the new programming scheme could accelerate the programming speed up to 742% and even reduce the bit error rate up to 471% for MLC pages.
Chien-Chung Ho, Yu-Ming Chang, Yuan-Hao Chang 0001, Tei-Wei Kuo
ACM Trans. Storage1
2017 Antiwear Leveling Design for SSDs With Hybrid ECC Capability
abstract
With the joint considerations of reliability and performance, hybrid error correction code (ECC) becomes an option in the designs of solid-state drives (SSDs). Unfortunately, wear leveling (WL) might result in the early performance degradation to SSDs, which is common with a limited number of P/E cycles, due to the efforts to delay the bit-error-rate growth. In this paper, an anti-WL design is proposed to avoid such a performance problem so that the performance of SSDs with hybrid ECC capability can be improved without sacrificing their reliability. The capability of the proposed design was evaluated by a series of experiments, for which it was shown that the proposed design could greatly improve the read and write performance of SSDs up to 50% without affecting the endurance of the investigated SSDs, compared with traditional approaches.
Chien-Chung Ho, Yu-Ping Liu, Yuan-Hao Chang 0001, Tei-Wei Kuo
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Enabling sub-blocks erase management to boost the performance of 3D NAND flash memory
abstract
3D NAND has been proposed to provide a large capacity storage with low-cost consideration due to its high density memory architecture. However, 3D NAND needs to consume enormous time for garbage collection because of live-page copying overhead and long block erase time. To alleviate the impact of live-page copying on the performance of 3D NAND, a sub-block erase design has been designed. With sub-block erase design, this paper proposes a performance booster strategy to extremely boost the performance of garbage collection. As experimental results shows, the proposed strategy has a significant improvement on the average response time.
Tseng-Yi Chen, Yuan-Hao Chang 0001, Chien-Chung Ho, Shuo-Han Chen
DAC3
2015 Access Pattern Reshaping for eMMC-enabled SSDs
abstract
The growing popularity of embedded Multi-Media Controllers (eMMCs) presents a unique opportunity to design commodity grade solid-state drives products. This work addresses the essential design issues of such drives and introduces a light-weight FTL design. In particular, access patterns to an eMMC-enabled solid-state drive are reshaped to create sequential access patterns and specific write sizes to better accommodate the characteristics of eMMCs, without resorting to the conventional address translation FTL design. At the same time, garbage collection overheads are minimized with reliability considerations, since eMMCs are usually not equipped with a powerful controller of a sophisticated design. The capability of the proposed design is evaluated by a series of experiments, for which we have very encouraging results.
Chien-Chung Ho, Yuan-Hao Chang 0001, Tei-Wei Kuo
ICCAD1
2014 Current-aware scheduling for flash storage devices
abstract
As NAND flash memory has become a major choice of storage media in diversified computing environments, the performance issue of flash memory has been extensively addressed in many excellent designs. Among them, an effective strategy is to adopt multiple channels and flash-memory chips to improve the performance on data accesses. However, the degree of data access parallelism cannot be increased by simply increasing the number of channels and chips in the storage device, because it is seriously limited by the maximum current constraint of the bus interface and affected by the access patterns of user data. As a consequence, to maximize the degree of access parallelism, it is of paramount significance to have a proper scheduling strategy to determine the order that read/write requests are served. In this paper, a current-aware scheduling strategy for read/write requests is proposed to maximize the read performance without violating the bus current constraint and without missing (the deadline of) written data. The proposed strategy is then evaluated through a series of experiments, in which the results are quite encouraging.
Tzu-Jung Huang, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001, Tei-Wei Kuo
RTCSA2