Liang-Chi Chen

dblp:55/5299 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
9since 2021 · last 2026
0000-0003-2579-4305ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 11 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LAMP: An Adaptive Near-Memory Processing System for High-Performance Long-Read Mapping
Jo-Ling Huang, Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
DATE2
2025 PIMDup: An Optimized Deduplication Design on a Real Processing-in-Memory System
abstract
Data deduplication enhances storage efficiency through non-destructive compression but is often hindered by the chunking process, which requires scanning the entire dataset. While traditional methods leveraging conventional architectures and hardware accelerators (e.g., GPUs and FPGAs) have been developed to address this issue, they continue to face challenges related to excessive data movement and associated performance degradation. These limitations stem from the von Neumann architecture, where computation and storage are separated in a processor-centric design, necessitating multiple memory hierarchy traversals and causing inefficiencies. To overcome these challenges, we explore UPMEM’s DPU, a processing-in-memory (PIM) technology that reduces data movement by performing computations directly within memory. However, designing a deduplication system for DPUs presents unique obstacles, including restricted inter-DPU data sharing, the absence of native multiplication support, and significant DPU-CPU communication overhead. In response, we propose PIMDup, a DPU-optimized deduplication system that addresses these constraints through efficient parallelization, DPU-friendly chunking techniques, and reduced data transfer volumes. Experimental results demonstrate that PIMDup improves chunking performance without compromising deduplication accuracy, achieving a $1.67 \times$ speedup over CPU-based systems while maintaining 100% result consistency.
Chun-Le Yeh, Liang-Chi Chen, Chien-Chung Ho, Yu-Ming Chang, Da-Wei Chang
DAC2
2025 AridWalk: Efficient Graph Random Walks on a Resource-Limited Computational Storage Device
abstract
The effective utilization of graph structures relies on obtaining high-quality graph embeddings. Traditional embedding algorithms, such as DeepWalk and Node2Vec, which rely on random walk sampling, encounter significant challenges when applied to large-scale graphs due to the substantial data transfer demands between storage and memory. To address these limitations, we propose AridWalk, which enables a Computational Storage Device (CSD) to perform random walks directly at the storage level, minimizing external data transfers by only transferring essential data. To address the constraints of limited computational resources in the CSD, AridWalk is designed to maximize DRAM utilization while significantly reducing internal data movements, specifically between internal DRAM and flash memory. Experimental results demonstrate that AridWalk substantially decreases internal data movement, providing an efficient and scalable solution for conducting in-storage random walks on large graphs.
Liang-Chi Chen, Chien-Chung Ho, Tei-Wei Kuo, Yuan-Hao Chang 0001
ISLPED1
2025 Accelerating RNA-Seq Quantification on a Real Processing-in-Memory System
abstract
Recently, with the growth of the required data size for emerging applications (e.g., graph processing and machine learning), the von Neumann bottleneck has become a main problem for restricting the throughput of the applications. To address the problem, an acceleration technique called Processing in Memory (PIM) has garnered attention due to its potential to reduce off-chip data movement between the processing unit (e.g., CPU) and memory device (e.g., DRAM). In 2019, UPMEM introduced the commercially available processing-in-memory product, the DRAM Processing Unit (DPU) [8], showing a new chance for accelerating data-intensive applications. Among data-intensive applications, RNA sequence (RNA-seq) quantification is used to measure the abundance of RNA sequences, and it also plays a critical role in the field of bioinformatics. We aim to leverage UPMEM DPU to accelerate RNA-seq Quantification. However, due to the DPU usage limitations caused by DPU hardware, there are some challenges to realizing RNA-seq Quantification on the DPU system. To overcome these challenges, we propose UpPipe, which consists of the DPU-friendly transcriptome allocation, the DPU-aware pipeline management, and the WRAM prefetching scheme. The UpPipe considers the hardware limitations of DPUs, enabling efficient sequence alignment even within the resource-constrained DPUs. The experimental results demonstrate the feasibility and efficiency of our proposed design. We also provide an evaluation study on the impact of data granularity selection on pipeline management and the optimal size for the WRAM prefetching scheme.
Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
IEEE Trans. Computers1
2025 A Survey on Flash-Memory Storage Systems: A Host-Side Perspective
abstract
NAND flash memory has become the dominant storage media choice in a vast majority of application scenarios. Compared to mechanical hard disks, flash offers better access performance, energy efficiency, and shock resistance. However, the unique hardware peculiarities of this technology require dedicated facilities to manage the flash space and data. The implementation of flash management facilities has alternatively been realized either at the device or host computer level. Managing flash on the device side eases integration/compatibility and increases performance in certain scenarios. However, the limited computing resources inherent to devices and the lack of higher-level file system/application information make these solutions suboptimal in many situations. Managing flash on the host allows leveraging its abundant resources, and host-side knowledge such as data access patterns can be exploited to optimize flash management, at the cost of increased host-side complexity. The pros and cons of each approach also led to the appearance of hybrid, cross-layer solutions, enabling the collaboration of different layers of the storage stack. Recently, the pressure on modern storage systems requires that an increasing amount of flash management responsibilities is offloaded to the host, and the development of application-specific cross-layer solutions: In that context, it is crucial to review these developments. In this article, we make a comprehensive survey of the host-side management technologies of flash memory, application-/system-level flash-friendly designs, and emergent applications based on flash memory.
Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun-Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001
ACM Trans. Storage4
2023 UpPipe: A Novel Pipeline Management on In-Memory Processors for RNA-seq Quantification
abstract
RNA sequence quantification is an important analysis method to measure transcript abundances. A key overhead in RNA-seq quantification is to map a set of RNA reads to multiple reference transcripts, i.e., transcriptome. Besides, the performance of RNA-seq quantification is strictly limited by the excessive amounts of data movement between CPU and memory, i.e., memory wall problem on the conventional architecture. As the first publicly commercial processing-in-memory (PIM) system, UPMEM DPU, is proposed, the PIM gradually becomes a promising solution to overcome the memory wall problem. DPUs show great potential to accelerate data-intensive workloads by minimizing off-chip data movement between CPU and memory. Thus, this paper aims to improve the performance of RNA-seq quantification by fully exploiting the strengths of DPU. To achieve that, we propose a novel DPU-aware pipeline design "UpPipe" built on the software layer to address the hardware constraints of DPU. To the best of our knowledge, this is the first work to enable pipeline management on the DPU system. The evaluation results demonstrate the feasibility of our proposed design and provide a comprehensive study on how to utilize the limited hardware resources of DPUs efficiently.
Liang-Chi Chen, Chien-Chung Ho, Yuan-Hao Chang 0001
DAC1
2023 Reaping Both Latency and Reliability Benefits With Elaborate Sanitization Design for 3D TLC NAND Flash
abstract
With the rising security concern on modern storage systems, the concept of data sanitization has been widely investigated recently. Among the existing works targeting data sanitization, an overwriting-based approach, namely one-shot sanitization, is one of the most efficient sanitization approaches. Nonetheless, we find that the one-shot sanitization approach would fail to achieve precise data sanitization for 3D TLC NAND flash, because of incurring undesired data errors. That is, how to simultaneously realize precise sanitization and high security with decent latency and reliability on emerging storage devices remains unsolved. This work proposes an elaborate sanitization design that skillfully manipulates the threshold voltage ($V_{t}$) distribution of sanitized pages. Not only does the proposed design achieve precise sanitization and high security, but it also enhances read performance and data reliability. Specifically, this work elaborately sanitizes data by merging specific$V_{t}$distributions of the target physical page on 3D TLC NAND flash. Besides, the proposed approach further takes lateral charge migration into consideration to improve data reliability. We conduct a series of experiments to evaluate our proposed approach on real 3D TLC NAND flash. The experiment results demonstrate the proposed approach can achieve elaborate data sanitization under various scenarios and improve read performance by 29%.
Wei-Chen Wang 0002, Chien-Chung Ho, Yung-Chun Li, Liang-Chi Chen, Yu-Ming Chang
IEEE Trans. Computers4
2023 WARM-tree: Making Quadtrees Write-efficient and Space-economic on Persistent Memories
abstract
Recently, the value of data has been widely recognized, which highlights the significance of data-centric computing in diversified application scenarios. In many cases, the data are multidimensional, and the management of multidimensional data often confronts greater challenges in supporting efficient data access operations and guaranteeing the space utilization. On the other hand, while many existing index data structures have been proposed for multidimensional data management, however, their designs are not fully optimized for modern nonvolatile memories, in particular the byte-addressable persistent memories. As a result, they might undergo serious access performance degradation or fail to guarantee space utilization. This observation motivates the redesigning of index data structures for multidimensional point data on modern persistent memories, such as the phase-change memory. In this work, we present the WARM-tree , a m ultidimensional t ree for r educing the w rite a mplification effect, for multidimensional point data. In our evaluation studies, as compared to the bucket PR quadtree and R*-tree, the WARM-tree can provide any worst-case space utilization guarantees in the form of \(\frac{m-1}{m}\) ( m ∈ ℤ^+) and effectively reduces the write traffic of key insertions by up to 48.10% and 85.86%, respectively, at the price of degraded average space utilization and prolonged latency of query operations. This suggests that the WARM-tree is a potential multidimensional index structure for insert-intensive workloads.
Shin-Ting Wu, Liang-Chi Chen, Po-Chun Huang, Yuan-Hao Chang 0001, Chien-Chung Ho, Wei-Kuan Shih
ACM Trans. Embed. Comput. Syst.2
2022 LongPhase: an ultra-fast chromosome-scale phasing algorithm for small and large variants
abstract
MOTIVATION: Long-read phasing has been used for reconstructing diploid genomes, improving variant calling and resolving microbial strains in metagenomics. However, the phasing blocks of existing methods are broken by large Structural Variations (SVs), and the efficiency is unsatisfactory for population-scale phasing. RESULTS: This article presents a novel algorithm, LongPhase, which can simultaneously phase single nucleotide polymorphisms (SNPs) and SVs of a human genome in 10-20 min, 10× faster than the state-of-the-art WhatsHap, HapCUT2 and Margin. In particular, co-phasing SNPs and SVs produces much larger haplotype blocks (N50 = 25 Mbp) than those of existing methods (N50 = 10-15 Mbp). We show that LongPhase combined with Nanopore ultra-long reads is a cost-effective and highly contiguous solution, which can produce between one and 26 blocks per chromosome arm without the need for additional trios, chromosome-conformation and strand-seq data. AVAILABILITYAND IMPLEMENTATION: LongPhase is freely available at https://github.com/twolinin/LongPhase/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jyun-Hong Lin, Liang-Chi Chen, Shu-Chi Yu, Yao-Ting Huang
Bioinform.2
2011 Transition test bring-up and diagnosis on UltraSPARCTM processors
abstract
We describe methods to use PLL-based transition test in support of chip bring-up. We used it to diagnose slow paths for performance improvement. During bring-up, often the issues of setup, design, scan patterns, and silicon slow paths are mixed together, making diagnosis more difficult. We discuss techniques used to understand and resolve these issues, and show examples of the benefit of transition test over functional and system test.
Liang-Chi Chen, Peter Dahlgren, Paul Dickinson, Scott Davidson 0001
ITC1
2009 Using transition test to understand timing behavior of logic circuits on UltraSPARCTM T2 family
abstract
Delay test is crucial for finding slow paths and slow ICs, both during bringup and during speed binning. Path delay test has traditionally been considered to be superior in finding slow paths. This paper describes our experiments indicating that this is not always the case. For the UltraSPARC T2 microprocessor series we found that transition delay test often ran slower, was more effective in finding the root cause of the slow path, and correlated well with functional diags also used for speed binning. Transition test does a better job finding delay issues related to the impact of simultaneous switching and coupling noise on chip speed. We used transition test to measure the impact on chip timing of voltage, temperature, and we also used it to confirm the results of improving slow paths.
Liang-Chi Chen, Paul Dickinson, Peter Dahlgren, Scott Davidson 0001, Olivier Caty, Kevin Wu
ITC1
2008 Transition Test on UltraSPARC- T2 Microprocessor
abstract
Sun's T2 processor transition test methodology, verification and silicon debug are described. We illustrate our test mechanism, test sequence, test development, and the verification of this mechanism in silicon. Methods to identify slow flops, diagnose gate dominated paths, excite long delay for wire dominated paths, and find essential patterns for speed characterization are presented. In addition, a method is developed for recovering devices with a fewer good processor cores. All these applications together make transition test a much more powerful tool.
Liang-Chi Chen, Paul Dickinson, Prasad Mantri, Murali M. R. Gala, Peter Dahlgren, Subhra Bhattacharya, Olivier Caty, Kevin Woodling, Thomas A. Ziaja, David Curwen, Wendy Yee, Ellen Su, Guixiang Gu, Tim Nguyen
ITC1
2002 TA-PSV - Timing Analysis for Partially Specified Vectors
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
J. Electron. Test.1
2001 A New Gate Delay Model for Simultaneous Switching and Its Applications
abstract
We present a new model to capture the delay phenomena associ-ated with simultaneous to-controlling transitions. The proposed delay model accurately captures the effect of the targeted delay phe-nomena over a wide range of transition times and skews. It also cap-tures the effects of more variables than table lookup methods can handle. The model helps improve the accuracy of static timing anal-ysis, incremental timing refinement, and timing-based ATPG.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
DAC1
2001 Crosstalk test generation on pseudo industrial circuits: a case study
abstract
In this paper, we present data that validates the viability of a university prototype crosstalk ATPG system, XGEN, on real designs. We remodeled Intel circuits and performed test generation using actual parasitic data. A crosstalk ATPG implementation flow was developed based on Intel tools. Validation results are shown for the modified circuits. Critical issues for preserving accurate timing information and capturing crosstalk effects are discussed.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
ITC1
2000 A new framework for static timing analysis, incremental timing refinement, and timing simulation
abstract
In this paper we present a framework that enables the computation of tight ranges of signal arrival, transition, and required times for rising and falling transitions at each circuit line, given an input sequence consisting of two partially specified vectors. At one extreme, when the vectors are completely unspecified, this framework becomes identical to static timing analysis (STA). At the other extreme, when the vectors are completely specified, this framework performs timing simulation (TS). Our key motivation for developing this framework was to reduce the amount of search required by a test generator that uses timing information. During test generation for a target fault, values are specified incrementally and this framework enables refinement of timing windows. We demonstrate that this approach significantly improves test generation efficiency. In this mode, the ATPG is said to be performing incremental timing refinement (ITR).
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
Asian Test Symposium1
1997 High Quality Robust Tests for Path Delay Faults
abstract
Detailed circuit simulations have demonstrated that a classical two-pattern robust test for a path delay fault may not excite the worst case delay of the target path. We have developed a new definition of robust test that maintains the desirable properties of classical robust tests while incorporating two additional considerations, namely side-fan-in transitions and pre-initialization, which are shown to have a significant impact on the delay of the target path. The associated test generation problem was formulated as a constrained optimization problem, and an ATPG system developed to generate three-pattern robust tests that excite the worst case delay of the target path. The ATPG works on a gate level model that is augmented to capture the necessary switch level details. Experimental results show that the quality of robust delay tests varies dramatically and that the proposed high quality robust delay tests are needed for improving test quality.
Liang-Chi Chen, Sandeep Gupta 0001, Melvin A. Breuer
VTS1
1993 On-chip test generation for combinational circuits by LFSR modification
abstract
A new on-chip test generation technique based on the built-in self test (BIST) and deterministic test generation concepts has been proposed. Given a test set, the test patterns can be regenerated on the chip and applied to the circuit under test without the use of any external test equipments. A systematic procedure for the modification of a basic linear feedback shift register (LFSR) to realize the on-chip test generation hardware is given. Since the delay introduced by the modification of the LFSR is only two gate delays, at-speed testing of circuits is feasible. Experiments are conducted and test application time and hardware overhead are compared with a known test technique under the same fault coverage conditions. It is shown that both test cost and test application time can be decreased significantly by using the proposed technique.
Shambhu J. Upadhyaya, Liang-Chi Chen
ICCAD2