VLDB 2026 Research / reviewers in the wild / expert
Chin-Fu Nien
dblp:268/1596
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0002-1892-6691ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tris-GCN: A 3D NAND Flash-based In-Storage Processing Architecture for GCN AccelerationabstractGraph convolutional networks (GCNs) excel in many applications, but scaling to large graphs is bottlenecked by heavy data movement. Existing in-storage processing (ISP) solutions offload I/O-intensive operations to the SSD controller to reduce PCIe traffic, but limited parallelism and flash bandwidth still constrain energy efficiency, even with accuracy-degrading neighbor sampling. We propose Tris-GCN, a 3D NAND-based ISP design that executes core GCN computations in situ by leveraging inherent flash computing capabilities, reducing channel traffic without accuracy loss. It incorporates mapping and scheduling optimizations for energy efficiency, alongside wear-leveling with selective recomputation for reliability. Results show that Tris-GCN achieves average 11.7× speedup (up to 41.0×) and 99.7% energy savings over CPU baselines, and 2.45× speedup and 63.8% energy savings over the SOTA ISP on the Amazon dataset. Yi-Wa Wu, Jia-You Li, Chi-Jung Chen, Ching (Ryan) Cheng, Chia-Chun Wang, Chin-Fu Nien, Hsiang-Yun Cheng |
ISLPED | 6 |
| 2025 | ReTAP: Processing-in-ReRAM Bitap Approximate String Matching Accelerator for Genomic AnalysisabstractRead mapping, which involves computationally intensive approximate string matching (ASM) on large datasets, is the primary performance bottleneck in genome sequence analysis. To accelerate read mapping, a processing-in-memory (PIM) architecture that conducts highly parallel computations within the memory to reduce energy-inefficient data movements can be a promising solution. In this paper, we present ReTAP, a processing-in-ReRAM Bitap accelerator for genomic analysis. Instead of using the intricate dynamic programming algorithm, our design incorporates the Bitap algorithm, which uses only simple bitwise operations to perform ASM. Additionally, we explore the opportunity to reduce redundant computations by dynamically adjusting the error tolerance of Bitap and co-design the hardware to enhance computation parallelism. Our evaluation demonstrates that ReTAP outperforms GenASM, the state-of-the-art Bitap accelerator, with a 5.74× throughput and 1.79× energy efficiency. Tsung-Yu Liu, Yen-An Lu, James Yu, Chin-Fu Nien, Hsiang-Yun Cheng |
ASP-DAC | 4 |
| 2025 | Efficient Power- and Area-Optimized 800-Spin Ising Chip for Solving Combinatorial Optimization Problems Using Multirun Decremental AnnealingabstractThis paper presents a very-large-scale-integration (VLSI)-based multi-run decremental annealing (MRDA) accelerator for solving combinatorial optimization problems (COPs). The proposed Ising chip, which features 800 spins, utilizes MRDA to minimize hardware area while maintaining high accuracy. The proposed method utilizes a local energy calculation approach, which replaces total energy computations and effectively reduces redundancy in such fully connected models. The multi-threaded design improves speed while lowering power and area overheads. Extensive tests on the Gset dataset and Max-Cut problems demonstrate that the proposed method achieves rapid convergence to near-optimal solutions, ensuring both high accuracy and low energy consumption. Fabricated using Taiwan Semiconductor Manufacturing Company (TSMC) 90 nm process the chip runs at 91MHz, with a core area of 2.4mm2 and power consumption of 9.4mW. This design demonstrates superior efficiency and accuracy in solving complex COPs. Yuan-Ho Chen, Yu-Jie Yen, Chin-Fu Nien, Chao-Sung Lai |
IEEE Internet Things J. | 3 |
| 2024 | ReTAP: Processing-in-ReRAM Bitap Approximate String Matching Accelerator for Genomic AnalysisabstractRead mapping, which involves computationally in-tensive approximate string matching (ASM) on large datasets, is the primary performance bottleneck in genome sequence analysis. To accelerate read mapping, a processing-in-memory (PIM) architecture that conducts highly parallel computations within the memory to reduce energy-inefficient data movements can be a promising solution. In this paper, we present ReTAP, a processing-in-ReRAM Bitap accelerator for genomic analysis. Instead of using the intricate dynamic programming algorithm, our design incorporates the Bitap algorithm, which uses only simple bitwise operations to perform ASM. Additionally, we explore the opportunity to reduce redundant computations by dynamically adjusting the error tolerance of Bitap and co-design the hardware to enhance computation parallelism. Our evaluation demonstrates that ReTAP outperforms GenASM, the state-of-the-art Bitap accelerator, with a 153.7 x higher throughput. Tsung-Yu Liu, Yen-An Lu, James Yu, Chin-Fu Nien, Hsiang-Yun Cheng |
DATE | 4 |
| 2024 | ReAIM: A ReRAM-based Adaptive Ising Machine for Solving Combinatorial Optimization ProblemsabstractRecently, in light of the success of quantum computers, research teams have actively developed quantum-inspired computers using classical computing technology. One notable success story is the Ising Machine (IM), which excels in efficiently solving NP-hard combinatorial optimization problems (COPs) in various domains such as finance, drug development, and logistics. However, IMs may encounter the von Neumann bottleneck due to significant data transfers between computing and memory units. To tackle this issue, processing-in-memory (PiM) leveraging resistive random-access memory (ReRAM) has emerged as a potential solution. Nonetheless, existing ReRAM-based IM accelerators face challenges stemming from non-ideal ReRAM devices and the limited flexibility in selecting solver algorithms. To overcome these limitations, we propose ReAIM, a ReRAMbased Adaptive Ising Machine, co-designed with an adaptive parameter search algorithm to dynamically select an appropriate solver algorithm and hardware/software parameters for the target COP. Based on our in-depth analysis, ReAIM accounts for the impact of both software parameters, such as the choice of local search algorithm and the number of spin flips, and hardware characteristics, including ReRAM-induced errors. Its reconfigurable architecture and optimized mapping/scheduling strategies enable efficient execution of the adaptive algorithm across various COPs, even for large-scale problems. Compared to the state-of-the-art SRAM-based IM accelerator, ReAIM achieves a $2.2 \times$ shorter time-to-solution (TTS), showcasing superior hardware performance without compromising solution quality. Hao-Wei Chiang, Chin-Fu Nien, Hsiang-Yun Cheng, Kuei-Po Huang |
ISCA | 2 |
| 2022 | RePAIR: A ReRAM-based Processing-in-Memory Accelerator for Indel RealignmentabstractGenomic analysis has attracted a lot of interest recently since it is the key to realizing precision medicine for diseases such as cancer. Among all the genomic analysis pipeline stages, Indel Realignment is the most time-consuming and induces intensive data movements. Thus, we propose RePAIR, the first ReRAM-based processing-in-memory accelerator targeting the Indel Realignment algorithm. To further increase the computation parallelism, we design several mapping and scheduling optimization schemes. RePAIR achieves 7443× speedup and is 27211× more energy efficient over the GATK3.8 running on a CPU server, significantly outperforming the state-of-the-art. Chin-Fu Nien, Kuang-Chao Chou, Hsiang-Yun Cheng |
DATE | 2 |
| 2022 | DL-RSIM: A Reliability and Deployment Strategy Simulation Framework for ReRAM-based CNN AcceleratorsabstractMemristor-based deep learning accelerators provide a promising solution to improve the energy efficiency of neuromorphic computing systems. However, the electrical properties and crossbar structure of memristors make these accelerators error-prone. In addition, due to the hardware constraints, the way to deploy neural network models on memristor crossbar arrays affects the computation parallelism and communication overheads. To enable reliable and energy-efficient memristor-based accelerators, a simulation platform is needed to precisely analyze the impact of non-ideal circuit/device properties on the inference accuracy and the influence of different deployment strategies on performance and energy consumption. In this paper, we propose a flexible simulation framework, DL-RSIM, to tackle this challenge. A rich set of reliability impact factors and deployment strategies are explored by DL-RSIM, and it can be incorporated with any deep learning neural networks implemented by TensorFlow. Using several representative convolutional neural networks as case studies, we show that DL-RSIM can guide chip designers to choose a reliability-friendly design option and energy-efficient deployment strategies and develop optimization techniques accordingly. Hsiang-Yun Cheng, Chia-Lin Yang, Meng-Yao Lin, Kai Lien, Han-Wen Hu, Hung-Sheng Chang, Hsiang-Pang Li, Meng-Fan Chang, Yen-Ting Tsou, Chin-Fu Nien |
ACM Trans. Embed. Comput. Syst. | 11 |
| 2021 | RePIM: Joint Exploitation of Activation and Weight Repetitions for In-ReRAM DNN AccelerationabstractEliminating redundant computations is a common approach to improve the performance of ReRAM-based DNN accelerators. While existing practical ReRAM-based accelerators eliminate part of the redundant computations by exploiting sparsity in inputs and weights or utilizing weight patterns of DNN models, they fail to identify all the redundancy, resulting in many unnecessary computations. Thus, we propose a practical design, RePIM, that is the first to jointly exploit the repetition of both inputs and weights. Our evaluation shows that RePIM is effective in eliminating unnecessary computations, achieving an average of $ 15.24\times$ speedup and 96.07% energy savings over the state-of-the-art practical ReRAM-based accelerator. Chen-Yang Tsai, Chin-Fu Nien, Tz-Ching Yu, Hung-Yu Yeh, Hsiang-Yun Cheng |
DAC | 2 |
| 2021 | ReSpar: Reordering Algorithm for ReRAM-based Sparse Matrix-Vector Multiplication AcceleratorabstractSparse matrix-vector multiplication (SpMV) serves as a crucial operation for several key application domains, such as graph analytics and scientific computing, in the era of big data. The performance of SpMV is bounded by the data transmissions across memory channels in conventional von Neumann systems. Emerging metal-oxide resistive random access memory (ReRAM) has shown its potential to address this memory wall challenge through performing SpMV directly within its crossbar arrays. However, due to the tightly coupled crossbar structure, it is unlikely to skip all redundant data loading and computations with zero-valued entries of the sparse matrix in such ReRAM-based processing-in-memory architecture. These unnecessary ReRAM writes and computations hurt the energy efficiency. As only the crossbar-sized sub-matrices with full-zero entries can be skipped, prior studies have proposed some matrix reordering methods to aggregate non-zero entries to few crossbar arrays, such that more full-zero crossbar arrays can be skipped. Nevertheless, the effectiveness of prior reordering methods is constrained by the original ordering of matrix rows. In this paper, we show that the amount of full-zero sub-matrices derived by these prior studies are less than a theoretical lower bound in some cases, indicating that there are still rooms for improvement. Hence, we propose a novel reordering algorithm, ReSpar, that aims to aggregate matrix rows with similar non-zero column entries together and concentrates the non-zeros columns to increase the zero-skipping opportunities. Results show that ReSpar achieves 1.68× and 1.37× more energy savings, while reducing the required number of crossbar loads by 40.4% and 27.2% on average. Yi-Jou Hsiao, Chin-Fu Nien, Hsiang-Yun Cheng |
ICCD | 2 |
| 2020 | GraphRSim: A Joint Device-Algorithm Reliability Analysis for ReRAM-based Graph ProcessingabstractGraph processing has attracted a lot of interests in recent years as it plays a key role to analyze huge datasets. ReRAM-based accelerators provide a promising solution to accelerate graph processing. However, the intrinsic stochastic behavior of ReRAM devices makes its computation results unreliable. In this paper, we build a simulation platform to analyze the impact of non-ideal ReRAM devices on the error rates of various graph algorithms. We show that the characteristic of the targeted graph algorithm and the type of ReRAM computations employed greatly affect the error rates. Using representative graph algorithms as case studies, we demonstrate that our simulation platform can guide chip designers to select better design options and develop new techniques to improve reliability. Chin-Fu Nien, Yi-Jou Hsiao, Hsiang-Yun Cheng, Cheng-Yu Wen, Ya-Cheng Ko, Che-Ching Lin |
DATE | 1 |