EDBT 2026 Demo / reviewers in the wild / expert
Fang-Yi Gu
dblp:320/9790
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-5523-470XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TT-RRAM: Joint Improvement of Sparsity and Repetition in Tensor-Train Inference With Fused Processing on RRAM AcceleratorsabstractDeep neural networks compressed using Tensor-Train decomposition , known as TT-DNNs, significantly reduce model size to decrease storage requirements but suffer from more frequent data movement during inference. To address this issue, RRAM-based DNN accelerators, by leveraging computing-in-memory, can mitigate the data movement and efficiently execute the VMM operations required during inference, making them an ideal candidate for accelerating the inference for TT-DNNs. However, RRAM-based accelerators face three challenges in TT-format DNN inference: (1) weights distribution on the RRAM crossbar with low sparsity and repetition; (2) low input sparsity; and (3) high data dependencies and storage overhead during inference. To address these challenges, we propose three corresponding methods: (1) Base-Offset Splitting (BOS) to improve the weights distribution on the crossbar, thereby enhancing performance potential; (2) Inverse Input Reusing (IIR) to enhance the sparsity of inputs during inference; and (3) Fused Scheduling (FS) to reorganize the computation order and leveraging the parallel processing capabilities of the RRAM-based accelerator to improve overall inference efficiency while also reducing data dependencies and storage overhead. Experimental results demonstrate that our proposed methods achieve a 3.44× to 5.53× performance improvement and 55.6% to 67.1% energy savings compared to the state-of-the-art RRAM-based accelerator, while also reducing storage overhead by 66.7% to 99.1%. Fang-Yi Gu, Pin-Hong Liu, Ing-Chao Lin, Bo Yuan 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Efficient Model Switching in RRAM-Based DNN AcceleratorsabstractResistive random access memory (RRAM) has emerged as a promising technology for deep neural network (DNN) accelerators, but programming every weight in a DNN onto RRAM cells for inference can be both time-consuming and energy-intensive, especially when switching between different DNN models. This article introduces a hardware-aware multimodel merging (HA3M) framework designed to minimize the need for reprogramming by maximizing weight reuse, while taking into account the hardware constraints of the accelerator. The framework includes three key approaches: 1) crossbar (XB)-aware model mapping (XAMM); 2) block-based layer matching (BLM); and 3) multimodel retraining (MMR). XAMM reduces the XB usage of the preprogrammed model on RRAM XBs while preserving the model’s structure. BLM reuses preprogrammed weights in a block-based manner, ensuring the inference process remains unchanged. MMR then equalizes the block-based matched weights across multiple models. Experimental results show that the proposed framework significantly reduces programming cycles in multi-DNN switching scenarios while maintaining or even enhancing accuracy, and eliminating the need for reprogramming. Fang-Yi Gu, Ing-Chao Lin, Bing Li 0005, Ulf Schlichtmann, Grace Li Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | A Hardware Friendly Variation-Tolerant Framework for RRAM-Based Neuromorphic ComputingabstractEmerging resistive random access memory (RRAM) attracts considerable interest in computing-in-memory by its high efficiency in multiply-accumulate operation, which is the key computation in the neural network (NN). However, due to the imperfect fabrication, RRAM cells suffer from the variations, which make the values in RRAM cells deviate from the target values so that the accuracy of the RRAM-based NN accelerator degrades significantly. Moreover, in a practical hardware design of RRAM-based NN accelerators, if the number of wordlines and bitlines in a crossbar array activated at the same time increases, ADCs with a high resolution are required and the power consumption of ADC increases. This paper proposes a novel methodology to mitigate the impact of variations in RRAM-based neural network accelerators. The methodology includes a unary-based non-uniform quantization method and a variation-aware operation unit (OU) based framework. The unary-based non-uniform quantization method equalizes the significance of weights stored in each RRAM cell to reduce the impact of variations. The variation-aware OU-based framework activates only RRAM cells in the same OU at the same time, which reduces the power consumption of ADCs. Additionally, the framework introduces three methods, including OU skipping, OU recombination, and OU compensation, to further mitigate the impact of variations. The experiments show that the proposed approach outperforms the state-of-the-art among four NN models on two datasets with 2-bit cell resolution. Fang-Yi Gu, Cheng-Han Yang, Ing-Chao Lin, Da-Wei Chang, Darsen D. Lu, Ulf Schlichtmann |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | A Novel and Efficient Block-Based Programming for ReRAM-Based Neuromorphic ComputingabstractReRAM-based accelerators have emerged as promising accelerators for deep neural networks (DNNs). How-ever, programming every ReRAM cell to its corresponding conductance before inference can be time-consuming and energy-intensive using existing one-by-one/row-by-row programming mechanisms. Although a two-phase multi-row programming scheme has been proposed to enhance programming efficiency, there are situations where multiple rows cannot be programmed together and only row-by-row programming can be employed. Therefore, this paper proposes a new block-based programming architecture for ReRAM crossbars that enables precise control of wordline and bitline transistors. In addition, a block-based programming framework, including the approximation phase and the fine-tuning phase, along with a multi-line programming algorithm and a programming-aware model retraining are proposed to reduce programming cycles and energy consumption. Experimental results demonstrate that our proposed method can reduce programming cycles and energy consumption by 46%-49 % and 63 % -64 %, respectively, compared to the state of the art. Additionally, the area and power overhead are negligible. Wei-Lun Chen, Fang-Yi Gu, Ing-Chao Lin, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann |
ICCAD | 2 |
| 2022 | WRAP: Weight RemApping and Processing in RRAM-based Neural Network Accelerators Considering Thermal EffectabstractResistive random-access memory (RRAM) has shown great potential for computing in memory (CIM) to support the requirements of high memory bandwidth and low power in neuromorphic computing systems. However, the accuracy of RRAM-based neural network (NN) accelerators can degrade significantly due to the intrinsic statistical variations of the resistance of RRAM cells, as well as the negative effects of high temperatures. In this paper, we propose a subarray-based thermal-aware weight remapping and processing framework (WRAP) to map the weights of a neural network model into RRAM subarrays. Instead of dealing with each weight individually, this framework maps weights into subarrays and performs subarray-based algorithms to reduce computational complexity while maintaining accuracy under thermal impact. Experimental results demonstrate that using our framework, inference accuracy losses of four DNN models are less than 2% compared to the ideal results and 1% with compensation applied even when the surrounding temperature is around 360K. Po-Yuan Chen, Fang-Yi Gu, Yu-Hong Huang, Ing-Chao Lin |
DATE | 2 |