Shu-Yin Ho

dblp:278/8904 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021
YearPublicationVenuePosition
2023 Reliable Brain-inspired AI Accelerators using Classical and Emerging Memories
abstract
By taking inspiration from the operation of biological brains, emerging brain-inspired hardware has the potential to revolutionize the way computations are performed. Brain-inspired computing can be realized using both classical CMOS and emerging beyond-CMOS technologies, whereas the latter holds the promise to provide substantial energy savings akin to the employment of non-volatile memories. One way to implement highly efficient brain-inspired AI applications is through analog computing schemes, such as Integrate-and-Fire (IF) Spiking Neural Networks (SNNs), which can be implemented using both CMOS and beyond-CMOS technologies as synaptic storage. However, managing the inherent degradation of computing accuracy in analog circuits and mitigating their effects on the predictive accuracy of AI systems remains a key challenge due to the inherent nature of analog computing.In this paper, we discuss how the aforementioned challenges can be addressed. In the first part, we present our SPICE-Torch, a framework that connects low-level SPICE simulations of circuits and memories performing analog computations with high-level accuracy evaluations of NN models based on PyTorch. Furthermore, we present an example of neuromorphic optimization using classical CMOS technology. In the second part, we introduce memristors as an emerging beyond-CMOS technology that can retain their state without any outside influence and are well-suited for brain-inspired neuromorphic hardware. We demonstrate that brain-inspired hardware, realized using classical CMOS or beyond-CMOS technologies, has the potential to revolutionize the way we process information and solve complex computation problems. Nevertheless, to harness its full potential, reliability issues have to be managed carefully and HW/SW codesign is key. Our presented framework SPICE-Torch, which connects low-level SPICE simulations of circuits performing analog computations with high-level accuracy evaluations of NN models based on PyTorch is available as open-source in https://github.com/myay/SPICE-Torch.
Mikail Yayla, Simon Thomann, Md. Mazharul Islam 0006, Ming-Liang Wei, Shu-Yin Ho, Ahmedullah Aziz, Chia-Lin Yang, Jian-Jia Chen, Hussam Amrouch
VTS5
2023 Impact of Non-Volatile Memory Cells on Spiking Neural Network Annealing Machine With In-Situ Synapse Processing
abstract
Solving constraint satisfaction problems (CSPs) is in high demand for various applications. SNN serves as a competitive annealing machine that can solve the CSP more efficiently than well-known Metropolis sampling and Hopfield networks. NVM-based crossbars with analog Integrate and Fire (IF) neurons can evolve the state of SNN to solve CSP more efficiently. However, analog computations inherently suffer from imprecisions in NVM cells, e.g., current variation, OFF-state leakage, and temperature-induced drift. We are the first to analyze the impacts of various memory technologies, including 2T-NOR, FeFET, WOx ReRAM, and HfOx ReRAM, on solving the Ising model, Sudoku, and Traveling-salesman-problem (TSP). The results show that both 2T-NOR Flash and FeFET with normalized standard deviation( ${\sigma}/{u}$ ) $<$ $5\%$ and ON-OFF ratio $>$ $1000$ are both ideal candidates as synapse devices at room temperature, while other devices suffer from the effects of current variation and OFF-state leakage, which would require the neuron circuits to have infeasible membrane capacitance size. However, the drift of cell current and the reduction of the ON-OFF ratio drops the success rate as the temperature increases. The success rate of solving TSP drops by 60 $\%$ and 90 $\%$ while the temperature increases from 300K to 358K for 2T-NOR and FeFET, respectively. Throughout the simulation, we show that the transistor-based memory is suggested to be a synapse device. Yet, we also find that the tolerance of temperature is inevitable under limited capacitance. Exploration of temperature-tolerated design of circuit and memory design is still in demand for future works.
Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Hussam Amrouch, Chia-Lin Yang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 Binarized SNNs: Efficient and Error-Resilient Spiking Neural Networks through Binarization
abstract
Spiking Neural Networks (SNNs) are considered the third generation of NNs and can reach similar accuracy as conventional deep NNs, but with a considerable improvement in efficiency. However, to achieve high accuracy, state-of-the-art SNNs employ stochastic spike coding of the inputs, requiring multiple cycles of computation. Because of this and due to the nature of analog computing, it is required to accumulate and hold the charges of multiple cycles, necessitating a large membrane capacitor. This results in high energy, long latency, and expensive area costs, constituting one of the major bottlenecks in analog SNN implementations. Membrane capacitor size determines the precision of the firing time. Hence reducing the capacitor size considerably degrades the inference accuracy. To alleviate this, we focus on bridging the gap between binarized NNs (BNNs) and SNNs. BNNs are rapidly emerging as an attractive alternative for NNs due to their high efficiency and error tolerance. In this work, we evaluate the impact of deploying error-resilient BNNs, i.e. BNNs that have been proactively trained in the presence of errors, on analog implementation of SNNs. We show that for BNNs, the capacitor size and latency can be reduced significantly compared to state-of-the-art SNNs, which employ multi-bit models. Our experiments demonstrate that when error-resilient BNNs are deployed on analog-based SNN accelerator, the size of the membrane capacitor is reduced by 50%, the inference latency is decreased by two orders of magnitude, and energy is reduced by 57% compared to the baseline 4-bit SNN implementation, under minimal accuracy cost.
Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Chia-Lin Yang, Hussam Amrouch
ICCAD3
2020 On Minimizing Analog Variation Errors to Resolve the Scalability Issue of ReRAM-Based Crossbar Accelerators
abstract
Crossbar accelerators with a resistive random-access memory (ReRAM) are a promising solution for accelerating neural network applications. The advantages of achieving high computation throughput per watt make ReRAM-based crossbar accelerators become a potential solution for accelerating inference operations in the Internet of Things and edge devices. Due to the analog variation errors, the launched ReRAM-based crossbar accelerators can only perform well when each ReRAM cell is used to represent a limited number of data bits. To make such ReRAM-based crossbar accelerators applicable in wide application scenarios, several proposed researches target at binary neural networks and focus on the chip designs in relieving the implementation challenges on computation accuracy for realizing single-bit ReRAM-based crossbar accelerators. Even though several small-sized ReRAM-based crossbar accelerators are announced, the scalability issue hinders ReRAM-based crossbar accelerators from being scaled up. That is, when there are more and more wordline in an ReRAM-based crossbar accelerator, the analog variation error is amplified and thus seriously degrades the computation accuracy. In this work, we propose an adaptive data manipulation strategy to substantially reduce analog variation errors so as to fill up the gap on scaling up the ReRAM-based crossbar accelerators. In particular, a weightrounding design is proposed to manipulate data to minimize overlapping variation so that the number of wordlines can be scaled up. In addition, an input subcycling design is proposed to further trade tolerable errors with neural networks' execution time. Moreover, a bitline redundant design is proposed to trade acceptable space overhead for eliminating the analog variation errors. The emulation experiments show that the proposed adaptive data manipulation strategy can improve the accuracy in running MNIST and CIFAR-10 by 1.3× and 2.6× with nearly no management penalty and hardware cost. The experimental results also show the close-to-ideal-case accuracy by substantially reducing analog variation errors.
Yao-Wen Kang, Chun-Feng Wu, Yuan-Hao Chang 0001, Tei-Wei Kuo, Shu-Yin Ho
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5