Pai-Yu Tan

dblp:260/6601 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-7131-888XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2023 A Low-Bitwidth Integer-STBP Algorithm for Efficient Training and Inference of Spiking Neural Networks
abstract
Spiking neural networks (SNNs) that enable energy-efficient neuromorphic hardware are receiving growing attention. Training SNNs directly with back-propagation has demonstrated accuracy comparable to deep neural networks (DNNs). However, previous direct-training algorithms require high-precision floating-point operations, which are not suitable for low-power end-point devices. The high-precision operations also require the learning algorithm to run on high-performance accelerator hardware. In this paper, we propose an improved approach that converts the high-precision floating-point operations to low-bitwidth integer operations for an existing direct-training algorithm, i.e., the Spatio-Temporal Back-Propagation (STBP) algorithm. The proposed low-bitwidth Integer-STBP algorithm requires only integer arithmetic for SNN training and inference, which greatly reduces the computational complexity. Experimental results show that the proposed STBP algorithm achieves comparable accuracy and higher energy efficiency than the original floating-point STBP algorithm. Moreover, it can be implemented on low-power end-point devices to provide learning capability during inference, which are mostly supported by fixed-point hardware.
Pai-Yu Tan, Cheng-Wen Wu
ASP-DAC1
2023 A 40-nm 1.89-pJ/SOP Scalable Convolutional Spiking Neural Network Learning Core With On-Chip Spatiotemporal Back-Propagation
abstract
In recent years, progress in spiking neural network (SNN) research has generated growing interest in specialized SNN hardware. However, most of the hardware studies are about inference-only engines, and the training process for low-power endpoint devices remains arduous due to high numerical precision requirements. In this article, we introduce a scalable convolutional SNN learning core for energy-efficient training utilizing the spatiotemporal back-propagation (STBP) algorithm. We modify the STBP algorithm with five hardware-based enhancing methods, which minimize the hardware implementation cost without losing accuracy. We propose a unified core architecture encompassing three data-flow modes for three training phases. It also offers multicore scalability for wider and deeper models. A 40-nm prototype chip has been implemented, achieving peak training and inference efficiencies of 3 and 7.7 TOPS/W, respectively, at 90% input sparsity and an energy per synaptic operation (SOP) metric of 1.89 pJ/SOP. Based on the chip, our multicore prototype system attains a competitive accuracy of 99.1% on MNIST. We also exhibit the first on-chip training results on SVHN and CIFAR10, with an average of$36\times $improvement in efficiency compared with a typical commercial GPU platform.
Pai-Yu Tan, Cheng-Wen Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2022 A Decision Tree-Based Screening Method for Improving Test Quality of Memory Chips
abstract
There is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, our heuristics can screen out 25% more bad dies, i.e., we can improve the FT quality without additional cost.
Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao
ITC-Asia2
2022 Weak Die Screening by Feature Prioritized Random Forest for Improving Semiconductor Quality and Reliability
abstract
with the increasing demand for safety-critical products, the quality and reliability of semiconductor components are among the top priorities. In recent years, test data analytics by machine learning (ML) algorithms are widely considered to have great potential for improving the quality and reliability of semiconductor chips. In this work, we inspect a typical test flow of advanced semiconductor products, and propose an ML-based weak die screening method for improving the quality and reliability of shipped products. We propose the feature prioritized random forest (FPRF) model, which can fit smoothly into the existing test flow. We perform experiments on an advanced SRAM product using the FPRF model. We perform feature analysis based on the test data obtained from the final test (FT). After the FPRF screening, we are able to screen out more bad dies from those that have passed the FT. For an overkill rate of 12.93 %, the bad die hit rate can be as high as 96.55%. One can explore the FPRF model for other products as well.
Shian-Yu Lin, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao
ITC-Asia2
2022 Improving Test Quality of Memory Chips by a Decision Tree-Based Screening Method
abstract
There is a growing demand for high-reliability and high-quality integrated circuit (IC) products, while their test costs should be kept as low as possible. We investigate the test process of advanced memory chips, where the high temperature operating life (HTOL) test has been used to determine their intrinsic reliability. This high temperature sampling test can run from 168 to 1,000 hours, so it is time-consuming and expensive. Recently, machine learning (ML) algorithms have been used to solve classification problems, so far as good training data can be obtained. In our case, there is already a large amount of parametric test data generated from the existing test flow. Therefore, in this work, we propose a decision tree (DT)-based screening method to predict weak (unreliable) dies that would fail the HTOL test. We show that experienced test engineers can prioritize the parametric test data for better use of the DT model. Finally, we take advantage of the high interpretability of DT to develop the multi-feature heuristics, which can be used to improve the quality of final test (FT). Keeping the overkill rate at 0%, we can screen out 25% more bad dies in the 5nm SRAM case with the heuristics, and in the 4nm case, we can screen out 14% more bad dies, i.e., we can improve the FT quality without additional cost.
Ya-Chi Cheng, Pai-Yu Tan, Cheng-Wen Wu, Ming-Der Shieh, Chien-Hui Chuang, Gordon Liao
ITC2
2021 An Improved STBP for Training High-Accuracy and Low-Spike-Count Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) that facilitate energy-efficient neuromorphic hardware are getting increasing attention. Directly training SNN with backpropagation has already shown competitive accuracy compared with Deep Neural Networks. Besides the accuracy, the number of spikes per inference has a direct impact on the processing time and energy once employed in the neuromorphic processors. However, previous direct-training algorithms do not put great emphasis on this metric. Therefore, this paper proposes four enhancing schemes for the existing direct-training algorithm, Spatio-Temporal Back-Propagation (STBP), to improve not only the accuracy but also the spike count per inference. We first modify the reset mechanism of the spiking neuron model to address the information loss issue, which enables the firing threshold to be a trainable variable. Then we propose two novel output spike decoding schemes to effectively utilize the spatio-temporal information. Finally, we reformulate the derivative approximation of the non-differentiable firing function to simplify the computation of STBP without accuracy loss. In this way, we can achieve higher accuracy and lower spike count per inference on image classification tasks. Moreover, the enhanced STBP is feasible for the on-line learning hardware implementation in the future.
Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu
DATE1
2020 A 90nm 103.14 TOPS/W Binary-Weight Spiking Neural Network CMOS ASIC for Real-Time Object Classification
abstract
This paper introduces a low-power 90nm CMOS binary weight spiking neural network (BW-SNN) ASIC for real-time image classification. The chip maximizes data reuse through systolic arrays that house the entire 5-layer BW-SNN, requiring a minimum off-chip bandwidth for data access. The chip achieves 97.57% accuracy for real-time bottled-drink recognition, consuming only 0.62uJ per inference. For comparison purpose, it achieves 98.73% accuracy for MNIST hand-written character recognition, consuming only 0.59uJ per inference. The bottled-drink recognition is demonstrated at 300 fps that is well enough for many other real-time applications. The peak efficiency point is 103.14TOPS/W at a voltage of 0.6V, which outperforms other designs so far as we know. By normalizing to the 28nm technology node, the proposed ASIC is about 5× more efficient and 7× lower hardware cost as compared with the state-of-the-art designs.
Po-Yao Chuang, Pai-Yu Tan, Cheng-Wen Wu, Juin-Ming Lu
DAC2