Yi-Ta Chen

dblp:249/2284 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-1724-0206ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object Detection
abstract
Object detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively.
Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu
ISCAS6
2023 Joint Optimization of Dimension Reduction and Mixed-Precision Quantization for Activation Compression of Neural Networks
abstract
Recently, deep convolutional neural networks (CNNs) have achieved eye-catching results in various applications. However, intensive memory access of activations introduces considerable energy consumption, resulting in a great challenge for deploying CNNs on resource-constrained edge devices. Existing research utilizes dimension reduction (DR) and mixed-precision (MP) quantization separately to reduce computational complexity without paying attention to their interaction. Such naïve concatenation of different compression strategies ends up with suboptimal performance. To develop a comprehensive compression framework, we propose an optimization system by jointly considering DR and MP quantization, which is enabled by independent groupwise learnable MP schemes. Group partitioning is guided by a well-designed automatic group partition mechanism that can distinguish compression priorities among channels, and it can deal with the tradeoff between model accuracy and compressibility. Moreover, to preserve model accuracy under low bit-width quantization, we propose a dynamic bit-width searching technique to enable continuous bit-width reduction. Our experimental results show that the proposed system reaches 69.03%/70.73% with average 2.16/2.61 bits per value on Resnet18/MobileNetV2, while introducing only approximately 1% accuracy loss of the uncompressed full-precision models. Compared with individual activation compression schemes, the proposed joint optimization system reduces 55%/9% (−2.62/−0.27 bits) memory access of DR and 55%/63% (−2.60/−4.52 bits) memory access of MP quantization, respectively, on Resnet18/ MobileNetV2 with comparable or even higher accuracy.
Yu-Shan Tai, Cheng-Yang Chang, Chieh-Fang Teng, Yi-Ta Chen, An-Yeu Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 S-QRD-ELM: Scalable QR-Decomposition-Based Extreme Learning Machine Engine Supporting Online Class-Incremental Learning for ECG-Based User Identification
abstract
User identification enables secure access to data and machines in smart factories. Compared with other modalities, ECG-based user identification is rising due to its intrinsic liveness proof and invulnerability to spoofing without contact. On the other hand, as new employees are registered at the factory, the ECG-based user identification system needs to be updated based on the new coming data. This scenario can be defined as an online class-incremental learning (O-CIL) problem. By exploiting hardware-software co- design, this work presents a Scalable QR-decomposition-based extreme learning machine (S-QRD-ELM) engine that can effectively and efficiently support O-CIL for ECG-based user identification. At the software level, we apply the concept of “the others” class and inversion-free QR-decomposition (QRD) recursive least squares to the S-QRD-ELM. This makes S-QRD-ELM achieve 79.7% higher accuracy in the O-CIL scenario compared with the neural network trained with back-propagation (BP-NN). At the hardware level, a one-dimensional diagonally-mapped linear array (1D-DMLA) is proposed to efficiently compute the QRD and back-substitution (BS) operations inside the S-QRD-ELM, reducing 98.5% of the silicon area. Moreover, the integrated processing element (PE) design with the unified COordinate Rotation DIgital Computer (u-CORDIC) further reduces 15.3% of the area and 22.4% of the power consumption. This engine is fabricated in 40nm CMOS technology with a$1.33\times 1.33$mm2 die area. The chip achieves$0.02\mu \text{J}$/sample and$2.47\mu \text{J}$/sample inferencing and learning energy efficiency, respectively, which is$6.4\times $and$28.5\times $than the state-of-the-art. To the best of our knowledge, the proposed highly energy-efficient S-QRD-ELM engine is the first chip to meet the requirements of O-CIL for ECG-based user identification.
Yi-Ta Chen, Yu-Chuan Chuang, Li-Sheng Chang, An-Yeu Wu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2022 An Effective Entropy-Assisted Mind-Wandering Detection System Using EEG Signals of MM-SART Database
abstract
Mind-wandering (MW), which is usually defined as a lapse of attention has negative effects on our daily life. Therefore, detecting when MW occurs can prevent us from those negative outcomes resulting from MW. In this work, we first collected a multi-modal Sustained Attention to Response Task (MM-SART) database for MW detection. Eighty-two participants' data were collected in our dataset. For each participant, we collected measures of 32-channels electroencephalogram (EEG) signals, photoplethysmography (PPG) signals, galvanic skin response (GSR) signals, eye tracker signals, and several questionnaires for detailed analyses. Then, we propose an effective MW detection system based on the collected EEG signals. To explore the non-linear characteristics of the EEG signals, we utilize entropy-based features. The experimental results show that we can reach 0.712 AUC score by using the random forest (RF) classifier with the leave-one-subject-out cross-validation. Moreover, to lower the overall computational complexity of the MW detection system, we propose correlation importance feature elimination (CIFE) along with AUC-based channel selection. By using two most significant EEG channels, we can reduce the training time of the classifier by 44.16%. By applying CIFE on the feature set, we can further improve the AUC score to 0.725 but with only 14.6% of the selection time compared with the recursive feature elimination (RFE). Finally, we can apply the current work to educational scenarios nowadays, especially in remote learning systems.
Yi-Ta Chen, Hsing-Hao Lee, Ching-Yen Shih, Zih-Ling Chen, Win-Ken Beh, Su-Ling Yeh, An-Yeu Wu
IEEE J. Biomed. Health Informatics1
2021 A Scalable Extreme Learning Machine (S-ELM) for Class-Incremental ECG-Based User Identification
abstract
User identification using electrocardiogram (ECG) is emerging due to the uniqueness and convenience of ECG signals. In addition, in real world applications, new subjects may enter the existed identification system and be authorized to access the private data. Therefore, we propose a scalable extreme learning machine (S-ELM) to meet the demand for class-incremental ECG-based user identification. We first prove that the output weight of the S-ELM learnt in a class-incremental manner is as same as that of a regular ELM which prior has the information of the number of total class. Therefore, in our experiment, S-ELM immunes from the catastrophic forgetting phenomenon, which is a common problem in class-incremental scenarios. Comparing to another class-incremental extreme learning machine such as progressive ELM, S-ELM outperforms progressive ELM by 7% accuracy in online dataset. Comparing to another commonly applied classifier, support vector machine (SVM) with linear and radial basis function (RBF) kernels, S-ELM shows its efficiency by 13.35% and 10.54% higher accuracy but only spends 5.09% and 3.48% of the inference time. Therefore, the proposed S-ELM is promising for the class-incremental ECG-based user identification.
Cheng-Lin Lee, Yi-Ta Chen, An-Yeu Wu
ISCAS2