Chih-Chyau Yang

dblp:38/283 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
5since 2021 · last 2024
0000-0001-6508-8160ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 LSvT: EEG Channel Localization and Selection via Training for BCI Applications
abstract
Electroencephalography (EEG) is widely utilized in neuroscience and clinical applications, serving as a noninvasive method for monitoring brain activity. However, despite its extensive use, the multitude of channels recorded by scalp electrodes presents challenges, including impractical usage, heightened model complexity, and potential overfitting issues during training. This paper addresses the challenges of high dimensionality in EEG data and introduces two innovative EEG channel selection algorithms, achieving significant reductions in channels, model size, and complexity while maintaining high classification accuracy. Validated through experiments on EEGNet and the MNE/BCI Competition IV 2a datasets, these algorithms prove valuable for practical and cost-efficient scenarios. LSvT-S emphasizes high channel reduction and low complexity, while LSvT-G targets faster channel selection, offering users choices. Experiments on MNE and BCI Competition IV 2a datasets show that LSvT-S achieves a remarkable 96.7% and 81.8% reduction in channels, along with 44% and 12.4% reductions in model size, and 96.2% and 80.8% in computation complexity, respectively. Meanwhile, LSvT-G achieves an additional speedup ranging from 32.8x to 7.6x.
Chun-Ming Huang, Wei-Lin Lai, Chih-Chyau Yang, Yi-Jie Hsieh, Chien-Ming Wu, Chu-Hui Lee
IJCNN3
2024 A 71.2-μW Speech Recognition Accelerator With Recurrent Spiking Neural Network
abstract
This paper introduces a 71.2-$\mu$W speech recognition accelerator designed for edge devices’ real-time applications, emphasizing an ultra low power design. Achieved through algorithm and hardware co-optimizations, we propose a compact recurrent spiking neural network with two recurrent layers, one fully connected layer, and a low time step (1 or 2). The 2.79-MB model undergoes pruning and 4-bit fixed-point quantization, shrinking it by 96.42% to 0.1 MB. On the hardware front, we take advantage ofmixed-level pruning,zero-skippingandmerged spiketechniques, reducing complexity by 90.49% to 13.86 MMAC/S. Theparallel time-step executionaddresses inter-time-step data dependencies and enables weight buffer power savings through weight sharing. Capitalizing on the sparse spike activity, an input broadcasting scheme eliminates zero computations, further saving power. Implemented on the TSMC 28-nm process, the design operates in real time at 100 kHz, consuming 71.2$\mu$W, surpassing state-of-the-art designs. At 500 MHz, it has 28.41 TOPS/W and 1903.11 GOPS/mm$^2$in energy and area efficiency, respectively.
Chih-Chyau Yang, Tian-Sheuan Chang
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 DLA-SP: A System Platform for Deep-Learning Accelerator Integration
abstract
Many deep-learning Accelerators (DLAs) are presented to meet high-performance needs in versatile deep-learning applications. However, they usually lack a flexible platform and flow for system integration. For assisting the professors and students in the academia of Taiwan to speed up their deep-learning accelerator implementation and verification of innovative system designs, the Taiwan Semiconductor Research Institute provides a new service platform named DLA-SP. The developed DLA-SP platform is highly flexible in contrast to existing DLA prototyping systems by providing three design resources: A parameterized DLA wrapper, an integration flow, and a hardware/software board support package. The parameterized DLA wrapper is capable of rate adaptation and format conversion between the DLA and the system. The number of DLA I/Os can be reduced to 29.7% to avoid the pad-limit DLA chip problem; The integration flow and methodology enable the DLA hardware system integration and DLA patterns reuse; The hardware/software board support package is capable of rapid deep-learning application development. The DLA-SP helps professors and students to concentrate their efforts on their deep-learning accelerators, and easily reuse system platforms, which greatly reduces the developing cycle of an embedded system. In this paper, a flexible system platform for deep-learning accelerators is presented, the accelerators can be easily integrated into our proposed DLA-SP system platform in both FPGA and chip ways for system verification and demonstration. A case study of a keyword-spotting accelerator is adopted as an example to illustrate how a DLA is integrated and works properly with our DLA-SP platform.
Chih-Chyau Yang, Fu-Chen Cheng 0001, Tsung-Jen Hsieh, Chien-Ming Wu, Chun-Ming Huang
IECON1
2023 A 1.6-mW Sparse Deep Learning Accelerator for Speech Separation
abstract
Low-power deep learning accelerators (DLAs) on the speech processing enable real-time applications on edge devices. However, most of the existing accelerators suffer from high-power consumption and focus on image applications only. This article presents a low-power accelerator for speech separation through algorithm and hardware optimizations. At the algorithm level, the model is compressed with structured sensitivity as well as unstructured pruning, and further quantized to the shifted 8-bit floating-point format instead of the 32-bit floating-point format. The computations with the zero kernel and zero activation values are skipped by decomposition of the dilated and transposed convolutions. At the hardware level, the compressed model is then supported by an architecture with eight independent multipliers and accumulators (MACs) with a simple zero-skipping hardware to take advantage of the activation sparsity and low-power processing. The proposed approach reduces the model size by 95.44% and computation complexity by 93.88%. The final implementation with the TSMC 40-nm process can achieve real-time speech separation and consumes 1.6-mW power when operated at 150 MHz. The normalized energy efficiency and area efficiency are 2.344 TOPS/W and 14.42 GOPS/mm2, respectively.
Chih-Chyau Yang, Tian-Sheuan Chang
IEEE Trans. Very Large Scale Integr. Syst.1
2022 A Real-Time 1280 × 720 Object Detection Chip With 585 MB/s Memory Traffic
abstract
Memory bandwidth has become the real-time bottleneck of current deep learning accelerators (DLAs), particularly for high definition (HD) object detection. Under resource constraints, this article proposes a low memory traffic DLA chip with joint hardware and software optimization. To maximize hardware utilization under memory bandwidth, we morph and fuse the object detection model into a group fusion-ready model to reduce intermediate data access. This reduces the YOLOv2’s feature memory traffic from 2.9 to 0.15 GB/s. To support group fusion, our previous DLA-based hardware employees a unified buffer with write-masking for simple layer-by-layer processing in a fusion group. When compared to our previous DLA with the same processing element (PE) numbers, the chip implemented in a 40-nm process supports$1280\times 720$at 30 frames per second (FPS) object detection and consumes$7.9\times $less external dynamic random access memory (DRAM) access energy, from 2607 to 327.6 mJ.
Kuo-Wei Chang, Hsu-Tung Shih, Tian-Sheuan Chang, Shang-Hong Tsai, Chih-Chyau Yang, Chien-Ming Wu, Chun-Ming Huang
IEEE Trans. Very Large Scale Integr. Syst.5
2019 A Smart Sensor Development Platform and Its System Demonstration
abstract
The smart sensor plays an important role for Internet of Things (IoT) development in the past few years. Due to the fast advance of IC fabrication and electronic design automation technologies, integrating a smart sensor design into a single chip has become practical. To assist the MEMS sensor teams in Taiwan academia to accelerate their smart sensor development, this paper presents a smart sensor development platform which consists of a common platform unit and a sensor unit. Our proposed smart sensor development platform provides the solutions of FPGA-based design, Smart Sensor on Chip (SSoC) design, and Smart Sensor in Package (SSiP) design for smart sensor development. The design flow and deliverables for the SSoC and SSiP smart sensor implementations are also presented in this paper. The proposed smart sensor common platform unit was taped out with UMC 0.18um process to perform the silicon proof. Moreover, this paper also presents a modularized wireless sensor system, MorSensor, to facilitate the smart sensor system demonstration. A case study for the FPGA-based smart sensor is also given in this paper. The experiment results show that the presented smart sensor platform is very suitable for smart sensor development, while MorSensor is suitable for the smart sensor system demonstration.
Chun-Ming Huang, Chih-Chyau Yang, Yi-Jie Hsieh, Chun-Wen Cheng, Yi-Jun Liu, Jia-Rong Chang, Yu-Tsang Chang, Chien-Ming Wu
ISCAS2
2017 A modular wireless sensor platform and its applications
abstract
In this paper, we propose a modular wireless sensor platform that consists of sensor modules. Each sensor module is a part of sensor system and in charge of one job in the system, such as computation, communication, output or sensing. Users can stack multiple modules together to build a unique sensor platform. Since users are able to easily replace one module with others, the proposed platform is highly extendable and reusable. Besides, we also design different kinds of mounts, so that sensors can be mounted on objects. This is especially helpful for some moving sensing applications such as attitude monitoring. To demonstrate the proposed platform, we show an alcohol detection application in the paper. The results show that the proposed platform is suitable for academic researches and industrial prototype verification.
Chun-Ming Huang, Yi-Jie Hsieh, Wei-Lin Lai, Yi-Jun Liu, Chun-Ying Juan, Ssu-Ying Chen, Jin-Ju Chue, Chih-Chyau Yang, Chien-Ming Wu
ISCAS9
2010 A packet-based emulating platform with serializer/deserializer interface for heterogeneous IP verification
abstract
This paper proposes a packet-based verification platform with serial link interface for emulating the hardware of the heterogeneous IPs before tape out. With the serial link interface Serializer/Deserializer (SerDes) added between IPs, significant amount of pin counts can be reduced in the platform. An adapter is inserted between IP and SerDes to convert parallel bus into packets and handle the handshaking. Under our proposed adapter architecture and handshaking scheme, the limitation on the number of the master adapter is eliminated compared with Bus-based Advanced High-performance Bus (AHB) architecture. Simulation results show the data transfer through our proposed architecture works correctly without the limitation on the number of masters. With the proposed adapter and SerDes architecture, the number of required signals in the interconnect is reduced from 79 to two for the AHB bus.
Chih-Hsing Lin, Yung-Chang Chang, Wen-Chih Huang, Wei-Chih Lai, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Chun-Ming Huang, Chih-Chyau Yang, Shih-Lun Chen
ISCAS9
2009 Implementation and Prototyping of a Complex Multi-project System-on-a-chip
abstract
A silicon prototyping methodology is presented for Multi-Project System-on-a-Chip (MP-SoC) implementation. A multi-projects platform was created for integrating heterogeneous SoC projects into a single chip. The total silicon prototyping cost of these projects can be greatly reduced by sharing a common platform. To demonstrate the effectiveness of the proposed methodology, a MP-SoC chip was implemented with eleven SoC projects sharing the common platform. The total silicon area is about 37.97 mm2in the TSMC 0.13 um CMOS generic logic process technology. Compared with the total chip area 129.39 mm2by implementing these projects separately, the results show that there are 91.42 mm2silicon areas reduced by the MP-SoC platform. In order to verify MP-SoC through silicon prototyping, a system modeling and hardware/ software co-design virtual platform were implemented. A configurable SoC prototyping system, namely CONCORD, is also created as a verification platform for emulating the hardware of MP-SoC before chip being taped out. The CONCORD system provides higher connection flexibility, modularization, and architecture consistence than conventional FPGA systems.
Chun-Ming Huang, Chien-Ming Wu, Chih-Chyau Yang, Wei-De Chien, Shih-Lun Chen, Chi-Shi Chen, Jiann-Jenn Wang, Chin-Long Wey
ISCAS3
2008 PrSoC: Programmable System-on-chip (SoC) for silicon prototyping
abstract
This paper presents a Programmable SoC (System- on-chip) design methodology which integrates multiple heterogeneous SoC design projects into a single chip such that the total silicon prototyping cost for these projects can be greatly reduced by sharing the common SoC platform. Results show that an integrated SoC platform is comprised of eight SoC projects. When these eight SoC projects are designed separately, the total area is approximately 143.03mm , while the area of the integrated platform is about 24.43mm . The area reduction is significant, so is the fabrication cost. Once the integrated platform chip is fabricated, three programming schemes are carried out to allow the integrated chip to act as the individual SoC design projects. A test chip is designed and implemented using the TSMC 0.13um CMOS generic logic process technology.
Chun-Ming Huang, Chien-Ming Wu, Chih-Chyau Yang, Chin-Long Wey
ISCAS3