Jianbiao Xiao

dblp:285/0974 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-7535-4378ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 An Energy-Efficient and End-to-End Keyword Spotting Processor Using 3D Broadcast Sparse Computing Array and Hybrid Sparsity Strategy Pruning
abstract
Keyword spotting (KWS) has been widely applied in various low-power offline voice wake-up applications, with core requirements for low energy consumption, low cost, and low latency. Neural network (NN)-based KWS processors can achieve state-of-the-art (SOTA) accuracy, but NNs also introduce considerable computation and parameter overhead. NN accelerators based on model pruning and sparse computing can leverage the sparsity of activations and weights to reduce storage, energy consumption, and latency, albeit at the cost of increased hardware area and dense computation energy. However, compared to large NN models, KWS NN models typically have far fewer parameters (often less than a million), resulting in limited redundancy for pruning and thus low sparsity. Consequently, traditional model pruning and sparse computing methods are not suitable for KWS processor design. To address this, we propose: 1) a hybrid sparse pruning architecture (HSPA) that maps joint weight-activation pruning (JWAP) to weight-only pruning (WP), avoiding high hardware overhead JWAP-specific circuits while maintaining high sparsity; 2) a 3D broadcast sparse computing array (3D-BSCA), a novel 3D computing array supporting unstructured WP that uses a broadcast strategy to replace complex data routing and reduce hardware costs; and 3) a joint activation-index computation engine (JAICE) that designs and reuses a reconfigurable computing engine to generate zero-value indices, eliminating additional index computation hardware overhead. This work was implemented with layout and simulation in a 65nm CMOS process, achieving 91.9% accuracy and 1.37 μJ energy consumption on KWS tasks. Additionally, HSPA enabled a high sparsity of up to 55%, reducing energy consumption by 43.9%, while 3D-BSCA and JAICE collectively reduced hardware overhead by 40% compared to conventional sparse computing.
Jianbiao Xiao, Hengxin Wang, Chiyu Zou, Yangning Hu, Jun Zhou 0017
ACM Great Lakes Symposium on VLSI1
2023 A High Accuracy and Low Power CNN-Based Environmental Sound Classification Processor
abstract
The environmental sound classification (ESC) has attracted increasing attention as the environmental sound contains a wealth of information that can be used to detect particular events. However, so far, most of the existing work in ESC still remains in the stage of algorithm design and the design of ESC processor has not been thoroughly investigated. The existing ESC processor designs have issues in meeting low power consumption and high accuracy simultaneously due to the lack of joint-optimization between algorithm and hardware, and very few work has demonstrated a complete ESC system containing all the necessary modules. In this work, a high accuracy and low power CNN-based ESC processor has been proposed, featuring: 1) a big-small CNN-based reconfigurable ESC processing hardware architecture to reduce the power consumption and hardware overhead while maintaining high classification accuracy. 2) a Mel feature adaptation engine reusing the neural network processing unit to further reduce the power consumption. 3) an event-driven ESC processing technique to reduce the inference time and the power consumption. The design has been implemented on a Kintex-7 FPGA and achieves low power consumption of 0.313W with high accuracy of 84.5% for the ESC-50 dataset, outperforming other state-of-the-art ESC processors.
Lujie Peng, Junyu Yang, Longke Yan, Xiben Jiao, Jianbiao Xiao, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.6
2022 An energy-efficient seizure detection processor using event-driven multi-stage CNN classification and segmented data processing with adaptive channel selection
abstract
Recently wearable EEG monitoring devices with seizure detection processor using convolutional neural network (CNN) have been proposed to detect the seizure onset of patients in real time for alert or stimulation purpose. High energy efficiency and accuracy are required for the seizure detection processor due to the tight energy constraint of wearable devices. However, the use of CNN and multi-channel processing nature of seizure detection result in significant energy consumption. In this work, an energy-efficient seizure detection processor is proposed, featuring multi-stage CNN classification, segmented data processing and adaptive channel selection to reduce the energy consumption while achieving high accuracy. The design has been fabricated and tested using a 55nm process technology. Compared with several state-of-the-art designs, the proposed design achieves the lowest energy per classification (0.32 μJ) with high sensitivity (97.78%) and low false positive rate per hour (0.5).
Jiahao Liu 0006, Zirui Zhong, Hui Qiu, Jianbiao Xiao, Jiajing Fan, Zhaomin Zhang, Sixu Li, Siqi Yang 0002, Weiwei Shan, Shuisheng Lin, Liang Chang 0002, Jun Zhou 0017
DAC5
2022 ULSED: An ultra-lightweight SED model for IoT devices
Lujie Peng, Junyu Yang, Jianbiao Xiao, Mingxue Yang, Yujiang Wang 0003, Haojie Qin, Xiaorong Li, Jun Zhou 0017
J. Parallel Distributed Comput.3
2022 ULECGNet: An Ultra-Lightweight End-to-End ECG Classification Neural Network
abstract
ECG classification is a key technology in intelligent electrocardiogram (ECG) monitoring. In the past, traditional machine learning methods such as support vector machine (SVM) and K-nearest neighbor (KNN) have been used for ECG classification, but with limited classification accuracy. Recently, the end-to-end neural network has been used for ECG classification and shows high classification accuracy. However, the end-to-end neural network has large computational complexity including a large number of parameters and operations. Although dedicated hardware such as field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) can be developed to accelerate the neural network, they result in large power consumption, large design cost, or limited flexibility. In this work, we have proposed an ultra-lightweight end-to-end ECG classification neural network that has extremely low computational complexity (∼8.2k parameters & ∼227k multiplication/addition operations) and can be squeezed into a low-cost microcontroller (MCU) such as MSP432 while achieving 99.1% overall classification accuracy. This outperforms the state-of-the-art ECG classification neural network. Implemented on MSP432, the proposed design consumes only 0.4 mJ and 3.1 mJ per heartbeat classification for normal and abnormal heartbeats respectively for real-time ECG classification.
Jianbiao Xiao, Jiahao Liu 0006, Huanqi Yang, Ning Wang 0070, Zhen Zhu 0005, Yu Long 0005, Liang Chang 0002, Jun Zhou 0017
IEEE J. Biomed. Health Informatics1
2021 Energy-efficient computing-in-memory architecture for AI processor: device, circuit, architecture perspective
Liang Chang 0002, Zhaomin Zhang, Jianbiao Xiao, Zhen Zhu 0005, Weihang Li, Zixuan Zhu 0001, Siqi Yang 0002, Jun Zhou 0017
Sci. China Inf. Sci.4
2021 A Fast and Energy-Efficient SNN Processor With Adaptive Clock/Event-Driven Computation Scheme and Online Learning
abstract
In the recent years, the spiking neural network (SNN) has attracted increasing attention due to its low energy consumption and online learning potential. However, the design of SNN processor has not been thoroughly investigated in the past, resulting in limited performance and energy consumption. In this work, a fast and energy-efficient SNN processor with adaptive clock/event-driven computation scheme and online learning capability has been proposed. Several techniques have been proposed to reduce the computation time and energy consumption, including Adaptive Clock- and Event-Driven Computing Scheme, Neighboring PE Borrowing Technique, Compressed Spike Routing Technique and Reconfigurable PE for Inference and Learning. Implemented on a Virtex-7 FPGA, the proposed design achieves computation time of 3.15 ms/image, inference energy consumption of$0.028~\mu $J/synapse/image and online learning energy consumption of$0.297~\mu $J/synapse/image for the MNIST 10-class dataset, which outperform several state-of-the-art SNN processors. The proposed SNN processor is suitable for real-time and energy-constrained applications.
Sixu Li, Zhaomin Zhang, Ruixin Mao, Jianbiao Xiao, Liang Chang 0002, Jun Zhou 0017
IEEE Trans. Circuits Syst. I Regul. Pap.4