EDBT 2026 Demo / reviewers in the wild / expert
Kangjun Bai
dblp:184/4445 · also Kang Jun Bai
· DBLP profile ↗
10ranked-venue papers
7as first author
4since 2021 · last 2023
0000-0003-4437-0006ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 7 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Enabling a New Methodology of Neural Coding: Multiplexing Temporal Encoding in Neuromorphic ComputingabstractFrom rate to temporal encoding, spiking information processing has demonstrated advantages across diverse neuromorphic applications. In the aspects of data capacity and robustness, multiplexing encoding outperforms alternative encoding schemes. In this work, we aim to implement a new class of multiplexing temporal encoders, patterning stimuli in multiple timescales to improve the information processing capability, and robustness of systems deployed in noisy environments. Benefitted by the internal reference frame using subthreshold membrane oscillation (SMO), the encoded spike patterns are less sensitive to the input noise, increasing the encoder’s robustness. Our design results in a tremendous saving on power consumption and silicon area compared with the power-hungry analog-to-digital converters. Furthermore, a working prototype of the multiplexing temporal encoder built based on an interspike interval (ISI) encoding scheme is implemented on a silicon chip using the standard 180-nm CMOS process. To the best of our knowledge, our introduced encoder demonstrates the first integrated circuit (IC) implementation of neural encoding with multiplexing topology. Finally, the accuracy and efficiency of our design are evaluated through standard machine learning benchmarks, including Modified National Institute of Standards and Technology (MNIST), Canadian Institute For Advanced Research (CIFAR)-10, Street View House Number (SVHN), and spectrum sensing in high-speed communication networks. While our multiplexing temporal encoder demonstrates a higher classification accuracy across all the benchmarks, the power consumption and dissipated energy per spike reach merely$2.6~\mu \text {W}$and 95 fJ/spike, respectively, with an effective frame rate of 300 MHz. Compared with alternative encoding schemes, our multiplexing temporal encoder achieves at most 100% higher data capacity, 11.4% more accurate in classification, and 25% more robust against noise. Compared with the state-of-the-art designs, our work achieves up to$105 \times $power efficiency without significantly increasing the silicon area. Honghao Zheng, Kangjun Bai, Yang Yi 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Diagnosing Clinical Diseases using an Edge-Enabled Deep Learning TechnologyabstractAlong with the development of high-speed communication networks, edge-enabled mobile devices have opened new possibilities for diagnosing health conditions or developing suitable treatment plans. While the latest deep learning technology has deployed to restructure and translate complex medical applications, the costly training operation using large-scale neural networks with tremendous amount of data remain the major challenge. In this work, we take advantages of reservoir computing to develop a reliable and low-cost medical diagnostic system for edge-enabled devices. Specifically, an echo state network (ESN) was trained to discover non-obvious correlation and likelihood from biomedical data with respect to various patients. Through the determination of cardiovascular and coronavirus diseases, numerical evaluations demonstrated advantage of ESN against the state-of-the-art. At particularly no computation overhead, ESN precisely described the prediction tasks of health conditions, offering improvements of up to 1000x in sample reduction, 175x in training speedup, and 15 percentage points in prediction accuracy. Kangjun Bai, Yang Yi 0002 |
SEC | 1 |
| 2021 | A Hybrid FPGA-ASIC Delayed Feedback Reservoir System to Enable Spectrum Sensing/Sharing for Low Power IoT Devices ICCAD Special Session PaperabstractThe delayed feedback reservoir (DFR) network is a delay-dynamic architecture that incorporates time in its training and inference. This quality enables DFR networks to proficiently model time series in a scalable architecture with only one nonlinear neuron. Previous studies have highlighted the accuracy and energy efficiency of DFR networks in ASIC implementations; however, these approaches are limited by hardcoded weights and static reservoir architectures. In this work, we introduce a hybrid FPGA-ASIC DFR system that combines the flexibility of a FPGA platform with the energy efficiency of an ASIC. To be specific, the FPGA allows for dynamic reconfiguration and training of the readout weights during runtime, while the ASIC provides an analog activation function for the single neuron. The accuracy and energy consumption of the introduced system is demonstrated for the applications of NARMA10 as well as MIMO spectrum sensing which is a critical component of dynamic spectrum sharing/access for 5G/beyond-5G systems. Results showcase the potential to enable on-board intelligence for future wireless systems, especially for Internet of Things (IoT) devices in low-power environments. Osaze Shears, Kangjun Bai, Lingjia Liu 0001, Yang Yi 0002 |
ICCAD | 2 |
| 2021 | Spatial-Temporal Hybrid Neural Network With Computing-in-Memory ArchitectureabstractDeep learning (DL) has gained unprecedented success in many real-world applications. However, DL poses difficulties for efficient hardware implementation due to the needs of a complex gradient-based learning algorithm and the required high memory bandwidth for synaptic weight storage, especially in today's data-intensive environment. Computing-in-memory (CIM) strategies have emerged as an alternative for realizing energy-efficient neuromorphic applications in silicon, reducing resources and energy required for neural computations. In this work, we exploit a CIM-based spatial-temporal hybrid neural network (STHNN) with a unique learning algorithm. To be specific, we integrate both multilayer perceptron and recurrent-based delay-dynamical system, making the network becomes linear separable while processing information in both spatial and temporal domains, better yet, reducing the memory bandwidth and hardware overhead through the CIM architecture. The prototype fabricated in 180 nm CMOS process is built of fully-analog components, yielding an average on-chip classification accuracy up to 86.9% on handprinted alphabet characters with a power consumption of 33 mW. Beyond that, through the handwritten digit database and the radio frequency fingerprinting dataset, software-based numerical evaluations offer 1.6 -to- 9.8 × and 1.9 -to- 4.4 × speedup, respectively, without significantly degrading its classification accuracy compared to the cutting-edge DL approaches. Kangjun Bai, Lingjia Liu 0001, Yang Yi 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Detection Through Deep Neural Networks: A Reservoir Computing Approach for MIMO-OFDM Symbol DetectionabstractThe Reservoir Computing, a neural computing framework suited for temporal information processing, utilizes a dynamic reservoir layer for high-dimensional encoding, enhancing the separability of the network. In this paper, we exploit a Deep Learning (DL)-based detection strategy for Multiple-input, Multiple-output Orthogonal Frequency-Division Multiplexing (MIMO-OFDM) symbol detection. To be specific, we introduce a Deep Echo State Network (DESN), a unique hierarchical processing structure with multiple time intervals, to enhance the memory capacity and accelerate the detection efficiency. The resulting hardware prototype with the hybrid memristor-CMOS co-design provides the in-memory computing and parallel processing capabilities, significantly reducing the hardware and power overhead. With the standard 180nm CMOS process and memristive synapses, the introduced DESN consumes merely 105mW of power consumption, exhibiting 16.7% power reduction compared to shallow ESN designs even with more dynamic layers and associated neurons. Furthermore, numerical evaluations demonstrate advantages of the DESN over state-of-the-art detection techniques in the literate for MIMO-OFDM systems even with a very limited training set, yielding a 47.8% improvement against conventional symbol detection techniques. Kangjun Bai, Lingjia Liu 0001, Zhou Zhou 0002, Yang Yi 0002 |
ICCAD | 1 |
| 2020 | A Training-Efficient Hybrid-Structured Deep Neural Network With Reconfigurable Memristive SynapsesabstractThe continued success in the development of neuromorphic computing has immensely pushed today's artificial intelligence forward. Deep neural networks (DNNs), a brainlike machine learning architecture, rely on the intensive vector-matrix computation with extraordinary performance in data-extensive applications. Recently, the nonvolatile memory (NVM) crossbar array uniquely has unvailed its intrinsic vector-matrix computation with parallel computing capability in neural network designs. In this article, we design and fabricate a hybrid-structured DNN (hybrid-DNN), combining both depth-in-space (spatial) and depth-in-time (temporal) deep learning characteristics. Our hybrid-DNN employs memristive synapses working in a hierarchical information processing fashion and delay-based spiking neural network (SNN) modules as the readout layer. Our fabricated prototype in 130-nm CMOS technology along with experimental results demonstrates its high computing parallelism and energy efficiency with low hardware implementation cost, making the designed system a candidate for low-power embedded applications. From chaotic time-series forecasting benchmarks, our hybrid-DNN exhibits 1.16×-13.77× reduction on the prediction error compared to the state-of-the-art DNN designs. Moreover, our hybrid-DNN records 99.03% and 99.63% testing accuracy on the handwritten digit classification and the spoken digit recognition tasks, respectively. Kangjun Bai, Qiyuan An, Lingjia Liu 0001, Yang Yi 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Deep-DFR: A Memristive Deep Delayed Feedback Reservoir Computing System with Hybrid Neural Network TopologyabstractDeep neural networks (DNNs), the brain-like machine learning architecture, have gained immense success in data-extensive applications. In this work, a hybrid structured deep delayed feedback reservoir (Deep-DFR) computing model is proposed and fabricated. Our Deep-DFR employs memristive synapses working in a hierarchical information processing fashion with DFR modules as the readout layer, leading our proposed deep learning structure to be both depth-in-space and depth-in-time. Our fabricated prototype along with experimental results demonstrate its high energy efficiency with low hardware implementation cost. With applications on the image classification, MNIST and SVHN, our Deep-DFR yields a 1.26~7.69X reduction on the testing error compared to state-of-the-art DNN designs. Kangjun Bai, Qiyuan An, Yang Yi 0002 |
DAC | 1 |
| 2018 | Enabling a new era of brain-inspired computing: energy-efficient spiking neural network with ring topologyabstractThe reservoir computing, an emerging computing paradigm, has proven its benefit to multifarious applications. In this work, we successfully designed and fabricated an analog delayed feedback reservoir (DFR) chip. Measurement results demonstrate its rich dynamic behaviors and high energy efficiency. System performance, as well as the robustness, are evaluated. The application of video frame recognition is investigated using a hybrid neural network, which employs the multilayer perceptron (MLP) training model as the readout layer of our designed DFR system, and yields 98% classification accuracy. Compared to results of using the MLP training only, our hybrid training model exhibits much higher recognition rate and accuracy. Kangjun Bai, Kian Hamedani, Yang Yi 0002 |
DAC | 1 |
| 2018 | DFR: An Energy-efficient Analog Delay Feedback Reservoir Computing System for Brain-inspired ComputingabstractNeuromorphic computing, which is built on a brain-inspired silicon chip, is uniquely applied to keep pace with the explosive escalation of algorithms and data density on machine learning. Reservoir computing, an emerging computing paradigm based on the recurrent neural network with proven benefits across multifaceted applications, offers an alternative training mechanism only at the readout stage. In this work, we successfully design and fabricate an energy-efficient analog delayed feedback reservoir (DFR) computing system, which is built upon a temporal encoding scheme, a nonlinear transfer function, and a dynamic delayed feedback loop. Measurement results demonstrate its high energy efficiency with rich dynamic behaviors, making the designed system a candidate for low power embedded applications. The system performance, as well as the robustness, are studied and analyzed through the Monte Carlo simulation. The chaotic time series prediction benchmark, NARMA10, is examined through the proposed DFR computing system, and exhibits a 36%−85% reduction on the error rate compared to state-of-the-art DFR computing system designs. To the best of our knowledge, our work represents the first analog integrated circuit (IC) implementation of the DFR computing system. Kangjun Bai, Yang Yi 0002 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2016 | Cyclical sensing integrate-and-fire circuit for memristor array based neuromorphic computingabstractThe brain-inspired, spike-based neuromorphic system is highly anticipated in the artificial intelligence community due to its high computational efficiency. The recently developed memristor-crossbar-array technology, which is able to efficiently emulate the plasticity of biological synapses and accommodate matrix multiplication, has demonstrated its potential for neuromorphic computing. To facilitate the computation, a high-speed integrate-and-fire circuit (IFC) and a counter were previously developed to efficiently convert the current from the memristor array into rate-coded spikes. However, the linear dynamic range of the circuit, which is limited by its responding speed, is challenged when the input intensity and the conductance of the memristor array are both high simultaneously. In this paper, a novel cyclical sensing scheme is developed that can significantly extend the linear dynamic range of the original IFC. Meanwhile, the power efficiency of the IFC can also be increased. The circuit simulation results indicated that the cyclical sensing IFC was able to efficiently and accurately facilitate the matrix multiplication when it was integrated with a 32×32 memristor crossbar array. With the optimized crossbar array structure and its peripheral circuits, the developed cyclical sensing IFC has shown great promise in accelerating matrix multiplication in spike-based computing systems. Hao Jiang 0014, Fu Luo, Kangjun Bai, J. Joshua Yang, Qiangfei Xia, Yiran Chen 0001, Qing Wu 0002 |
ISCAS | 4 |