Yuan Gao 0011

dblp:76/2452-11 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0002-0832-2982ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 2 first-author · 28 since 2021
YearPublicationVenuePosition
2026 A Two-Stage Machine Learning Assisted Calibration Scheme Achieving 73.3dB SNDR and 87dB SFDR for A 14-bit Pipelined NS-SAR ADC
Jiaju Lu, Xinzhe Xie, Wang Ling Goh, Jinhai Hu, Yuan Gao 0011
ISCAS6
2026 Low-Power Learnable Digital Audio Feature Extractor for Always-on Keyword Spotting in Edge Devices
abstract
This paper presents a low-power, learnable digital audio feature extractor (AuFEx) for always-on keyword spotting (KWS) in edge devices. A noise-aware design flow is introduced to integrate the tuning of AuFEx parameters directly into the neural network classifier’s training process. This co-design approach enables the joint optimization of feature extractor and classifier within a unified framework. By incorporating noise during the training process, the system becomes more robust to variations in signal-to-noise ratio (SNR), maintaining high inference accuracy even with lightweight neural network classifiers. This design flow supports the design of both time-domain AuFEx (TD-AuFEx) and frequency-domain AuFEx (FD-AuFEx) for single-keyword wake-word detection (WWD) and 10-keyword KWS tasks, respectively. Implemented in a 40nm CMOS process, both TD-AuFEx and FD-AuFEx achieve over 2% accuracy improvement with the smallest size backend classifier. Specifically, the TD-AuFEx for WWD task achieves classifier size of 3.1k parameters with 494 nW power consumption and$375~\mu $s latency. The accuracy is maintained between 96.2% – 97.9% for SNR in the range of 5 – 20 dB. The FD-AuFEx for 10-keyword KWS achieves classifier size of 6.39k parameters with$1.258~\mu $W power consumption and 34.625 ms latency. The accuracy is maintained between 87.5% – 92.2% for SNR in the range of 5 dB - 20 dB, which is one of the highest compared to the other state-of-the-art designs.
Jinhai Hu, Wang Ling Goh, Yi Sheng Chong, Anh-Tuan Do, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.6
2026 FPGA-Based Real-Time ECG Classification System Using Quantized Inception-ResNeXt Neural Network and CWT Approximation
abstract
This article presents a software–hardware codesigned field-programmable gate array (FPGA)-based real-time electrocardiogram (ECG) classification system that combines methodological and practical innovations to achieve state-of-the-art performance with an ultracompact model. On the software side, we introduce a hardware-adaptive, configurable quantization-aware training (QAT) framework that enables layerwise precision assignment and flexible quantization, ensuring that the trained model is highly accurate and hardware-friendly even at ultralow bit widths. On the hardware side, we propose a resource-efficient FPGA accelerator featuring a streaming architecture and a cosine-approximated continuous wavelet transform (CWT) module, optimized for low-power and real-time inference. Implemented in an FPGA, we demonstrate that a six-layer Inception-ResNeXt (IRN) network can achieve 99.5% inference accuracy on the MIT-BIH ECG dataset with 200-mW dynamic power and 0.0767-mJ/inference energy efficiency.
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011, Hen-Wei Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2026 FPGA Implementation of PoolFormer Network Using Python-Driven High-Level Synthesis Framework for Edge-AIoT Speech Recognition
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Very Large Scale Integr. Syst.5
2025 A 0.58mW dB-Linear Time Gain Compensation Amplifier with ±0.5-dB Gain Error for Imaging Applications
abstract
This paper presents a low-power time gain compensation (TGC) amplifier that provides accurate dB-linear voltage gain for imaging applications. Different from the conventional amplifier, the TGC amplifier generates exponentially variable gain over time to compensate the increasing signal attenuation along the ultrasound propagation path. This TGC amplifier employs a current-steering architecture to control the bias currents for the core cells. Implemented in a 55nm CMOS BCDLITE process, this design occupies an active chip area of 0.11mm2and consumes a total power of 0.58mW. Compared to the state-of-the-art TGC designs, the proposed TGC amplifier demonstrated the capability to maintain a dB-linear gain up to 14dB with less than ±0.5 dB gain error. This circuit is well-suited for integration with other high-voltage (HV) blocks, enabling a monolithic transceiver implementation.
Zhaoyang Cao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS4
2025 Neuromorphic FeRAM-Based Co-Design for Imaging Enhancement in Handheld Photoacoustic Systems
abstract
This paper introduces a novel platform designed to enhance the imaging quality of handheld photoacoustic imaging (PAI) systems, addressing the limitations of current portable PAI devices. The platform integrates the MultiResU-Net imaging enhancement algorithm with a Ferroelectric random-access memory (FeRAM) crossbar array, enabling efficient in-memory computing that is highly suitable for deep neural networks involving extensive matrix multiplications. The hardware implementation is optimized for low-power operation on edge devices, and a specifically designed algorithmic strategy is introduced to accurately simulate hardware variations with a time complexity of O(mn). The feasibility and effectiveness of this approach are demonstrated through simulations using synthesized and in vivo data, showing a more than tenfold improvement in imaging resolution. The neural network inference is significantly accelerated, completing within microseconds, thereby fully supporting real-time imaging. The entire platform is compact, with dimensions of 25×25×20 cm3, making it a portable, high-resolution, real-time imaging solution for personalized healthcare.
Tiancheng Cao, Zhengyuan Zhang 0002, Shuailin Tao, Chen Liu 0009, Wang Ling Goh, Yuanjing Zheng, Yuan Gao 0011
ISCAS7
2025 A Digital Compute-in-Memory Macro Featuring Two's Complement Multiplication for LSTM-based Biomedical Signal Classification
abstract
This paper presents a digital compute-in-memory (DCIM) macro that supports two’s complement multiplication, specifically designed for processing electrocardiogram (ECG) signals using a Long Short-Term Memory (LSTM) neural network. Two distinct bitcell computing mechanisms are introduced: one for two’s complement bit-serial recurrent inputs using a 6T SRAM bitcell with two transmission gates (TGs) for outputting a weight bit or its complement, and another for encoded one-hot ECG inputs using an 8T bitcell to output weight values based on "1" detection in the input. Each column of bitcells performs multiply-and-accumulate operations, computing bitwise vector-matrix multiplication between inputs and SRAM-stored weights. Partial sums generated by columns of DCIM cells are processed through an adder tree controlled by a shift register, yielding the final LSTM gate-sum result via a parallel adder. The proposed DCIM macro enhances hardware efficiency by reducing transistor count and supports precise two’s complement multiplication. It achieves 96.9% accuracy on a 5-class classification task, using 32-level one-hot ECG input and an INT5 quantized LSTM neural network.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2025 A Gait Data Compression and Reconstruction Framework for Edge Device using Low-Dimensional Attention Model with Autoencoder
abstract
This paper presents a gait data compression and reconstruction framework based on a low-dimensional attention model with autoencoder. By reducing the size of the attention filter to match the maximum matrix rank, the dimensionality of the attention filter can be reduced to enhance the compression ratio. Extensive evaluations using MHEALTH dataset demonstrated that the proposed method can achieve compression ratio of 24 with low reconstruction error of Percent Root Mean Square Difference (PRD) of 0.0323, Correlation Coefficient (CC) of 0.9510, and Signal-to-Noise Ratio Loss (SNRL) of 1.51 dB. The proposed compression model is implemented in hardware using microcontroller. Fixed-point quantization and optimized Softmax layer representation are performed to reduce the hardware resources requirement.
Shuailin Tao, Wang Ling Goh, Tiancheng Cao, Yuan Gao 0011
ISCAS4
2025 A Low-Noise Class-F23 VCO With Harmonic Resonance Expansion and 2nd/3rd-Harmonic Outputs for Multiband mm-Wave Applications
abstract
This paper presents a low-noise class-F23voltage-controlled oscillator (VCO) with harmonic resonance expansion and$2^{\mathrm {nd}}$/$3^{\mathrm {rd}}$-harmonic outputs for multiband mm-wave applications. By using a single four-coil transformer to extend the common-mode (CM) and differential-mode (DM) harmonic resonance bandwidths, the$2^{\mathrm {nd}}$- and$3^{\mathrm {rd}}$-harmonic resonances can be acquired without additional frequency alignment calibration. Meanwhile, benefiting from the favorable differential response at the$2^{\mathrm {nd}}$- and$3^{\mathrm {rd}}$-harmonic frequencies, the corresponding harmonic frequency outputs are extracted, simultaneously. Fabricated in 65-nm CMOS process, the proposed VCO achieves a frequency tuning range (FTR) of 20.8% from 10.71 GHz to 13.20GHz with 1-MHz offset phase noise (PN) from -117.4 to -114.5 dBc/Hz, while consuming 7.8-9.8 mW at 0.6 V. The VCO core area is only 0.054 mm2and the flicker noise corner is 310-450 kHz. The figure-of-merit (FoM) at 10-MHz offset scores 189.8-191.7 dBc/Hz. The harmonic outputs achieve a FTR of 21.42-26.40 GHz and 32.13-39.60 GHz, with a 1-MHz-offset PN from –110.5 to –107.5 dBc/Hz and –107.6 to –104.8 dBc/Hz.
Yuan Gao 0011, Depeng Sun, Feng Bu, Bowen Wang 0001, Zhicheng Dong 0002, Xiaoteng Zhao, Tao Zhang 0086, Ruixue Ding, Shubin Liu 0001, Zhangming Zhu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2025 LearnAFE: Circuit-Algorithm Co-Design Framework for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper. Instead, this paper proposes a joint optimization of the backend classifier with the AFE’s transfer function to achieve system-level optimum. More specifically, the transfer function parameters of an analog bandpass filter (BPF) bank are tuned in a signal-to-noise ratio (SNR)-aware training loop for the classifier. Using a co-design loss function LBPF, this work shows superior optimization of both the filter bank and the classifier. Implemented in open-source SKY130 130nm CMOS process, the optimized design achieved 90.5%–94.2% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5 dB to 20 dB, with only 22k classifier parameters. Compared to conventional approach, the proposed audio AFE achieves 8.7% and 12.9% reduction in power and capacitor area respectively.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 A 310 nA DCM Hysteretic Buck Converter With 95.8% Peak Efficiency and Greater Than 90% Efficiency in 10 μA-100 mA Load Range for Battery Powered IoT Sensors
abstract
This paper presents a discontinuous conduction mode (DCM) hysteretic buck converter for battery-powered Internet of Things (IoT) sensors. To reduce the quiescent power consumption, a duty-cycled comparator with delay compensation feedback loop is used in the zero current detector (ZCD). This feedback loop generates a self-calibrated comparator input offset to compensate for the comparator propagation delay. Additionally, a dynamic biasing scheme is applied to the hysteretic comparator to enhance the conversion efficiency under light load. The proposed buck converter was implemented in a 55-nm CMOS BCDLite process with an active area of 0.76mm$ \times 0.69$mm. With 310 nA quiescent current, this buck converter supports load currents in the range of$1~{\mu }$A - 300 mA with output voltage 0.8 V - 2.2 V. In particular, greater than 90% efficiency is maintained for load current from$10~{\mu }$A to 100 mA with 95.8% peak efficiency.
Wang Ling Goh, Liter Siek, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A 7.4-9.2-GHz Fractional-N Differential Sampling PLL Based on Phase-Domain and Voltage-Domain Hybrid Calibration
abstract
This brief proposes a 7.4–9.2-GHz low-noise fractional-N differential sampling phase-locked loop (DSPLL), which features doubled phase detector (PD) gain. By using the phase-domain and voltage-domain hybrid calibration, the accumulated quantization error (Q-error) of the delta-sigma modulator (DSM) is compensated, and the locking problem caused by large sampling voltage fluctuation is solved. Meanwhile, a voltage shifting technique is introduced to adjust the locked voltage region of differential sampling PD (DSPD), which can improve the linearity of DSPLL for better calibration. Fabricated in 65-nm CMOS process, the presented DSPLL achieves measured integrated jitter of 69.09 and 73.26 fs for integer-N and fractional-N modes, respectively. The reference spur is −72.96 dBc, and the worst fractional spur is −55.26 dBc. The total power consumption is 19.2 mW at a 1.2-V supply, achieving a figure of merit jitter (FOMJ) of −249.9 dB.
Feng Bu, Ruixue Ding, Depeng Sun, Yuan Gao 0011, Xiaoteng Zhao, Lisheng Chen, Shubin Liu 0001, Zhangming Zhu
IEEE Trans. Very Large Scale Integr. Syst.5
2025 Edge PoolFormer: Modeling and Training of PoolFormer Network on RRAM Crossbar for Edge-AI Applications
abstract
PoolFormer is a subset of Transformer neural network with a key difference of replacing computationally demanding token mixer with pooling function. In this work, a memristor-based PoolFormer network modeling and training framework for edge-artificial intelligence (AI) applications is presented. The original PoolFormer structure is further optimized for hardware implementation on RRAM crossbar by replacing the normalization operation with scaling. In addition, the nonidealities of RRAM crossbar from device to array level as well as peripheral readout circuits are analyzed. By integrating these factors into one training framework, the overall neural network performance is evaluated holistically and the impact of nonidealities to the network performance can be effectively mitigated. Implemented in Python and PyTorch, a 16-block PoolFormer network is built with$64\times 64$four-level RRAM crossbar array model extracted from measurement results. The total number of the proposed Edge PoolFormer network parameters is 0.246 M, which is at least one order smaller than the conventional CNN implementation. This network achieved inference accuracy of 88.07% for CIFAR-10 image classification tasks with accuracy degradation of 1.5% compared to the ideal software model with FP32 precision weights.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
IEEE Trans. Very Large Scale Integr. Syst.3
2024 Late Breaking Results: Circuit-Algorithm Co-design for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable audio analog front-end (AFE) which includes an analog filterbank for feature extraction and a classifier based on Depthwise Separable Convolutional Neural Network (DSCNN). Instead of the traditional approach to design the analog filterbank and digital classifier separately, a learnable filterbank is proposed and its source-follower bandpass filter (SF-BPF) parameters are optimized together with the neural network classifier in a signal-to-noise ratio (SNR)-aware training process. A new system criterion function (Lbpf) is proposed to include classification loss and filter performance into the training process. The optimized audio AFE achieves 10.6% and 11.7% reduction in BPF power and chip area, respectively. Meanwhile, this approach achieved 88.6%--94.5% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5dB to 20dB, with only 16k trainable parameters.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
DAC5
2024 A 30.5-to-31 GHz Sampling PLL With Double-Edge Sampling PD and Implict Common-Mode VCO Scoring 39.69-fs RMS Jitter and -253.6-dB FoM in a 0.047mm2 Area
abstract
This paper presents an integer-N sampling phase-locked loop (S-PLL) characterized by both low jitter and low spur. The design integrates a high-gain double-edge sampling phase detector (PD) and a self-retimed multi-modulus divider (MMD) aimed at mitigating the in-band noise. Furthermore, it features a compact implicit common-mode voltage-controlled oscillator (VCO) with a second harmonic tuning tailored for noise reduction. The proposed S-PLL, fabricated in a 28-nm CMOS technology, operates at 31 GHz with a 250-MHz reference. The measured RMS jitter is 39.69 fs integrated from 10 kHz to 100 MHz, with a reference spur of −63.2 dBc. The proposed PLL achieves a figure-of-merit (FoM) of −253.6 dB with a 28-mW power and a 0.047-mm2area.
Zhicheng Dong 0002, Xiaoteng Zhao, Weitan Huang, Yuan Gao 0011, Depeng Sun, Shubin Liu 0001, Lihong Yang, Zhangming Zhu
ISCAS4
2024 A 99.8-dB SNDR 10kHz-BW Second-Order DT Delta-Sigma Modulator with Single OTA and Enhanced Noise-Coupling
abstract
This paper presents a low-power second-order discrete-time (DT) Delta-Sigma Modulator (DSM) for Internet of Things (IoT) applications. The conventional noise coupling (NC) topology is improved with additional dual feedback paths to suppress the high-frequency gain of the noise transfer function (NTF), thereby reducing the internal voltage swing and relaxing the requirement of the quantizer resolution. In addition, stage-sharing technique is applied to reuse the OTA for both the integrator and the NC adder. Hence, second-order noise shaping is achieved with a single OTA. Implemented in a 130nm CMOS process, the proposed design demonstrated a simulated SNDR of 99.8dB in a 10kHz bandwidth with a 5.12MS/s sampling rate, consuming 200μW. State-of-the-art Schreier FoM (SNDR) and Walden FoM of 176.8dB and 125fJ/conv-step are achieved.
Jiaju Lu, Wang Ling Goh, Yuan Gao 0011
ISCAS4
2024 High Accuracy and Low Latency Mixed Precision Neural Network Acceleration for TinyML Applications on Resource-Constrained FPGAs
abstract
Recent advancements in mixed precision quantization algorithms hold significant promise for Tiny Machine Learning (TinyML) applications that face tight resource constraints. However, to leverage the benefits of mixed precision NNs, specialized hardware is required. The reconfigurability of Field Programmable Gate Arrays (FPGAs) allows tensor computations of any precision. In this work, we propose a streaming architecture with layer-specific processing elements to support interlayer mixed precision NN computations on FPGAs. Additionally, we extend the supported quantization scheme to intralayer level by introducing multiply-accumulation (MAC) units that support intralayer dynamic precision quantization. Bit-serial MAC units are chosen to accommodate more PEs on resource-constrained FPGAs. We successfully implemented a mixed quantized NN for the keyword spotting task on a FPGA-only Digilent A7-35T platform with 62% reduction in resource consumption when compared to the previous FPGA-only implementations. Our implementation also achieves 2 times speedup as compared to the state-of-the-art ASIC implementation while meeting the accuracy requirement of 90% set by MLPerf®Tiny Benchmark.
Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2024 Squeeze-Excite Fusion Based Multimodal Neural Network for Sleep Stage Classification with Flexible EEG/ECG Signal Acquisition Circuit
abstract
This paper presents a multimodal fusion strategy for sleep stage classification using polysomnography (PSG) with electroencephalogram (EEG) and Electrocardiogram (ECG) data. The Squeeze-Excite (SE) Fusion mechanism is implemented to enhance the collaborative impact of EEG and ECG signals on neural network classification. To address the challenges of imbalance in the dataset, a balanced sampler is used. Improved feature extraction is achieved through Linear-frequency cepstral coefficients (LFCC) applied to the EEG signal. A recurrent convolutional neural network (RCNN) reduces model parameters and optimizes architecture, while quantizing the network weight down to INT4 ensures hardware compatibility, especially for edge devices. Applying these methodologies to signals, this optimized approach achieves a significant validation accuracy of 77.6% with a compact 23.5KB weight memory size on the MIT-BIH dataset, covering six distinct classification categories.
Shuailin Tao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS4
2024 A Nanowatt Area-Efficient 16-Channel Bandpass Filterbank with Floating Active Capacitance Multiplier for Acoustic Signal Processing
abstract
This paper introduces an area-efficient nanowatt 16-channel bandpass filterbank tailored for acoustic signal processing in Artificial Internet of Things (AIoT) sensor systems. Integration of a floating active capacitance multiplier (FACM) adeptly addresses the inherent challenge of increased area in source-follower-based filters due to their larger capacitance. Compared to conventional gm-C filters, this work offers reduced power consumption, a more straightforward structure, and tunability. In a standard 0.13-μm CMOS process, simulations reveal that the filterbank's frequency range spans from 92 Hz to 5060 Hz with a gain of 17 dB, and a total power consumption of 97 nW for the 16 channels. When benchmarked against recent bandpass filters, this work achieves an impressive area efficiency of 0.038 mm2/channel and superior linearity. Such a compact and energy-efficient design offers a competitive solution for analog feature extraction in edge AI applications.
Wang Ling Goh, Yuan Gao 0011
ISCAS3
2023 RRAM-PoolFormer: A Resistive Memristor-based PoolFormer Modeling and Training Framework for Edge-AI Applications
abstract
PoolFormer is a type of neural network architecture that is abstracted from Transformer where the computationally heavy token mixer module is replaced with simple pooling function. This paper presents a memristor-based PoolFormer modeling and training framework for edge-AI applications. To fit for implementation on resistive crossbar array, original PoolFormer structure is further optimized by replacing normalization operation with hardware friendly scaling operation. In addition, the non-idealities of RRAM crossbar from device to array level as well as peripheral readout circuits are also included. By incorporating these elements under a single framework for network training, their impact to the network performance can be effectively mitigated. This framework is implemented in a combination of Python and PyTorch. A 16-block PoolFormer network is designed and optimized for CIFAR-10 image classification tasks using measured$\mathbf{64}\times \mathbf{64}$RRAM crossbar array results. The total network weights are only 0.26M, which is at least one order of magnitude smaller than that of the conventional DNN implementation. When compared to the ideal model with FP64 weight bit-length, 85.86% inference accuracy is reached with only 4-level weight resolutions and less than 4% accuracy loss.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
ISCAS3
2023 Classification of ECG Anomaly with Dynamically-biased LSTM for Continuous Cardiac Monitoring
abstract
This paper presents an electrocardiogram (ECG) signal classification model based on dynamically-biased Long Short-Term Memory (DB-LSTM) network. Compared to conventional LSTM networks, DB-LSTM introduces a set of parameters$C$which save the previous time-step cell gate states of the unit cell. Hence, more feature information is preserved and a smaller size network is required for the classification task. Comprehensive simulations using MIT-BIH ECG datasets show that this model can perform ECG feature classification with shorter time window, faster training convergence while achieving comparable training and classification accuracy with much lower weigh resolution. Compared to the other state-of- art ECG analysis algorithms, this model only requires 4 layers, and it achieved 96.74% accuracy when weights are truncated from FP32 to INT4 with only 2.4% accuracy degradation. Implemented on Xilinx Artix-7 FPGA, the proposed design is estimated to consume only 40μW dynamic power, which is a promising candidate for resource constrained edge devices.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2023 Sparsity Through Spiking Convolutional Neural Network for Audio Classification at the Edge
abstract
Convolutional neural networks (CNNs) have shown to be effective for audio classification. However, deep CNNs can be computationally heavy and unsuitable for edge intelligence as embedded devices are generally constrained by memory and energy requirements. Spiking neural networks (SNNs) offer potential as energy-efficient networks but typically underperform typical deep neural networks in accuracy. This paper proposes a spiking convolutional neural network (SCNN) that exhibits excellent accuracy of above 98 % on a multi-class audio classification task. Accuracy remains high with weight quantization to INT8-precision. Additionally, this paper examines the role of neuron parameters in co-optimizing activation sparsity and accuracy.
Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2023 A 110nW Always-on Keyword Spotting Chip using Spiking CNN in 40nm CMOS
abstract
This paper presents an ultra-low power keyword spotting (KWS) chip for Artificial Intelligence of Things (AIoT) device's always-on ambient sensing function. The core KWS engine is based on a spiking convolutional neural network (SCNN) model for its attractive features of sparse activation and addition-only operations inside the spiking neurons. The proposed SCNN model improves the existing framewise incremental computation flow by adding a spike processing unit (SPU) to reduce the computing cycles. The power and latency of the whole system are reduced by 16.5% and 43.2% respectively. Extensive network quantization reduces the weight bit-length to 4-bit and only 1-bit activation is required. The chip also supports power gating by an energy-based voice activity detection (VAD) module to further reduce power consumption in random and sparse event (RSE) scenarios. Full chip simulation results show that the chip consumes only 110nW with 2.15% False alarm rate and 3.00% False reject rate in a 10% voice event stream test. It achieves state-of-art recognition accuracy of 99% and 96% for one and two keyword detection tasks.
Junran Pu, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do, Yuan Gao 0011
ISCAS8
2023 A Nanowatt Temperature-Independent Tunable Active Capacitance Multiplier with DC Compensation in $0.13-\mu\mathrm{m}$ CMOS
abstract
This paper presents a nanowatt active capacitance multiplier (ACM) with enhanced tunable capacitance multiplication factor and temperature-independent current control for Artificial Internet of Things (AIoT) sensor applications. The proposed circuit is based on second-generation voltage conveyor (VCII) topology and biased in subthreshold region for high energy efficiency. An improved stacked translinear loop is designed to achieve temperature-independent bias with reduced power consumption. A DC compensation circuit is incorporated to compensate the output DC offset due to active bias circuit current, so that ACM can directly interface with other ultra-low power circuits. The proposed circuit is implemented in a standard$0.13-\mu\mathrm{m}$CMOS process. Simulation results show that the 3-dB bandwidth is from 0.04 Hz to 8 kHz and the multiply factor can be tuned from 337 to 561 with power consumption in the range of 41.6 nW to 46.4 nW. The achieved figure-of-merits (FoMs) is compared favorably with the other state-of-the-arts.
Wang Ling Goh, Yuan Gao 0011
ISCAS5
2022 A 1800μm2, 953Gbps/W AES Accelerator for IoT Applications in 40nm CMOS
abstract
A compact and energy-efficient AES accelerator for area and power-constrained IoT applications was fabricated in a 40nm CMOS process. By eliminating the need of intermediate data registers for MixColumns and ShiftRows in our proposed AES accelerator, we were able to reduce the total flip-flops to only 269 bits. Further, by reusing functional blocks and swapping the D flip-flops in data storage with scan flip-flops, our chip occupies only a tiny area of $1800 \mu \text{m}^{2}$ with an extremely low number of 657 gates. In addition, clock gating method and near-threshold voltage were used in our design. Thus, our accelerator consumes only $3.2 \mu \text{W}$ with an operation efficiency of 953 Gbps/W using a 0.48 V supply voltage. Compared with prior arts, our design has savings of 53% on area and 55% on the number of gates. When operated with a supply voltage of 0.48 V at 25°C, we can also achieve lower energy efficiency.
Jingjing Lan, Vishnu P. Nambiar, Ming Ming Wong, Fei Li 0015, Yuan Gao 0011, Kevin Tshun Chuan Chai, Anh-Tuan Do
ISCAS5
2022 A Nanowatt Comparator with Feedforward Slew Rate Enhancement and PVT-Insensitive Bias for Always-on MEMS Switch Wake-up Sensor
abstract
In this paper, a nanowatt comparator for always-on microelectromechanical systems (MEMS) switch wake-up sensor is presented. To achieve continuous event detection with nanowatt ultralow power (ULP) consumption, the comparator building blocks are designed to work in subthreshold region. Various circuit techniques including self-cascode transconductance, feedforward slew rate enhancement are adopted to optimize the circuit performance under nA power consumption. Programmable comparator hysteresis allows the flexibility to conFigure the comparator to operate in different environments. PVT-insensitive voltage/current bias circuits are included to provide stable bias to ensure robust operation across the process corners. The proposed circuit is designed and implemented in a standard 1P8M0.13$-\mu$m CMOS process. The core circuit area is 220$\mu$m× 60$\mu$m. Post layout simulations show that the overall comparator power consumption is in the range of 5. 2-7.8nW under 0. 8V-1.2V supply, the corresponding work bandwidth is up to 6kHz.
Jinhen Lee, Jianming Zhao, Yuan Gao 0011
ISCAS3
2021 Design of Fully Differential Energy-Efficient Inverter-Based Low-Noise Amplifier for Ultrasound Imaging
abstract
Inverter-based low-noise amplifier (LNA) offers an elegant solution in terms of power and area efficiency, which is favorable for ultrasound receivers. Nevertheless, it imposes challenges on circuit designs regarding the biasing circuit, the limited gain, robustness over process, supply voltage and temperature (PVT) variations. This paper presents an in-depth study of various inverter-based LNAs for ultrasound receiver design. In addition, three fully differential inverter-based LNAs are designed and optimized with minimum power and area penalty in the common-mode feedback (CMFB) circuit. The LNAs are implemented in a standard 0.18 - $\mu m$ CMOS process, with the identical power consumption, simulation results show that LNA with single CMFB achieves lower distortion, i.e. a third harmonic distortion (HD3) of $\lt -51.5$ dB with 35 mVp-p input, LNA with dual CMFB obtains better noise performance, i.e. $4.85 nV / \sqrt{Hz}$ at 3 MHz.
Zhaoyang Cao, Yuan Gao 0011, Wang Ling Goh
VLSI-SoC3
2021 A 13.56 MHz Active Rectifier with PMOS AC-DC Interface for Wireless Powered Medical Implants
abstract
A 13.56 MHz active rectifier for wireless powered medical implants is presented in this paper. Only PMOS power transistors are used in the AC-DC interface in the active rectifier to prevent the inherent parasitic diode conduction that may lead to potential latch-up during the rectifier start-up period. To eliminate the reverse current introduced by the delay of comparator signal, a delay compensation loop is designed to dynamically adjust the rectifier at various DC output. A fully onchip $2 \times$ charge pump and capacitor assisted driver/sampling circuits are proposed to turn on/off the power PMOS in the rectifier. The rectifier is implemented in a standard 0.18-$\mu$ m CMOS process with a core active area of 0.4 mm $\times0.9$ mm. The measurement results show that this PMOS rectifier achieves peak power efficiency of 85% when load resistor is $500 \Omega$ and the DC output voltage range is 2V--3.5V.
Jianming Zhao, Yuan Gao 0011
VLSI-SoC2
2020 Design of Pyroelectric Infrared Detector And Micropower CMOS Integrated Circuitry Towards a Monolithic Gas Sensor
abstract
With the emerging trend in miniaturization of complete optical sensor on a silicon platform for gas sensing applications, this paper presents a promising candidate towards monolithic gas sensor with features of 1) a CMOS compatible Aluminum Nitride (AlN)-based pyroelectric infrared detector (PID); 2) customized micropower high-precision CMOS integrated readout circuits. In the PID design, the sensitivity and device integrity are improved with optimized sensing area and modified layer structure. In the readout circuits, to amplify the pico-amperes current from detector, a tunable resistive transimpedance amplifier (TIA) with robust pseudo-resistor in series is proposed to achieve reduced variations against process-voltage-temperature (PVT). A micropower incremental analog-to-digital converter (ADC) is designed to digitize the voltage output from TIA. In addition, to compensate the drifts induced by local temperature change around the detector, an accurate BJT-based temperature sensor is integrated in the circuits as well. Fabricated in an 8-inch wafer, the AlN-based PID with sensing area of 0.29 mm2achieves the detectivity of 6.04× 106cm√ Hz/W. The readout circuits are implemented in a 40nm-CMOS process, consuming total power of 24.5 μW under 1.2-V supply. Simulation results show that the TIA gain variation is11-bit resolution over 1 kHz bandwidth, and the temperature error over -20 °C - 80 °C is ±0.65 °C (3σ) without any calibration.
Doris Keh Ting Ng, Yuan Gao 0011
IECON4
2019 A 1.6MHz Swing-Boosted Relaxation Oscillator with ±0.15%/V 23.4ppm/°C Frequency Inaccuracy using Voltage-to-Delay Feedback
abstract
This paper presents a relaxation oscillator for on-chip clock reference or sensor interface application. To improve the frequency stability over temperature and supply variations, a voltage-to-delay feedback is proposed to compensate the circuit delay variation. In addition, a switch-capacitor swing boosting (SCSB) circuit is proposed to enhance the output swing for phase noise reduction. Implemented in 0.18-μm CMOS process, the proposed oscillator shows a 1.6MHz output frequency, with low frequency inaccuracy of 23.4ppm/°C across 0°C-90°C and ±0.15%/V over 1.2V-1.52V. The measured phase noise is -118.6dBc/Hz at 100kHz offset, corresponding to 156dBc/Hz FOM. The oscillator consumes 51.4μW under 1.3V supply voltage.
Wei Zhou 0036, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2018 A Rectifier-less Energy Harvesting Interface Circuit for Low-Voltage Piezoelectric Transducers
abstract
This paper proposes a rectifier-less, topology for harvesting energy from low-voltage piezoelectric transducers. The proposed system is based on Synchronous Electric Charge Extraction and utilizes a bi-directional switching converter which inherently produces both positive and negative voltages that enables rectifier-free operation. A PCB prototype implemented using discrete components is tested with a commercially available transducer. The measurement results confirm the effectiveness of the topology in harvesting energy from low voltages. The bi-directional converter achieves a peak power conversion efficiency of 73 % while the control circuits consume only 7.5 μW at 3 V.
Arish Shareef, Wang Ling Goh, Srikanth Narasimalu, Yuan Gao 0011
ISCAS4
2018 A 16.6 μW 3.12 MHz RC Relaxation Oscillator with 160.3 dBc/Hz FOM
abstract
This paper presents a new RC relaxation oscillator for biomedical sensor interface circuit. A novel switch-capacitor based RC charging/discharging circuit is proposed to effectively improve the oscillator phase noise and power performance. The inverter-based comparator with replica biasing is employed and optimized to enhance the phase noise performance and to lower output dependence on the supply voltage variation. The oscillator's temperature insensitivity is also improved by resistor temperature compensation. The prototype RC relaxation oscillator circuit is designed in a commercial 65nm CMOS process. The post-layout simulation results showed 3.12 MHz output frequency, -112dBc/Hz phase noise at 100 kHz offset, and 16.6 μW power consumption under 1 V supply voltage. The frequency variation is ±0.294%/V for supply within 1 V to 1.6 V, and 11.31 ppm/°C for temperature across -40°C to 100°C. The overall circuit performance is compared favorably to the state-of-art designs, with an outstanding Figure of Merit (FOM) of 160.03 dBc/Hz at 100 kHz.
Wei Zhou 0036, Wang Ling Goh, Jia Hao Cheong, Yuan Gao 0011
ISCAS4
2015 An output feedback-based start-up technique with automatic disabling for battery-less energy harvesters
abstract
This paper presents a start-up circuit based upon output voltage feedback for thermal energy harvesting. The feedback loop consisting of an output load, a boost converter core, a charge-pump-based doubler and a CMOS inverter switch to track the output voltage and generate a reset signal to kickstart the boost converter. The proposed start-up circuit shows a remarkably improved start-up time of 35μs at the input voltage of 290mV. The start-up circuit is automatically disabled after the boost converter reaches steady-state. Only one off-chip component is required, improving the cost efficiency. The proposed start-up circuit consumes 3.2nW during the steady state, aiding the boost converter to achieve an efficiency of 77% at input voltage of 290mV. It was designed in standard 65-nm CMOS process technology and occupies the area of 0.072mm2.
Abhik Das, Yuan Gao 0011, Tony Tae-Hyoung Kim
ISCAS2
2011 An integrated beamformer for IR-UWB receiver in 0.18-µm CMOS
abstract
This paper presents a fully integrated 2-channel beamformer for 3-5 GHz impulse-based ultra-wideband (IR-UWB) receiver. A true time delay (TTD) element with optimized cutoff frequency and active loss compensation is proposed to provide large continuous time delay tuning with reduced chip size and power consumption. Each beamforming channel has a measured peak gain of 18.5 dB and a continuous time delay between 0-255 ps. The proposed beamformer can provide maximum scan angle of up to ±60° for two antennas with 8 cm spacing. The whole beamformer consumes 38 mA from a 1.8V DC power supply.
Yuan Gao 0011, Yuanjin Zheng, Shengxi Diao, Yao Zhu 0005, Chun-Huat Heng
ISCAS1