Jinhai Hu

dblp:22/10209 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0003-2758-2349ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 5 first-author · 9 since 2021
YearPublicationVenuePosition
2026 A Reconfigurable Audio Analog Front-End with Hardware-Aware Transfer Learning for Task-Adaptive Keyword Spotting
Jinhai Hu, Wang Ling Goh
ISCAS1
2026 A Two-Stage Machine Learning Assisted Calibration Scheme Achieving 73.3dB SNDR and 87dB SFDR for A 14-bit Pipelined NS-SAR ADC
Jiaju Lu, Xinzhe Xie, Wang Ling Goh, Jinhai Hu, Yuan Gao 0011
ISCAS5
2026 Low-Power Learnable Digital Audio Feature Extractor for Always-on Keyword Spotting in Edge Devices
abstract
This paper presents a low-power, learnable digital audio feature extractor (AuFEx) for always-on keyword spotting (KWS) in edge devices. A noise-aware design flow is introduced to integrate the tuning of AuFEx parameters directly into the neural network classifier’s training process. This co-design approach enables the joint optimization of feature extractor and classifier within a unified framework. By incorporating noise during the training process, the system becomes more robust to variations in signal-to-noise ratio (SNR), maintaining high inference accuracy even with lightweight neural network classifiers. This design flow supports the design of both time-domain AuFEx (TD-AuFEx) and frequency-domain AuFEx (FD-AuFEx) for single-keyword wake-word detection (WWD) and 10-keyword KWS tasks, respectively. Implemented in a 40nm CMOS process, both TD-AuFEx and FD-AuFEx achieve over 2% accuracy improvement with the smallest size backend classifier. Specifically, the TD-AuFEx for WWD task achieves classifier size of 3.1k parameters with 494 nW power consumption and$375~\mu $s latency. The accuracy is maintained between 96.2% – 97.9% for SNR in the range of 5 – 20 dB. The FD-AuFEx for 10-keyword KWS achieves classifier size of 6.39k parameters with$1.258~\mu $W power consumption and 34.625 ms latency. The accuracy is maintained between 87.5% – 92.2% for SNR in the range of 5 dB - 20 dB, which is one of the highest compared to the other state-of-the-art designs.
Jinhai Hu, Wang Ling Goh, Yi Sheng Chong, Anh-Tuan Do, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 A 0.58mW dB-Linear Time Gain Compensation Amplifier with ±0.5-dB Gain Error for Imaging Applications
abstract
This paper presents a low-power time gain compensation (TGC) amplifier that provides accurate dB-linear voltage gain for imaging applications. Different from the conventional amplifier, the TGC amplifier generates exponentially variable gain over time to compensate the increasing signal attenuation along the ultrasound propagation path. This TGC amplifier employs a current-steering architecture to control the bias currents for the core cells. Implemented in a 55nm CMOS BCDLITE process, this design occupies an active chip area of 0.11mm2and consumes a total power of 0.58mW. Compared to the state-of-the-art TGC designs, the proposed TGC amplifier demonstrated the capability to maintain a dB-linear gain up to 14dB with less than ±0.5 dB gain error. This circuit is well-suited for integration with other high-voltage (HV) blocks, enabling a monolithic transceiver implementation.
Zhaoyang Cao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2025 A Digital Compute-in-Memory Macro Featuring Two's Complement Multiplication for LSTM-based Biomedical Signal Classification
abstract
This paper presents a digital compute-in-memory (DCIM) macro that supports two’s complement multiplication, specifically designed for processing electrocardiogram (ECG) signals using a Long Short-Term Memory (LSTM) neural network. Two distinct bitcell computing mechanisms are introduced: one for two’s complement bit-serial recurrent inputs using a 6T SRAM bitcell with two transmission gates (TGs) for outputting a weight bit or its complement, and another for encoded one-hot ECG inputs using an 8T bitcell to output weight values based on "1" detection in the input. Each column of bitcells performs multiply-and-accumulate operations, computing bitwise vector-matrix multiplication between inputs and SRAM-stored weights. Partial sums generated by columns of DCIM cells are processed through an adder tree controlled by a shift register, yielding the final LSTM gate-sum result via a parallel adder. The proposed DCIM macro enhances hardware efficiency by reducing transistor count and supports precise two’s complement multiplication. It achieves 96.9% accuracy on a 5-class classification task, using 32-level one-hot ECG input and an INT5 quantized LSTM neural network.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS1
2025 LearnAFE: Circuit-Algorithm Co-Design Framework for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper. Instead, this paper proposes a joint optimization of the backend classifier with the AFE’s transfer function to achieve system-level optimum. More specifically, the transfer function parameters of an analog bandpass filter (BPF) bank are tuned in a signal-to-noise ratio (SNR)-aware training loop for the classifier. Using a co-design loss function LBPF, this work shows superior optimization of both the filter bank and the classifier. Implemented in open-source SKY130 130nm CMOS process, the optimized design achieved 90.5%–94.2% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5 dB to 20 dB, with only 22k classifier parameters. Compared to conventional approach, the proposed audio AFE achieves 8.7% and 12.9% reduction in power and capacitor area respectively.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 Late Breaking Results: Circuit-Algorithm Co-design for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable audio analog front-end (AFE) which includes an analog filterbank for feature extraction and a classifier based on Depthwise Separable Convolutional Neural Network (DSCNN). Instead of the traditional approach to design the analog filterbank and digital classifier separately, a learnable filterbank is proposed and its source-follower bandpass filter (SF-BPF) parameters are optimized together with the neural network classifier in a signal-to-noise ratio (SNR)-aware training process. A new system criterion function (Lbpf) is proposed to include classification loss and filter performance into the training process. The optimized audio AFE achieves 10.6% and 11.7% reduction in BPF power and chip area, respectively. Meanwhile, this approach achieved 88.6%--94.5% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5dB to 20dB, with only 16k trainable parameters.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
DAC1
2024 Squeeze-Excite Fusion Based Multimodal Neural Network for Sleep Stage Classification with Flexible EEG/ECG Signal Acquisition Circuit
abstract
This paper presents a multimodal fusion strategy for sleep stage classification using polysomnography (PSG) with electroencephalogram (EEG) and Electrocardiogram (ECG) data. The Squeeze-Excite (SE) Fusion mechanism is implemented to enhance the collaborative impact of EEG and ECG signals on neural network classification. To address the challenges of imbalance in the dataset, a balanced sampler is used. Improved feature extraction is achieved through Linear-frequency cepstral coefficients (LFCC) applied to the EEG signal. A recurrent convolutional neural network (RCNN) reduces model parameters and optimizes architecture, while quantizing the network weight down to INT4 ensures hardware compatibility, especially for edge devices. Applying these methodologies to signals, this optimized approach achieves a significant validation accuracy of 77.6% with a compact 23.5KB weight memory size on the MIT-BIH dataset, covering six distinct classification categories.
Shuailin Tao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2023 Classification of ECG Anomaly with Dynamically-biased LSTM for Continuous Cardiac Monitoring
abstract
This paper presents an electrocardiogram (ECG) signal classification model based on dynamically-biased Long Short-Term Memory (DB-LSTM) network. Compared to conventional LSTM networks, DB-LSTM introduces a set of parameters$C$which save the previous time-step cell gate states of the unit cell. Hence, more feature information is preserved and a smaller size network is required for the classification task. Comprehensive simulations using MIT-BIH ECG datasets show that this model can perform ECG feature classification with shorter time window, faster training convergence while achieving comparable training and classification accuracy with much lower weigh resolution. Compared to the other state-of- art ECG analysis algorithms, this model only requires 4 layers, and it achieved 96.74% accuracy when weights are truncated from FP32 to INT4 with only 2.4% accuracy degradation. Implemented on Xilinx Artix-7 FPGA, the proposed design is estimated to consume only 40μW dynamic power, which is a promising candidate for resource constrained edge devices.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS1