Wang Ling Goh

dblp:22/4018 · DBLP profile ↗
← Back
51ranked-venue papers
1as first author
33since 2021 · last 2026
0000-0001-7466-8941ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 33 since 2021Artificial intelligence and machine learning · 4Graphics, computer vision, multimedia, augmented reality and games · 3Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Reconfigurable Audio Analog Front-End with Hardware-Aware Transfer Learning for Task-Adaptive Keyword Spotting
Jinhai Hu, Wang Ling Goh
ISCAS3
2026 A Two-Stage Machine Learning Assisted Calibration Scheme Achieving 73.3dB SNDR and 87dB SFDR for A 14-bit Pipelined NS-SAR ADC
Jiaju Lu, Xinzhe Xie, Wang Ling Goh, Jinhai Hu, Yuan Gao 0011
ISCAS4
2026 Low-Power Learnable Digital Audio Feature Extractor for Always-on Keyword Spotting in Edge Devices
abstract
This paper presents a low-power, learnable digital audio feature extractor (AuFEx) for always-on keyword spotting (KWS) in edge devices. A noise-aware design flow is introduced to integrate the tuning of AuFEx parameters directly into the neural network classifier’s training process. This co-design approach enables the joint optimization of feature extractor and classifier within a unified framework. By incorporating noise during the training process, the system becomes more robust to variations in signal-to-noise ratio (SNR), maintaining high inference accuracy even with lightweight neural network classifiers. This design flow supports the design of both time-domain AuFEx (TD-AuFEx) and frequency-domain AuFEx (FD-AuFEx) for single-keyword wake-word detection (WWD) and 10-keyword KWS tasks, respectively. Implemented in a 40nm CMOS process, both TD-AuFEx and FD-AuFEx achieve over 2% accuracy improvement with the smallest size backend classifier. Specifically, the TD-AuFEx for WWD task achieves classifier size of 3.1k parameters with 494 nW power consumption and$375~\mu $s latency. The accuracy is maintained between 96.2% – 97.9% for SNR in the range of 5 – 20 dB. The FD-AuFEx for 10-keyword KWS achieves classifier size of 6.39k parameters with$1.258~\mu $W power consumption and 34.625 ms latency. The accuracy is maintained between 87.5% – 92.2% for SNR in the range of 5 dB - 20 dB, which is one of the highest compared to the other state-of-the-art designs.
Jinhai Hu, Wang Ling Goh, Yi Sheng Chong, Anh-Tuan Do, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.3
2026 FPGA-Based Real-Time ECG Classification System Using Quantized Inception-ResNeXt Neural Network and CWT Approximation
abstract
This article presents a software–hardware codesigned field-programmable gate array (FPGA)-based real-time electrocardiogram (ECG) classification system that combines methodological and practical innovations to achieve state-of-the-art performance with an ultracompact model. On the software side, we introduce a hardware-adaptive, configurable quantization-aware training (QAT) framework that enables layerwise precision assignment and flexible quantization, ensuring that the trained model is highly accurate and hardware-friendly even at ultralow bit widths. On the hardware side, we propose a resource-efficient FPGA accelerator featuring a streaming architecture and a cosine-approximated continuous wavelet transform (CWT) module, optimized for low-power and real-time inference. Implemented in an FPGA, we demonstrate that a six-layer Inception-ResNeXt (IRN) network can achieve 99.5% inference accuracy on the MIT-BIH ECG dataset with 200-mW dynamic power and 0.0767-mJ/inference energy efficiency.
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011, Hen-Wei Huang
IEEE Trans. Very Large Scale Integr. Syst.3
2026 FPGA Implementation of PoolFormer Network Using Python-Driven High-Level Synthesis Framework for Edge-AIoT Speech Recognition
Tiancheng Cao, Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A 0.58mW dB-Linear Time Gain Compensation Amplifier with ±0.5-dB Gain Error for Imaging Applications
abstract
This paper presents a low-power time gain compensation (TGC) amplifier that provides accurate dB-linear voltage gain for imaging applications. Different from the conventional amplifier, the TGC amplifier generates exponentially variable gain over time to compensate the increasing signal attenuation along the ultrasound propagation path. This TGC amplifier employs a current-steering architecture to control the bias currents for the core cells. Implemented in a 55nm CMOS BCDLITE process, this design occupies an active chip area of 0.11mm2and consumes a total power of 0.58mW. Compared to the state-of-the-art TGC designs, the proposed TGC amplifier demonstrated the capability to maintain a dB-linear gain up to 14dB with less than ±0.5 dB gain error. This circuit is well-suited for integration with other high-voltage (HV) blocks, enabling a monolithic transceiver implementation.
Zhaoyang Cao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2025 Neuromorphic FeRAM-Based Co-Design for Imaging Enhancement in Handheld Photoacoustic Systems
abstract
This paper introduces a novel platform designed to enhance the imaging quality of handheld photoacoustic imaging (PAI) systems, addressing the limitations of current portable PAI devices. The platform integrates the MultiResU-Net imaging enhancement algorithm with a Ferroelectric random-access memory (FeRAM) crossbar array, enabling efficient in-memory computing that is highly suitable for deep neural networks involving extensive matrix multiplications. The hardware implementation is optimized for low-power operation on edge devices, and a specifically designed algorithmic strategy is introduced to accurately simulate hardware variations with a time complexity of O(mn). The feasibility and effectiveness of this approach are demonstrated through simulations using synthesized and in vivo data, showing a more than tenfold improvement in imaging resolution. The neural network inference is significantly accelerated, completing within microseconds, thereby fully supporting real-time imaging. The entire platform is compact, with dimensions of 25×25×20 cm3, making it a portable, high-resolution, real-time imaging solution for personalized healthcare.
Tiancheng Cao, Zhengyuan Zhang 0002, Shuailin Tao, Chen Liu 0009, Wang Ling Goh, Yuanjing Zheng, Yuan Gao 0011
ISCAS5
2025 A Digital Compute-in-Memory Macro Featuring Two's Complement Multiplication for LSTM-based Biomedical Signal Classification
abstract
This paper presents a digital compute-in-memory (DCIM) macro that supports two’s complement multiplication, specifically designed for processing electrocardiogram (ECG) signals using a Long Short-Term Memory (LSTM) neural network. Two distinct bitcell computing mechanisms are introduced: one for two’s complement bit-serial recurrent inputs using a 6T SRAM bitcell with two transmission gates (TGs) for outputting a weight bit or its complement, and another for encoded one-hot ECG inputs using an 8T bitcell to output weight values based on "1" detection in the input. Each column of bitcells performs multiply-and-accumulate operations, computing bitwise vector-matrix multiplication between inputs and SRAM-stored weights. Partial sums generated by columns of DCIM cells are processed through an adder tree controlled by a shift register, yielding the final LSTM gate-sum result via a parallel adder. The proposed DCIM macro enhances hardware efficiency by reducing transistor count and supports precise two’s complement multiplication. It achieves 96.9% accuracy on a 5-class classification task, using 32-level one-hot ECG input and an INT5 quantized LSTM neural network.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2025 Qubit-State Discrimination using Neural Networks with Rapid and Energy-Efficient Compute Arrays
abstract
Neural networks (NNs) implemented on field-programmable gate arrays (FPGAs) provide fast, high-fidelity solutions for processing readout signals from quantum information processors. However, application-specific integrated circuits (ASICs) instead of FPGAs hold the potential for improved performance, a largely unexplored path. This work proposes specialized hardware for NN-based qubit-state discrimination. We optimize the NN architecture to minimize resource requirements by reducing the layer width, employing linear activation functions, and weight quantization. Quantization-aware training is used to preserve accuracy despite these optimizations. Next, a compute array employing output stationary dataflow is chosen to process the NN workload. The compute array with abundant multipliers and adders can complete one NN inference in 63 ns, which makes it a good candidate for real-time qubit-state discrimination.
Yuntian Liu, Yi Sheng Chong, Benjamin Lienhard, Minghao Fan, Wang Ling Goh, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS5
2025 A Gait Data Compression and Reconstruction Framework for Edge Device using Low-Dimensional Attention Model with Autoencoder
abstract
This paper presents a gait data compression and reconstruction framework based on a low-dimensional attention model with autoencoder. By reducing the size of the attention filter to match the maximum matrix rank, the dimensionality of the attention filter can be reduced to enhance the compression ratio. Extensive evaluations using MHEALTH dataset demonstrated that the proposed method can achieve compression ratio of 24 with low reconstruction error of Percent Root Mean Square Difference (PRD) of 0.0323, Correlation Coefficient (CC) of 0.9510, and Signal-to-Noise Ratio Loss (SNRL) of 1.51 dB. The proposed compression model is implemented in hardware using microcontroller. Fixed-point quantization and optimized Softmax layer representation are performed to reduce the hardware resources requirement.
Shuailin Tao, Wang Ling Goh, Tiancheng Cao, Yuan Gao 0011
ISCAS2
2025 Design and Analysis of Latch-Type Comparators for Cryogenic Operations
abstract
A low-power, high-speed latched comparator is the most crucial component in high-speed analog-to-digital converters (ADCs). With recent advances in quantum sensing and quantum computing, these ADCs need to operate at cryogenic temperatures. However, the performance of existing latched-comparators at cryogenic temperatures has not been thoroughly evaluated. In this paper, we address this gap by characterizing and comparing five representative comparator architectures. We elucidate their operational principles at both cryogenic (4K) and room (300K) temperatures, utilizing TSMC 28nm CMOS technology. Our study showed that, the triple latch design achieves the best performance at cryogenic temperature, while at room temperature, the modified strong-regeneration design has the best performance.
Qibang Zang, Wang Ling Goh, Andres Brito, Anh-Tuan Do
ISCAS2
2025 LearnAFE: Circuit-Algorithm Co-Design Framework for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper. Instead, this paper proposes a joint optimization of the backend classifier with the AFE’s transfer function to achieve system-level optimum. More specifically, the transfer function parameters of an analog bandpass filter (BPF) bank are tuned in a signal-to-noise ratio (SNR)-aware training loop for the classifier. Using a co-design loss function LBPF, this work shows superior optimization of both the filter bank and the classifier. Implemented in open-source SKY130 130nm CMOS process, the optimized design achieved 90.5%–94.2% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5 dB to 20 dB, with only 22k classifier parameters. Compared to conventional approach, the proposed audio AFE achieves 8.7% and 12.9% reduction in power and capacitor area respectively.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 A 310 nA DCM Hysteretic Buck Converter With 95.8% Peak Efficiency and Greater Than 90% Efficiency in 10 μA-100 mA Load Range for Battery Powered IoT Sensors
abstract
This paper presents a discontinuous conduction mode (DCM) hysteretic buck converter for battery-powered Internet of Things (IoT) sensors. To reduce the quiescent power consumption, a duty-cycled comparator with delay compensation feedback loop is used in the zero current detector (ZCD). This feedback loop generates a self-calibrated comparator input offset to compensate for the comparator propagation delay. Additionally, a dynamic biasing scheme is applied to the hysteretic comparator to enhance the conversion efficiency under light load. The proposed buck converter was implemented in a 55-nm CMOS BCDLite process with an active area of 0.76mm$ \times 0.69$mm. With 310 nA quiescent current, this buck converter supports load currents in the range of$1~{\mu }$A - 300 mA with output voltage 0.8 V - 2.2 V. In particular, greater than 90% efficiency is maintained for load current from$10~{\mu }$A to 100 mA with 95.8% peak efficiency.
Wang Ling Goh, Liter Siek, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 Edge PoolFormer: Modeling and Training of PoolFormer Network on RRAM Crossbar for Edge-AI Applications
abstract
PoolFormer is a subset of Transformer neural network with a key difference of replacing computationally demanding token mixer with pooling function. In this work, a memristor-based PoolFormer network modeling and training framework for edge-artificial intelligence (AI) applications is presented. The original PoolFormer structure is further optimized for hardware implementation on RRAM crossbar by replacing the normalization operation with scaling. In addition, the nonidealities of RRAM crossbar from device to array level as well as peripheral readout circuits are analyzed. By integrating these factors into one training framework, the overall neural network performance is evaluated holistically and the impact of nonidealities to the network performance can be effectively mitigated. Implemented in Python and PyTorch, a 16-block PoolFormer network is built with$64\times 64$four-level RRAM crossbar array model extracted from measurement results. The total number of the proposed Edge PoolFormer network parameters is 0.246 M, which is at least one order smaller than the conventional CNN implementation. This network achieved inference accuracy of 88.07% for CIFAR-10 image classification tasks with accuracy degradation of 1.5% compared to the ideal software model with FP32 precision weights.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
IEEE Trans. Very Large Scale Integr. Syst.7
2024 Late Breaking Results: Circuit-Algorithm Co-design for Learnable Audio Analog Front-End
abstract
This paper presents a circuit-algorithm co-design framework for learnable audio analog front-end (AFE) which includes an analog filterbank for feature extraction and a classifier based on Depthwise Separable Convolutional Neural Network (DSCNN). Instead of the traditional approach to design the analog filterbank and digital classifier separately, a learnable filterbank is proposed and its source-follower bandpass filter (SF-BPF) parameters are optimized together with the neural network classifier in a signal-to-noise ratio (SNR)-aware training process. A new system criterion function (Lbpf) is proposed to include classification loss and filter performance into the training process. The optimized audio AFE achieves 10.6% and 11.7% reduction in BPF power and chip area, respectively. Meanwhile, this approach achieved 88.6%--94.5% accuracy for 10-keyword classification task across a wide range of input signal SNR from 5dB to 20dB, with only 16k trainable parameters.
Jinhai Hu, Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
DAC4
2024 Quantum Readout Processing Accelerator with a CORDIC Core at Cryogenic Temperature
abstract
Quantum computing has been the most promising in addressing computational challenging problems beyond classical computers’ capabilities. Recent work focus on developing systemon-chip (SoC) for quantum control and readout, yet there are scarce literature discussing how the measurement blocks are developed in details. This paper presents a quantum readout processing accelerator that operates at cryogenic temperature to process measurement data directly to obtain the qubit state. The proposed accelerator includes a coordinate rotation digital computer (CORDIC) based direct digital frequency synthesizer (DDFS) and a qubit state decision unit. To enable compact processing accelerator, the error of the qubit state prediction is analyzed when choosing different bitwidths and number of CORDIC iterations, aiming to balance the accuracy and hardware cost. The proposed accelerator, implemented using a 28-nm CMOS, consumes 36mW of power at room temperature and 8.9mW at −196°C. The processing latency and total energy consumption when determining a qubit state are reduced by 6 and 3 orders of magnitude respectively, when compared to a general purpose processor.
Yi Sheng Chong, Hongyu Cao, Wang Ling Goh, Patrick Bore, Yuanzheng Paul Tan, Yung Szen Yap, Rainer Dumke, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS3
2024 A 99.8-dB SNDR 10kHz-BW Second-Order DT Delta-Sigma Modulator with Single OTA and Enhanced Noise-Coupling
abstract
This paper presents a low-power second-order discrete-time (DT) Delta-Sigma Modulator (DSM) for Internet of Things (IoT) applications. The conventional noise coupling (NC) topology is improved with additional dual feedback paths to suppress the high-frequency gain of the noise transfer function (NTF), thereby reducing the internal voltage swing and relaxing the requirement of the quantizer resolution. In addition, stage-sharing technique is applied to reuse the OTA for both the integrator and the NC adder. Hence, second-order noise shaping is achieved with a single OTA. Implemented in a 130nm CMOS process, the proposed design demonstrated a simulated SNDR of 99.8dB in a 10kHz bandwidth with a 5.12MS/s sampling rate, consuming 200μW. State-of-the-art Schreier FoM (SNDR) and Walden FoM of 176.8dB and 125fJ/conv-step are achieved.
Jiaju Lu, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2024 High Accuracy and Low Latency Mixed Precision Neural Network Acceleration for TinyML Applications on Resource-Constrained FPGAs
abstract
Recent advancements in mixed precision quantization algorithms hold significant promise for Tiny Machine Learning (TinyML) applications that face tight resource constraints. However, to leverage the benefits of mixed precision NNs, specialized hardware is required. The reconfigurability of Field Programmable Gate Arrays (FPGAs) allows tensor computations of any precision. In this work, we propose a streaming architecture with layer-specific processing elements to support interlayer mixed precision NN computations on FPGAs. Additionally, we extend the supported quantization scheme to intralayer level by introducing multiply-accumulation (MAC) units that support intralayer dynamic precision quantization. Bit-serial MAC units are chosen to accommodate more PEs on resource-constrained FPGAs. We successfully implemented a mixed quantized NN for the keyword spotting task on a FPGA-only Digilent A7-35T platform with 62% reduction in resource consumption when compared to the previous FPGA-only implementations. Our implementation also achieves 2 times speedup as compared to the state-of-the-art ASIC implementation while meeting the accuracy requirement of 90% set by MLPerf®Tiny Benchmark.
Wei Soon Ng, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2024 3881 Gbps/W, 3005 µm AES Core with State Based Clock Gating for IoT applications
abstract
An efficient and extremely low-energy AES-256 accelerator was implemented in a 40nm CMOS process for space and energy-limited IoT applications, supporting encryption, decryption, ECB and CBC modes. By reusing functional block hardware and utilizing scan flip-flops, the proposed AES design occupies merely 3005 µm of silicon area. Additionally, clock gating circuits were extensively used at circuit level to halt essential flip-flops during the ShiftRows and its reversed operation for aggressive power saving. The post-layout simulation showed that our proposed design achieved the lowest power consumption of 1.83 µW at 0.3 V, 17 MHz, with energy efficiency as high as 3881 Gbps/W.
Zhangyi Pei, Vishnu P. Nambiar, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do
ISCAS4
2024 Squeeze-Excite Fusion Based Multimodal Neural Network for Sleep Stage Classification with Flexible EEG/ECG Signal Acquisition Circuit
abstract
This paper presents a multimodal fusion strategy for sleep stage classification using polysomnography (PSG) with electroencephalogram (EEG) and Electrocardiogram (ECG) data. The Squeeze-Excite (SE) Fusion mechanism is implemented to enhance the collaborative impact of EEG and ECG signals on neural network classification. To address the challenges of imbalance in the dataset, a balanced sampler is used. Improved feature extraction is achieved through Linear-frequency cepstral coefficients (LFCC) applied to the EEG signal. A recurrent convolutional neural network (RCNN) reduces model parameters and optimizes architecture, while quantizing the network weight down to INT4 ensures hardware compatibility, especially for edge devices. Applying these methodologies to signals, this optimized approach achieves a significant validation accuracy of 77.6% with a compact 23.5KB weight memory size on the MIT-BIH dataset, covering six distinct classification categories.
Shuailin Tao, Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS3
2024 A Nanowatt Area-Efficient 16-Channel Bandpass Filterbank with Floating Active Capacitance Multiplier for Acoustic Signal Processing
abstract
This paper introduces an area-efficient nanowatt 16-channel bandpass filterbank tailored for acoustic signal processing in Artificial Internet of Things (AIoT) sensor systems. Integration of a floating active capacitance multiplier (FACM) adeptly addresses the inherent challenge of increased area in source-follower-based filters due to their larger capacitance. Compared to conventional gm-C filters, this work offers reduced power consumption, a more straightforward structure, and tunability. In a standard 0.13-μm CMOS process, simulations reveal that the filterbank's frequency range spans from 92 Hz to 5060 Hz with a gain of 17 dB, and a total power consumption of 97 nW for the 16 channels. When benchmarked against recent bandpass filters, this work achieves an impressive area efficiency of 0.038 mm2/channel and superior linearity. Such a compact and energy-efficient design offers a competitive solution for analog feature extraction in edge AI applications.
Wang Ling Goh, Yuan Gao 0011
ISCAS2
2023 RRAM-PoolFormer: A Resistive Memristor-based PoolFormer Modeling and Training Framework for Edge-AI Applications
abstract
PoolFormer is a type of neural network architecture that is abstracted from Transformer where the computationally heavy token mixer module is replaced with simple pooling function. This paper presents a memristor-based PoolFormer modeling and training framework for edge-AI applications. To fit for implementation on resistive crossbar array, original PoolFormer structure is further optimized by replacing normalization operation with hardware friendly scaling operation. In addition, the non-idealities of RRAM crossbar from device to array level as well as peripheral readout circuits are also included. By incorporating these elements under a single framework for network training, their impact to the network performance can be effectively mitigated. This framework is implemented in a combination of Python and PyTorch. A 16-block PoolFormer network is designed and optimized for CIFAR-10 image classification tasks using measured$\mathbf{64}\times \mathbf{64}$RRAM crossbar array results. The total network weights are only 0.26M, which is at least one order of magnitude smaller than that of the conventional DNN implementation. When compared to the ideal model with FP64 weight bit-length, 85.86% inference accuracy is reached with only 4-level weight resolutions and less than 4% accuracy loss.
Tiancheng Cao, Weihao Yu 0001, Yuan Gao 0011, Chen Liu 0009, Shuicheng Yan, Wang Ling Goh
ISCAS6
2023 Classification of ECG Anomaly with Dynamically-biased LSTM for Continuous Cardiac Monitoring
abstract
This paper presents an electrocardiogram (ECG) signal classification model based on dynamically-biased Long Short-Term Memory (DB-LSTM) network. Compared to conventional LSTM networks, DB-LSTM introduces a set of parameters$C$which save the previous time-step cell gate states of the unit cell. Hence, more feature information is preserved and a smaller size network is required for the classification task. Comprehensive simulations using MIT-BIH ECG datasets show that this model can perform ECG feature classification with shorter time window, faster training convergence while achieving comparable training and classification accuracy with much lower weigh resolution. Compared to the other state-of- art ECG analysis algorithms, this model only requires 4 layers, and it achieved 96.74% accuracy when weights are truncated from FP32 to INT4 with only 2.4% accuracy degradation. Implemented on Xilinx Artix-7 FPGA, the proposed design is estimated to consume only 40μW dynamic power, which is a promising candidate for resource constrained edge devices.
Jinhai Hu, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2023 Sparsity Through Spiking Convolutional Neural Network for Audio Classification at the Edge
abstract
Convolutional neural networks (CNNs) have shown to be effective for audio classification. However, deep CNNs can be computationally heavy and unsuitable for edge intelligence as embedded devices are generally constrained by memory and energy requirements. Spiking neural networks (SNNs) offer potential as energy-efficient networks but typically underperform typical deep neural networks in accuracy. This paper proposes a spiking convolutional neural network (SCNN) that exhibits excellent accuracy of above 98 % on a multi-class audio classification task. Accuracy remains high with weight quantization to INT8-precision. Additionally, this paper examines the role of neuron parameters in co-optimizing activation sparsity and accuracy.
Cong Sheng Leow, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2023 A 110nW Always-on Keyword Spotting Chip using Spiking CNN in 40nm CMOS
abstract
This paper presents an ultra-low power keyword spotting (KWS) chip for Artificial Intelligence of Things (AIoT) device's always-on ambient sensing function. The core KWS engine is based on a spiking convolutional neural network (SCNN) model for its attractive features of sparse activation and addition-only operations inside the spiking neurons. The proposed SCNN model improves the existing framewise incremental computation flow by adding a spike processing unit (SPU) to reduce the computing cycles. The power and latency of the whole system are reduced by 16.5% and 43.2% respectively. Extensive network quantization reduces the weight bit-length to 4-bit and only 1-bit activation is required. The chip also supports power gating by an energy-based voice activity detection (VAD) module to further reduce power consumption in random and sparse event (RSE) scenarios. Full chip simulation results show that the chip consumes only 110nW with 2.15% False alarm rate and 3.00% False reject rate in a 10% voice event stream test. It achieves state-of-art recognition accuracy of 99% and 96% for one and two keyword detection tasks.
Junran Pu, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do, Yuan Gao 0011
ISCAS5
2023 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network Processing
abstract
Compute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures.
Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do
ISCAS2
2023 A Nanowatt Temperature-Independent Tunable Active Capacitance Multiplier with DC Compensation in $0.13-\mu\mathrm{m}$ CMOS
abstract
This paper presents a nanowatt active capacitance multiplier (ACM) with enhanced tunable capacitance multiplication factor and temperature-independent current control for Artificial Internet of Things (AIoT) sensor applications. The proposed circuit is based on second-generation voltage conveyor (VCII) topology and biased in subthreshold region for high energy efficiency. An improved stacked translinear loop is designed to achieve temperature-independent bias with reduced power consumption. A DC compensation circuit is incorporated to compensate the output DC offset due to active bias circuit current, so that ACM can directly interface with other ultra-low power circuits. The proposed circuit is implemented in a standard$0.13-\mu\mathrm{m}$CMOS process. Simulation results show that the 3-dB bandwidth is from 0.04 Hz to 8 kHz and the multiply factor can be tuned from 337 to 561 with power consumption in the range of 41.6 nW to 46.4 nW. The achieved figure-of-merits (FoMs) is compared favorably with the other state-of-the-arts.
Wang Ling Goh, Yuan Gao 0011
ISCAS4
2022 0.08mm2 128nW MFCC Engine for Ultra-low Power, Always-on Smart Sensing Applications
abstract
Mel frequency cepstral coefficient (MFCC) features are widely used in applications such as keyword spotting, bearing fault detection and heart sound classification. This work proposes a low power MFCC engine that enables its use for battery-powered edge applications. Three hardware algorithm co-optimizations were adopted to achieve energy efficient MFCC hardware implementation. The approximated MFCC features due to the optimizations still allows good accuracy when deployed in several applications such as keyword spotting and bearing fault detection, reporting negligible accuracy drop of $\le 1.5$%. The proposed MFCC hardware consumes only 128nW at 0.3V supply and occupies only 0.08mm2in 40nm CMOS technology, which are $5 \times $ and $2.75 \times $ power and area reduction respectively when compared to the prior arts.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS2
2022 Recovering Accuracy of RRAM-based CIM for Binarized Neural Network via Chip-in-the-loop Training
abstract
Resistive random access memory (RRAM) based computing-in-memory (CIM) is attractive for edge artificial intelligence (AI) applications, thanks to its excellent energy efficiency, compactness and high parallelism in matrix vector multiplication (MatVec) operations. However, existing RRAM-based CIM designs often require complex programming scheme to precisely control the RRAM cells to reach the desired resistance states so that the neural network classification accuracy is maintained. This leads to large area and energy overhead as well as low RRAM area utilization. Hence, compact RRAM-based CIM with simple pulse-based programming scheme is thus more desirable. To achieve this, we propose a chip-in-the-loop training approach to compensate for the network performance drop due to the stochastic behavior of the RRAM cells. Note that, although the target RRAM cell here is a two-state RRAM (i.e binary, having only high and low resistance states), their inherent analog resistance values are used in the CIM operation. Our experiment using a 4-layer fully-connected binary neural network (BNN) showed that after retraining, the RRAM-based network accuracy can be recovered, regardless of the RRAM resistance distribution and $\frac{\text{R}_{\text{HRS}}}{\text{R}_{\text{LRS}}}$ resistance ratio.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS2
2021 An Energy-Efficient Convolution Unit for Depthwise Separable Convolutional Neural Networks
abstract
High performance but computationally expensive Convolutional Neural Networks (CNNs) require both algorithmic and custom hardware improvement to reduce model size and to improve energy efficiency for edge computing applications. Recent CNN architectures employ depthwise separable convolution to reduce the total number of weights and MAC operations. However, depthwise separable convolution workload does not run efficiently in existing CNN accelerators. This paper proposes an energy-efficient CONV unit for pointwise and depthwise operation. The CONV unit utilizes weight stationary to enable high efficiency. The row partial sum reduction is engaged to increase parallelism in pointwise convolution thereby lightening the memory requirements on output partial sums. Our design achieves a maximum efficiency of 3.17 TOPS/W at 0.85V/40nm CMOS which is well-suited for energy constrained edge computing applications.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS2
2021 Design of Fully Differential Energy-Efficient Inverter-Based Low-Noise Amplifier for Ultrasound Imaging
abstract
Inverter-based low-noise amplifier (LNA) offers an elegant solution in terms of power and area efficiency, which is favorable for ultrasound receivers. Nevertheless, it imposes challenges on circuit designs regarding the biasing circuit, the limited gain, robustness over process, supply voltage and temperature (PVT) variations. This paper presents an in-depth study of various inverter-based LNAs for ultrasound receiver design. In addition, three fully differential inverter-based LNAs are designed and optimized with minimum power and area penalty in the common-mode feedback (CMFB) circuit. The LNAs are implemented in a standard 0.18 - $\mu m$ CMOS process, with the identical power consumption, simulation results show that LNA with single CMFB achieves lower distortion, i.e. a third harmonic distortion (HD3) of $\lt -51.5$ dB with 35 mVp-p input, LNA with dual CMFB obtains better noise performance, i.e. $4.85 nV / \sqrt{Hz}$ at 3 MHz.
Zhaoyang Cao, Yuan Gao 0011, Wang Ling Goh
VLSI-SoC4
2021 Efficient Implementation of Activation Functions for LSTM accelerators
abstract
Activation functions such as hyperbolic tangent (tanh) and logistic sigmoid (sigmoid) are critical computing elements in a long short term memory (LSTM) cell and network. These activation functions are non-linear, leading to challenges in their hardware implementations. Area-efficient and high performance hardware implementation of these activation functions thus becomes crucial to allow high throughput in a LSTM accelerator. In this work, we propose an approximation scheme which is suitable for both tanh and sigmoid functions. The proposed hardware for sigmoid function is 8.3 times smaller than the state-of-the-art, while for tanh function, it is the second smallest design. When applying the approximated tanh and sigmoid of 2% error in a LSTM cell computation, its final hidden state and cell state record errors of 3.1% and 5.8% respectively. When the same approximated functions are applied to a single layer LSTM network of 64 hidden nodes, the accuracy drops by 2.8% only. This proposed small yet accurate activation function hardware is promising to be used in Internet of Things (IoT) applications where accuracy can be traded off for ultra-low power consumption.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
VLSI-SoC2
2021 A 5.28-mm² 4.5-pJ/SOP Energy-Efficient Spiking Neural Network Hardware With Reconfigurable High Processing Speed Neuron Core and Congestion-Aware Router
abstract
In recent years, fast computation, low power, and small footprint are the key motivations for building SNN hardware. The unique features of SNN hardware have not been fully exploited, where the computation speed and energy efficiency of the SNN hardware can be improved according to the sparse spiking and non-uniform traffic of SNN. In this paper, we propose a 5.28-mm$^{2}~4096$-neuron 1M-synapse energy-efficient digital SNN hardware that can achieve ultra-low energy per synaptic operation of 4.5 pJ. The proposed neuron computing unit is implemented in pipeline architecture to achieve high synaptic processing speed. The proposed spike processing unit can significantly increase the processing speed of the neuron core by$1.9\times $and$9.4\times $when the spike injection rate is 50% and 10%, respectively. Besides, the increase in the processing speed of the neuron core leads to a reduction in energy consumption of up to 81.5%. An event-driven clock gating circuit that can reduce the power consumption of the proposed neuron block by more than 70% is proposed in this paper. This paper proposes a supervised STDP+ algorithm for SNN training, and the classification accuracy of the MNIST digits is 89.6% with 73.6% weight sparsity of the output layer.
Junran Pu, Wang Ling Goh, Vishnu P. Nambiar, Ming Ming Wong, Anh-Tuan Do
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Scalable Block-Based Spiking Neural Network Hardware with a Multiplierless Neuron Model
abstract
This paper proposes a scalable hardware architecture for block-based spiking neural networks utilizing a multiplierless spiking neuron model. These blocks were implemented as a neurocore mesh generated from an interconnect algorithm, allowing for seamless scalability of the network size while mitigating connectivity errors. The routing fabric asynchronous protocol allows for critical timing paths between blocks to be relaxed. The proposed neuron model consumed less logic compared to standard models with multipliers, reducing up to 16% of the neurocore logic cell area. The network was implemented alongside a computing subsystem as an FPGA-based system-on-chip, communicating via a fabric interconnect bridge. Experimental results validate the functionality of the proposed system, and achieved comparable classification accuracy to existing works.
Vishnu P. Nambiar, Eng-Kiat Koh, Junran Pu, Aarthy Mani, Ming Ming Wong, Wang Ling Goh, Anh-Tuan Do
ISCAS7
2020 Bottom-Up Scene Text Detection with Markov Clustering Networks
Zichuan Liu, Guosheng Lin, Wang Ling Goh
Int. J. Comput. Vis.3
2019 Towards Robust Curve Text Detection With Conditional Spatial Expansion
abstract
It is challenging to detect curve texts due to their irregular shapes and varying sizes. In this paper, we first investigate the deficiency of the existing curve detection methods and then propose a novel Conditional Spatial Expansion (CSE) mechanism to improve the performance of curve text detection. Instead of regarding the curve text detection as a polygon regression or a segmentation problem, we treat it as a region expansion process. Our CSE starts with a seed arbitrarily initialized within a text region and progressively merges neighborhood regions based on the extracted local features by a CNN and contextual information of merged regions. The CSE is highly parameterized and can be seamlessly integrated into existing object detection frameworks. Enhanced by the data-dependent CSE mechanism, our curve text detection system provides robust instance-level text region extraction with minimal post-processing. The analysis experiment shows that our CSE can handle texts with various shapes, sizes, and orientations, and can effectively suppress the false-positives coming from text-like textures or unexpected texts included in the same RoI. Compared with the existing curve text detection algorithms, our method is more robust and enjoys a simpler processing flow. It also creates a new state-of-art performance on curve text benchmarks with Fscore of up to 78.4%.
Zichuan Liu, Guosheng Lin, Sheng Yang 0006, Fayao Liu, Weisi Lin, Wang Ling Goh
CVPR6
2019 Block-Based Spiking Neural Network Hardware with Deme Genetic Algorithm
abstract
Hardware implementation of spiking neural networks (SNN) has been the focus of many previous works due to its higher execution speed. A block-based SNN architecture with a simple spiking neuron model is proposed in this paper. Compared to traditional spiking neuron models, the proposed model simplifies the equation of the membrane potential for ease of hardware implementation. The block-based SNN architecture also makes the hardware implementation more scalable and simplifies floorplanning. Deme genetic algorithm (GA) was applied for training the SNN model, and a population encoding scheme was used for spike time conversion. Two case studies were carried out to verify the functionality of the proposed model, namely number recognition and Fisher Iris classification. Experimental results showed that the proposed SNN model with deme GA was able to achieve comparable or higher classification accuracy than previous works.
Junran Pu, Vishnu P. Nambiar, Anh-Tuan Do, Wang Ling Goh
ISCAS4
2019 A 1.6MHz Swing-Boosted Relaxation Oscillator with ±0.15%/V 23.4ppm/°C Frequency Inaccuracy using Voltage-to-Delay Feedback
abstract
This paper presents a relaxation oscillator for on-chip clock reference or sensor interface application. To improve the frequency stability over temperature and supply variations, a voltage-to-delay feedback is proposed to compensate the circuit delay variation. In addition, a switch-capacitor swing boosting (SCSB) circuit is proposed to enhance the output swing for phase noise reduction. Implemented in 0.18-μm CMOS process, the proposed oscillator shows a 1.6MHz output frequency, with low frequency inaccuracy of 23.4ppm/°C across 0°C-90°C and ±0.15%/V over 1.2V-1.52V. The measured phase noise is -118.6dBc/Hz at 100kHz offset, corresponding to 156dBc/Hz FOM. The oscillator consumes 51.4μW under 1.3V supply voltage.
Wei Zhou 0036, Wang Ling Goh, Yuan Gao 0011
ISCAS2
2018 Feasibility of dictionary-based sparse coding for data compression in machine condition-based monitoring
abstract
On-time machine maintenance and upkeep are critical to ensure on-time production. In recent years, Condition-Based Monitoring (CBM) has been regarded as one of the most state-of- the-art machine maintenance techniques that can significantly lower unscheduled maintenance cost and provide greater efficiency. This however means that more data needs to be streamed and stored for CBM to work properly, which might prove costly in the long run in terms of data storage. This paper therefore focuses on the feasibility of using Sparse Coding as a method for signal compression for CBM purposes. Specifically, the feasibility of the method will be measured for vibration signals coming from the spindle component of machines.
Wang Ling Goh, Nicholas Sadjoli
ICIS1
2018 SqueezedText: A Real-Time Scene Text Recognition by Binary Convolutional Encoder-Decoder Network
abstract
A new approach for real-time scene text recognition is proposed in this paper. A novel binary convolutional encoder-decoder network (B-CEDNet) together with a bidirectional recurrent neural network (Bi-RNN). The B-CEDNet is engaged as a visual front-end to provide elaborated character detection, and a back-end Bi-RNN performs character-level sequential correction and classification based on learned contextual knowledge. The front-end B-CEDNet can process multiple regions containing characters using a one-off forward operation, and is trained under binary constraints with significant compression. Hence it leads to both remarkable inference run-time speedup as well as memory usage reduction. With the elaborated character detection, the back-end Bi-RNN merely processes a low dimension feature sequence with category and spatial information of extracted characters for sequence correction and classification. By training with over 1,000,000 synthetic scene text images, the B-CEDNet achieves a recall rate of 0.86, precision of 0.88 and F-score of 0.87 on ICDAR-03 and ICDAR-13. With the correction and classification by Bi-RNN, the proposed real-time scene text recognition achieves state-of-the-art accuracy while only consumes less than 1-ms inference run-time. The flow processing flow is realized on GPU with a small network size of 1.01 MB for B-CEDNet and 3.23 MB for Bi-RNN, which is much faster and smaller than the existing solutions.
Zichuan Liu, Yixing Li, Fengbo Ren, Wang Ling Goh, Hao Yu 0001
AAAI4
2018 Learning Markov Clustering Networks for Scene Text Detection
abstract
A novel framework named Markov Clustering Network (MCN) is proposed for fast and robust scene text detection. MCN predicts instance-level bounding boxes by firstly converting an image into a Stochastic Flow Graph (SFG) and then performing Markov Clustering on this graph. Our method can detect text objects with arbitrary size and orientation without prior knowledge of object size. The stochastic flow graph encode objects' local correlation and semantic information. An object is modeled as strongly connected nodes, which allows flexible bottom-up detection for scale-varying and rotated objects. MCN generates bounding boxes without using Non-Maximum Suppression, and it can be fully parallelized on GPUs. The evaluation on public benchmarks shows that our method outperforms the existing methods by a large margin in detecting multioriented text objects. MCN achieves new state-of-art performance on challenging MSRA-TD500 dataset with precision of 0.88, recall of 0.79 and F-score of 0.83. Also, MCN achieves realtime inference with frame rate of 34 FPS, which is 1.5× speedup when compared with the fastest scene text detection algorithm.
Zichuan Liu, Guosheng Lin, Sheng Yang 0006, Jiashi Feng, Weisi Lin, Wang Ling Goh
CVPR6
2018 A Rectifier-less Energy Harvesting Interface Circuit for Low-Voltage Piezoelectric Transducers
abstract
This paper proposes a rectifier-less, topology for harvesting energy from low-voltage piezoelectric transducers. The proposed system is based on Synchronous Electric Charge Extraction and utilizes a bi-directional switching converter which inherently produces both positive and negative voltages that enables rectifier-free operation. A PCB prototype implemented using discrete components is tested with a commercially available transducer. The measurement results confirm the effectiveness of the topology in harvesting energy from low voltages. The bi-directional converter achieves a peak power conversion efficiency of 73 % while the control circuits consume only 7.5 μW at 3 V.
Arish Shareef, Wang Ling Goh, Srikanth Narasimalu, Yuan Gao 0011
ISCAS2
2018 A 16.6 μW 3.12 MHz RC Relaxation Oscillator with 160.3 dBc/Hz FOM
abstract
This paper presents a new RC relaxation oscillator for biomedical sensor interface circuit. A novel switch-capacitor based RC charging/discharging circuit is proposed to effectively improve the oscillator phase noise and power performance. The inverter-based comparator with replica biasing is employed and optimized to enhance the phase noise performance and to lower output dependence on the supply voltage variation. The oscillator's temperature insensitivity is also improved by resistor temperature compensation. The prototype RC relaxation oscillator circuit is designed in a commercial 65nm CMOS process. The post-layout simulation results showed 3.12 MHz output frequency, -112dBc/Hz phase noise at 100 kHz offset, and 16.6 μW power consumption under 1 V supply voltage. The frequency variation is ±0.294%/V for supply within 1 V to 1.6 V, and 11.31 ppm/°C for temperature across -40°C to 100°C. The overall circuit performance is compared favorably to the state-of-art designs, with an outstanding Figure of Merit (FOM) of 160.03 dBc/Hz at 100 kHz.
Wei Zhou 0036, Wang Ling Goh, Jia Hao Cheong, Yuan Gao 0011
ISCAS2
2017 A passively compensated capacitive sensor readout with biased varactor temperature compensation and temperature coherent quantization
abstract
This paper presents a frequency-mode capacitive sensor readout front-end operating in a wide temperature range. First, a biased varactor temperature compensation (BVTC) is proposed to compensate the aggregate temperature gradients from the sensor and the oscillator circuit, achieving a nullified temperature coefficient for the oscillation frequency. Second, a temperature coherent quantization (TCQ) approach is proposed to enhance the sensor' sensitivity and to provide a self-referenced clock for digitization, whereby the influence from the temperature effect of the clock is minimized through a hybrid down-conversion time-to-digital converter (DTDC). A prototype chip was fabricated using the 0.18-μm CMOS process and it was verified using a commercial 5.92-to-6.53-pF capacitive pressure sensor over a temperature range of −20°C to 120°C. Proven by experiments, the prototype presented as good as ±0.4% full-scale pressure error in the 140°C temperature range.
Wang Ling Goh, Kevin Tshun Chuan Chai, Xin Lou 0001, Wen Bin Ye 0001
ISCAS3
2016 A 13.5-MHz relaxation oscillator with ±0.5% temperature stability for RFID application
abstract
This paper presents a 13.5-MHz low-power, RC on-chip relaxation oscillator with split-capacitor technique for RFID application. This oscillator implements only one comparator and one reference voltage to minimize power consumption and silicon area. A loop delay variation cancellation technique that employs an integrator loop and a split-capacitor architecture helps attained the temperature stability in the proposed oscillator. The proposed design is fabricated in 0.18-μm CMOS process. The relaxation oscillator consumes 48.8 μW at 1.8-V power supply. The measurement results show that the circuit can generate a stable frequency of 13.5 MHz. The output frequency variation is less than ±0.5% of temperature range from -30° C to 120° C, and the supply voltage variation coefficient is 0.5%/V across 1.5 V to 2.1V supply voltage.
Jiacheng Wang 0001, Wang Ling Goh
ISCAS2
2016 A Two-Stage Large-Capacitive-Load Amplifier With Multiple Cross-Coupled Small-Gain Stages
abstract
A two-stage large-capacitive-load amplifier with multiple cross-coupled small-gain stages is proposed in this paper. The cross-coupled structure of the small-gain stages augments the large-signal responses, providing significant improvement in the effective output-stage transconductance and, hence, the gain- bandwidth product (GBW). Implemented in a standard 0.13-μm CMOS technology and powered by a 0.7 V supply with a current consumption of 20 μA, the proposed amplifier achieves the GBW of 1.17 MHz and the phase margin of 74.8° while driving a capacitive load of 9.5 nF. The average slew rate is 0.3679 V/μs. The on-chip compensation capacitor is only 1.62 pF. The active chip area is 0.0056 mm2.
Marco Ho, Jianping Guo 0004, Tin Wai Mui, Kai Ho Mak, Wang Ling Goh, Hiu Ching Poon, Shi Bu, Ming Wai Lau, Ka Nang Leung
IEEE Trans. Very Large Scale Integr. Syst.5
2016 Asymmetrical Dead-Time Control Driver for Buck Regulator
abstract
This brief presents an asymmetrical dead-time control driver (ASDTCD) for synchronous buck converter operating in the continuous conduction mode. Dead-time control is an important metric for improving the efficiency of switching mode power regulator. Without an additional circuit, the proposed ASDTCD can generate dead time by controlling the slope for the output signal of the driver. The proposed ASDTCD utilizes the transition between triode region and saturation region for the power transistor to avoid body-diode conduction and shoot-through current while minimizing the switching loss. Thus, high-speed body-diode conduction sensor is avoided; thereby, reducing the power consumption and saving silicon area. In addition, the body-diode conduction time control accuracy is also enhanced. Less than 1-ns body-diode conduction time has been achieved without bringing in shoot-through current across 10-450-mA load range. With less than 0.5% of the total input power consumed, the proposed ASDTCD takes less than 1% of the power transistor area. This design is implemented in the 0.18-μm CMOS process.
Chundong Wu, Wang Ling Goh, Chiang Liang Kok, Liter Siek, Yat-Hei Lam, Ravinder Pal Singh
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Design considerations of STCB OTA in CMOS 65nm with large capacitive loads
abstract
A modified structure of OTA in CMOS 65-nm with signal- and transient-current boosting is presented in this paper. The structure uses simple cascode current mirrors to overcome channel-modulation effect of the 65-nm MOSFETs and to maintain low-error current matching. Simulations show that the parasitic poles of the OTA in CMOS 65-nm are located at very high frequencies and the achievable bandwidth is much increased with sufficient phase margin to maintain closed-loop stability.
Kai Ho Mak, Marco Ho, Ka Nang Leung, Wang Ling Goh
ISCAS4
2011 Power-Efficient Explicit-Pulsed Dual-Edge Triggered Sense-Amplifier Flip-Flops
abstract
A novel explicit-pulsed dual-edge triggered sense-amplifier flip-flop (DET-SAFF) for low-power and high-performance applications is presented in this paper. By incorporating the dual-edge triggering mechanism in the new fast latch and employing conditional precharging, the DET-SAFF is able to achieve low-power consumption that has small delay. To further reduce the power consumption at low switching activities, a clock-gated sense-amplifier (CG-SAFF) is engaged. Extensive post-layout simulations proved that the proposed DET-SAFF exhibits both the low-power and high-speed properties, with delay and power reduction of up to 43.3% and 33.5% of those of the prior art, respectively. When the switching activity is less than 0.5, the proposed CG-SAFF demonstrates its superiority in terms of power reduction. During zero input switching activity, CG-SAFF can realize up to 86% in power saving. Lastly, a modification to the proposed circuit has led to an improved common-mode rejection ratio (CMRR) DET-SAFF.
Myint Wai Phyu, Kangkang Fu, Wang Ling Goh, Kiat Seng Yeo
IEEE Trans. Very Large Scale Integr. Syst.3
2010 Design of Low-Power High-Speed Truncation-Error-Tolerant Adder and Its Application in Digital Signal Processing
abstract
In modern VLSI technology, the occurrence of all kinds of errors has become inevitable. By adopting an emerging concept in VLSI design and test, error tolerance (ET), a novel error-tolerant adder (ETA) is proposed. The ETA is able to ease the strict restriction on accuracy, and at the same time achieve tremendous improvements in both the power consumption and speed performance. When compared to its conventional counterparts, the proposed ETA is able to attain more than 65% improvement in the Power-Delay Product (PDP). One important potential application of the proposed ETA is in digital signal processing systems that can tolerate certain amount of errors.
Wang Ling Goh, Weija Zhang, Kiat Seng Yeo, Zhi-Hui Kong
IEEE Trans. Very Large Scale Integr. Syst.2
2009 A Low-Noise Multi-GHz CMOS Multiloop Ring Oscillator With Coarse and Fine Frequency Tuning
abstract
A 7-GHz CMOS voltage controlled ring oscillator that employs multiloop technique for frequency boosting is presented in this paper. The circuit permits lower tuning gain through the use of coarse/fine frequency control. The lower tuning gain also translates into a lower sensitivity to the voltage at the control lines. Fabricated in a standard 0.13-mum CMOS process, the proposed voltage-controlled ring oscillator exhibits a low phase noise of -103.4 dBc/Hz at 1 MHz offset from the center frequency of 7.64 GHz, while consuming a current of 40 mA excluding the buffer.
Hai Qi Liu, Wang Ling Goh, Liter Siek, Wei Meng Lim, Yue Ping Zhang
IEEE Trans. Very Large Scale Integr. Syst.2