Yi Sheng Chong

dblp:279/1473 · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-4136-6570ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-author · 13 since 2021
YearPublicationVenuePosition
2026 Low-Power Learnable Digital Audio Feature Extractor for Always-on Keyword Spotting in Edge Devices
abstract
This paper presents a low-power, learnable digital audio feature extractor (AuFEx) for always-on keyword spotting (KWS) in edge devices. A noise-aware design flow is introduced to integrate the tuning of AuFEx parameters directly into the neural network classifier’s training process. This co-design approach enables the joint optimization of feature extractor and classifier within a unified framework. By incorporating noise during the training process, the system becomes more robust to variations in signal-to-noise ratio (SNR), maintaining high inference accuracy even with lightweight neural network classifiers. This design flow supports the design of both time-domain AuFEx (TD-AuFEx) and frequency-domain AuFEx (FD-AuFEx) for single-keyword wake-word detection (WWD) and 10-keyword KWS tasks, respectively. Implemented in a 40nm CMOS process, both TD-AuFEx and FD-AuFEx achieve over 2% accuracy improvement with the smallest size backend classifier. Specifically, the TD-AuFEx for WWD task achieves classifier size of 3.1k parameters with 494 nW power consumption and$375~\mu $s latency. The accuracy is maintained between 96.2% – 97.9% for SNR in the range of 5 – 20 dB. The FD-AuFEx for 10-keyword KWS achieves classifier size of 6.39k parameters with$1.258~\mu $W power consumption and 34.625 ms latency. The accuracy is maintained between 87.5% – 92.2% for SNR in the range of 5 dB - 20 dB, which is one of the highest compared to the other state-of-the-art designs.
Jinhai Hu, Wang Ling Goh, Yi Sheng Chong, Anh-Tuan Do, Yuan Gao 0011
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Qubit-State Discrimination using Neural Networks with Rapid and Energy-Efficient Compute Arrays
abstract
Neural networks (NNs) implemented on field-programmable gate arrays (FPGAs) provide fast, high-fidelity solutions for processing readout signals from quantum information processors. However, application-specific integrated circuits (ASICs) instead of FPGAs hold the potential for improved performance, a largely unexplored path. This work proposes specialized hardware for NN-based qubit-state discrimination. We optimize the NN architecture to minimize resource requirements by reducing the layer width, employing linear activation functions, and weight quantization. Quantization-aware training is used to preserve accuracy despite these optimizations. Next, a compute array employing output stationary dataflow is chosen to process the NN workload. The compute array with abundant multipliers and adders can complete one NN inference in 63 ns, which makes it a good candidate for real-time qubit-state discrimination.
Yuntian Liu, Yi Sheng Chong, Benjamin Lienhard, Minghao Fan, Wang Ling Goh, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS2
2024 PACE: A Scalable and Energy Efficient CGRA in a RISC-V SoC for Edge Computing Applications
abstract
▪Coarse-grained reconfigurable arrays (CGRAs) deliver high energy efficiency while maintaining the programmability advantages. ▪CGRA is the ideal candidate for efficiently handling loop kernels, which allows it to offload repetitive looping functions such as vector multiplication or hashing algorithms from CPUs. ▪It relies on a compiler to convert a given workload into a data flow graph (DFG) which is then mapped onto the hardware in a manner that achieves the highest possible energy efficiency.
Vishnu P. Nambiar, Yi Sheng Chong, Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Zhaoying Li 0004, Rohan Juneja, Li-Shiuan Peh, Tulika Mitra, Anh-Tuan Do
HCS2
2024 Quantum Readout Processing Accelerator with a CORDIC Core at Cryogenic Temperature
abstract
Quantum computing has been the most promising in addressing computational challenging problems beyond classical computers’ capabilities. Recent work focus on developing systemon-chip (SoC) for quantum control and readout, yet there are scarce literature discussing how the measurement blocks are developed in details. This paper presents a quantum readout processing accelerator that operates at cryogenic temperature to process measurement data directly to obtain the qubit state. The proposed accelerator includes a coordinate rotation digital computer (CORDIC) based direct digital frequency synthesizer (DDFS) and a qubit state decision unit. To enable compact processing accelerator, the error of the qubit state prediction is analyzed when choosing different bitwidths and number of CORDIC iterations, aiming to balance the accuracy and hardware cost. The proposed accelerator, implemented using a 28-nm CMOS, consumes 36mW of power at room temperature and 8.9mW at −196°C. The processing latency and total energy consumption when determining a qubit state are reduced by 6 and 3 orders of magnitude respectively, when compared to a general purpose processor.
Yi Sheng Chong, Hongyu Cao, Wang Ling Goh, Patrick Bore, Yuanzheng Paul Tan, Yung Szen Yap, Rainer Dumke, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS1
2024 A 420 GOPS/W CGRA with a Configurable MAC and Dynamic Truncation
abstract
Edge devices demand for highly efficient yet flexible processing capability to handle dynamic real-time workloads. Coarse grain reconfigurable architecture (CGRA) emerges as a suitable accelerator candidate in edge devices, because they are as flexible as general purpose processors and offer high efficiency close to that of domain specific accelerators. However, a typical CGRA requires two cycles for a multiply-and-accumulate (MAC) operation, and workloads such as neural network inference and signal processing involve many MAC operations, resulting in long CGRA processing time. This work proposes a CGRA that has configurable MAC units in the processing elements (PEs) that can perform an addition (ADD) or multiplication (MUL) or a MAC by using the same multiplier and adder, in a single cycle. The readout precision of MAC result can be adjusted by a truncation block. The proposed CGRA is implemented with 40nm CMOS technology. It attains an energy efficiency of 420.6GOPS/W operating at supply of 0.6V and frequency of 21MHz, which is 1.4 times higher than the state-of-the-art.
Yi Sheng Chong, Rakshith Harish, Rajesh Chandrasekhara Panicker, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS1
2024 3881 Gbps/W, 3005 µm AES Core with State Based Clock Gating for IoT applications
abstract
An efficient and extremely low-energy AES-256 accelerator was implemented in a 40nm CMOS process for space and energy-limited IoT applications, supporting encryption, decryption, ECB and CBC modes. By reusing functional block hardware and utilizing scan flip-flops, the proposed AES design occupies merely 3005 µm of silicon area. Additionally, clock gating circuits were extensively used at circuit level to halt essential flip-flops during the ShiftRows and its reversed operation for aggressive power saving. The post-layout simulation showed that our proposed design achieved the lowest power consumption of 1.83 µW at 0.3 V, 17 MHz, with energy efficiency as high as 3881 Gbps/W.
Zhangyi Pei, Vishnu P. Nambiar, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do
ISCAS3
2024 1.63 pJ/SOP Neuromorphic Processor With Integrated Partial Sum Routers for In-Network Computing
abstract
Neuromorphic computing is promising to achieve unprecedented energy efficiency by emulating the human brain’s mechanism. Conventional neuromorphic accelerators employ split-and-merge method to map spiking neural networks’ inputs to surpass the fan-in capabilities of a single neuron core. However, this approach gives rise to the risk of accuracy compromise and extra core usage for the merging process. Moreover, it requires excessive data movement and clock cycles to aggregate spikes generated by partial sums instead of total sums obtained from different cores with substantial power and energy overhead. This work presents a novel approach to addressing the challenges imposed by the split-and-merge method. We propose an energy-efficient, reconfigurable neuromorphic processor that leverages several key techniques to mitigate the above issues. First, we introduce a partial sum router circuitry that enables in-network computing (INC), eliminating the need for extra merge cores. Second, we adopt software-defined Networks-on-Chip (NoCs) by leveraging predefined, efficient routing, eliminating power-hungry routing computation. At last, we incorporate fine-grained power gating and clock gating techniques for further power reduction. Experimental results from our test chip demonstrate the lossless mapping of the algorithm and exceptional energy efficiency, achieving an energy consumption of 1.63 pJ/SOP at 0.48 V. This energy efficiency represents a 22.4% improvement compared to the state-of-the-art results. Our proposed neuromorphic processor provides an efficient and flexible solution for neural network processing, mitigating the limitations of the traditional split-and-merge approach while delivering superior energy efficiency.
Dongrui Li, Ming Ming Wong, Yi Sheng Chong, Jun Zhou 0014, Mohit Upadhyay, Ananta Narayanan Balaji, Aarthy Mani, Weng-Fai Wong, Li-Shiuan Peh, Anh-Tuan Do, Bo Wang 0020
IEEE Trans. Very Large Scale Integr. Syst.3
2023 1.7pJ/SOP Neuromorphic Processor with Integrated Partial Sum Routers for In-Network Computing
abstract
Conventional neuromorphic accelerators primarily leverage split-merge method to accommodate a neural network that is beyond a single core's size, leading to possible accuracy loss, extra core usage and significant power and energy overhead. This work presents an energy-efficient, reconfigurable neuro-morphic processor to address the problem by (i) a partial sum router circuitry that enables in-network computing to remove the need of extra merge cores; (ii) software-defined Networks-on-Chip that eliminates the power-hungry routing compute and (iii) fine-grained power gating and clock gating technique for power reduction. Our test chip achieves lossless mapping as the algorithm and an energy efficiency of 1.7pJ/SOP at 0.5V, 19% lower than state-of-the-art result.
Bo Wang 0020, Ming Ming Wong, Dongrui Li, Yi Sheng Chong, Jun Zhou 0014, Weng-Fai Wong, Li-Shiuan Peh, Aarthy Mani, Mohit Upadhyay, Ananta Narayanan Balaji, Anh-Tuan Do
ISCAS4
2023 A 110nW Always-on Keyword Spotting Chip using Spiking CNN in 40nm CMOS
abstract
This paper presents an ultra-low power keyword spotting (KWS) chip for Artificial Intelligence of Things (AIoT) device's always-on ambient sensing function. The core KWS engine is based on a spiking convolutional neural network (SCNN) model for its attractive features of sparse activation and addition-only operations inside the spiking neurons. The proposed SCNN model improves the existing framewise incremental computation flow by adding a spike processing unit (SPU) to reduce the computing cycles. The power and latency of the whole system are reduced by 16.5% and 43.2% respectively. Extensive network quantization reduces the weight bit-length to 4-bit and only 1-bit activation is required. The chip also supports power gating by an energy-based voice activity detection (VAD) module to further reduce power consumption in random and sparse event (RSE) scenarios. Full chip simulation results show that the chip consumes only 110nW with 2.15% False alarm rate and 3.00% False reject rate in a 10% voice event stream test. It achieves state-of-art recognition accuracy of 99% and 96% for one and two keyword detection tasks.
Junran Pu, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do, Yuan Gao 0011
ISCAS3
2022 0.08mm2 128nW MFCC Engine for Ultra-low Power, Always-on Smart Sensing Applications
abstract
Mel frequency cepstral coefficient (MFCC) features are widely used in applications such as keyword spotting, bearing fault detection and heart sound classification. This work proposes a low power MFCC engine that enables its use for battery-powered edge applications. Three hardware algorithm co-optimizations were adopted to achieve energy efficient MFCC hardware implementation. The approximated MFCC features due to the optimizations still allows good accuracy when deployed in several applications such as keyword spotting and bearing fault detection, reporting negligible accuracy drop of $\le 1.5$%. The proposed MFCC hardware consumes only 128nW at 0.3V supply and occupies only 0.08mm2in 40nm CMOS technology, which are $5 \times $ and $2.75 \times $ power and area reduction respectively when compared to the prior arts.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS1
2022 Recovering Accuracy of RRAM-based CIM for Binarized Neural Network via Chip-in-the-loop Training
abstract
Resistive random access memory (RRAM) based computing-in-memory (CIM) is attractive for edge artificial intelligence (AI) applications, thanks to its excellent energy efficiency, compactness and high parallelism in matrix vector multiplication (MatVec) operations. However, existing RRAM-based CIM designs often require complex programming scheme to precisely control the RRAM cells to reach the desired resistance states so that the neural network classification accuracy is maintained. This leads to large area and energy overhead as well as low RRAM area utilization. Hence, compact RRAM-based CIM with simple pulse-based programming scheme is thus more desirable. To achieve this, we propose a chip-in-the-loop training approach to compensate for the network performance drop due to the stochastic behavior of the RRAM cells. Note that, although the target RRAM cell here is a two-state RRAM (i.e binary, having only high and low resistance states), their inherent analog resistance values are used in the CIM operation. Our experiment using a 4-layer fully-connected binary neural network (BNN) showed that after retraining, the RRAM-based network accuracy can be recovered, regardless of the RRAM resistance distribution and $\frac{\text{R}_{\text{HRS}}}{\text{R}_{\text{LRS}}}$ resistance ratio.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS1
2021 An Energy-Efficient Convolution Unit for Depthwise Separable Convolutional Neural Networks
abstract
High performance but computationally expensive Convolutional Neural Networks (CNNs) require both algorithmic and custom hardware improvement to reduce model size and to improve energy efficiency for edge computing applications. Recent CNN architectures employ depthwise separable convolution to reduce the total number of weights and MAC operations. However, depthwise separable convolution workload does not run efficiently in existing CNN accelerators. This paper proposes an energy-efficient CONV unit for pointwise and depthwise operation. The CONV unit utilizes weight stationary to enable high efficiency. The row partial sum reduction is engaged to increase parallelism in pointwise convolution thereby lightening the memory requirements on output partial sums. Our design achieves a maximum efficiency of 3.17 TOPS/W at 0.85V/40nm CMOS which is well-suited for energy constrained edge computing applications.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
ISCAS1
2021 Efficient Implementation of Activation Functions for LSTM accelerators
abstract
Activation functions such as hyperbolic tangent (tanh) and logistic sigmoid (sigmoid) are critical computing elements in a long short term memory (LSTM) cell and network. These activation functions are non-linear, leading to challenges in their hardware implementations. Area-efficient and high performance hardware implementation of these activation functions thus becomes crucial to allow high throughput in a LSTM accelerator. In this work, we propose an approximation scheme which is suitable for both tanh and sigmoid functions. The proposed hardware for sigmoid function is 8.3 times smaller than the state-of-the-art, while for tanh function, it is the second smallest design. When applying the approximated tanh and sigmoid of 2% error in a LSTM cell computation, its final hidden state and cell state record errors of 3.1% and 5.8% respectively. When the same approximated functions are applied to a single layer LSTM network of 64 hidden nodes, the accuracy drops by 2.8% only. This proposed small yet accurate activation function hardware is promising to be used in Internet of Things (IoT) applications where accuracy can be traded off for ultra-low power consumption.
Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do
VLSI-SoC1
2020 Area and Energy Efficient 2D Max-Pooling For Convolutional Neural Network Hardware Accelerator
abstract
In the era of digital world, Convolutional Neural-network is widely adopted. To keep the network in a reasonable size, down-sampling is required. Max-pooling is a common solution for down-sampling. High throughput and area-efficient -hardware design of 2D max-pooling is essential for CNN accelerator. In this paper, we propose a highly efficient hardware design of max-pooling for 2D data input. The proposed design can be integrated into the neural-network accelerator as it can process the 2D input feature map data in parallel. The proposed solution, which is supporting up to 31 parallel data input, has been synthesized using at 40nm process node with area of 90×100 μm2, and the estimated power is 2.9 mW @200Mhz @1.1V power supply.
Yi Sheng Chong, Anh-Tuan Do
IECON2