VLDB 2026 Research / reviewers in the wild / expert
Vishnu P. Nambiar
dblp:53/8665
· DBLP profile ↗
18ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0001-5570-5911ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Qubit-State Discrimination using Neural Networks with Rapid and Energy-Efficient Compute ArraysabstractNeural networks (NNs) implemented on field-programmable gate arrays (FPGAs) provide fast, high-fidelity solutions for processing readout signals from quantum information processors. However, application-specific integrated circuits (ASICs) instead of FPGAs hold the potential for improved performance, a largely unexplored path. This work proposes specialized hardware for NN-based qubit-state discrimination. We optimize the NN architecture to minimize resource requirements by reducing the layer width, employing linear activation functions, and weight quantization. Quantization-aware training is used to preserve accuracy despite these optimizations. Next, a compute array employing output stationary dataflow is chosen to process the NN workload. The compute array with abundant multipliers and adders can complete one NN inference in 63 ns, which makes it a good candidate for real-time qubit-state discrimination. Yuntian Liu, Yi Sheng Chong, Benjamin Lienhard, Minghao Fan, Wang Ling Goh, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 6 |
| 2024 | PACE: A Scalable and Energy Efficient CGRA in a RISC-V SoC for Edge Computing Applicationsabstract▪Coarse-grained reconfigurable arrays (CGRAs) deliver high energy efficiency while maintaining the programmability advantages. ▪CGRA is the ideal candidate for efficiently handling loop kernels, which allows it to offload repetitive looping functions such as vector multiplication or hashing algorithms from CPUs. ▪It relies on a compiler to convert a given workload into a data flow graph (DFG) which is then mapped onto the hardware in a manner that achieves the highest possible energy efficiency. Vishnu P. Nambiar, Yi Sheng Chong, Thilini Kaushalya Bandara, Dhananjaya Wijerathne, Zhaoying Li 0004, Rohan Juneja, Li-Shiuan Peh, Tulika Mitra, Anh-Tuan Do |
HCS | 1 |
| 2024 | Quantum Readout Processing Accelerator with a CORDIC Core at Cryogenic TemperatureabstractQuantum computing has been the most promising in addressing computational challenging problems beyond classical computers’ capabilities. Recent work focus on developing systemon-chip (SoC) for quantum control and readout, yet there are scarce literature discussing how the measurement blocks are developed in details. This paper presents a quantum readout processing accelerator that operates at cryogenic temperature to process measurement data directly to obtain the qubit state. The proposed accelerator includes a coordinate rotation digital computer (CORDIC) based direct digital frequency synthesizer (DDFS) and a qubit state decision unit. To enable compact processing accelerator, the error of the qubit state prediction is analyzed when choosing different bitwidths and number of CORDIC iterations, aiming to balance the accuracy and hardware cost. The proposed accelerator, implemented using a 28-nm CMOS, consumes 36mW of power at room temperature and 8.9mW at −196°C. The processing latency and total energy consumption when determining a qubit state are reduced by 6 and 3 orders of magnitude respectively, when compared to a general purpose processor. Yi Sheng Chong, Hongyu Cao, Wang Ling Goh, Patrick Bore, Yuanzheng Paul Tan, Yung Szen Yap, Rainer Dumke, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 8 |
| 2024 | A 420 GOPS/W CGRA with a Configurable MAC and Dynamic TruncationabstractEdge devices demand for highly efficient yet flexible processing capability to handle dynamic real-time workloads. Coarse grain reconfigurable architecture (CGRA) emerges as a suitable accelerator candidate in edge devices, because they are as flexible as general purpose processors and offer high efficiency close to that of domain specific accelerators. However, a typical CGRA requires two cycles for a multiply-and-accumulate (MAC) operation, and workloads such as neural network inference and signal processing involve many MAC operations, resulting in long CGRA processing time. This work proposes a CGRA that has configurable MAC units in the processing elements (PEs) that can perform an addition (ADD) or multiplication (MUL) or a MAC by using the same multiplier and adder, in a single cycle. The readout precision of MAC result can be adjusted by a truncation block. The proposed CGRA is implemented with 40nm CMOS technology. It attains an energy efficiency of 420.6GOPS/W operating at supply of 0.6V and frequency of 21MHz, which is 1.4 times higher than the state-of-the-art. Yi Sheng Chong, Rakshith Harish, Rajesh Chandrasekhara Panicker, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 4 |
| 2024 | Exploring Error Correction Circuits on RISC-V based Systems for Space ApplicationsabstractRISC-V systems are becoming increasingly adopted in space applications. SRAM (Static Random-Access Memory) data memory is a critical component that occupies a large portion of the processor peripheral system. SRAM is vulnerable to single-event upsets (SEUs). Existing studies mainly considered Hamming error correction codes (ECCs) for memory protection in RISC-V processor. In this paper, we explore different ECCs as well as the triple modular redundancy (TMR) as solutions to mitigate SEU effects on SRAMs for RISC-V system. To overcome the area and power overhead of TMR with higher fault tolerance than ECCs, we propose a dual modular redundancy (DMR) with ECC memory protection scheme. We conduct a comprehensive error analysis to evaluate different fault-tolerant designs under various radiation attack scenarios, utilizing real-world data to devise the fault-injection campaign. An efficient scheduling for the self-refresh operation of SRAMs is proposed to prevent the error accumulation. The proposed DMR with ECC design reduces the power and area overhead of TMR by 28% and 11% respectively and improve the error resilience of the SRAM significantly compared with Hsiao and Hamming ECC schemes. Nazim Altar Koca, Chip-Hong Chang, Anh-Tuan Do, Vishnu P. Nambiar |
ISCAS | 4 |
| 2024 | 3881 Gbps/W, 3005 µm AES Core with State Based Clock Gating for IoT applicationsabstractAn efficient and extremely low-energy AES-256 accelerator was implemented in a 40nm CMOS process for space and energy-limited IoT applications, supporting encryption, decryption, ECB and CBC modes. By reusing functional block hardware and utilizing scan flip-flops, the proposed AES design occupies merely 3005 µm of silicon area. Additionally, clock gating circuits were extensively used at circuit level to halt essential flip-flops during the ShiftRows and its reversed operation for aggressive power saving. The post-layout simulation showed that our proposed design achieved the lowest power consumption of 1.83 µW at 0.3 V, 17 MHz, with energy efficiency as high as 3881 Gbps/W. Zhangyi Pei, Vishnu P. Nambiar, Yi Sheng Chong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 2 |
| 2022 | 0.08mm2 128nW MFCC Engine for Ultra-low Power, Always-on Smart Sensing ApplicationsabstractMel frequency cepstral coefficient (MFCC) features are widely used in applications such as keyword spotting, bearing fault detection and heart sound classification. This work proposes a low power MFCC engine that enables its use for battery-powered edge applications. Three hardware algorithm co-optimizations were adopted to achieve energy efficient MFCC hardware implementation. The approximated MFCC features due to the optimizations still allows good accuracy when deployed in several applications such as keyword spotting and bearing fault detection, reporting negligible accuracy drop of $\le 1.5$%. The proposed MFCC hardware consumes only 128nW at 0.3V supply and occupies only 0.08mm2in 40nm CMOS technology, which are $5 \times $ and $2.75 \times $ power and area reduction respectively when compared to the prior arts. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 4 |
| 2022 | Recovering Accuracy of RRAM-based CIM for Binarized Neural Network via Chip-in-the-loop TrainingabstractResistive random access memory (RRAM) based computing-in-memory (CIM) is attractive for edge artificial intelligence (AI) applications, thanks to its excellent energy efficiency, compactness and high parallelism in matrix vector multiplication (MatVec) operations. However, existing RRAM-based CIM designs often require complex programming scheme to precisely control the RRAM cells to reach the desired resistance states so that the neural network classification accuracy is maintained. This leads to large area and energy overhead as well as low RRAM area utilization. Hence, compact RRAM-based CIM with simple pulse-based programming scheme is thus more desirable. To achieve this, we propose a chip-in-the-loop training approach to compensate for the network performance drop due to the stochastic behavior of the RRAM cells. Note that, although the target RRAM cell here is a two-state RRAM (i.e binary, having only high and low resistance states), their inherent analog resistance values are used in the CIM operation. Our experiment using a 4-layer fully-connected binary neural network (BNN) showed that after retraining, the RRAM-based network accuracy can be recovered, regardless of the RRAM resistance distribution and $\frac{\text{R}_{\text{HRS}}}{\text{R}_{\text{LRS}}}$ resistance ratio. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 4 |
| 2022 | A 1800μm2, 953Gbps/W AES Accelerator for IoT Applications in 40nm CMOSabstractA compact and energy-efficient AES accelerator for area and power-constrained IoT applications was fabricated in a 40nm CMOS process. By eliminating the need of intermediate data registers for MixColumns and ShiftRows in our proposed AES accelerator, we were able to reduce the total flip-flops to only 269 bits. Further, by reusing functional blocks and swapping the D flip-flops in data storage with scan flip-flops, our chip occupies only a tiny area of $1800 \mu \text{m}^{2}$ with an extremely low number of 657 gates. In addition, clock gating method and near-threshold voltage were used in our design. Thus, our accelerator consumes only $3.2 \mu \text{W}$ with an operation efficiency of 953 Gbps/W using a 0.48 V supply voltage. Compared with prior arts, our design has savings of 53% on area and 55% on the number of gates. When operated with a supply voltage of 0.48 V at 25°C, we can also achieve lower energy efficiency. Jingjing Lan, Vishnu P. Nambiar, Ming Ming Wong, Fei Li 0015, Yuan Gao 0011, Kevin Tshun Chuan Chai, Anh-Tuan Do |
ISCAS | 2 |
| 2021 | An Energy-Efficient Convolution Unit for Depthwise Separable Convolutional Neural NetworksabstractHigh performance but computationally expensive Convolutional Neural Networks (CNNs) require both algorithmic and custom hardware improvement to reduce model size and to improve energy efficiency for edge computing applications. Recent CNN architectures employ depthwise separable convolution to reduce the total number of weights and MAC operations. However, depthwise separable convolution workload does not run efficiently in existing CNN accelerators. This paper proposes an energy-efficient CONV unit for pointwise and depthwise operation. The CONV unit utilizes weight stationary to enable high efficiency. The row partial sum reduction is engaged to increase parallelism in pointwise convolution thereby lightening the memory requirements on output partial sums. Our design achieves a maximum efficiency of 3.17 TOPS/W at 0.85V/40nm CMOS which is well-suited for energy constrained edge computing applications. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
ISCAS | 4 |
| 2021 | Efficient Implementation of Activation Functions for LSTM acceleratorsabstractActivation functions such as hyperbolic tangent (tanh) and logistic sigmoid (sigmoid) are critical computing elements in a long short term memory (LSTM) cell and network. These activation functions are non-linear, leading to challenges in their hardware implementations. Area-efficient and high performance hardware implementation of these activation functions thus becomes crucial to allow high throughput in a LSTM accelerator. In this work, we propose an approximation scheme which is suitable for both tanh and sigmoid functions. The proposed hardware for sigmoid function is 8.3 times smaller than the state-of-the-art, while for tanh function, it is the second smallest design. When applying the approximated tanh and sigmoid of 2% error in a LSTM cell computation, its final hidden state and cell state record errors of 3.1% and 5.8% respectively. When the same approximated functions are applied to a single layer LSTM network of 64 hidden nodes, the accuracy drops by 2.8% only. This proposed small yet accurate activation function hardware is promising to be used in Internet of Things (IoT) applications where accuracy can be traded off for ultra-low power consumption. Yi Sheng Chong, Wang Ling Goh, Yew-Soon Ong, Vishnu P. Nambiar, Anh-Tuan Do |
VLSI-SoC | 4 |
| 2021 | A 5.28-mm² 4.5-pJ/SOP Energy-Efficient Spiking Neural Network Hardware With Reconfigurable High Processing Speed Neuron Core and Congestion-Aware RouterabstractIn recent years, fast computation, low power, and small footprint are the key motivations for building SNN hardware. The unique features of SNN hardware have not been fully exploited, where the computation speed and energy efficiency of the SNN hardware can be improved according to the sparse spiking and non-uniform traffic of SNN. In this paper, we propose a 5.28-mm$^{2}~4096$-neuron 1M-synapse energy-efficient digital SNN hardware that can achieve ultra-low energy per synaptic operation of 4.5 pJ. The proposed neuron computing unit is implemented in pipeline architecture to achieve high synaptic processing speed. The proposed spike processing unit can significantly increase the processing speed of the neuron core by$1.9\times $and$9.4\times $when the spike injection rate is 50% and 10%, respectively. Besides, the increase in the processing speed of the neuron core leads to a reduction in energy consumption of up to 81.5%. An event-driven clock gating circuit that can reduce the power consumption of the proposed neuron block by more than 70% is proposed in this paper. This paper proposes a supervised STDP+ algorithm for SNN training, and the classification accuracy of the MNIST digits is 89.6% with 73.6% weight sparsity of the output layer. Junran Pu, Wang Ling Goh, Vishnu P. Nambiar, Ming Ming Wong, Anh-Tuan Do |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Post-Silicon Validation Methodology for Resource-Constrained Neuromorphic HardwareabstractWith the semiconductor industry moving towards smaller and more advanced process nodes, the reliability of the fabricated system-on-chips (SoCs) has become a significant problem, although it can be mitigated by proper verification and validation methodologies. Typical neuromorphic SoC chips are resource constrained and have very low power envelopes, pushing designers to forego the integration of onchip debug circuits. Futhermore, these specialized neuromorphic SoCs may not even be capable of supporting memory debug circuits due to the inclusion of novel memory types. Hence, it is a huge challenge to execute a comprehensive system level post-silicon validation suite on such chips, while aiding debugging efforts reliably. This paper presents a post-silicon validation methodology and algorithm to test all the neurocores within our custom designed neuromorphic processing unit (NPU) chip, focused on pinpointing the faults on the synaptic weight storage elements. The described methodology does not require any specialized built-in debug circuits to infer and detect faults. Based on the test results, the custom NPU chip was able to respond accordingly to all performed tests under optimal voltage and frequency conditions. Yun Kwan Lee, Vishnu P. Nambiar, Kim Seng Goh, Anh-Tuan Do |
IECON | 2 |
| 2020 | Scalable Block-Based Spiking Neural Network Hardware with a Multiplierless Neuron ModelabstractThis paper proposes a scalable hardware architecture for block-based spiking neural networks utilizing a multiplierless spiking neuron model. These blocks were implemented as a neurocore mesh generated from an interconnect algorithm, allowing for seamless scalability of the network size while mitigating connectivity errors. The routing fabric asynchronous protocol allows for critical timing paths between blocks to be relaxed. The proposed neuron model consumed less logic compared to standard models with multipliers, reducing up to 16% of the neurocore logic cell area. The network was implemented alongside a computing subsystem as an FPGA-based system-on-chip, communicating via a fabric interconnect bridge. Experimental results validate the functionality of the proposed system, and achieved comparable classification accuracy to existing works. Vishnu P. Nambiar, Eng-Kiat Koh, Junran Pu, Aarthy Mani, Ming Ming Wong, Wang Ling Goh, Anh-Tuan Do |
ISCAS | 1 |
| 2019 | Block-Based Spiking Neural Network Hardware with Deme Genetic AlgorithmabstractHardware implementation of spiking neural networks (SNN) has been the focus of many previous works due to its higher execution speed. A block-based SNN architecture with a simple spiking neuron model is proposed in this paper. Compared to traditional spiking neuron models, the proposed model simplifies the equation of the membrane potential for ease of hardware implementation. The block-based SNN architecture also makes the hardware implementation more scalable and simplifies floorplanning. Deme genetic algorithm (GA) was applied for training the SNN model, and a population encoding scheme was used for spike time conversion. Two case studies were carried out to verify the functionality of the proposed model, namely number recognition and Fisher Iris classification. Experimental results showed that the proposed SNN model with deme GA was able to achieve comparable or higher classification accuracy than previous works. Junran Pu, Vishnu P. Nambiar, Anh-Tuan Do, Wang Ling Goh |
ISCAS | 2 |
| 2014 | Optimization of structure and system latency in evolvable block-based neural networks using genetic algorithm
Vishnu P. Nambiar, Mohamed Khalil Hani, Muhammad N. Marsono, Chen Wei Sia |
Neurocomputing | 1 |
| 2014 | Hardware implementation of evolvable block-based neural networks utilizing a cost efficient sigmoid-like activation function
Vishnu P. Nambiar, Mohamed Khalil Hani, Riadh Sahnoun, Muhammad N. Marsono |
Neurocomputing | 1 |
| 2013 | Co-simulation methodology for improved design and verification of hardware neural networksabstractThis paper presents a methodology for speeding up the design and verification of process of artificial neural networks (ANNs) in system-on-chip (SoC) hardware with the help of co-simulation. Application of advanced design methodologies for complex designs such as ANNs are important in todays fast moving hardware design industry. However, it is difficult to fully verify the functionality of ANNs when designed in hardware. Most forms of ANN require the use of complex training algorithms, which are difficult to implement in a testbench even with the help of modern interfaces such as SystemVerilog's DPI-C. The neural network topology selected as the case study for this paper are evolvable block-based neural networks (BbNNs). The case studies employed during the verification process are the XOR problem, driver drowsiness classification, and heart arrhythmia classification. The proposed methodology significantly reduces the verification time required for the design of hardware-based neural networks. This allows complex models applicable a variety of applications to be quickly designed, such as industrial motor controllers or fuzzy systems. Mohamed Khalil Hani, Vishnu P. Nambiar, Muhammad N. Marsono |
IECON | 2 |