EDBT 2026 Demo / reviewers in the wild / expert
Ajay Kumar Singh
dblp:06/7373
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analog and Temporary On-chip Memory for ANN Training and InferenceabstractOn-chip training at the edge becomes a primary requisite for real-time and security-sensitive artificial neural network (ANN) applications. In-memory computation (IMC) techniques have been proposed to facilitate data-intensive computational operations in ANNs. IMC-based multiply-accumulate (MAC) accelerates ANN training but suffers from significant communication overhead between the MAC engine and the off-chip storage for the intermediate data. This article proposes an analog temporary on-chip memory (ATOM) to store this intermediate data during ANN training. The ANN training architecture with the proposed ATOM has two significant advantages. First, the energy required to store intermediate data is scaled down by \(\sim\) 40 \(\times\) due to the on-chip and analog nature of the memory. Second, the proposed architecture avoids power and area-consuming analog-to-digital converters (ADCs) between neural network stages. The ATOM cell measurements are carried out from 20 fabricated chips, and the impact of ATOM characteristics on ANN system performance accuracy is analyzed. This article shows significant latency improvement of \(\sim\) 9 \(\times\) and area savings of \(\sim\) 5 \(\times\) for intermediate data storage compared to the on-chip SRAM during ANN training’s forward and backward pass operations. An improvement in the area and latency will be beneficial to instrument the area- and energy-efficient hardware system for on-chip ANN applications. Shreyas Deshmukh, Raghav Singhal, Shruti Landge, Vivek Saraswat, Anmol Biswas, Abhishek Kadam, Ajay Kumar Singh, Sreenivas Subramoney, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2024 | A Compact Low Power Multi-mode Spiking Neuron using Band to Band TunnelingabstractEfficient and compact neurons with low power consumption are crucial when designing large-scale spiking neural networks (SNNs) for hardware implementation. Many architectures in the literature showcase different spike patterns associated with biological neurons. However, using bulky capacitors to generate the different time constants related to complex neuron patterns makes these circuits area inefficient. This paper presents a band-to-band-tunneling (BTBT) based energy-efficient and compact neuron capable of producing various spike patterns. The BTBT region’s extremely low current enables different time constants while eliminating the need of bulky capacitors. The circuit is based on the Izhikevich neuron model. The proposed circuit is designed in Silicon on Insulator technology to exhibit important firing patterns observed in the biological cortex, viz. regular spiking, fast-spiking, and chattering, and it is fine-tuned for efficient operation at low subthreshold voltages. This circuit utilizes only 129 μm2area and consumes only 6.7 fJ energy per spike ( approximately 40% lower area and energy per spike than state-of-the-art multi-mode neurons) in G45RFSOI technology. Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 2 |
| 2024 | A Compact 140nW/input Winner-Take-All Circuit for Spiking Neural NetworksabstractSolving classification problems using Spiking Neural Networks (SNNs) involves determining the most active neuron in the output layer. Scalable, low-power and low-area hardware solutions for such decision-making are vital for neuromorphic edge applications to meet power and space constraints. In this work, we propose a low-power, compact Winner-Take-All (WTA) circuit, a multi-input multi-output dynamic threshold comparator that simultaneously compares multiple analog voltage inputs and provides a one-hot-encoded digital output vector indicating the result of the classification. The design eliminates the need for cascading and a dedicated feedback circuit. A spike integrator stage captures the temporal activity of a set of neurons, and these activities are compared and digitized by the proposed WTA comparator stage. The proposed WTA designed in GF45RFSOI technology, exhibits self-excitation and global-inhibition properties, offers scalability, consumes 44% less power (140 nW ) and occupies a 40% lower area (166 μm2), compared to state-of-the-art. Gaurav R, Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 3 |
| 2024 | A sub-100 nW Power, Compact CTDSM with a Band-To-Band Tunnelling Loop FilterabstractThis work presents a continuous-time delta-sigma modulator (CTDSM) deploying an experimentally demonstrated band-to-band-tunelling (BTBT) SOI MOSFET-based loop filter. With a compact, low-pass filter circuit and extremely low current in the BTBT regime, a loop filter implementation will provide optimality in terms of area and power performance. In literature for moderate-resolution CTDSMs, traditional loop filters are implemented with either fully passive, active, or hybrid integrators. These designs have a tight tradeoff in terms of area and power. The passive integrators have optimal power but suboptimal area, while the active integrators have optimal area and sub-optimal power consumption. The proposed work tries to break this tradeoff using BTBT regime loop filters. The CTDSM was designed in a GF45RFSOI technology and achieves a peak SNR/SNDR of 48.41 dB/47.94 dB for a 5 kHz bandwidth. The power consumption is 76.3 nW, with an area of 102.7 μm2— more than 100x area reduction over previous state-of-the-art moderate-precision CTDSM designs. This makes the proposed CTDSM extremely compact and power-efficient compared to traditional state-of-the-art moderate-resolution DSMs. Atharva Raut, Abhishek Kadam, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
ISCAS | 3 |
| 2023 | Real-world Performance Estimation of Liquid State Machines for Spoken Digit ClassificationabstractLiquid State Machine (LSM) is a brain-inspired neural network architecture for solving temporal classification problems like speech recognition. The simple structure of LSM with a reservoir and single-layer classifier is attractive from a hardware implementation perspective. When the LSM is considered for low-power hardware implementation in real-world command word recognition tasks, challenges like nonidealities in sensor filter response and ambient noise become critical concerns. In this work, we evaluate the performance of LSM based on two aspects (1) ambient noise and (2) sensor/preprocessing circuit nonidealities. For Ambient noise, we use additive white gaussian noise (AWGN) and ambient noise using the iNoise Indian Noise dataset that covers various natural indoor, outdoor, and travel-related environmental sounds. To understand the impact of input hardware nonidealities, we analyzed the impact of the audio preprocessing filter's quality factor, order, center frequency variations, and output nonlinearity on LSM performance. We use the spoken digits classification in the TI-46 dataset. This paper's findings present design guidelines for the system designers intending to use liquid-state machines for speech classification tasks. In terms of filter design, first, there is a broad Q, order space for filter design where performance is high. We use the hardware-friendly parallel 4th order Butterworth bandpass filter model to provide a baseline 98% accuracy in speech classification tasks. Second, the performance of LSM degrades proportionally to the variation in the center frequency of the bandpass filters in the filter bank. Third, nonlinearity with the third-order harmonic of 50 dBc can be tolerated. Regarding ambient noise, our study shows that a 40 dB SNR for AWGN is sufficient for ideal performance. Second, the best case of “home” noise leads to a performance of 91.4%. Outdoor and travel noise reduce the classification performance to 78.8% and 62.4%, respectively. However, ideal performance is recovered if the signal to noise ratio (SNR) is increased, particularly by 10 dB in indoor conditions and 30 dB in outdoor conditions. Thus, our study presents an engineering evaluation for real-world spoken digit recognition using LSMs. Abhishek Kadam, Anmol Biswas, Vivek Saraswat, Ajay Kumar Singh, Laxmeesha Somappa, Maryam Shojaei Baghini, Udayan Ganguly |
IJCNN | 4 |
| 2022 | Quantum Tunneling Based Ultra-Compact and Energy Efficient Spiking Neuron Enables Hardware SNNabstractLow-power and low-area neurons are essential for hardware implementation of large-scale SNNs. Various novel-physics-based leaky-integrate-and-fire (LIF) neuron architectures have been proposed with low power and area, but are not compatible with CMOS technology to enable brain scale implementation of SNN. In this paper, for the first time, we demonstrate hardware implementation of recurrent SNN using proposed low-power, low-area, and low-leakage band-to-band-tunneling (BTBT) based neurons. A low-power thresholding circuit is proposed. We further propose a predistortion technique to linearize a nonlinear neuron without any area and power overhead. We establish the equivalence of the proposed neuron with the ideal LIF neuron to demonstrate its versatility. The tunneling regime enables a high input impedance in the BTBT neurons (few$\text{G}\Omega$) to enable a voltage input without loading the synaptic array. To verify the effect of the proposed neuron, a 36-neuron recurrent SNN is fabricated in GF-45nm PDSOI technology. We achieved 5000x lower energy-per-spike at a similar area and 10x lower standby power at a similar area and energy-per-spike. Such overall performance improvement enables brain scale computing. Ajay Kumar Singh, Vivek Saraswat, Maryam Shojaei Baghini, Udayan Ganguly |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Frequency Estimation for Resonant MEMS SensorsabstractResonant MEMS sensors have potential of providing highly precise measurements. The accurate processing of the output relies on precise frequency estimation techniques especially in the context of portable sensors for particulate matter. This paper investigates five commonly used single-tone frequency estimation techniques (3-point DFT interpolation, parabolic interpolation of periodogram peak, Prony's method, modified Pisarenko and zero crossing method) with respect to the estimation accuracy, memory requirement and computational complexity. The effect of noise and harmonics on estimation accuracy of these five techniques are analyzed and validated through simulation. The experimental data is acquired from a resonant MEMS sensor with a center frequency of 3.15 MHz. The output is sampled at 100MS/s using a 12-bit ADC. These five techniques are applied to the various data sets acquired from an experimental setup. The comparison results along with the analysis are presented. Ajay Kumar Singh, Laxmeesha Somappa, Malar Chellasivalingam, Ashwin A. Seshia, Maryam Shojaei Baghini |
ISCAS | 1 |
| 2004 | TCP-ADA: TCP with adaptive delayed acknowledgement for mobile ad hoc networksabstractEarlier studies have shown that TCP suffers from performance degradation over mobile ad hoc networks due to poor wireless channel characteristics and host mobility. This paper shows that generating acknowledgement for each data packet also deteriorates TCP throughput over these networks more significantly than in wired networks. In these networks, acknowledgement packets share the path with data packets. This creates the contention and collision between ACK and DATA packets, resulting in reduced TCP throughput. Decreasing the number of ACKs enhances the performance of TCP as it reduces the contention and collision with DATA packets. We analyse this enhancement mathematically and derive the relationship between throughput and number of DATA packets covered in one ACK. The study shows that maximum throughput is achieved when one ACK acknowledges full congestion window of packets. Based on this analysis, we devise and propose TCP with adaptive delayed acknowledgement (TCP-ADA), which tries to decrease the number of ACKs to one per congestion window, adapting the ACK generation time. Simulation analysis is performed to investigate the enhancement achievable in practical scenario and these results corroborate our mathematical study. Ajay Kumar Singh, Kishore Kankipati |
WCNC | 1 |