VLDB 2026 Research / reviewers in the wild / expert
Chetan Singh Thakur
dblp:149/4625
· DBLP profile ↗
31ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0002-1240-6214ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Margin Propagation Based XOR-SAT Solvers for Decoding of LDPC CodesabstractDecoding of Low-Density Parity Check (LDPC) codes can be viewed as a special case of XOR-SAT problems, for which low-computational complexity bit-flipping algorithms have been proposed in the literature. However, a performance gap exists between the bit-flipping LDPC decoding algorithms and the benchmark LDPC decoding algorithms, such as the Sum-Product Algorithm (SPA). In this paper, we propose an XOR-SAT solver using log-sum-exponential functions and demonstrate its advantages for LDPC decoding. This is then approximated using the Margin Propagation formulation to attain a low-complexity LDPC decoder. The proposed algorithm uses soft information to decide the bit-flips that maximize the number of parity check constraints satisfied over an optimization function. The proposed solver can achieve results that are within 0.1dB of the Sum-Product Algorithm for the same number of code iterations. It is also at least$10 \times $lower than other Gradient-Descent Bit Flipping decoding algorithms, which are also bit-flipping algorithms based on optimization functions. The approximation using the Margin Propagation formulation does not require any multipliers, resulting in significantly lower computational complexity than other soft-decision Bit-Flipping LDPC decoders. Ankita Nandi, Shantanu Chakrabartty, Chetan Singh Thakur |
IEEE Trans. Commun. | 3 |
| 2024 | Live Demonstration: Real-time audio and visual inference on the RAMAN TinyML acceleratorabstractThe setup includes a host PC, camera and microphone sensors, and a Pynq-Z2 FPGA board. The neuromorphic cochlear model and RAMAN accelerator for neural network inference are deployed on the FPGA. The ARM processor on the FPGA sends the image received, cochleagram and the classified outputs to the PC to be visualized. Adithya Krishna, Ashwin Rajesh, Hitesh Pavan Oleti, Anand Chauhan, Shankaranarayanan H, André van Schaik, Mahesh Mehendale, Chetan Singh Thakur |
ISCAS | 8 |
| 2024 | tinyRadar: LSTM-based Real-time Multi-target Human Activity Recognition for Edge ComputingabstractWireless Human Activity Recognition (HAR) has emerged as a vital technology with wide-ranging applications, including healthcare, aged care, and child monitoring. Radar-based HAR systems, grounded in electromagnetic principles, offer resilience to lighting variations and uphold user privacy by efficiently processing sparse point cloud data. These systems demonstrate robust performance even in obstructed environments. Nonetheless, existing radar-based HAR methods face a limitation in relying on fixed time windows for data classification. This approach may not be the most adaptable, especially when monitoring individuals of various ages, from children to the elderly, who perform activities at different speeds. This paper introduces "tinyRadar," a novel system that capitalizes on the capabilities of the Texas Instruments IWR6843 radar for sensing and the Raspberry Pi 4 for executing Long Short-Term Memory (LSTM) inference. tinyRadar is trained on activities of varying durations, enabling it to cater to different human activity speeds. Remarkably, tinyRadar achieves 93% real-time inference accuracy in recognizing eight distinct activity classes, classifying each activity frame within 10 ms, with a compact model size of 311 KB. Satyapreet Singh Yadav, Shreyansh Anand, Adithya M. D, Dasari Sai Nikitha, Chetan Singh Thakur |
ISCAS | 5 |
| 2024 | RAMAN: A Reconfigurable and Sparse tinyML Accelerator for Inference on EdgeabstractDeep Neural Network (DNN) based inference at the edge is challenging as these compute, and data-intensive algorithms need to be implemented at low cost and low power while meeting the latency constraints of the target applications. Sparsity, in both activations and weights inherent to DNNs, is a key knob to leverage. In this paper, we present RAMAN, a Re-configurable and spArse tinyML Accelerator for infereNce on edge, architected to exploit the sparsity to reduce area (storage), power as well as latency. RAMAN can be configured to support a wide range of DNN topologies -consisting of different convolution layer types and a range of layer parameters (feature-map size and the number of channels). RAMAN can also be configured to support accuracy vs. power/latency tradeoffs using techniques deployed at compile-time and run-time. We present the salient features of the architecture, provide implementation results and compare the same with the state-of-the-art. RAMAN employs novel dataflow inspired by Gustavson’s algorithm that has optimal input activation (IA) and output activation (OA) reuse to minimize memory access and the overall data movement cost. The dataflow allows RAMAN to locally reduce the partial sum (Psum) within a processing element array to eliminate the Psum writeback traffic. Additionally, we suggest a method to reduce peak activation memory by overlapping IA and OA on the same memory space, which can reduce storage requirements by up to 50%. RAMAN was implemented on a low-power and resource-constrained Efinix Ti60 FPGA with 37.2K LUTs and 8.6K register utilization. RAMAN processes all layers of the MobileNetV1 model at 98.47 GOp/s/W and the DS-CNN model at 79.68 GOp/s/W by leveraging both weight and activation sparsity. Adithya Krishna, Srikanth Rohit Nudurupati, Chandana D. G, Pritesh Dwivedi, André van Schaik, Mahesh Mehendale, Chetan Singh Thakur |
IEEE Internet Things J. | 7 |
| 2024 | Beyond supervision: An unsupervised spatio-temporal point cloud noise modeling for event vision sensor
Lakshmi Annamalai, Chetan Singh Thakur |
Pattern Recognit. Lett. | 2 |
| 2024 | ARYABHAT: A Digital-Like Field Programmable Analog Computing Array for Edge AIabstractRecent advances in margin-propagation (MP) based approximate computing have resulted in analog computing circuits that exhibit scaling properties similar to that of digital computing circuits. MP-based circuits allow trading off energy-efficiency with speed and precision, endow robustness to temperature variations, and make the design portable across different process nodes. In this work, We leverage these scaling properties to design ARYABHAT, a field-programmable analog machine learning processor that can be synthesized like digital field-programmable gate arrays (FPGAs). ARYABHAT features a fully reconfigurable tile-based modular analog architecture with adjustable throughput and configurable energy requirements, making it suitable for various machine-learning computations. The architecture can perform computations at variable accuracy and different power-performance specifications and can simultaneously leverage near-memory computing paradigms to improve computational throughput. We also present a complete programming and test ecosystem for ARYABHAT called ARYAFlow and ARYATest. As proof of concept, we showcase the implementation of machine learning algorithms at different performance specifications. Pratik Kumar, Ankita Nandi, Ayan Saha, Kurupati Sai Pruthvi Teja, Ratul Das, Shantanu Chakrabartty, Chetan Singh Thakur |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Process, Bias, and Temperature Scalable CMOS Analog Computing Circuits for Machine LearningabstractAnalog computing is attractive compared to digital computing due to its potential for achieving higher computational density and higher energy efficiency. However, unlike digital circuits, conventional analog computing circuits cannot be easily mapped across different process nodes due to differences in transistor biasing regimes, temperature variations and limited dynamic range. In this work, we generalize the previously reported margin-propagation-based analog computing framework for designing novel shape-based analog computing (S-AC) circuits that can be easily cross-mapped across different process nodes. Similar to digital designs S-AC designs can also be scaled for precision, speed, and power. As a proof-of-concept, we show several examples of S-AC circuits implementing mathematical functions that are commonly used in machine learning architectures. Using circuit simulations we demonstrate that the circuit input/output characteristics remain robust when mapped from a planar CMOS 180nm process to a FinFET 7nm process. Also, using benchmark datasets we demonstrate that the classification accuracy of a S-AC based neural network remains robust when mapped across the two processes and to changes in temperature. Pratik Kumar, Ankita Nandi, Shantanu Chakrabartty, Chetan Singh Thakur |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | An Always-On tinyML Acoustic Classifier for Ecological ApplicationsabstractLong-term monitoring and tracking of wildlife and endangered species in their natural environment is challenging due to human factors and logistical limitations. We present a light-weight, always-on acoustic classification system that can identify the density of specific wildlife species in an ecological environment where human presence may be undesirable. The system uses a template-based support-vector-machine (SVM) classifier that combines acoustic filtering and classification into an in-filter computing and a hardware-friendly platform. We demonstrate the system’s capabilities for identifying the density of different bird species using ARM Cortex-M4 based AudioMoth hardware. The embedded software, designed specifically for the AudioMoth hardware, can generate the programmable parameters, given limited training samples corresponding to different wildlife species. We show that the system can identify four different bird species with an accuracy of more than 95% and consumes a memory footprint of 14 KB SRAM and 149 KB Flash memory that can run for 48 days on battery without any human intervention. Hemanth Reddy Sabbella, Abhishek Ramdas Nair, V. Gumme, Satyapreet Singh Yadav, Shantanu Chakrabartty, Chetan Singh Thakur |
ISCAS | 6 |
| 2022 | tinyRadar: mmWave Radar based Human Activity Classification for Edge ComputingabstractThe rising need for elderly care, child care, and intrusion detection challenges the sustainability of traditional systems that depend on in-person monitoring and surveillance. The current state-of-the-art technology heavily relies on InfraRed (IR) and camera-based systems, which often require cloud computing. It can lead to higher latency, data theft, and privacy issues of being continuously monitored. This paper proposes a novel tiny-ML-based single-chip radar solution for on-edge sensing and detection of human activity. Edge computing within a small form factor solves the issue of data theft and privacy concerns as radar provides point cloud information. Also, it can operate in adverse environmental conditions like fog, dust, and low light. This work used the Texas Instruments IWR6843 millimeter wave (mmWave) radar board to implement signal processing and Convolutional Neural Network (CNN) for human activity classification. A dataset for four different human activities generalized over six subjects was collected to train the 8-bit quantized CNN model. The real-time inference engine implemented on Cortex®-R4F using CMSIS-NN framework has a model size of 1.44 KB, gives the classification result after every 120 ms, and has an overall subject-independent accuracy of 96.43%. Satyapreet Singh Yadav, Radha Agarwal, Kola Bharath, Sandeep Rao, Chetan Singh Thakur |
ISCAS | 5 |
| 2022 | In-Filter Computing for Designing Ultralight Acoustic Pattern RecognizersabstractWe present a novel in-filter computing framework that can be used for designing ultralight acoustic classifiers for use in the smart Internet of Things (IoT). Unlike a conventional acoustic pattern recognizer, where the feature extraction and classification are designed independently, the proposed architecture integrates the convolution and nonlinear filtering operations directly into the kernels of a support vector machine (SVM). The result of this integration is a template-based SVM whose memory and computational footprint (training and inference) is light enough to be implemented on a field-programmable gate array (FPGA)-based IoT platform. While the proposed in-filter computing framework is general enough, in this article, we demonstrate this concept using a cascade of an asymmetric resonator with inner hair cells (CAR-IHCs)-based acoustic feature extraction algorithm. The complete system has been optimized using time-multiplexing and parallel-pipeline techniques for a Xilinx Spartan 7 series FPGA. We show that the system can achieve robust classification performance on benchmark sound recognition tasks using only 1.5k lookup tables (LUTs) and 2.8k flip-flops (FFs), a significant improvement over other approaches. Abhishek Ramdas Nair, Shantanu Chakrabartty, Chetan Singh Thakur |
IEEE Internet Things J. | 3 |
| 2022 | Neuromorphic Time-Multiplexed Reservoir Computing With On-the-Fly Weight Generation for Edge DevicesabstractThe human brain has evolved to perform complex and computationally expensive cognitive tasks, such as audio-visual perception and object detection, with ease. For instance, the brain can recognize speech in different dialects and perform other cognitive tasks, such as attention, memory, and motor control, with just 20 W of power consumption. Taking inspiration from neural systems, we propose a low-power neuromorphic hardware architecture to perform classification on temporal data at the edge. The proposed architecture uses a neuromorphic cochlea model for feature extraction and reservoir computing (RC) framework as a classifier. In the proposed hardware architecture, the RC framework is modified for on-the-fly generation of reservoir connectivity, along with binary feedforward and reservoir weights. Also, a large reservoir is split into multiple small reservoirs for efficient use of hardware resources. These modifications reduce the computational and memory resources required, thereby resulting in a lower power budget. The proposed classifier is validated for speech and human activity recognition (HAR) tasks. We have prototyped our hardware architecture using Intel's cyclone-10 low-power series field-programmable gate array (FPGA), consuming only 4790 logic elements (LEs) and 34.9-kB memory, making it a perfect candidate for edge computing applications. Moreover, we have implemented a complete system for speech recognition with the feature extraction block (cochlea model) and the proposed classifier, utilizing 15 532 LEs and 38.4-kB memory. By using the proposed idea of multiple small reservoirs along with on-the-fly generation of reservoir binary weights, our architecture can reduce the power consumption and memory requirement by order of magnitude compared to existing FPGA models for speech recognition tasks with similar complexity. Sarthak Gupta, Satrajit Chakraborty, Chetan Singh Thakur |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Multiplierless MP-Kernel Machine for Energy-Efficient Edge DevicesabstractWe present a novel framework for designing multiplierless kernel machines that can be used on resource-constrained platforms such as intelligent edge devices. The framework uses a piecewise linear (PWL) approximation based on a margin propagation (MP) technique and uses only addition/subtraction, shift, comparison, and register underflow/overflow operations. We propose a hardware-friendly MP-based inference and online training algorithm that has been optimized for a field-programmable gate array (FPGA) platform. Our FPGA implementation eliminates the need for digital signal processor (DSP) units and reduces the number of Look-Up Tables (LUTs). By reusing the same hardware for inference and training, we show that the platform can overcome classification errors and local minima artifacts that result from MP approximation. The implementation of this proposed multiplierless MP-kernel machine on FPGA results in an estimated energy consumption of 13.4 pJ and power consumption of 107 mW with ~9 k LUTs and Flip Flops (FFs) each for a 256 $\times $ 32 sized kernel making it superior in terms of power, performance, and area compared with other comparable implementations. Abhishek Ramdas Nair, Pallab Kumar Nath, Shantanu Chakrabartty, Chetan Singh Thakur |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Real-Time Object Detection and Localization in Compressive Sensed VideoabstractTypically a 1-2MP CCTV camera generates around 7-12GB of data per day. Frame-by-frame processing of such an enormous amount of data requires hefty computational resources. In recent years, compressive sensing approaches have shown impressive compression results by reducing the sampling bandwidth. Different sampling mechanisms were developed to incorporate compressive sensing in image and video acquisition. Though all-CMOS [1], [2] sensor cameras that perform compressive sensing can help save a lot of bandwidth on sampling and minimize the memory required to store videos, the traditional signal processing, and deep learning models can realize operations only on the reconstructed data. To realize the original uncompressed domain, most reconstruction techniques are computationally expensive and time-consuming. To bridge this gap, we propose a novel task of detection and localization of objects directly on the compressed frames. Thereby mitigating the need to reconstruct the frames and reducing the search rate up to $20 \times$ (compression rate). We achieved an accuracy of 46.27% mAP with the proposed model on a GeForce GTX 1080 Ti. We were also able to show real-time inference on an NVIDIA TX2 embedded board with 45.11% mAP, thereby achieving the best balance between the accuracy, inference time, and memory constraints. Yeshwanth Bethi, Sathyaprakash Narayanan, Venkat Rangan, Anirban Chakraborty 0001, Chetan Singh Thakur |
ICIP | 5 |
| 2021 | Bayesian Source Localization Using Stochastic ComputationabstractBayesian models are challenging to implement on hardware with the conventional design methodologies due to their high computational complexity. Conventional digital architectures are designed for deterministic computation and are not optimal for implementing probabilistic algorithms on hardware. In this work, we propose an alternative method to implement the probabilistic algorithms such as Bayesian models on hardware using a stochastic computation (SC) framework. This framework leverages on the probabilistic nature of the Bayesian models and facilitates the implementation of complex probabilistic models using simple logic gates. From an application standpoint, we propose a novel Bayesian source localization model (BSLM) that estimates a source's position in a noisy environment by solving the Bayesian recursive equation implemented on Field Programmable Gate Array (FPGA) with low resource utilization. The proposed SC design framework will pave the way to build complex probabilistic algorithms for real-time edge computing applications. Adithya Krishna, Chetan Singh Thakur |
ISCAS | 2 |
| 2021 | Biomimetic FPGA-based spatial navigation model with grid cells and place cells
Adithya Krishna, Divyansh Mittal, Siri Garudanagiri Virupaksha, Abhishek Ramdas Nair, Rishikesh Narayanan, Chetan Singh Thakur |
Neural Networks | 6 |
| 2020 | Implementation of Bayesian Fly Tracking Model using Analog Neuromorphic CircuitsabstractThere is a growing body of evidence that suggests that the neurons in the brain calculate the posterior probability of states and events based on observations provided by the sensory neurons. Based on this hypothesis, a neuromorphic framework is proposed, where the sensory neurons of the dragonfly make noisy observations of the fruit fly and uses the underlying Hidden Markov Model (HMM) to track the fruit fly in two dimensional space. The dragonfly estimates the target position by solving the Bayesian recursive equations online. This work presents a novel approach for implementing probabilistic networks using sub-threshold analog neuromorphic circuits, with the ability to perform the computation in real-time. This framework will pave the way to build complex probabilistic algorithms based on HMMs for low power real-time applications. Alin Thomas Tharakan, Dheeraj Bhaskar, Chetan Singh Thakur |
ISCAS | 3 |
| 2020 | QUICKSAL: A small and sparse visual saliency model for efficient inference in resource constrained hardwareabstractVisual saliency is an important problem in the field of cognitive science and computer vision with applications such as surveillance, adaptive compressing, detecting unknown objects and scene understanding. In this paper, we propose a small and sparse neural network model for performing salient object segmentation that is suitable for use in mobile and embedded applications. Our model is built using depthwise separable convolutions and bottleneck inverted residuals which have been proven to perform very memory efficient inference and can be easily implemented using standard functions available in all deep learning frameworks. The multiscale features extracted along the layers with deep residuals allow our network to learn high quality saliency maps. We present the quantitative results of our QUICKSAL model with multiple levels of model sparsity ranging from 0% to ~96%, with the non-zero parameter count varying from ~3.3M to ~0.14M respectively - on publicly available benchmark datasets - showing that our highly constrained approach is comparable to other state-of-the-art approaches (parameter count ~35M). We also present qualitative results on camouflage images and show that our model can successfully distinguish between the salient and non-salient parts even when both seem blended together. Vignesh Ramanathan, Pritesh Dwivedi, Bharath Katabathuni, Anirban Chakraborty 0001, Chetan Singh Thakur |
WACV | 5 |
| 2020 | Neuromorphic Fringe Projection ProfilometryabstractWe address the problem of 3-D reconstruction using neuromorphic cameras (also known as event-driven cameras), which are a new class of vision-inspired imaging devices. Neuromorphic cameras are becoming increasingly popular for solving image processing and computer vision problems as they have significantly lower data rates than conventional frame-based cameras. We develop a neuromorphic-camera-based Fringe Projection Profilometry (FPP) system. We use the Dynamic Vision Sensor (DVS) in the DAVIS346 neuromorphic camera for acquiring measurements. Neuromorphic FPP is faster than a single-line-scanning method. Also, unlike frame-based FPP, the efficacy of the proposed method is not limited by the background while acquiring measurements. The working principle of the DVS also allows one to efficiently handle shadows thereby preventing ambiguities during 2-D phase unwrapping. Ashish Rao Mangalore, Chandra Sekhar Seelamantula, Chetan Singh Thakur |
IEEE Signal Process. Lett. | 3 |
| 2019 | SAMIR: Sparsity Amplified Iteratively-reweighted Beamforming for High-rsolution Ultrasound ImagingabstractIn ultrasound imaging, one typically employs delay-and-sum (DAS) beamformers for image reconstruction. An apodization window is used to suppress the side-lobes of an array beam pattern. The application of an apodization window to suppress the side-lobes widens the main-lobe width. We consider a statistical beamformer and present two variants. The signal of interest is modeled as a Laplacian-distributed random variable and additive interference components as Gaussian distributed. The resultant LASSO formulation is known to suffer from underestimation of large signal amplitudes due to the ℓ1-norm regularization. In the first variant, we reformulate the LASSO problem with a minimax-concave penalty (called Sparsity AMplified (SAM)) to contain the bias, thereby enhancing the beamformed image. A closed-form pointwise estimator is obtained for the optimization problem. In the second variant, we propose Sparsity AMplified Iteratively-Reweighted (SAMIR) beamforming algorithm, which leverages the properties of an apodization function. In SAMIR beamforming, we jointly optimize the cost over the signal of interest and the extrinsic apodization weights. This beamformer results in high-resolution ultrasound images, especially in the lateral direction. The proposed methods are compared with the standard DAS and a recently proposed statistically-modeled beamformer, iMAP, for a different number of plane-wave insonifications. Amol G. Mahurkar, Praveen Kumar Pokala, Chetan Singh Thakur, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2019 | Low Power Neuromorphic Analog System Based on Sub-Threshold Current Mode CircuitsabstractHardware implementation of brain-inspired algorithms such as reservoir computing, neural population coding and deep learning (DL) networks is useful for edge computing devices. The need for hardware implementation of neural network algorithms arises from the high resource utilization in form of processing and power requirements, making them difficult to integrate with edge devices. In this paper, we propose a non-spiking four quadrant current mode neuron model that has a generalized design to be used for population coding, echo-state networks (uses reservoir network), and DL networks. The model is implemented in analog domain with transistors in sub-threshold region for low power consumption and simulated using 180nm technology. The proposed neuron model is configurable and versatile in terms of non-linearity, which empowers the design of a system with different neurons having different activation functions. The neuron model is more robust in case of population coding and echo-state networks (ESNs) as we use random device mismatches to our advantage. The proposed model is current input and current output, hence, easily cascaded together to implement deep layers. The system was tested using the classic XOR gate classification problem, exercising 10 hidden neurons with population coding architecture. Further, derived activation functions of the proposed neuron model have been used to build a dynamical system, input controlled oscillator, using ESNs. Sarthak Gupta, Pratik Kumar, Satrajit Chakraborty, Chetan Singh Thakur |
ISCAS | 5 |
| 2019 | Live Demonstration: Real-Time Implementation of Proto-Object Based Visual Saliency ModelabstractWe will demonstrate a real-time implementation of a protoobject based neuromorphic visual saliency model [1] on an embedded processing board. Visual saliency models are difficult to implement in hardware for real-time applications due to their computational complexity. The conventional implementation is not optimal because of the requirement of a large number of convolution operations for filtering on several feature channels across multiple image pyramids. Our current implementation considers the dynamic temporal motion change by convoluting along time efficiently by parallelly processing them. We have implemented the model on an NVIDIA Jetson TX1 board (Fig. 1), which has NVIDIA Maxwell GPU with 256 NVIDIA CUDA Cores, hosted on an Ubuntu environment. The board has a 5 MP fixed focus MIPI CSI camera through which the frames are fetched using a Quad-core ARM Cortex-A57 MPCore Processor with 4 GB LPDDR4 Memory. The camera module fetches the frames to the application for processing, and the result is then displayed through the HDMI port. The application is written in Tensorflow, Cuda, and Python and uses several Python libraries. For further analysis, the user can also save the output onto a file. Sathyaprakash Narayanan, Yeshwanth Bethi, Jamal Molin, Ernst Niebur, Ralph Etienne-Cummings, Chetan Singh Thakur |
ISCAS | 6 |
| 2019 | N-HAR: A Neuromorphic Event-Based Human Activity Recognition System using Memory SurfacesabstractIn recent years, a new generation of low-power, neuromorphic, event-based vision sensors has been gaining popularity for their very low latency and data sparsity. Though the conventional frame-based cameras have advanced in a lot of ways, they suffer from data redundancy and temporal latency. The bio-inspired artificial retinas eliminate the data redundancy by capturing only the change in illumination at each pixel and asynchronously communicating in binary spikes. In this work, we propose a system to achieve the task of human activity recognition based on the event-based camera data. We show that such tasks, which generally need high frame rate sensors for accurate predictions, can be achieved by adapting existing computer vision techniques to the spiking domain. We used event memory surfaces to make the sparse event data compatible with deep convolutional neural networks (CNNs). We leverage upon the recent advances in deep convolutional networks based video analysis and adapt such frameworks onto the neuromorphic domain. We also provide the community with a new dataset consisting of five categories of human activities captured in real world without any simulations. We achieved an accuracy of 94.3% using event memory surfaces on our activity recognition dataset. Bibrat Ranjan Pradhan, Yeshwanth Bethi, Sathyaprakash Narayanan, Anirban Chakraborty 0001, Chetan Singh Thakur |
ISCAS | 5 |
| 2019 | Real-Time Image Segmentation using Neuromorphic Pixel ArrayabstractImage segmentation is a critical step in achieving computer vision. For fast and near real time image segmentation it is important to use dedicated hardware as it works much faster as compared to software based solutions. This paper presents a CMOS based real time image segmentation analog circuit using semi-supervised learning scheme. The proposed circuit ensures minimal operating power required when compared to its digital hardware implementation due to sub-threshold region operation of MOSFET and much lesser number of transistors being used. The circuit operation is based upon an anisotropic current spreading in non linear circuits. An analog pixel array is created which takes in weights and initial seed as input and segments the image based on them. Weights are generated from an image based on the feature similarity and are then fed to the pixel array. These weights dictate the anisotropic diffusion of current in the analog array and thus required segmentation is achieved. The proposed analog circuit based segmentation scheme is highly power efficient as compared to digital hardware implementation. Thus this scheme can be used in a power constrained environment as in case of Internet of Things (IoTs) network with independent battery-less nodes. Our approach will also be suitable for applications that require highperformance computing to run in real time, such as biomedical image segmentation for image-guided surgery. Raja Sharma, Sarthak Gupta, Pratik Kumar, Chetan Singh Thakur |
ISCAS | 5 |
| 2019 | Analog Neuromorphic System Based on Multi Input Floating Gate MOS Neuron ModelabstractThis paper introduces a novel implementation of the low-power analog artificial neural network (ANN) using Multiple Input Floating Gate MOS (MIFGMOS) transistor for machine learning applications. The number of inputs to a neuron in an ANN is the major bottleneck in building a large scale analog system. The proposed MIFGMOS transistor enables to build a large scale system by combining multiple inputs in a single transistor with a small silicon footprint. Here, we show the MIFGMOS based implementation of the Extreme Learning Machine (ELM) architecture using the receptive field approach with transistor operating in the sub-threshold region. The MIFGMOS produces output current as a function of the weighted combination of the voltage applied to its gate terminals. In the ELM architecture, the weights between the input and the hidden layer are random and this allows exploiting the random device mismatch due to the fabrication process, for building Integrated Circuits (IC) based on ELM architecture. Thus, we use implicit random weights present due to device mismatch, and there is no need to store the input weights. We have verified our architecture using circuit simulations on regression and various classification problems such as on the MNIST data-set and a few UCI data-sets. The proposed MIFGMOS enables combining multiple inputs in a single transistor and will thus pave the way to build large scale deep learning neural networks. Ankit Tripathi, Mehdi Arabizadeh, Sourabh Khandelwal, Chetan Singh Thakur |
ISCAS | 4 |
| 2017 | Low-power, low-mismatch, highly-dense array of VLSI Mihalas-Niebur neuronsabstractWe present an array of Mihalas-Niebur neurons with dynamically reconfigurable synapses implemented in 0.5 μm CMOS technology optimized for low-power, low-mismatch, and high-density. This neural array has two modes of operation: one is each cell in the array operates as independent leaky integrate-and-fire neurons, and the second is two cells work together to model the Mihalas-Niebur neuron dynamics. Depending on the mode of operation, this implementation consists of 2040 Mihalas-Niebur neurons or 4080 I&F neurons within a 3mm χ 3mm area. Each I&F neuron cell consumes an area of 1495μm2 and the neural array dissipates 360pJ of energy per synaptic event measured at 5.0V power supply (∼14pJ at 1.0V estimated from SPICE simulation). Jamal Molin, Adebayo Eisape, Chetan Singh Thakur, Vigil Varghese, Christian Brandli, Ralph Etienne-Cummings |
ISCAS | 3 |
| 2017 | Neuromorphic visual saliency implementation using stochastic computationabstractVisual saliency models are difficult to implement in hardware for real time applications due to their computational complexity. The conventional digital implementation is not optimal because of the requirement of a large number of convolution operations for filtering on several feature channels across multiple image pyramids [1], [2]. Here, we propose an alternative approach to implement a neuromorphic visual saliency algorithm [3] in digital hardware using stochastic computation, which can achieve very low power and small area. We show the real time implementation of important building blocks of the system and compare the overall system with its software implementation. Our implementation will be useful for facilitating high-fidelity selective rendering in computer graphics applications using the output of the saliency model, and for communications, where the non-salient parts of an image can be compressed more heavily than the salient parts. Our implementation will find several applications as a frontend co-processor for information triaging, compression and analysis in computer vision tasks. Our proposed SC-based convolution circuit could be a potential building block for implanting deep convolutional neural networks (CNN) on hardware. Chetan Singh Thakur, Jamal Molin, Jie Zhang 0063, Ernst Niebur, Ralph Etienne-Cummings |
ISCAS | 1 |
| 2017 | Live demonstration: A compact all-CMOS spatiotemporal compressed sensing video cameraabstractA compact all-CMOS spatiotemporal compressed sensing (CS) video camera is demonstrated. This CS-based framework [1], implemented on integrated circuits, is able to achieve 20-fold reduction in the readout speed and consumes only 14μW to provide 100 fps videos. Taking advantage of dictionary learning and sparse recovery, this prototype image sensor (127×90 pixels) can reconstruct 100 fps videos from the coded images sampled at 5 fps. Jie Zhang 0063, Chetan Singh Thakur, John M. Rattray, Sang (Peter) Chin, Trac D. Tran, Ralph Etienne-Cummings |
ISCAS | 3 |
| 2016 | A stochastic approach to STDPabstractWe present a digital implementation of the Spike Timing Dependent Plasticity (STDP) learning rule. The proposed digital implementation consists of an exponential decay (exp-decay) generator array and a STDP adaptor array. The weight values are stored in a digital memory, and the STDP adaptor w ill send these values to the exp-decay generator using a digital spike of which the duration is modulated according to these values. The exp-decay generator will then generate an exponential decay, which will be used by the STDP adaptor for performing the weight adaption. The exponential decay, which is computational expensive, is efficiently implemented by using a novel stochastic approach. This stochastic approach was fully analysed and characterised. We use a time multiplexing approach to achieve 8192 (8k) virtual STDP adaptors and exp-decay generators with only one physical adaptor and exp-decay generator respectively. We have validated our stochastic STDP approach with measurement results of a balanced excitation experiment. In that experiment, the competition (induced by STDP) between the synapses can establish a bimodal distribution of the synaptic weights: either towards zero (weak) or the maximum (strong) values. Our stochastic approach is therefore ideal for implementing the STDP learning rule in large-scale spiking neural networks running in real time. Runchun Wang, Chetan Singh Thakur, Tara J. Hamilton, Jonathan Tapson, André van Schaik |
ISCAS | 2 |
| 2015 | A neuromorphic hardware framework based on population codingabstractIn the biological nervous system, large neuronal populations work collaboratively to encode sensory stimuli. These neuronal populations are characterised by a diverse distribution of tuning curves, ensuring that the entire range of input stimuli is encoded. Based on these principles, we have designed a neuromorphic system called a Trainable Analogue Block (TAB), which encodes given input stimuli using a large population of neurons with a heterogeneous tuning curve profile. Heterogeneity of tuning curves is achieved using random device mismatches in VLSI (Very Large Scale Integration) process and by adding a systematic offset to each hidden neuron. Here, we present measurement results of a single test cell fabricated in a 65nm technology to verify the TAB framework. We have mimicked a large population of neurons by re-using measurement results from the test cell by varying offset. We thus demonstrate the learning capability of the system for various regression tasks. The TAB system may pave the way to improve the design of analogue circuits for commercial applications, by rendering circuits insensitive to random mismatch that arises due to the manufacturing process. Chetan Singh Thakur, Tara J. Hamilton, Runchun Wang, Jonathan Tapson, André van Schaik |
IJCNN | 1 |
| 2014 | FPGA implementation of the CAR Model of the cochleaabstractThe front end of the human auditory system, the cochlea, converts sound signals from the outside world into neural impulses transmitted along the auditory pathway for further processing. The cochlea senses and separates sound in a nonlinear active fashion, exhibiting remarkable sensitivity and frequency discrimination. Although several electronic models of the cochlea have been proposed and implemented, none of these are able to reproduce all the characteristics of the cochlea, including large dynamic range, large gain and sharp tuning at low sound levels, and low gain and broad tuning at intense sound levels. Here, we implement the `Cascade of Asymmetric Resonators' (CAR) model of the cochlea on an FPGA. CAR represents the basilar membrane filter in the `Cascade of Asymmetric Resonators with Fast-Acting Compression' (CAR-FAC) cochlear model. CAR-FAC is a neuromorphic model of hearing based on a pole-zero filter cascade model of auditory filtering. It uses simple nonlinear extensions of conventional digital filter stages that are well suited to FPGA implementations, so that we are able to implement up to 1224 cochlear sections on Virtex-6 FPGA to process sound data in real time. The FPGA implementation of the electronic cochlea described here may be used as a front-end sound analyser for various machine-hearing applications. Chetan Singh Thakur, Tara J. Hamilton, Jonathan Tapson, André van Schaik, Richard F. Lyon |
ISCAS | 1 |
| 2014 | Live demonstration: FPGA implementation of the CAR model of the cochleaabstractWe will demonstrate a 100-CAR-section cochlear model running in real time on an FPGA. Although our result suggests that an electronic cochlea with 1224 cochlear sections can be implemented on an average FPGA [1], the data rate limit of USB 2.0 does not permit us to implement more than 100 filter sections and display the output on a PC. Future work will explore alternatives to increase the bandwidth such as a PCI interface or USB 3.0 that will enable us to implement more filter sections. Nonetheless, our work demonstrates the capability of the CAR model to process sound in real-time. Chetan Singh Thakur, James Wright, Tara J. Hamilton, Jonathan Tapson, André van Schaik |
ISCAS | 1 |