Arindam Basu

dblp:95/2418 · DBLP profile ↗
← Back
79ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0003-1035-8770ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 57 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CADC: Crossbar-Aware Dendritic Convolution for Efficient In-memory Computing
abstract
Convolutional neural networks (CNNs) are computationally intensive and often accelerated using crossbar-based in-memory computing (IMC) architectures. However, large convolutional layers must be partitioned across multiple crossbars, generating numerous partial sums (psums) that require additional buffer, transfer, and accumulation, thus introducing significant system-level overhead. Inspired by dendritic computing principles from neuroscience, we propose crossbar-aware dendritic convolution (CADC), a novel approach that dramatically increases sparsity in psums by embedding a nonlinear dendritic function (zeroing negative values) directly within crossbar computations. Experimental results demonstrate that CADC significantly reduces psums, eliminating 80% in LeNet-5 on MNIST, 54% in ResNet-18 on CIFAR-10, 66% in VGG-16 on CIFAR-100, and up to 88% in spiking neural networks (SNN) on the DVS Gesture dataset. The induced sparsity from CADC provides two key benefits: (1) enabling zero-compression and zero-skipping, thus reducing buffer and transfer overhead by 29.3%, and accumulation overhead by 47.9%; (2) minimizing ADC quantization noise accumulation, resulting in small accuracy degradation-only 0.01% for LeNet-5, 0.1% for ResNet-18, 0.5% for VGG-16, and 0.9% for SNN. Compared to vanilla convolution (vConv), CADC exhibits accuracy changes ranging from +0.11 % to + 0. 1 9 % for LeNet-5, - 0. 0 4 % t - 0. 2 7 % for ResNet-18, +0. 9 9 % to + 1. 6 0 % for VGG-16, and - 0. 5 7 % to + 1. 3 2 % for SNN, across crossbar sizes from $64 \times 64$ to $256 \times 256$. Ultimately, a SRAM-based IMC implementation of CADC achieves 2.15 TOPS and 40.8 TOPS/W for ResNet-18 (4/2/4b), realizing a $11 \times-18 \times$ speedup and $1.9 \times-22.9 \times$ improvement in energy efficiency compared to existing IMC accelerators.
Ye Ke, Hongyang Shang, Arindam Basu
ASP-DAC5
2026 Near-Memory Architecture for Threshold-Ordinal Surface-Based Corner Detection of Event Cameras
abstract
Event-based Cameras (EBCs) are widely utilized in surveillance and autonomous driving applications due to their high speed and low power consumption. Corners are essential low-level features in event-driven computer vision, and novel algorithms utilizing event-based representations, such as Threshold-Ordinal Surface (TOS), have been developed for corner detection. However, the implementation of these algorithms on resource-constrained edge devices is hindered by significant latency, undermining the advantages of EBCs. To address this challenge, a near-memory architecture for efficient TOS updates (NM-TOS) is proposed. This architecture employs a read-write decoupled 8T SRAM cell and optimizes patch update speed through pipelining. Hardware-software co-optimized peripheral circuits and dynamic voltage and frequency scaling (DVFS) enable power and latency reductions. Compared to traditional digital implementations, our architecture reduces latency/energy by 24.7×/1.2× at Vdd=1.2 V or 1.93×/6.6× at Vdd=0.6 V based on 65nm CMOS process. Monte Carlo simulations confirm robust circuit operation, demonstrating zero bit error rate at operating voltages above 0.62 V, with only 0.2% at 0.61 V and 2.5% at 0.6 V. Corner detection evaluation using precision-recall area under curve (AUC) metrics reveals minor AUC reductions of 0.027 and 0.015 at 0.6 V for two popular EBC datasets.
Hongyang Shang, An Guo 0001, Ye Ke, Arindam Basu
DATE6
2026 Live Demonstration: An Event-Driven E-Skin System with Dynamic Binary Scanning and real time SNN Classification
Zhengnan Fu, Li Gaishan, Anubhab Tripathi, Arindam Basu
ISCAS4
2026 An Event-Driven E-Skin System with Dynamic Binary Scanning and real time SNN Classification
Li Gaishan, Zhengnan Fu, Anubhab Tripathi, Arindam Basu
ISCAS5
2026 SRAM-Based Compute-in-Memory Accelerator for Linear-decay Spiking Neural Networks
Hongyang Shang, Yahan Yang, Arindam Basu
ISCAS6
2026 3D Stack In-Sensor-Computing (3DS-ISC): Accelerating Time-Surface Construction for Neuromorphic Event Cameras
abstract
This work proposes a 3D Stack In-Sensor-Computing (3DS-ISC) architecture for efficient event-based vision processing. A real-time normalization method using an exponential decay function is introduced to construct the time-surface, reducing hardware usage while preserving temporal information. The circuit design utilizes the leakage characterization of Dynamic Random Access Memory(DRAM) for timestamp normalization. Custom interdigitated metal-oxide-metal capacitor (MOMCAP) is used to store the charge and low leakage switch (LL switch) is used to extend the effective charge storage time. The 3DS-ISC architecture integrates sensing, memory, and computation to overcome the memory wall problem, reducing power, latency, and reducing area by$69\times $,$2.2\times $and$1.9\times $, respectively, compared with its 2D counterpart. Moreover, compared to works using a 16-bit SRAM to store timestamps, the ISC analog array can reduce power consumption by three orders of magnitude. In real computer vision (CV) tasks, we applied the spatial-temporal correlation filter (STCF) for denoise, and 3D-ISC achieved almost equivalent accuracy compared to the digital implementation using high precision timestamps. As for the image classification, time-surface constructed by 3D-ISC is used as the input of GoogleNet, achieving 99% on N-MNIST, 85% on N-Caltech101, 78% on CIFAR10-DVS, and 97% on DVS128 Gesture, comparable with state-of-the-art results on each dataset. Additionally, the 3D-ISC method is also applied to image reconstruction using the DAVIS240C dataset, achieving the highest average SSIM (0.62) among three methods. This work establishes a foundation for real-time, resource-efficient event-based processing and points to future integration of advanced computational circuits for broader applications.
Hongyang Shang, Ye Ke, Arindam Basu
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 A Real-Time End-to-End Event-Based Tactile Sensing System
abstract
Event-based tactile sensing offers a promising alternative to frame-based approaches by reducing data redundancy, yet existing systems often lack end-to-end hardware support and remain power-inefficient. This work presents an energy-efficient event-based tactile sensing system that codesigns algorithms and hardware for real-time perception. At the front end, a leakage-compensated event-driven readout circuit integrates multiple tactile sensors into a single node, minimizing wiring and static power. At the algorithm level, a 2-D convolutional neural network (CNN) reconstructs event frames for handwritten digit recognition with strong robustness to varying interaction durations, while a graph neural network (GNN) processes irregular tactile layouts for object classification. A customized event-based tactile processor (ETP) supports end-to-end tactile processing, incorporating a dual-mode event frame builder (EFB) and a neural network processing unit (NPU) with an optimized data-reuse scheme. Evaluated using 28-nm CMOS, the ASIC performs handwritten digit recognition and object classification in 0.44 and 0.37 ms, respectively. The system consumes the total power of 5.2 mW with only 0.11-mW dynamic power, demonstrating an efficient solution for always-on tactile perception.
Yuncheng Lu, Kiho Seong, Si En Timothy Ng, Shibi Varku, Arindam Basu, Nripan Mathews, Tony Tae-Hyoung Kim
IEEE Trans. Very Large Scale Integr. Syst.5
2025 High Energy-efficiency and Low latency In-Memory Computing using Analog Accumulator and In-Memory ADC with shared References
abstract
This article proposes a $256 \times 128$ in-memory computing array using reconfigurable in-memory analog-to-digital conversion with shared references for high area efficiency (area overhead of 3% is $9 X$ better than traditional). A dual-8T SRAM bitcell is used to achieve read-write decoupling and store ternary weights. Read World Line Under Drive enabled Cascode helps to minimize current variations producing high linearity. Multi-bit input is handled with low latency and high energyefficiency by using bit-slicing (BS) with near-memory chargesharing based binary weighted accumulator (CHA). Using noise resilient training, we show software comparable performance for a MLP on MNIST, VGG-8 on CIFAR-10, and graph attention network on Cora, with respective accuracy reductions of only $0.1 \%, 0.8 \%$ and 0.5% due to non-idealities. The proposed macro demonstrates high energy/area efficiency (1146 TOPS/W, 27 TOPS $/ \mathbf{m m}^{\mathbf{2}}$ at $\mathbf{1} \boldsymbol{/} \mathbf{2} \boldsymbol{/} \mathbf{1 b}$) in $\mathbf{6 5 ~ n m}$ CMOS. It increases throughput (by $1.9 X$) and linearity (by $23 X$) compared to input pulse-width modulation by using BS and CHA. Compared to conventional BS with digital accumulation after ADC, this method has $1.7 X / 6.6 X$ better energy-efficiency/throughput by reducing ADC operations.
Zhengnan Fu, Hongyang Shang, Arindam Basu
DAC5
2025 Event-based Neural Spike Detection Using Spiking Neural Networks for Neuromorphic iBMI Systems
abstract
Implantable brain-machine interfaces (iBMIs) are evolving to record from thousands of neurons wirelessly but face challenges in data bandwidth, power consumption, and implant size. We propose a novel Spiking Neural Network Spike Detector (SNN-SPD) that processes event-based neural data generated via delta modulation and pulse count modulation, converting signals into sparse events. By leveraging the temporal dynamics and inherent sparsity of spiking neural networks, our method improves spike detection performance while maintaining low computational overhead suitable for implantable devices. Our experimental results demonstrate that the proposed SNN-SPD achieves an accuracy of 95.72% at high noise levels (standard deviation 0.2), which is about 2% higher than the existing Artificial Neural Network Spike Detector (ANN-SPD). Moreover, SNN-SPD requires only 0.41% of the computation and about 26.62% of the weight parameters compared to ANN-SPD, with zero multiplications. This approach balances efficiency and performance, enabling effective data compression and power savings for next-generation iBMIs.
Chanwook Hwang, Biyan Zhou, Ye Ke, Vivek Mohan, Jong Hwan Ko, Arindam Basu
ISCAS6
2025 Architectural Exploration of Hybrid Neural Decoders for Neuromorphic Implantable BMI
abstract
This work presents an efficient decoding pipeline for neuromorphic implantable brain-machine interfaces (Neu-iBMI), leveraging sparse neural event data from an event-based neural sensing scheme. We introduce a tunable event filter (EvFilter), which also functions as a spike detector (EvFilter-SPD), significantly reducing the number of events processed for decoding by 192× and 554×, respectively. The proposed pipeline achieves high decoding performance, up to R2= 0.73, with ANN- and SNN-based decoders, eliminating the need for signal recovery, spike detection, or sorting, commonly performed in conventional iBMI systems. The SNN-Decoder reduces computations and memory required by 5 − 23× compared to NN-, and LSTM-Decoders, while the ST-NN-Decoder delivers similar performance to an LSTM-Decoder requiring 2.5× fewer resources. This streamlined approach significantly reduces computational and memory demands, making it ideal for low-power, on-implant, or wearable iBMIs.
Vivek Mohan, Biyan Zhou, Zhou Wang 0005, Anil A. Bharath, Emmanuel M. Drakakis, Arindam Basu
ISCAS6
2025 SPACE: SPike-Aware Consistency Enhancement for Test-Time Adaptation in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs), as a biologically plausible alternative to Artificial Neural Networks (ANNs), have demonstrated advantages in terms of energy efficiency, temporal processing, and biological plausibility. However, SNNs are highly sensitive to distribution shifts, which can significantly degrade their performance in real-world scenarios. Traditional test-time adaptation (TTA) methods designed for ANNs often fail to address the unique computational dynamics of SNNs, such as sparsity and temporal spiking behavior. To address these challenges, we propose SPike-Aware Consistency Enhancement (SPACE), the first source-free and single-instance TTA method specifically designed for SNNs. SPACE leverages the inherent spike dynamics of SNNs to maximize the consistency of spike-behavior-based local feature maps across augmented versions of a single test sample, enabling robust adaptation without requiring source data. We evaluate SPACE on multiple datasets. Furthermore, SPACE exhibits robust generalization across diverse network architectures, consistently enhancing the performance of SNNs on CNNs, Transformer, and ConvLSTM architectures. Experimental results show that SPACE outperforms state-of-the-art ANN methods while maintaining lower computational cost, highlighting its effectiveness and robustness for SNNs in real-world settings. The code will be available at https://github.com/ethanxyluo/SPACE.
Xinyu Luo, Kecheng Chen, Pao-Sheng Sun, Chris Xing Tian, Arindam Basu, Haoliang Li
NeurIPS5
2025 A 22-nm 64-kB lightning-like hybrid computing-in-memory macro with a compressed adder tree and analog-storage quantizers for transformer and CNNs
An Guo 0001, Xi Chen 0107, Fangyuan Dong, Jinwu Chen, Zhihang Yuan, Xing Hu 0010, Guangyu Sun 0003, Arindam Basu, Jun Yang 0006, Xin Si
Sci. China Inf. Sci.9
2025 Topkima-Former: Low-Energy, Low-Latency Inference for Transformers Using Top-k In-Memory ADC
abstract
Transformer has emerged as a leading architecture in neural language processing (NLP) and computer vision (CV). However, the extensive use of nonlinear operations, like softmax, poses a performance bottleneck during transformer inference and comprises up to 40% of the total latency. Hence, we propose innovations at the circuit, algorithm and architecture levels to accelerate the transformer. At the circuit level, we propose Topkima—combining top-kactivation selection with in-memory ADC (IMA) to implement efficient softmax without any sorting overhead. Only theklargest activations are sent to softmax calculation block, reducing the huge computational cost of softmax. At the algorithmic level, a modified training scheme utilizes top-kactivations only during the forward pass, combined with a sub-top-kmethod to address the crossbar size limitation by aggregating each sub-top-kvalues as global top-k. At the architecture level, we introduce a fine pipeline for efficiently scheduling data flows and an improved scale-free technique for removing scaling cost. The combined system, dubbed Topkima-Former, enhances$1.8\times -84\times $speedup and$1.2\times -36\times $energy efficiency (EE) over prior In-memory computing (IMC) accelerators. Compared to a conventional softmax macro and a digital top-k(Dtopk) softmax macro, our proposed Topkima softmax macro achieves about$15\times $and$8\times $faster speed respectively. Experimental evaluations demonstrate minimal (0.42% to 1.60%) accuracy loss for different models in both vision and NLP tasks.
Xiaoqi Peng, Hongyang Shang, Ye Ke, Xiaofeng Yang 0004, Arindam Basu
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 HAST: A Hardware-Efficient Spatio-Temporal Correlation Near-Sensor Noise Filter for Dynamic Vision Sensors
abstract
The Dynamic Vision Sensor (DVS) is a bio-inspired image sensor which has many advantages such as high dynamic range, high bandwidth, high temporal resolution and low power consumption for Internet of Video Things and Edge Computing applications. However, spuriously generated Background Activity (BA) noise events can significantly degrade the quality of DVS output and cause unnecessary computations throughout the image processing chain, reducing its energy efficiency. Near-sensor filters can mitigate this problem by preventing the BA noise events from reaching downstream stages. In this paper, we propose a novel, hardware-efficient, spatio-temporal correlation filter (HAST) for near-sensor BA noise filtering. It uses compact two-dimensional binary arrays along with simple, arithmetic-free hash-based functions for storage and retrieval operations. This approach eliminates the need to use timestamps for determining the chronological order of events. HAST uses much lower memory and energy compared to other hardware-friendly filters (BAF/STCF) while matching their performance in simulations with standard datasets; for a sensor of resolution$346\times 260$pixels, it requires only 5–18% of their memory, and about 15% of their energy per event for correlation time$\tau $ranging from 1 to 50 ms. The memory and energy gains of the filter increase with sensor resolution. In FPGA implementation, HAST achieves about 29% higher throughput than BAF/STCF while utilizing only about 5% of their memory. The filter parameter values can be chosen by Design Space Exploration (DSE) for optimized performance-resource trade-offs based on application requirements.
Pradeep Kumar Gopalakrishnan, Chip-Hong Chang, Arindam Basu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 Older and Wiser: The Marriage of Device Aging and Intellectual Property Protection of DNNs
abstract
Deep neural networks (DNNs), such as the widely-used GPT-3 with billions of parameters, are often kept secret due to high training costs and privacy concerns surrounding the data used to train them. Previous approaches to securing DNNs typically require expensive circuit redesign, resulting in additional overheads such as increased area, energy consumption, and latency. To address these issues, we propose a novel hardware-software co-design approach for DNN intellectual property (IP) protection that capitalizes on the inherent aging characteristics of circuits and a novel differential orientation fine-tuning (DOFT) to ensure effective protection.
Ning Lin, Shaocong Wang 0001, Yue Zhang 0011, Yangu He, Kwunhang Wong, Arindam Basu, Dashan Shang, Xiaoming Chen 0003
DAC6
2024 Memory Efficient Corner Detection for Event-Driven Dynamic Vision Sensors
abstract
Event cameras offer low-latency and data compression for visual applications, through event-driven operation, that can be exploited for edge processing in tiny autonomous agents. Robust, accurate and low latency extraction of highly informative features such as corners is key for most visual processing. While several corner detection algorithms have been proposed, state-of-the-art performance is achieved by "luvHarris". However, this algorithm requires a high number of memory accesses per event, making it less-than ideal for low-latency, low-energy implementation in tiny edge processors. In this paper, we propose a new event-driven corner detection implementation tailored for edge computing devices, which requires much lower memory access than lu-vHarris while also improving accuracy. Our method trades computation for memory access, which is more expensive for large memories. For a DAVIS346 camera, our method requires ≈ 3.8X less memory, ≈ 36.6X less memory accesses with only ≈ 2.3X more computes.
Pao-Sheng Sun, Arren Glover, Chiara Bartolozzi, Arindam Basu
ICASSP4
2024 A Low-Power Spike Detector Using In-Memory Computing for Event-based Neural Frontend
abstract
With the sensor scaling of next-generation Brain-Machine Interface (BMI) systems, the massive A/D conversion and analog multiplexing at the neural frontend poses a challenge in terms of power and data rates for wireless and implantable BMIs. While previous works have reported the neuromorphic compression of neural signal, further compression requires integration of spike detectors on chip. In this work, we propose an efficient HRAM-based spike detector using In-memory computing for compressive event-based neural frontend. Our proposed method involves detecting spikes from event pulses without reconstructing the signal and uses a 10T hybrid in-memory computing bitcell for the accumulation and thresholding operations. We show that our method ensures a spike detection accuracy of 92-99% for neural signal inputs while consuming only 13.8 nW per channel in 65 nm CMOS.
Ye Ke, Arindam Basu
ISCAS2
2024 Hybrid Event-Frame Neural Spike Detector for Neuromorphic Implantable BMI
abstract
This work introduces two novel neural spike detection schemes intended for use in next-generation neuromorphic brain-machine interfaces (iBMIs). The first, an Event-based Spike Detector (Ev-SPD) which examines the temporal neighborhood of a neural event for spike detection, is designed for in-vivo processing and offers high sensitivity and decent accuracy (94-97%). The second, Neural Network-based Spike Detector (NN-SPD) which operates on hybrid temporal event frames, provides an off-implant solution using shallow neural networks with impressive detection accuracy (96-99%) and minimal false detections. These methods are evaluated using a synthetic dataset with varying noise levels and validated through comparison with ground truth data. The results highlight their potential in next-gen neuromorphic iBMI systems and emphasize the need to explore this direction further to understand their resource-efficient and high-performance capabilities for practical iBMI settings.
Vivek Mohan, Wee-Peng Tay, Arindam Basu
ISCAS3
2024 A Hybrid Neuromorphic Object Tracking and Classification Framework for Real-Time Systems
abstract
Deep learning inference that needs to largely take place on the "edge" is a highly computational and memory intensive workload, making it intractable for low-power, embedded platforms such as mobile nodes and remote security applications. To address this challenge, this article proposes a real-time, hybrid neuromorphic framework for object tracking and classification using event-based cameras that possess desirable properties such as low-power consumption (5-14 mW) and high dynamic range (120 dB). Nonetheless, unlike traditional approaches of using event-by-event processing, this work uses a mixed frame and event approach to get energy savings with high performance. Using a frame-based region proposal method based on the density of foreground events, a hardware-friendly object tracking scheme is implemented using the apparent object velocity while tackling occlusion scenarios. The frame-based object track input is converted back to spikes for TrueNorth (TN) classification via the energy-efficient deep network (EEDN) pipeline. Using originally collected datasets, we train the TN model on the hardware track outputs, instead of using ground truth object locations as commonly done, and demonstrate the ability of our system to handle practical surveillance scenarios. As an alternative tracker paradigm, we also propose a continuous-time tracker with C++ implementation where each event is processed individually, which better exploits the low latency and asynchronous nature of neuromorphic vision sensors. Subsequently, we extensively compare the proposed methodologies to state-of-the-art event-based and frame-based methods for object tracking and classification, and demonstrate the use case of our neuromorphic approach for real-time and embedded applications without sacrificing performance. Finally, we also showcase the efficacy of the proposed neuromorphic system to a standard RGB camera setup when simultaneously evaluated over several hours of traffic recordings.
Andrés Ussa, Chockalingam Senthil Rajen, Tarun Pulluri, Deepak Singla, Jyotibdha Acharya, Gideon Fu Chuanrong, Arindam Basu, Bharath Ramesh 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 Architectural Exploration of Neuromorphic Compression based Neural Sensing for Next-Gen Wireless implantable-BMI
abstract
This work explores the architectural trade-offs and implications of a neuromorphic compression based neural sensing architecture with address-event representation inspired readout protocol for massively parallel, next-gen wireless iBMI. We use quantitative metrics such as root-mean-square error and correlation coefficient between the original and recovered signal to assess the effect of neuromorphic compression on spike shape, and spike detection accuracy, sensitivity, and false detection rate to understand the effect of compression on downstream iBMI tasks. We demonstrate that a data compression ratio of$> 50$can be achieved by selective transmission of event pulses generated in different modes for large electrode arrays with a correlation coefficient of$\approx 0.9$and a spike detection accuracy of over 90%.
Vivek Mohan, Wee-Peng Tay, Arindam Basu
ISCAS3
2023 Reconfigurable Leakage-based Weak PUF in 65nm CMOS with 0.63% instability
abstract
Reliability of hardware security devices is of paramount importance when deployed in a System-On-Chip along-with the IoT sensor nodes. A bit flip could cause the node to be unrecognizable by the trusted source and increase costs due to replacement and re-deployment. To tackle these problems we fabricate a weak PUF in 65nm CMOS which can be used for chip ID applications. The design is based on “OFF” devices which consume only leakage current, and more importantly, provide large drain current mismatch and therefore output voltage dispersion. This large dispersion results in lower instability and Bit Error Rate (BER), making the design ideal for integration with IoT devices. To this end we measure instability of 2.81% over 2000 evaluations, and native BER of 0.48%. Native BER is below 2.5% over both temperature (−40-80° C) and voltage range (0.8-1.2V). Additionally, reconfiguration aids in reducing the instability/BER to 0.63%/0.047%, implying excellent reproducibility of the key. High throughput of 9.6 Gb/s is measured by implementing SRAM-style array for parallel read-out. Core energy/bit is measured to be 5.99 fJ/bit.
Nimesh Shah, Arindam Basu
ISCAS2
2022 EBBINNOT: A Hardware-Efficient Hybrid Event-Frame Tracker for Stationary Dynamic Vision Sensors
abstract
As an alternative sensing paradigm, dynamic vision sensors (DVSs) have been recently explored to tackle scenarios where conventional sensors result in high data rate and processing time. This article presents a hybrid event-frame approach for detecting and tracking objects recorded by a stationary neuromorphic sensor, thereby exploiting the sparse DVS output in a low-power setting for traffic monitoring. Specifically, we propose a hardware-efficient processing pipeline that optimizes memory and computational needs that enable long-term battery-powered usage for Internet of Things applications. To exploit the background removal property of a static DVS, we propose an event-based binary image creation that signals presence or absence of events in a frame duration. This reduces memory requirement and enables the usage of simple algorithms like median filtering and connected component labeling for denoise and region proposal (RP), respectively. To overcome the fragmentation issue, a YOLO-inspired neural network-based detector and classifier to merge fragmented RPs has been proposed. Finally, a new overlap-based tracker was implemented, exploiting overlap between detections and tracks is proposed with heuristics to overcome occlusion. The proposed pipeline is evaluated with more than 5 h of traffic recording spanning three different locations on two different neuromorphic sensors (DVS and CeleX) and demonstrates similar performance. Compared to existing event-based feature trackers, our method provides similar accuracy while needing$\approx {6}\times $less computes. To the best of our knowledge, this is the first time a stationary DVS-based traffic monitoring solution is extensively compared to simultaneously recorded RGB frame-based methods while showing tremendous promise by outperforming state-of-the-art deep learning solutions. The traffic data set is publicly made available at:https://nusneuromorphic.github.io/dataset/index.html.
Vivek Mohan, Deepak Singla, Tarun Pulluri, Andrés Ussa, Pradeep Kumar Gopalakrishnan, Pao-Sheng Sun, Bharath Ramesh 0001, Arindam Basu
IEEE Internet Things J.8
2021 DeepFreeze: Cold Boot Attacks and High Fidelity Model Recovery on Commercial EdgeML Device
abstract
EdgeML accelerators like Intel Neural Compute Stick 2 (NCS) can enable efficient edge-based inference with complex pre-trained models. The models are loaded in the host (like Raspberry Pi) and then transferred to NCS for inference. In this paper, we demonstrate practical and low-cost cold boot based model recovery attacks on NCS to recover the model architecture and weights, loaded from the Raspberry Pi. The architecture is recovered with 100% success and weights with an error rate of 0.04%. The recovered model reports maximum accuracy loss of 0.5% as compared to original model and allows high fidelity transfer of adversarial examples. We further extend our study to other cold boot attack setups reported in the literature with higher error rates leading to accuracy loss as high as 70%. We then propose a methodology based on knowledge distillation to correct the erroneous weights in recovered model, even without access to original training data. The proposed attack remains unaffected by the model encryption features of the OpenVINO and NCS framework.
Yoo-Seung Won, Dirmanto Jap, Arindam Basu, Shivam Bhasin
ICCAD4
2021 Power-efficient Spike Sorting Scheme Using Analog Spiking Neural Network Classifier
abstract
The method to map the neural signals to the neuron from which it originates is spike sorting. A low-power spike sorting system is presented for a neural implant device. The spike sorter constitutes a two-step trainer module that is shared by the signal acquisition channel associated with multiple electrodes. A low-power Spiking Neural Network (SNN) module is responsible for assigning the spike class. The two-step shared supervised on-chip training module is presented for improved training accuracy for the SNN. Post implant, the relatively power-hungry training module can be activated conditionally based on a statistics-driven retraining algorithm that allows on the fly training and adaptation. A low-power analog implementation for the SNN classifier is proposed based on resistive crossbar memory exploiting its approximate computing nature. Owing to the direct mapping of SNN functionality using physical characteristics of devices, the analog mode implementation can achieve ∼21 × lower power than its fully digital counterpart. We also incorporate the effect of device variation in the training process to suppress the impact of inevitable inaccuracies in such resistive crossbar devices on the classification accuracy. A variation-aware, digitally calibrated analog front-end is also presented, which consumes less than ∼50 nW power and interfaces with the digital training module as well as the analog SNN spike sorting module. Hence, the proposed scheme is a low-power, variation-tolerant, adaptive, digitally trained, all-analog spike sorter device, applicable to implantable and wearable multichannel brain-machine interfaces.
Anand Kumar Mukhopadhyay, Indrajit Chakrabarti, Arindam Basu, Mrigank Sharad
ACM J. Emerg. Technol. Comput. Syst.4
2021 A 0.11-0.38 pJ/cycle Differential Ring Oscillator in 65 nm CMOS for Robust Neurocomputing
abstract
This paper presents a low-area and low-power consumption CMOS differential current controlled oscillator (CCO) for neuromorphic applications. The oscillation frequency is improved over the conventional one by reducing the number of MOS transistors thus lowering the load capacitor in each stage. The analysis shows that for the same power consumption, the oscillation frequency can be increased about 11% compared with the conventional one without degrading the phase noise. Alternatively, the power consumption can be reduced 15% at the same frequency. The prototype structures are fabricated in a standard 65 nm CMOS technology and measurements demonstrate that the proposed CCO operates from 0.7 - 1.2 V supply with maximum frequencies of 80 MHz and energy/cycle ranging from 0.11 - 0.38 pJ over the tuning range. Further, system level simulations show that the nonlinearity in current-frequency conversion by the CCO does not affect its use as a neuron in a Deep Neural Network if accounted for during training.
Xueyong Zhang, Jyotibdha Acharya, Arindam Basu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2020 A 75kb SRAM in 65nm CMOS for In-Memory Computing Based Neuromorphic Image Denoising
abstract
This paper presents an in-memory computing (IMC) architecture for image denoising. The proposed SRAM based inmemory processing framework works in tandem with approximate computing on a binary image generated from neuromorphic vision sensors. Implemented in TSMC 65nm process, the proposed architecture enables ≈ 2000X energy savings ( ≈222X from IMC) compared to a digital implementation when tested with the video recordings from a DAVIS sensor and achieves a peak throughput of 1.25 - 1.66 frames/μs.
Sumon Kumar Bose, Vivek Mohan, Arindam Basu
ISCAS3
2020 Reducing Temperature Induced Unreliability in Sub-Threshold Strong PUFs through Circuit Modeling
abstract
Due to increased adoption of IoT devices, Physically Unclonable Function (PUF) circuits are essential as a lightweight security primitive. However existing PUFs suffer from imperfect reliability due to environmental variation. In this paper, we propose a PUF-server paradigm in which the server hosts a temperature model of the PUF, and the PUF informs the server of its response, and temperature through an integrated sensor. The model can predict responses for two different strong PUFs across -45°C to 90°C and thus avoids losing challenge-response pairs(CRPs) owing to low mismatch. Simulations show the model-based reliability improves to 97.9% - 99.5% compared with conventional reliability of 83.9% - 86.3%. Conversely, a loss of 25% - 40% of the CRPs is observed in order to improve the conventional reliability to match the model-based reliability. These results are also shown to be robust against quantization errors of a temperature sensor.
Nimesh Shah, Sumon Kumar Bose, Chip-Hong Chang, Arindam Basu
ISCAS4
2020 Towards Autonomous Intra-Cortical Brain Machine Interfaces: Applying Bandit Algorithms for Online Reinforcement Learning
abstract
This paper presents application of Banditron - an online reinforcement learning algorithm (RL) in a discrete state intra-cortical Brain Machine Interface (iBMI) setting. We have analyzed two datasets from non-human primates (NHPs) - NHP A and NHP B each performing a 4-option discrete control task over a total of 8 days. Results show average improvements of 15%, 6% in NHP A and 15%, 21% in NHP B over state of the art algorithms - Hebbian Reinforcement Learning (HRL) and Attention Gated Reinforcement Learning (AGREL) respectively. Apart from yielding a superior decoding performance, Banditron is also the most computationally friendly as it requires two orders of magnitude less multiply-and-accumulate operations than HRL and AGREL. Furthermore, Banditron provides average improvements of at least 40%, 15% in NHPs A, B respectively compared to popularly employed supervised methods - LDA, SVM across test days. These results pave the way towards an alternate paradigm of temporally robust hardware friendly reinforcement learning based iBMIs.
Shoeb Shaikh, Rosa Q. So, Tafadzwa Sibindi, Camilo Libedinsky, Arindam Basu
ISCAS5
2020 HyNNA: Improved Performance for Neuromorphic Vision Sensor Based Surveillance using Hybrid Neural Network Architecture
abstract
Applications in the Internet of Video Things (IoVT) domain have very tight constraints with respect to power and area. While neuromorphic vision sensors (NVS) may offer advantages over traditional imagers in this domain, the existing NVS systems either do not meet the power constraints or have not demonstrated end-to-end system performance. To address this, we improve on a recently proposed hybrid event-frame approach by using morphological image processing algorithms for region proposal and address the low-power requirement for object detection and classification by exploring various convolutional neural network (CNN) architectures. Specifically, we compare the results obtained from our object detection framework against the state-of-the-art low-power NVS surveillance system and show an improved accuracy of 82.16% from 63.1%. Moreover, we show that using multiple bits does not improve accuracy, and thus, system designers can save power and area by using only single bit event polarity information. In addition, we explore the CNN architecture space for object classification and show useful insights to trade-off accuracy for lower power using lesser memory and arithmetic operations.
Deepak Singla, Lavanya Ramapantulu, Andrés Ussa, Bharath Ramesh 0001, Arindam Basu
ISCAS6
2020 ADIC: Anomaly Detection Integrated Circuit in 65-nm CMOS Utilizing Approximate Computing
abstract
In this article, we present a low-power (LP) anomaly detection integrated circuit (ADIC) based on a one-class classifier (OCC) neural network. The ADIC achieves LP operation through a combination of: 1) careful choice of algorithm for online learning and 2) approximate computing techniques to lower average energy. In particular, online pseudoinverse update method (OPIUM) is used to train a randomized neural network for quick and resource-efficient learning. An additional 42% energy saving can be achieved when a lighter version of OPIUM method is used for training with the same number of data samples lead to no significant compromise on the quality of inference. Instead of a single classifier with large number of neurons, an ensemble of K base learner (BL) approach is chosen to reduce learning memory by a factor of K. This also enables approximate computing by dynamically varying the neural network size based on anomaly detection. Fabricated in 65-nm CMOS, the ADIC has K = 7 BLs with 32 neurons in each BL and dissipates 11.87 and 3.35 pJ/OP during learning and inference, respectively, at Vdd= 0.75 V when all seven BLs are enabled. Furthermore, evaluated on the NASA bearing data set, approximately 80% of the chip can be shut down for 99% of the lifetime leading to an energy efficiency of 0.48 pJ/OP, an 18.5× reduction over full-precision computing running at Vdd= 1.2 V throughout the lifetime.
Bapi Kar, Pradeep Kumar Gopalakrishnan, Sumon Kumar Bose, Mohendra Roy, Arindam Basu
IEEE Trans. Very Large Scale Integr. Syst.5
2019 ADEPOS: anomaly detection based power saving for predictive maintenance using edge computing
abstract
In Industry 4.0, predictive maintenance (PdM) is one of the most important applications pertaining to the Internet of Things (IoT). Machine learning is used to predict the possible failure of a machine before the actual event occurs. However, main challenges in PdM are: (a) lack of enough data from failing machines, and (b) paucity of power and bandwidth to transmit sensor data to cloud throughout the lifetime of the machine. Alternatively, edge computing approaches reduce data transmission and consume low energy. In this paper, we propose Anomaly Detection based Power Saving (ADEPOS) scheme using approximate computing through the lifetime of the machine. In the beginning of the machine's life, low accuracy computations are used when machine is healthy. However, on detection of anomalies as time progresses, system is switched to higher accuracy modes. We show using the NASA bearing dataset that using ADEPOS, we need 8.8X less neurons on average and based on post-layout results, the resultant energy savings are 6.4--6.65X.
Sumon Kumar Bose, Bapi Kar, Mohendra Roy, Pradeep Kumar Gopalakrishnan, Arindam Basu
ASP-DAC5
2019 A 0.16pJ/bit recurrent neural network based PUF for enhanced machine learning attack resistance
abstract
Physically Unclonable Function (PUF) circuits are finding wide-spread use due to increasing adoption of IoT devices. However, the existing strong PUFs such as Arbiter PUFs (APUF) and its compositions are susceptible to machine learning (ML) attacks because the challenge-response pairs have a linear relationship. In this paper, we present a Recurrent-Neural-Network PUF (RNN-PUF) which uses a combination of feedback and XOR function to significantly improve resistance to ML attack, without significant reduction in the reliability. ML attack is also partly reduced by using a shared comparator with offset-cancellation to remove bias and save power. From simulation results, we obtain ML attack accuracy of 62% for different ML algorithms, while reliability stays above 93%. This represents a 33.5% improvement in our Figure-of-Merit. Power consumption is estimated to be 12.3μW with energy/bit of ≈ 0.16pJ.
Nimesh Shah, Manaar Alam, Durga Prasad Sahoo, Debdeep Mukhopadhyay, Arindam Basu
ASP-DAC5
2019 A Current Mirror Cross Bar Based 2.86-TOPS/W Machine Learner and PUF with <2.5% BER in 65nm CMOS for IoT Application
abstract
Energy-efficient machine-learning and physical unclonable function (PUF) becomes popular in Internet-of-Things (IoT) applications for saliency detection and privacy protection at sensor node. A machine-learning and PUF engine for IoT applications is presented in this work with a current mirror cross-bar (CMCB) array being a shared core, reducing silicon area. A novel dimension expansion technique is proposed to increase weight matrix dimension beyond the physically implemented array with small hardware and energy overhead. A signed multiply-and-accumulation is realized in CMCB with differential current path and 2-phase conversion. The proposed engine achieves an error rate of 6.34% on MNIST digit recognition task with an energy efficiency of 2.86 TOPS/W. The PUF achieves a native BER of 2.3% across corners and extremely low area/CRP of 4.17×10-59μm2/CRP.
Yi Chen 0012, Zheng Wang 0027, Aakash Patil, Arindam Basu
ISCAS4
2019 Spiking Neural Network Based Region Proposal Networks for Neuromorphic Vision Sensors
abstract
This paper presents a three layer spiking neural network based region proposal network operating on data generated by neuromorphic vision sensors. The proposed architecture consists of refractory, convolution and clustering layers designed with bio-realistic leaky integrate and fire (LIF) neurons and synapses. The proposed algorithm is tested on traffic scene recordings from a DAVIS sensor setup. The performance of the region proposal network has been compared with event based mean shift algorithm and is found to be far superior (≈ 50% better) in recall for similar precision (≈ 85%). Computational and memory complexity of the proposed method are also shown to be similar to that of event based mean shift.
Jyotibdha Acharya, Vandana Padala, Arindam Basu
ISCAS3
2019 Live Demonstration: Autoencoder-Based Predictive Maintenance for IoT
abstract
This live demo aims to show the performance of a two-layer neural network applied to predictive maintenance. The first layer encodes features based on prior knowledge, while the second layer is trained online to detect anomalies. The system is implemented on an FPGA, acquiring real-time data from sensors attached to a motor. Faults can be triggered artificially in real-time to demonstrate anomaly detection.
Pradeep Kumar Gopalakrishnan, Bapi Kar, Sumon Kumar Bose, Mohendra Roy, Arindam Basu
ISCAS5
2018 Domain Wall Motion-based XOR-like Activation Unit With A Programmable Threshold
abstract
Spintronic devices promise an excellent opportunity for implementing ultra-low power neuromorphic platforms due to the inherent correspondence between their physical characteristics and the required neuronal, synaptic functionalities. Neuromorphic circuits using domain wall motion-based threshold neurons have been demonstrated in previous studies. However, threshold neurons are unable to realize linearly inseparable functions. Our work addresses this challenge by proposing a new domain wall motion-based neural activation unit with XOR-like activation function. We also develop a new learning algorithm for neurons with this activation unit. Offline training is performed on real-world datasets from the UCI machine learning repository. Neuromorphic circuits corresponding to these datasets are also simulated. The results suggest femto-Joule range energy consumption of a neuron with the proposed activation unit and 1.08×-1.82× lower misclassification rate (MCR) of the proposed algorithm in comparison to the traditional perceptron learning algorithm.
Tarun Vatwani, Anupam Chattopadhyay, Arindam Basu, Xuanyao Fong
IJCNN4
2018 Spiking Neural Classifier with Lumped Dendritic Nonlinearity and Binary Synapses: A Current Mode VLSI Implementation and Analysis
abstract
We present a neuromorphic current mode implementation of a spiking neural classifier with lumped square law dendritic nonlinearity. It has been shown previously in software simulations that such a system with binary synapses can be trained with structural plasticity algorithms to achieve comparable classification accuracy with fewer synaptic resources than conventional algorithms. We show that even in real analog systems with manufacturing imperfections (CV of 23.5% and 14.4% for dendritic branch gains and leaks respectively), this network is able to produce comparable results with fewer synaptic resources. The chip fabricated in [Formula: see text]m complementary metal oxide semiconductor has eight dendrites per cell and uses two opposing cells per class to cancel common-mode inputs. The chip can operate down to a [Formula: see text] V and dissipates 19 nW of static power per neuronal cell and [Formula: see text] 125 pJ/spike. For two-class classification problems of high-dimensional rate encoded binary patterns, the hardware achieves comparable performance as software implementation of the same with only about a 0.5% reduction in accuracy. On two UCI data sets, the IC integrated circuit has classification accuracy comparable to standard machine learners like support vector machines and extreme learning machines while using two to five times binary synapses. We also show that the system can operate on mean rate encoded spike patterns, as well as short bursts of spikes. To the best of our knowledge, this is the first attempt in hardware to perform classification exploiting dendritic properties and binary synapses.
Aritra Bhaduri, Amitava Banerjee, Subhrajit Roy, Sougata Kumar Kar, Arindam Basu
Neural Comput.5
2017 Unsupervised learning of event-based image recordings using spike-timing-dependent plasticity
abstract
The remarkable versatility and efficiency of the brain makes it important to understand its principles in order to push the boundaries of modern computing. Many models in neuroscience effectively model the detailed biological properties of the brain - however, they do not exhibit good performance on any benchmark task. Other systems have good performance, but do not work in the manner that the brain does, thus they do not give any insight into the mechanisms of the brain. A recent model by Diehl and Cook [1] is a neuromorphic system modelled by spiking neural networks (SNN) and spike-timing-dependant plasticity (STDP) that is able to achieve very good results on the MNIST dataset, a benchmark dataset of handwritten numbers for classification algorithms. However, in this system the manner in which MNIST dataset is converted into spikes does not mimic the manner in which neuromorphic sensors record information. Therefore, in this paper, we test the method proposed by Diehl and Cook on the N-MNIST (neuromorphic MNIST) where the MNIST images were recorded using a silicon retina, a DVS sensor that was moved similar to the saccade-like movements of the retina. The output of the camera consists of events at pixels similar to the retinal spikes. In this paper we examined how the original system had to be changed to accommodate the new dataset and noted that with modified parameters that respond to characteristics of the new dataset the system had better accuracy. The 400 neuron system has an accuracy of 93.68% for a smaller 3-class training dataset, and entire 10-class N-MNIST dataset on an 800 neuron network has 80.63% accuracy. We noted that output neurons form clusters that respond to a particular class, making it suitable to stack the system in a hierarchical manner.
Laxmi R. Iyer, Arindam Basu
IJCNN2
2017 Current mirror array: A novel lightweight strong PUF topology with enhanced reliability
abstract
In this work, we present a novel design of physical unclonable function (PUF) based on the topology of current mirror array (CMA). The proposed strong PUF exploits the randomness in the current mirror transistors and generates the response bit by comparing the accumulated current values. Thresholding and reference current are used to increase the reliability of the proposed PUF without requiring additional hardware resources. The proposed PUF structure is also analysed in terms of its difficulty of model building for measurement-prediction attack. Measurement results on 0.35μm test chips demonstrate that the proposed PUF outperforms other state-of-the-art designs with smaller area/bit of 9 × 10−36μm2 and lower native bit error rate (BER) of 0.16%.
Zheng Wang 0027, Yi Chen 0012, Aakash Patil, Chip-Hong Chang, Arindam Basu
ISCAS5
2017 Hardware architecture for large parallel array of Random Feature Extractors applied to image recognition
Aakash Patil, Shanlan Shen, Enyi Yao, Arindam Basu
Neurocomputing4
2017 Triplet Spike Time-Dependent Plasticity in a Floating-Gate Synapse
abstract
Synapse plays an important role in learning in a neural network; the learning rules that modify the synaptic strength based on the timing difference between the pre- and postsynaptic spike occurrence are termed spike time-dependent plasticity (STDP) rules. The most commonly used rule posits weight change based on time difference between one presynaptic spike and one postsynaptic spike and is hence termed doublet STDP (D-STDP). However, D-STDP could not reproduce results of many biological experiments; a triplet STDP (T-STDP) that considers triplets of spikes as the fundamental unit has been proposed recently to explain these observations. This paper describes the compact implementation of a synapse using a single floating-gate (FG) transistor that can store a weight in a nonvolatile manner and demonstrates the T-STDP learning rule by modifying drain voltages according to triplets of spikes. We describe a mathematical procedure to obtain control voltages for the FG device for T-STDP and also show measurement results from an FG synapse fabricated in TSMC 0.35-μm CMOS process to support the theory. Possible very large scale integration implementation of drain voltage waveform generator circuits is also presented with the simulation results.
Roshan Gopalakrishnan, Arindam Basu
IEEE Trans. Neural Networks Learn. Syst.2
2017 An Online Unsupervised Structural Plasticity Algorithm for Spiking Neural Networks
abstract
In this paper, we propose a novel winner-take-all (WTA) architecture employing neurons with nonlinear dendrites and an online unsupervised structural plasticity rule for training it. Furthermore, to aid hardware implementations, our network employs only binary synapses. The proposed learning rule is inspired by spike-timing-dependent plasticity but differs for each dendrite based on its activation level. It trains the WTA network through formation and elimination of connections between inputs and synapses. To demonstrate the performance of the proposed network and learning rule, we employ it to solve two-class, four-class, and six-class classification of random Poisson spike time inputs. The results indicate that by proper tuning of the inhibitory time constant of the WTA, a tradeoff between specificity and sensitivity of the network can be achieved. We use the inhibitory time constant to set the number of subpatterns per pattern we want to detect. We show that while the percentages of successful trials are 92%, 88%, and 82% for two-class, four-class, and six-class classification when no pattern subdivisions are made, it increases to 100% when each pattern is subdivided into 5 or 10 subpatterns. However, the former scenario of no pattern subdivision is more jitter resilient than the later ones.
Subhrajit Roy, Arindam Basu
IEEE Trans. Neural Networks Learn. Syst.2
2017 VLSI Extreme Learning Machine: A Design Space Exploration
abstract
In this paper, we describe a compact low-power high-performance hardware implementation of extreme learning machine for machine learning applications. Mismatches in current mirrors are used to perform the vector-matrix multiplication that forms the first stage of this classifier and is the most computationally intensive. Both regression and classification (on UCI data sets) are demonstrated and a design space tradeoff between speed, power, and accuracy is explored. Our results indicate that for a wide set of problems, σ VTin the range of 15-25 mV gives optimal results. An input weight matrix rotation method to extend the input dimension and hidden layer size beyond the physical limits imposed by the chip is also described. This allows us to overcome a major limit imposed on most hardware machine learners. The chip is implemented in a 0.35-μm CMOS process and occupies a die area of around 5 mm × 5 mm. Operating from a 1 V power supply, it achieves an energy efficiency of 0.47 pJ/MAC at a classification rate of 31.6 kHz.
Enyi Yao, Arindam Basu
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Pulse-based feature extraction for hardware-efficient neural recording systems
abstract
Current brain-machine interfaces have two machine learners-one for spike sorting and the second for intention decoding that acts on the sorted spatio-temporal spike train. In this paper, we propose a pulse-based feature extractor that can enable these two machine learners to be combined into one. We show from simulations and measurements that the information about the spike shape is still retained in the pulse counts-hence, the circuit can also be used as a traditional feature extractor. The proposed circuit also has the advantage of sharing several blocks with spike detector designs reducing system level cost. Fabricated in 65nm CMOS and operating from Vdd = 1V, the feature extractor dissipates roughly 2μW of power for an input spike rate of 100Hz.
Aritra Bhaduri, Enyi Yao, Arindam Basu
ISCAS3
2016 Morphological learning in multicompartment neuron model with binary synapses
abstract
There is a vast amount of neurobiological evidence supporting the role of dendritic processing in neural computation. However, most of the neuromorphic chips designed have overlooked these research findings. Here, we present a neuron model with multiple nonlinear spatially sensitive dendrites. Location-dependent processing occurs in multiple dendritic compartments, where sparse binary synapses are formed by utilizing structural plasticity rule. This rule is a correlation-based learning scheme inspired by the Tempotron learning rule to form synaptic connections on the dendritic compartments. The performance of the model is compared with a lumped dendrite model as well as with other classifiers for spatiotemporal spike patterns. The results indicate that our biologically realistic multicompartment model using low resolution weights achieves about 4-5% higher accuracy than the lumped scheme and about 2% less accuracy than the Tempotron using high resolution weights, while with 4-bit weights the Tempotron accuracy drops by 5% of our proposed method.
Shaista Hussain, Arindam Basu
ISCAS2
2016 A low-voltage, low power STDP synapse implementation using domain-wall magnets for spiking neural networks
abstract
Online, real-time learning in neuromorphic circuits have been implemented through variants of Spike Time Dependent Plasticity (STDP). Current implementations have used either floating-gate devices or memristors to implement such learning synapses together with non-volatile storage. However, these approaches require high voltages (≈ 3-12V) for weight update and entail high energy for learning (≈ 4-30pJ/write). We present a domain wall memory based low-voltage, low-energy STDP synapse that can operate with a power supply as low as 0.8V and update the weight at ≈ 40fJ/write. Device level simulations are performed to prove its feasibility. Its use in associative learning is also demonstrated by using neurons with dendritic branches to classify spike patterns from MNIST dataset.
Govind Narasimman, Subhrajit Roy, Xuanyao Fong, Kaushik Roy 0001, Chip-Hong Chang, Arindam Basu
ISCAS6
2016 An Online Structural Plasticity Rule for Generating Better Reservoirs
abstract
In this letter, we propose a novel neuro-inspired low-resolution online unsupervised learning rule to train the reservoir or liquid of liquid state machines. The liquid is a sparsely interconnected huge recurrent network of spiking neurons. The proposed learning rule is inspired from structural plasticity and trains the liquid through formating and eliminating synaptic connections. Hence, the learning involves rewiring of the reservoir connections similar to structural plasticity observed in biological neural networks. The network connections can be stored as a connection matrix and updated in memory by using address event representation (AER) protocols, which are generally employed in neuromorphic systems. On investigating the pairwise separation property, we find that trained liquids provide 1.36 0.18 times more interclass separation while retaining similar intraclass separation as compared to random liquids. Moreover, analysis of the linear separation property reveals that trained liquids are 2.05 0.27 times better than random liquids. Furthermore, we show that our liquids are able to retain the generalization ability and generality of random liquids. A memory analysis shows that trained liquids have 83.67 5.79 ms longer fading memory than random liquids, which have shown 92.8 5.03 ms fading memory for a particular type of spike train inputs. We also throw some light on the dynamics of the evolution of recurrent connections within the liquid. Moreover, compared to separation-driven synaptic modification', a recently proposed algorithm for iteratively refining reservoirs, our learning rule provides 9.30%, 15.21%, and 12.52% more liquid separations and 2.8%, 9.1%, and 7.9% better classification accuracies for 4, 8, and 12 class pattern recognition tasks, respectively.
Subhrajit Roy, Arindam Basu
Neural Comput.2
2016 Learning Spike Time Codes Through Morphological Learning With Binary Synapses
abstract
In this brief, a neuron with nonlinear dendrites (NNLDs) and binary synapses that is able to learn temporal features of spike input patterns is considered. Since binary synapses are considered, learning happens through formation and elimination of connections between the inputs and the dendritic branches to modify the structure or morphology of the NNLD. A morphological learning algorithm inspired by the tempotron, i.e., a recently proposed temporal learning algorithm is presented in this brief. Unlike tempotron, the proposed learning rule uses a technique to automatically adapt the NNLD threshold during training. Experimental results indicate that our NNLD with 1-bit synapses can obtain accuracy similar to that of a traditional tempotron with 4-bit synapses in classifying single spike random latency and pairwise synchrony patterns. Hence, the proposed method is better suited for robust hardware implementation in the presence of statistical variations. We also present results of applying this rule to real-life spike classification problems from the field of tactile sensing.
Subhrajit Roy, Phyo Phyo San, Shaista Hussain, Wang Wei Lee, Arindam Basu
IEEE Trans. Neural Networks Learn. Syst.5
2015 A current-mode spiking neural classifier with lumped dendritic nonlinearity
abstract
We present the current mode implementation of a spiking neural classifier with lumped square law dendritic nonlinearity. It has been shown earlier that such a system with binary synapses can be trained with structural plasticity algorithms to achieve comparable classification accuracy with less synaptic resources than conventional algorithms. Hence, in our address event based implementation, we save 2-12X memory resources in storing connectivity information. The chip fabricated in 0.35μm CMOS has 8 dendrites per cell and uses two opposing cells per class to cancel common mode inputs. Preliminary results show the chip is functional and dissipates 30nW of static power per neuronal cell and 422pJ/spike.
Amitava Banerjee, Sougata Kumar Kar, Subhrajit Roy, Aritra Bhaduri, Arindam Basu
ISCAS5
2015 A 128 channel 290 GMACs/W machine learning based co-processor for intention decoding in brain machine interfaces
abstract
A machine learning co-processor in 0.35μm CMOS for motor intention decoding in the brain-machine interfaces is presented in this paper. Using Extreme Learning Machine algorithm, time delayed sample based feature dimension enhancement, low-power analog processing and massive parallelism, it achieves an energy efficiency of 290 GMACs/W at a classification rate of 50 Hz. A portable external unit based on the proposed co-processor is verified with neural data recorded in monkey finger movements experiment, achieving a decoding accuracy of 99.3%. With time-delayed feature dimension enhancement, the classification accuracy can be increased by 5% with limited number of input channels.
Yi Chen 0012, Enyi Yao, Arindam Basu
ISCAS3
2015 Triplet spike time dependent plasticity in a floating-gate synapse
abstract
Synapses plays an important role of learning in a neural network; the learning rules which modify the synaptic strength based on the timing difference between the pre- and post-synaptic spike occurrence is termed as Spike Time Dependent Plasticity (STDP). This paper describes the compact implementation of a synapse using single floating-gate (FG) transistor (and two additional high voltage transistors) that can store a weight in a non-volatile manner and demonstrate the triplet STDP (T-STDP) learning rule developed to explain biologically observed plasticity. We describe a mathematical procedure to obtain control voltages for the FG device for T-STDP and also show measurement results, from a FG synapse fabricated in TSMC 0.35μm CMOS process to support the theory.
Roshan Gopalakrishnan, Arindam Basu
ISCAS2
2015 A 1 V, compact, current-mode neural spike detector with detection probability estimator in 65 nm CMOS
abstract
In this paper, we describe a novel low power, compact, current-mode spike detector circuit for real-time neural recording systems where neural spikes or action potentials (AP) are of interest. Such a circuit can enable massive compression of data facilitating wireless transmission. This design operates by approximating the popularly used nonlinear energy operator (NEO) through standard current mode analog blocks that can operate at low voltages. To reduce sensitivity of threshold setting, this work uses a current-mode oscillator based detection probability estimator (DPE) to reject false positives caused by the background noise. The circuit is implemented in a 65 nm CMOS process and occupies 200 μm × 150 μm of chip area. Operating from a 1 V power supply, it consumes about 88 nW of static power and 10 nJ of dynamic energy per input spike.
Enyi Yao, Arindam Basu
ISCAS2
2015 Hardware-Amenable Structural Learning for Spike-Based Pattern Classification Using a Simple Model of Active Dendrites
abstract
This letter presents a spike-based model that employs neurons with functionally distinct dendritic compartments for classifying high-dimensional binary patterns. The synaptic inputs arriving on each dendritic subunit are nonlinearly processed before being linearly integrated at the soma, giving the neuron the capacity to perform a large number of input-output mappings. The model uses sparse synaptic connectivity, where each synapse takes a binary value. The optimal connection pattern of a neuron is learned by using a simple hardware-friendly, margin-enhancing learning algorithm inspired by the mechanism of structural plasticity in biological neurons. The learning algorithm groups correlated synaptic inputs on the same dendritic branch. Since the learning results in modified connection patterns, it can be incorporated into current event-based neuromorphic systems with little overhead. This work also presents a branch-specific spike-based version of this structural plasticity rule. The proposed model is evaluated on benchmark binary classification problems, and its performance is compared against that achieved using support vector machine and extreme learning machine techniques. Our proposed method attains comparable performance while using 10% to 50% less in computational resource than the other reported techniques.
Shaista Hussain, Shih-Chii Liu, Arindam Basu
Neural Comput.3
2015 On the Non-STDP Behavior and Its Remedy in a Floating-Gate Synapse
abstract
This brief describes the neuromorphic very large scale integration implementation of a synapse utilizing a single floating-gate (FG) transistor that can be used to store a weight in a nonvolatile manner and demonstrate biological learning rules such as spike-timing-dependent plasticity (STDP). The experimental STDP plot (change in weight against ∆t=tpost - tpre ) of a traditional FG synapse from previous studies shows a depression instead of potentiation at some range of positive values of ∆t -we call this non-STDP behavior. In this brief, we first analyze theoretically the reason for this anomaly and then present a simple solution based on changing control gate waveforms of the FG device to make the weight change conform closely to biological observations over a wide range of parameters. The experimental results from an FG synapse fabricated in AMS 0.35- μ m CMOS process design are also presented to justify the claim. Finally, we present the simulation results of a circuit designed to create the modified gate voltage waveform.
Roshan Gopalakrishnan, Arindam Basu
IEEE Trans. Neural Networks Learn. Syst.2
2014 Robust doublet STDP in a floating-gate synapse
abstract
Learning in a neural network typically happens with the modification or plasticity of synaptic weight. Thus the plasticity rule which modifies the synaptic strength based on the timing difference between the pre- and post-synaptic spike occurrence is termed as Spike Time Dependent Plasticity (STDP). This paper describes the neuromorphic VLSI implementation of a synapse utilizing a single floating-gate (FG) transistor that can be used to store a weight in a nonvolatile manner and demonstrate biological learning rules such as Long-Term Potentiation (LTP), Long-Term Depression (LTD) and STDP. The experimental STDP plot of a FG synapse (change in weight against Δt = tpost- tpre) from previous studies shows a depression instead of potentiation at some range of positive values of Δt for a wide set of parameters. In this paper, we present a simple solution based on changing control gate waveforms of the FG device that makes the weight change conform closely with biological observations over a wide range of parameters. We show results from a theoretical model to illustrate the effects of the modified waveform. The experimental results from a FG synapse fabricated in AMS 0.35μm CMOS process design are also presented to justify the claim.
Roshan Gopalakrishnan, Arindam Basu
IJCNN2
2014 Spike-timing dependent morphological learning for a neuron with nonlinear active dendrites
abstract
It has been shown earlier that simple abstraction of a neuron with nonlinear active dendrites and binary synapses has a higher computational power than a neuron with linearly summing dendrites. However, it has only been used to classify high dimensional binary patterns of mean spike rates. In this paper, a nonlinear dendritic (NLD) neuron equipped with binary synapses that is able to learn temporal features of spike input patterns is presented. Since the synapses are binary, learning happens through formation and elimination of connections between the inputs and the dendritic branches thus modifying the structure or "morphology" of the cell. A morphological learning algorithm inspired by the `Tempotron'-a recently proposed temporal learning algorithm-is presented in this work. Experimental results indicate that our neuron with NLD with 1-bit synapses can obtain similar accuracy as a traditional Tempotron with 4-bit synapses in classifying a population of single spike latency patterns. Hence, the proposed method is better suited for robust hardware implementation in the presence of statistical variations.
Phyo Phyo San, Shaista Hussain, Arindam Basu
IJCNN3
2014 Improved margin multi-class classification using dendritic neurons with morphological learning
abstract
We present an architecture of a spike based multiclass classifier using neurons with non-linear dendrites and sparse synaptic connectivity where each synapse takes a binary value. The learning in this model happens not through weight updates but through structural changes, i.e. a change of connectivity between inputs and dendrites. Hence, it is well suited for implementation in neuromorphic systems using address event representation (AER). We present a new learning rule that allows better generalization of the system to noisy testing data making it feasible to transfer learnt weights in software to a hardware device interfacing with noisy spiking sensors. The new rule improves testing accuracy by 7 - 10% compared to earlier versions. We also present preliminary results for multi-class classification on handwritten digits from the MNIST database and show that our system can attain comparable performance (≈ 3% more error) with other reported spike based classifiers while using at least 50% less synaptic resources.
Shaista Hussain, Shih-Chii Liu, Arindam Basu
ISCAS3
2014 Delay learning architectures for memory and classification
Shaista Hussain, Arindam Basu, Runchun Wang, Tara J. Hamilton
Neurocomputing2
2014 A 0.7 V low-power fully programmable Gaussian function generator for brain-inspired Gaussian correlation associative memory
Chip-Hong Chang, Arindam Basu, Liter Siek
Neurocomputing3
2014 Speech Processing on a Reconfigurable Analog Platform
abstract
We describe architectures for audio classification front ends on a reconfigurable analog platform. Real-time implementation of audio processing algorithms involving discrete-time signals tend to be power-intensive. We present an alternate continuous-time system implementation of a noise-suppression algorithm on our reconfigurable chip, while detailing the design considerations. We also describe a framework that enables future implementations of other speech processing algorithms, classifier front ends, and hearing aids.
Shubha Ramakrishnan, Arindam Basu, Leung Kin Chiu, Jennifer Hasler, David V. Anderson, Stephen Brink
IEEE Trans. Very Large Scale Integr. Syst.2
2013 Improving energy gains of inexact DSP hardware through reciprocative error compensation
abstract
We present a zero hardware-overhead design approach called reciprocative error compensation (REC) that significantly enhances the energy-accuracy trade-off gains in inexact signal processing datapaths by using a two-pronged approach: (a) deliberately redesigning the basic arithmetic blocks to effectively compensate for each other's (expected) error through inexact logic minimization, and (b) "reshaping" the response waveforms of the systems being designed to further reduce any residual error. We apply REC to several DSP primitives such as the FFT and FIR filter blocks, and show that this approach delivers 2-3 orders of magnitude lower (expected) error and more than an order of magnitude lesser Signal-to-Noise Ratio (SNR) loss (in dB) over the previously proposed inexact design techniques, while yielding similar energy gains. Post-layout comparisons in the 65nm process technology show that our REC approach achieves upto 73% energy savings (with corresponding delay and area savings of upto 16% and 62% respectively) when compared to an existing exact DSP implementation while trading a relatively small loss in SNR of less than 1.5 dB.
Lingamneni Avinash, Arindam Basu, Christian C. Enz, Krishna V. Palem, Christian Piguet
DAC2
2013 Morphological learning: Increased memory capacity of neuromorphic systems with binary synapses exploiting AER based reconfiguration
abstract
Spiking neurons with lumped nonlinearity representing active dendrites can perform a larger number of input-output mappings than is possible by a neuron with linear synaptic summation of its currents. This is possible due to the additional degree of freedom in such cells-its `morphology' reflected in the number of dendrites and the choice of which inputs form synapses on the same dendrite. We present a hardware friendly algorithm for learning such optimal morphologies utilizing correlations between inputs and dendritic branch activations. We demonstrate the increased memory capacity of neurons with nonlinear dendrites and binary synapses over typically used linearly summing cells with high resolution weights. We have shown that a neuron model with a fixed number of binary weights performs much worse on a pattern classification task when it uses traditional linear dendrites than when it utilizes nonlinear dendrites (19% compared to 9% errors for 1000 patterns). This method allows to trade-off weight resolution, a problem in most current neuromorphic systems, with configurability that is the strength of address event representation (AER) based systems which can store configuration details in an off-chip memory. On a fundamental level, it points to the need of having a higher ratio of nonlinear to linear operations in spiking neural networks than is typically used.
Shaista Hussain, Roshan Gopalakrishnan, Arindam Basu, Shih-Chii Liu
IJCNN3
2013 Silicon spiking neurons for hardware implementation of extreme learning machines
Arindam Basu, Sun Shuo, Hongming Zhou, Meng-Hiot Lim, Guang-Bin Huang
Neurocomputing1
2013 Models for characterizing noise based PCMOS circuits
abstract
Quick and accurate error-rate prediction of Probabilistic CMOS (PCMOS) circuits is crucial for their systematic design and performance evaluation. While still in the early stage of research, PCMOS has shown potential to drastically reduce energy consumption at a cost of increased errors. Recently, a methodology has been proposed which could predict the error rates of cascade structures of blocks in PCMOS. This methodology requires error rates of unique blocks to predict the error rates of multiblock cascade structures composed of these unique blocks. In this article we present a new model for characterization of probabilistic circuits/blocks and present a procedure to find and characterize unique circuits/blocks. Unlike prior approaches, our new model distinguishes distinct filtering effects per output, thereby improving prediction accuracy by an average of 95% over the prior art by Palem and coauthors. Furthermore, we show two models where our new model with three stages is 18% more accurate, on average, than our simpler two-stage model. We apply our proposed models to Ripple Carry Adders and Wallace Tree Multipliers and show that using our models, the methodology of cascade structures can predict error rates of PCMOS circuits with reasonable accuracy (within 9%) in PCMOS for uniform voltages as well as multiple voltages. Finally, our approach takes seconds of simulation time whereas using HSPICE would take days of simulation time.
Anshul Singh, Arindam Basu, Keck Voon Ling, Vincent John Mooney III
ACM Trans. Embed. Comput. Syst.2
2011 A Fully Integrated Architecture for Fast and Accurate Programming of Floating Gates Over Six Decades of Current
abstract
This paper presents an on-chip system with digital serial peripheral interface (SPI) interface that enables accurate programming of floating gate arrays at a high speed. The main component allowing this speedup is a floating point current measuring analog-to-digital convertor (ADC). The ADC comprises a wide range logarithmic transimpedance amplifier (TIA) followed by a linear ramp ADC. The TIA operates over seven decades of current going down to sub-pA levels. It incorporates an adaptive biasing scheme to save power. The topology provides a relatively temperature independent measurement of the floating-gate voltage. The TIA-ADC combination operates over six decades at a thermal noise limited accuracy of 9.5 bits when average conversion time is around 500 μs. The system features level-shifters and selection circuitry at the periphery of the floating gate array, current-steering digital-to-analog converters (DACs) to set gate and drain voltages, and SPI for a microprocessor or field-programmable gate array (FPGA). Algorithms using either pulse-width modulation or drain voltage modulation can be implemented on this platform. We present data for this system from 0.5 μm AMI and 0.35 μ m TSMC processes.
Arindam Basu, Paul E. Hasler
IEEE Trans. Very Large Scale Integr. Syst.1
2010 Neural dynamics in reconfigurable silicon
abstract
A neuromorphic analog chip is presented that is capable of implementing massively parallel neural computations while retaining the programmability of digital systems. We show measurements from neurons with Hopf bifurcations and integrate and fire neurons, excitatory and inhibitory synapses, passive dendrite cables and central pattern generators implemented on the chip. This chip provides a platform for not only simulating detailed neuron dynamics but also using the same to interface with actual cells in applications like a dynamic clamp. The programmability is achieved using floating gate transistors with on-chip programming control. The switch matrix for interconnecting the components also consists of floating-gate transistors. Massive computational area efficiency is obtained by using the reconfigurable interconnect as synaptic weights.
Arindam Basu, Shubha Ramakrishnan, Paul E. Hasler
ISCAS1
2010 Live demonstration: Hardware and software infrastructure for a family of floating-gate based FPAAs
abstract
Analog circuits and systems research and education can benefit from the flexibility provided by large-scale Field Programmable Analog Arrays (FPAAs). This demonstration will present visitors with the hardware and software infrastructure supporting the use of a family of floating-gate based FPAAs being developed at Georgia Tech. A picture of the programming and control hardware that will be demonstrated is found in Figure la. Figure lb shows the software flow that will be demonstrated. The infrastructure is compact and portable and provides the user with a comprehensive set of tools for custom analog circuit design and implementation. The infrastructure includes the FPAA integrated circuit (IC); discrete analog to digital converters (ADC), digital to analog converters (DAC) and amplifier ICs; a 32-Bit ARM based microcontroller (μC) for interfacing the FPAA with a laptop computer; and Matlab and targeting software. The FPAA hardware communicates with Matlab over a Universal Serial Bus (USB) connection. The USB connection also provides the hardware's power. The software tools in the demonstration include three major systems: a Matlab Simulink FPAA program, a SPICE to FPAA compiler called GRASPER, and a visualization tool called RAT. Figure 2 shows a block diagram of the Programming and Control board and demonstration setup.
Scott Koziol, Craig Schlottmann, Arindam Basu, Stephen Brink, Csaba Petre, Brian P. Degnan, Shubha Ramakrishnan, Paul E. Hasler, Aurele Balavoine
ISCAS3
2010 Hardware and software infrastructure for a family of floating-gate based FPAAs
abstract
Analog circuits and systems research and education can benefit from the flexibility provided by large-scale Field Programmable Analog Arrays (FPAAs). This paper presents the hardware and software infrastructure supporting the use of a family of floating-gate based FPAAs being developed at Georgia Tech. This infrastructure is compact and portable and provides the user with a comprehensive set of tools for custom analog circuit design and implementation. The infrastructure includes the FPAA IC, discrete ADC, DAC and amplifier ICs, a 32-Bit ARM based microcontroller for interfacing the FPAA with the user's computer, and Matlab and targeting software. The FPAA hardware communicates with Matlab over a USB connection. The USB connection also provides the hardware's power. The software tools include three major systems: a Matlab Simulink FPAA program, a SPICE to FPAA compiler called GRASPER, and a visualization tool called RAT. The hardware consists of two custom PCB designs which include a main board used to program and control an FPAA IC and an FPAA IC adaptor board used to interface a QFP packaged FPAA IC with the 100 pin ZIF socket on the main programming and control board.
Scott Koziol, Craig Schlottmann, Arindam Basu, Stephen Brink, Csaba Petre, Brian P. Degnan, Shubha Ramakrishnan, Paul E. Hasler, Aurele Balavoine
ISCAS3
2010 Integrated low voltage and low power CMOS circuits for optical sensing of diffraction based micromachined microphone
abstract
We present CMOS electronics for optical sensing of diffraction based micromachined microphone. Peak detector and track and hold circuits are used for continuous time approach with only 58 μA of current for 65dB peak SNR (lPa at 1kHz). We also present a 1-bit sigma delta interface ADC for directly digitizing the signal for digital processing with 63 μA of current for 62dB peak SNR (lPa @ 1kHz). The electronics are designed in 0.35um standard digital CMOS process with reduced supply of 1.5Volts. All the analog front end circuits operate in weak inversion with rail-to-rail wide input linear range.
Muhammad Shakeel Qureshi, Arindam Basu, Baris Bicen, Levent Degertekin, Paul E. Hasler
ISCAS2
2009 A learning digital computer
abstract
The concept of learning digital hardware is presented here. A proof of concept of a circuit that can arbitrarily control the current, and thus the switching speed and power consumption, of a digital circuit is given. This control of current is directly tuned by the feedback from the digital circuit itself, thus a learning digital computer. An argument for a completely new paradigm in digital computing follows whereby an entire system of learning digital circuits is proposed.
Bo Marr, Arindam Basu, Stephen Brink, Paul E. Hasler
DAC2
2009 A Large-scale Reconfigurable Smart Sensory Chip
abstract
The Reconfigurable Smart Sensory Chip (RSSC) is a powerful tool for fast prototyping sensory microsystems. Innovative design ideas can be quickly realized and tested in hardware without doing time-consuming and expensive silicon fabrication. The RSSC is a large-scale floating-gate based IC containing 8 universal sensor interface blocks, each of which can be configured for voltage sensing, capacitive sensing, or current sensing, and 28 configurable analog blocks. The outputs of the interface circuits can be multiplexed out in a time-division sequence or can be routed to the configurable analog blocks for further analog signal processing or data conversion. With more than 50,000 programmable elements and on-chip programming circuitry, RSSC is an extremely powerful tool to develop and test a great variety of smart sensory microsystems in minutes.
Sheng-Yu Peng, Gokce Gurun, Christopher M. Twigg, Muhammad Shakeel Qureshi, Arindam Basu, Stephen Brink, Paul E. Hasler, Levent Degertekin
ISCAS5
2008 Bifurcations in a silicon neuron
abstract
In this paper, we describe the bifurcations occurring in a silicon neuron with one sodium and one potassium channel. The channels are designed to model the physics of ion flow in actual biological channels instead of modeling a particular set of equations. We show a pair of subcritical Hopf-bifurcation with increase in current stimulus which is characteristic of class 2 excitability in Hodgkin-Huxley neurons. Theoretical analysis of the bifurcations lead to conditions for designing and biasing the circuit. The circuit is very compact, comprising six transistors and three capacitors, lending itself to easy integration. The parameters are set using floating-gate transistors and can be programmed as desired. We hope to study more complicated dynamics of large networks of these neurons, a task which might be beyond a typical digital computer.
Arindam Basu, Csaba Petre, Paul E. Hasler
ISCAS1
2007 A Fully Integrated Architecture for Fast Programming of Floating Gates
abstract
We present an on-chip system that enables programming floating gate arrays at a high speed. The main component allowing this speedup is a floating point current measuring ADC operating over 4 decades at 10bit accuracy or 7decades at 7 bit accuracy. The conversion time is around 200μs till around 30 pA of current. The gate and drain voltages are set by on-chip DACs. The digital words for the DAC are sent by an FPGA through an SPI interface. The controller for sequencing the operations as well as the look-up-table with characterization data are on the FPGA. Algorithms using either pulse-width modulation or drain voltage modulation can be implemented.
Arindam Basu, Paul E. Hasler
ISCAS1
2007 Dynamics of a Logarithmic Transimpedance Amplifier
abstract
We analyze the dynamics of a logarithmic transimpedance amplifier that are not explainable from its small-signal model. The amplifier was used both to measure pixel currents in an imager as well as currents from a floating gate array for accurate programming. We explain the marked asymmetry between the up-going and down-going current steps intuitively using circuit concepts and then based on phase-plane analysis. This analysis also helps develop a method of reducing the huge delay to down-going current steps.
Arindam Basu, Kofi M. Odame, Paul E. Hasler
ISCAS1
2007 A Low-Power, Compact, Adaptive Logarithmic Transimpedance Amplifier Operating over Seven Decades of Current
abstract
This paper describes a transimpedance amplifier using logarithmic compression of the input current for wide dynamic range current sensing applications. Measurements demonstrate operation over currents ranging from 200fA to 2μA with an average error of 0.8%. It is analytically shown that the power dissipation of the non-adaptive structure varies linearly with dynamic range. This amplifier alleviates this strong dependance on dynamic range and achieves low-power operation by adapting the bias current of the amplifier depending on input current thus burning an average power of 3.45μW per decade of current. Either SNR or bandwidth can be made to trade-off with the input current depending on application. If the bandwidth is limited to 5kHz, it achieves an average SNR of 65dB.
Arindam Basu, Ryan W. Robucci, Paul E. Hasler
ISCAS1
2007 Above Threshold pFET InjectionModeling intended for ProgrammingFloating-Gate Systems
abstract
We present a first-order model for pFET hotelectron injection that is consistent for subthreshold and above threshold current levels. Injection is a critical phenomena for high-precision programming of floating-gate devices, and accurate modeling fuels continued improvement of programming techniques. Previous work has shown good modeling for subthreshold operation; in this work we extend the modeling throughout the region, enabling improved programming algorithms for floating-gate switch elements, resistors, and highperformance circuit elements. We discuss the implementation of this model in CADENCE's version of SPICE.
Paul E. Hasler, Arindam Basu, Sctt Kozil
ISCAS2
2007 Transistor Channel Dendrites implementing HMM classifiers
abstract
Recently the authors presented transistor channel models of biological channels and the resulting implementation towards building spiking nodes, synapses, and dendrites. The authors also discussed how to build reconfigurable dendrites using programmable analog techniques. With all of this technology components available, the authors begin to address the question of the computation model possible using a dendrite element, as well as a network of dendrite elements. The authors discuss the connection between a dendrite element and a hidden Markov model (HMM) classifier branch, as well as a network of dendrites and somas to create an HMM classifier typical of what is used in speech recognition systems. The authors present simulation and experimental results for the branch elements; the authors also present initial results for a small dendrite based classifier structure to show the similarities to the HMM paradigm.
Paul E. Hasler, Scott Koziol, Ethan Farquhar, Arindam Basu
ISCAS4
2007 Studying Nonlinear Dynamical Systems on a Reconfigurable Analog Platform
abstract
We have developed a field programmable analog array (FPAA) that can be configured to synthesis and analyze a vast variety of circuits. This FPAA is a valuable platform for studying nonlinear dynamics in circuits, as it offers close to the flexibility of computer simulation, but with actual experimental results. We present data from a current mirror, a peak detector, and a second-order section
Kofi M. Odame, Christopher M. Twigg, Arindam Basu, Paul E. Hasler
ISCAS3
2004 A generalized analog architecture for DCT, DST and its inverse
abstract
This paper describes a sampled analog architecture, for computing DCT or DST, using the switched capacitor principle with capacitance switching. The input sample stream is applied to an array of capacitors and multiplied by all the DCT/DST coefficients concurrently using capacitor ratios. These capacitors are switched concurrently with the help of a switching matrix, to realize switched capacitor integrators for performing the necessary addition/subtraction. The architecture may also be used for computing inverse DCT and DST transforms. The proposed architecture is simple, regular and can be used for on-line computations, with good accuracy.
Ashis Kumar Mal, Arindam Basu
ICASSP (5)2