EDBT 2026 Demo / reviewers in the wild / expert
Dhireesha Kudithipudi
dblp:24/748
· DBLP profile ↗
33ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-4462-5224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 5 since 2021Artificial intelligence and machine learning · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Device-Algorithm Co-Design with FeFETs for On-device Continual Learning
Fatima Tuz Zohora, Nicolas Ramos, Abinidhi Geethaikrishnan, Hai Li 0001, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | Minion gated recurrent unit for continual learning
Abdullah M. Zyarah, Dhireesha Kudithipudi |
Neurocomputing | 2 |
| 2025 | Time-Series Forecasting and Sequence Learning Using Memristor-based Reservoir SystemabstractPushing the frontiers of time-series information processing in the ever-growing domain of edge devices with stringent resources has been impeded by the systems’ ability to process information and learn locally on the device. Local processing and learning of time-series information typically demand intensive computations and massive storage as the process involves retrieving information and tuning hundreds of parameters back in time. In this work, we developed a memristor-based echo state network accelerator that features efficient temporal data processing and in situ online learning. The proposed design is benchmarked using various datasets involving real-world tasks, such as forecasting the load energy consumption and weather conditions. The experimental results illustrate that the hardware model experiences a marginal degradation in performance as compared to the software counterpart. This is mainly attributed to the limited precision and dynamic range of network parameters when emulated using memristor devices. The proposed system is evaluated for lifespan, robustness, and energy-delay product. It is observed that the system demonstrates reasonable robustness for device failure below 10%, which may occur due to stuck-at faults. Furthermore, 247× reduction in energy consumption is achieved when compared to a custom CMOS digital design implemented at the same technology node. Abdullah M. Zyarah, Dhireesha Kudithipudi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2024 | PositCL: Compact Continual Learning with Posit Aware QuantizationabstractNeural network models catastrophically forget previously learned information while acquiring new knowledge, requiring a fundamental change in learning models and architectures. These enhancements to architecture structures and training mechanisms lead to an increase in memory and computational resources, making it difficult to deploy models on resource-constrained edge devices. To enhance both memory and computational efficiency, we propose a model compression approach for spiking continual learning models, where the model parameters are quantized with varying precision according to their weight distribution. Vedant Karia, Abdullah M. Zyarah, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | A domain-agnostic approach for characterization of lifelong learning systems
Megan M. Baker, Alexander New, Mario Aguilar-Simon, Ziad Al-Halah, Sébastien M. R. Arnold, Eseoghene Benjamin, Andrew P. Brna, Ethan Brooks, Ryan C. Brown, Zachary A. Daniels, Anurag Reddy Daram, Fabien Delattre, Ryan Dellana, Eric Eaton, Haotian Fu, Kristen Grauman, Jesse Hostetler, Shariq Iqbal, Cassandra Kent, Nicholas Ketz, Soheil Kolouri, George Dimitri Konidaris, Dhireesha Kudithipudi, Erik G. Learned-Miller, Michael L. Littman, Sandeep Madireddy, Jorge A. Mendez, Eric Q. Nguyen, Christine D. Piatko, Praveen K. Pilly, Aswin Raghavan, Abrar Rahman, Santhosh K. Ramakrishnan, Neale Ratzlaff, Andrea Soltoggio, Peter Stone 0001, Indranil Sur, Zhipeng Tang, Saket Tiwari, Kyle Vedder, Felix Wang, Zifan Xu, Angel Yanguas-Gil, Harel Yedidsion, Shangqun Yu, Gautam K. Vallabha |
Neural Networks | 23 |
| 2022 | SCOLAR: A Spiking Digital Accelerator with Dual Fixed Point for Continual LearningabstractSpiking neural network models when deployed in dynamic environments, catastrophically forget previously learned tasks. In this paper, we propose a reconfigurable spiking digital accelerator, which uses activity-dependent metaplasticity to mitigate catastrophic forgetting. The proposed accelerator has a custom low precision dual fixed point representation for network parameters. The custom precision leads to lower quantization error and higher accuracy. We evaluate the proposed accelerator on split-MNIST continual learning benchmark. Analysis shows that representing network parameters with 8-bit dual fixed point numbers reduces the memory footprint compared to 16-bit fixed point numbers, while maintaining comparable continual learning ability. Vedant Karia, Fatima Tuz Zohora, Nicholas Soures, Dhireesha Kudithipudi |
ISCAS | 4 |
| 2021 | MetaplasticNet: Architecture with Probabilistic Metaplastic Synapses for Continual LearningabstractMetaplasticity, the activity-dependent modification of synaptic plasticity, is an important technique for mitigating catastrophic forgetting in neural networks. Often, continual learning models with metaplasticity require compute-intensive training. In this research, we propose a probabilistic metaplastic synapse with discrete hidden states that alleviates the computational cost. We implement a digital architecture of the network with on-chip training to achieve further power savings. Results show upto ~ 22% and ~ 21% improvement in mean accuracy for Split-MNIST and sequential MNIST-FMNIST benchmarks respectively, compared to previous metaplasticity models. Simulations of the full digital architecture show ~ 53× lower power consumption per weight update with similar accuracy as gradient-based network counterparts. Fatima Tuz Zohora, Vedant Karia, Anurag Reddy Daram, Abdullah M. Zyarah, Dhireesha Kudithipudi |
ISCAS | 5 |
| 2020 | Metaplasticity in Multistate Memristor Synaptic NetworksabstractRecent studies have shown that metaplastic synapses can retain information longer than simple binary synapses and are beneficial for continual learning. In this paper, we explore the multistate metaplastic synapse characteristics in the context of high retention and reception of information. Inherent behavior of a memristor emulating the multistate synapse is employed to capture the metaplastic behavior. An integrated neural network study for learning and memory retention is performed by integrating the synapse in a 5 × 3 crossbar at the circuit level and 128 × 128 network at the architectural level. An on-device training circuitry ensures the dynamic learning in the network. In the 128 × 128 network, it is observed that the number of input patterns the multistate synapse can classify is ≃ 2.1× that of a simple binary synapse model, at a mean accuracy of ≥ 75%. Fatima Tuz Zohora, Abdullah M. Zyarah, Nicholas Soures, Dhireesha Kudithipudi |
ISCAS | 4 |
| 2020 | Neuromorphic System for Spatial and Temporal Information ProcessingabstractNeuromorphic systems that learn and predict from streaming inputs hold significant promise in pervasive edge computing and its applications. In this article, a neuromorphic system that processes spatio-temporal information on the edge is proposed. Algorithmically, the system is based on hierarchical temporal memory that inherently offers online learning, resiliency, and fault tolerance. Architecturally, it is a full custom mixed-signal design with an underlying digital communication scheme and analog computational modules. Therefore, the proposed system features reconfigurability, real-time processing, low power consumption, and low-latency processing. The proposed architecture is benchmarked to predict on real-world streaming data. The network's mean absolute percentage error on the mixed-signal system is 1.129 X lower compared to its baseline algorithm model. This reduction can be attributed to device non-idealities and probabilistic formation of synaptic connections. We demonstrate that the combined effect of Hebbian learning and network sparsity also plays a major role in extending the overall network lifespan. We also illustrate that the system offers 3.46 X reduction in latency and 77.02 X reduction in power consumption when compared to a custom CMOS digital design implemented at the same technology node. By employing specific low power techniques, such as clock gating, we observe 161.37 X reduction in power consumption. Abdullah M. Zyarah, Kevin Gomez, Dhireesha Kudithipudi |
IEEE Trans. Computers | 3 |
| 2019 | Deep Positron: A Deep Neural Network Using the Posit Number SystemabstractThe recent surge of interest in Deep Neural Networks (DNNs) has led to increasingly complex networks that tax computational and memory resources. Many DNNs presently use 16-bit or 32-bit floating point operations. Significant performance and power gains can be obtained when DNN accelerators support low-precision numerical formats. Despite considerable research, there is still a knowledge gap on how low-precision operations can be realized for both DNN training and inference. In this work, we propose a DNN architecture, Deep Positron, with posit numerical format operating successfully at ≤8 bits for inference. We propose a precision-adaptable FPGA soft core for exact multiply-and-accumulate for uniform comparison across three numerical formats, fixed, floating-point and posit. Preliminary results demonstrate that 8-bit posit has better accuracy than 8-bit fixed or floating-point for three different low-dimensional datasets. Moreover, the accuracy is comparable to 32-bit floating-point on a Xilinx Virtex-7 FPGA device. The trade-offs between DNN performance and hardware resources, i.e. latency, power, and resource utilization, show that posit outperforms in accuracy and latency at 8-bit and below. Zachariah Carmichael, Hamed Fatemi Langroudi, Char Khazanov, Jeffrey Lillie, John L. Gustafson, Dhireesha Kudithipudi |
DATE | 6 |
| 2019 | Exploiting Randomness in Deep Learning AlgorithmsabstractThe recent surge of interest in using deep neural networks for real-world tasks has led to training complex networks with billions of parameters that use enormous amounts of training data. Performing backpropagation in these deep networks is time consuming and requires large amount of resources that is usually limited by the underlying hardware. In order to move towards agile deep learning, we are motivated to exploit randomness in the networks. In this work, we explore the effects of utilizing random weights in convolutional neural networks. This is achieved through random initialization of weights and by freezing them. The training occurs only in the output layer. We also propose a novel weight distribution method based on the sum of sinusoids for random convolutional neural networks. Our experiments show that by leaving the weights random in convolutional neural networks relatively high performance can be achieved for MSTAR and CIFAR-10 datasets. Hamed Fatemi Langroudi, Cory E. Merkel, Humza Syed, Dhireesha Kudithipudi |
IJCNN | 4 |
| 2019 | Neuromemristive Multi-Layer Random Projection Network with On-Device LearningabstractThis paper proposes a neuromemristive multi-layer neural network with on-device learning. The proposed system is studied within the context of a feedforward multi-layer random projection network, where the core learning is modeled by a stochastic gradient descent simplified for memristor crossbar integration. Two random projection network topologies are explored for binomial and multinomial datasets. A detailed study on the resiliency of the networks in the presence of device failure is performed. The topology with softmax output layer exhibits stability and better resiliency in performance after experiencing a device failure. It is shown that this topology can regain full performance after experiencing 30% stuck-at-faults, with 2x increase in the hidden layer neurons. Abdullah M. Zyarah, Dhireesha Kudithipudi |
IJCNN | 2 |
| 2019 | Neuromemrisitive Architecture of HTM with On-Device Learning and NeurogenesisabstractHierarchical temporal memory (HTM) is a biomimetic sequence memory algorithm that holds promise for invariant representations of spatial and spatio-temporal inputs. This article presents a comprehensive neuromemristive crossbar architecture for the spatial pooler (SP) and the sparse distributed representation classifier, which are fundamental to the algorithm. There are several unique features in the proposed architecture that tightly link with the HTM algorithm. A memristor that is suitable for emulating the HTM synapses is identified and a new Z-window function is proposed. The architecture exploits the concept of synthetic synapses to enable potential synapses in the HTM. The crossbar for the SP avoids dark spots caused by unutilized crossbar regions and supports rapid on-chip training within two clock cycles. This research also leverages plasticity mechanisms such as neurogenesis and homeostatic intrinsic plasticity to strengthen the robustness and performance of the SP. The proposed design is benchmarked for image recognition tasks using Modified National Institute of Standards and Technology (MNIST) and Yale faces datasets, and is evaluated using different metrics including entropy, sparseness, and noise robustness. Detailed power analysis at different stages of the SP operations is performed to demonstrate the suitability for mobile platforms. Abdullah M. Zyarah, Dhireesha Kudithipudi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | On-Device Learning in Memristor Spiking Neural NetworksabstractIn this paper, a memristor spiking neuron and synaptic trace circuits for efficient on device learning are presented. A key feature of these circuits is the use of memristors to emulate the membrane potential of spiking neurons, as opposed to the conventional use of a capacitor. The circuits are designed in IBM 65nm technology node and validated on a small-scale spiking neural network. It was observed that a 3×3 spiking neural network consumes 19.1 μW of power at 100 MHz. Abdullah M. Zyarah, Nicholas Soures, Dhireesha Kudithipudi |
ISCAS | 3 |
| 2018 | Semi-Trained Memristive Crossbar Computing Engine with In Situ Learning AcceleratorabstractOn-device intelligence is gaining significant attention recently as it offers local data processing and low power consumption. In this research, an on-device training circuitry for threshold-current memristors integrated in a crossbar structure is proposed. Furthermore, alternate approaches of mapping the synaptic weights into fully trained and semi-trained crossbars are investigated. In a semi-trained crossbar, a confined subset of memristors are tuned and the remaining subset of memristors are not programmed. This translates to optimal resource utilization and power consumption, compared to a fully programmed crossbar. The semi-trained crossbar architecture is applicable to a broad class of neural networks. System level verification is performed with an extreme learning machine for binomial and multinomial classification. The total power for a single 4 × 4 layer network, when implemented in IBM 65nm node, is estimated to be ≈42.16μW and the area is estimated to be 26.48μm × 22.35μm. Abdullah M. Zyarah, Dhireesha Kudithipudi |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | A penalized maximum likelihood approach to the adaptive learning of the spatial pooler permanenceabstractHierarchical Temporal Memory is a machine learning algorithm for spatio-temporal information processing. One of the key functional units in this algorithm is the spatial pooler, which has been demonstrated to be efficient in classification, dimensionality reduction and for preprocessing non-spatial inputs. Formalization of the spatial pooler is proposed in recent literature. In this work, we present a principled theoretical formulation of the spatial poolers underlying learning scheme. Constraints from the active connected and emmeshed columns in a spatial pooler are included in the analysis. It has been observed that locally adaptive learning enhanced the performance of the spatial pooler as a feature selector. Ernest Fokoué, Lakshmi Ravi, Dhireesha Kudithipudi |
IJCNN | 3 |
| 2017 | Robustness of a memristor based liquid state machineabstractThe ability to learn from noisy and incomplete information is highly desired in cognitive systems. When these cognitive systems are realized in hardware, such as neuromemristive systems, an added constraint is how the algorithms adapt to the inherent noise and variability from the devices. In this work, we explore the robustness of the reservoir computing algorithm, specifically a liquid state machine, when realized as a mixed signal neuromemristive system. The study focuses on robustness of the liquid state machine under different manifestations of memristor read and write noise. A high-level analysis on the liquid state machine's accuracy for simultaneous occurrence of multiple sources of noise is also investigated. For analysis, TIMIT speech recognition and spoken arabic digit datasets were used. The results support that the neuromemristive liquid state machine has high immunity to a variety of non-ideal device effects. There is 22% degradation in the classification accuracy of the Liquid state machine, even in cases where 15% of the neurons in the liquid are faulty. Nicholas Soures, Lydia Hays, Dhireesha Kudithipudi |
IJCNN | 3 |
| 2017 | Extreme learning machine as a generalizable classification engineabstractExtreme learning machine is an emerging neural network architecture that offers fast learning and generalization for multiple tasks. In this work, a scalable digital architecture for multi-classifier extreme learning machine (MT-ELM) is proposed. The proposed architecture performs multiple classification tasks without reconfiguring the network. The design is validated with MNIST dataset and it is shown that the proposed model achieves an accuracy of 91.7% for classifying numbers in the MNIST dataset and an accuracy of 90.35% for categorizing number parity. The design is synthesized on a TSMC-65nm technology node and the power dissipation is 13.6 mW for MT-ELM network with 80 hidden neurons and 12 output neurons. Abdullah M. Zyarah, Dhireesha Kudithipudi |
IJCNN | 2 |
| 2017 | Ziksa: On-chip learning accelerator with memristor crossbars for multilevel neural networksabstractMemristor crossbars support efficient realizations of spiking and non-spiking neural networks designs. In most of these designs off-chip/ex-situ training is used to set/update the state of the memrisitve devices. However, there is a growing need to design an efficient on-chip/in-situ learning for mobile autonomous systems. In this research, we propose an on-chip learning accelerator, known as Ziksa, that is integrated with the memristor crossbars. We demonstrate how regression and back-propagation in multi-level networks can be realized through Ziksa. The proposed accelerator is evaluated on a fabricated TiN-TaOx-TaTiN memristor crossbar. A 3-layer feedforward network was tested using Ziksa for classification. An accuracy of 95.3% was achieved on Wisconsin breast cancer dataset. The proposed learning accelerator can be envisioned as a core building block in a wide-range of cognitive algorithms that rely on on-chip online learning. Abdullah M. Zyarah, Nicholas Soures, Lydia Hays, Robin Jacobs-Gedrim, Sapan Agarwal, Matthew J. Marinella, Dhireesha Kudithipudi |
ISCAS | 7 |
| 2017 | Stochastic CBRAM-Based Neuromorphic Time Series Prediction SystemabstractIn this research, we present a Conductive-Bridge RAM (CBRAM)-based neuromorphic system which efficiently addresses time series prediction. We propose a new (i) voltage-mode, stochastic, multiweight synapse circuit based on experimental bi-stable CBRAM devices, (ii) a voltage-mode neuron circuit based on the concept of charge sharing, and (iii) an optimized training methodology powered by a stochastic implementation of the Least-Mean-Squares (SLMS) training rule. To validate the proposed design, we use time series prediction for short-term electrical load forecasting in smart grids. Our system is able to forecast hourly electrical loads with a mean accuracy of 96%, an estimated power dissipation of 15 μW, and area of 14.5 μm 2 at 65 nm CMOS technology. Cory E. Merkel, Dhireesha Kudithipudi, Manan Suri, Bryant T. Wysocki |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2017 | Editorial: A Successful Year and Looking Forward to 2017 and BeyondabstractThis issue marks the first anniversary issue since I was honored to serve as the Editor-in-Chief (EiC) of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS). I am happy to report that we had a very successful year and here are a few highlights that I would like to share with the community.•The latest impact factor of TNNLS is 4.854 according to the Journal Citation Reports. This marks a record high impact factor for our journal and places TNNLS as the number one scholarly publication in Computer Science (Hardware & Architecture), number three in Computer Science (Theory & Methods), and number ten in Electrical and Electronic Engineering journals. Haibo He, Barbara Hammer, Daniel W. C. Ho, Fakhri Karray, Dhireesha Kudithipudi, José Antonio Lozano 0001, Teresa Bernarda Ludermir, Jacek Mandziuk, Stefano Melacci, Antonio Paiva, Hong Qiao, Alain Rakotomamonjy, Shiliang Sun, Johan A. K. Suykens |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2016 | A design of HTM spatial pooler for face recognition using memristor-CMOS hybrid circuitsabstractHierarchical Temporal Memory (HTM) is a machine learning algorithm that is inspired from the working principles of the neocortex, capable of learning, inference, and prediction for bit-encoded inputs. Spatial pooler is an integral part of HTM that is capable of learning and classifying visual data such as objects in images. In this paper, we propose a memristor-CMOS circuit design of spatial pooler and exploit memristors capabilities for emulating the synapses, where the strength of the weights is represented by the state of the memristor. The proposed design is validated on a challenging application of single image per person face recognition problem using AR database resulting in a recognition accuracy of 80%. Timur Ibrayev, Alex James 0001, Cory E. Merkel, Dhireesha Kudithipudi |
ISCAS | 4 |
| 2015 | Design and analysis of neuromemristive echo state networks with limited-precision synapsesabstractEcho state networks (ESNs) are gaining popularity as a method for recognizing patterns in time series data. ESNs are random, recurrent neural network topologies that are able to integrate temporal data over short time windows by operating on the edge of chaos. In this paper, we explore the design of a hardware ESN with bi-stable memristor-based synapses. Hybrid CMOS/memristor hardware implementations of ESNs are able to exploit non-linear device physics, improving power consumption, and boosting performance over software approaches. However, the digital nature of most experimental memristors places a limit on the precision of weight states in the ESN's readout layer. In spite of this, we show that ESNs with only 5 different readout layer weight states can acheive 67% accuracy in spoken digit recognition tasks. Colin Donahue, Cory E. Merkel, Qutaiba Saleh, Levs Dolgovs, Yu Kee Ooi, Dhireesha Kudithipudi, Bryant T. Wysocki |
CISDA | 6 |
| 2015 | Memristive computational architecture of an echo state network for real-time speech-emotion recognitionabstractEcho state neural networks (ESNs) provide an efficient classification technique for spatiotemporal signals. The feedback connections in the ESN topology enable feature extraction of both spatial and temporal components in time series data. This property has been used in several application domains such as image and video analysis, anomaly detection, and speech recognition. In this research, we explore a hardware architecture for realizing ESN efficiently in power-constrained devices. Specifically, we propose a scalable computational architecture applied to speech-emotion recognition. Two different topologies are explored, with memristive synapses. The simulation results are promising with a classification accuracy of ≈ 96% for two distinct emotion statuses. Qutaiba Saleh, Cory E. Merkel, Dhireesha Kudithipudi, Bryant T. Wysocki |
CISDA | 3 |
| 2014 | A current-mode CMOS/memristor hybrid implementation of an extreme learning machineabstractIn this work, we propose a current-mode CMOS/memristor hybrid implementation of an extreme learning machine (ELM) architecture. We present novel circuit designs for linear, sigmoid,and threshold neuronal activation functions, as well as memristor-based bipolar synaptic weighting. In addition, this work proposes a stochastic version of the least-mean-squares (LMS) training algorithm for adapting the weights between the ELM's hidden and output layers. We simulated our top-level ELM architecture using Cadence AMS Designer with 45 nm CMOS models and an empirical piecewise linear memristor model based on experimental data from an HfOx device. With 10 hidden node neurons, the ELM was able to learn a 2-input XOR function after 150 training epochs. Cory E. Merkel, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2014 | Temperature Sensing RRAM Architecture for 3-D ICsabstract3-D integrated circuits, or 3-D ICs, have gained significant attention in the research community over the past few years. This has primarily been motivated by their enhanced power, performance, and functionality over planar CMOS ICs. However, thermal management remains a key challenge in these devices due to the impedance of heat flow that results from die stacking. In this paper, we address this challenge by utilizing a temperature sensing resistive random access memory (TSRRAM) which can generate accurate thermal profiles to gauge the heat distribution within the 3-D IC. The architecture enables each RRAM switching element in the memory die to be used both as a memory bit and a temperature sensor. We simulated our TSRRAM design as an L2 cache for an alpha 21364 processor. We used a customized simulation framework to test the design accuracy and performance over several SPEC2000 CPU benchmarks, and achieved a 2.14 K mean error and an eight-cycle performance overhead with a 4-kB L2 cache size. Furthermore, we show that active sensing methods can be employed to achieve 100% coverage of global hot spot temperatures. Cory E. Merkel, Dhireesha Kudithipudi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Periodic activation functions in memristor-based analog neural networksabstractThis work explores the use of periodic activation functions in memristor-based analog neural networks. We propose a hardware neuron based on a folding amplifier that produces a periodic output voltage. Furthermore, the amplifier's fold factor be adjusted to change the number of low-to-high or high-to-low output voltage transitions. We also propose a memristor-based synapse circuit and training circuitry for realizing the Perceptron learning rule. Behavioral models of our circuits were developed for simulating a single-layer, single-output feedforward neural network. The network was trained to detect the edges of a grayscale image. Our results show that neurons with a single fold-with an activation function similar to a sigmoidal activation function-perform the worst for this application, since they are unable to learn functions with multiple decision boundaries. Conversely, the 4-fold neuron performs the best (up to ≈65% better than the 1-fold neuron), as its activation function is periodic, and it is able to learn functions with four decision boundaries. Cory E. Merkel, Dhireesha Kudithipudi, Nick Sereni |
IJCNN | 2 |
| 2013 | Memristor-Based Neural Logic Blocks for Nonlinearly Separable FunctionsabstractNeural logic blocks (NLBs) enable the realization of biologically inspired reconfigurable hardware. Networks of NLBs can be trained to perform complex computations such as multilevel Boolean logic and optical character recognition (OCR) in an area- and energy-efficient manner. Recently, several groups have proposed perceptron-based NLB designs with thin-film memristor synapses. These designs are implemented using a static threshold activation function, limiting the set of learnable functions to be linearly separable. In this work, we propose two NLB designs-robust adaptive NLB (RANLB) and multithreshold NLB (MTNLB)-which overcome this limitation by allowing the effective activation function to be adapted during the training process. Consequently, both designs enable any logic function to be implemented in a single-layer NLB network. The proposed NLBs are designed, simulated, and trained to implement ISCAS-85 benchmark circuits, as well as OCR. The MTNLB achieves 90 percent improvement in the energy delay product (EDP) over lookup table (LUT)-based implementations of the ISCAS-85 benchmarks and up to a 99 percent improvement over a previous NLB implementation. As a compromise, the RANLB provides a smaller EDP improvement, but has an average training time of only ≈ 4 cycles for 4-input logic functions, compared to the MTNLBs ≈ 8-cycle average training time. Michael Soltiz, Dhireesha Kudithipudi, Cory E. Merkel, Garrett S. Rose, Robinson E. Pino |
IEEE Trans. Computers | 2 |
| 2012 | Design-time performance evaluation of thermal management policies for SRAM and RRAM based 3D MPSoCsabstract3D-ICs hold significant promise for future generation multi processor systems-on-chip due to their potential for increased performance, decreased power, heterogeneous integration, and reduced cost over planar ICs. However, the vertical integration of these structures exacerbates the heat dissipation and run-time thermal management issues. There have been a number of design- and run-time thermal management policies proposed, but few focus on examining overall system performance. Additionally, the heterogeneity of 3D-ICs allows for the integration of novel technologies, such as resistive random access memories (RRAMs), which offer higher density and lower power than traditional CMOS memory technologies. Our work presents a flexible design-time simulation framework to evaluate system performance and thermal profiles of 3D MPSoCs. We utilize this framework to study the effect of three dynamic thermal management policies (air-cooled load balancing, liquid-cooled load balancing, and air-cooled DVFS) on system performance and die temperature for multi-tiered 3D MPSoCs utilizing SRAM and RRAM-based L2 caches. We find that RRAM-based caches lower overall average maximum temperatures by 120 K and 24 K for air and liquid cooling systems, respectively (when compared to SRAM-based caches), at a worst-case performance delay of 47% and best-case delay of 13% for the parallel shared-memory benchmarks studied. David Brenner, Cory E. Merkel, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 3 |
| 2011 | Reconfigurable N-level memristor memory designabstractMemristive devices have gained significant research attention lately because of their unique properties and wide application spectrum. In particular, memristor-based resistive random access memory (RRAM) offers the high density, low power, and low volatility required for next-generation non-volatile memory. The ability to program memristive devices into several different resistance states has also led to the proposal of multilevel RRAM. This work analyzes the application of thinfilm memristors as N-level RRAM elements. The tradeoffs between the number of memory levels and each RRAM element's reliability will be discussed. A metric is proposed to rate each RRAM element in the presence of process variations. A memory architecture is also presented which allows the number of memory levels to be reconfigured based on different application characteristics. The proposed architecture can achieve a write time speedup of 5.9 over other memristor memory architectures with 80% ion mobility degradation. Cory E. Merkel, Nakul Nagpal, Sindhura Mandalapu, Dhireesha Kudithipudi |
IJCNN | 4 |
| 2010 | Performance enhancement of subthreshold circuits using substrate biasing and charge-boosting buffersabstractSubthreshold circuits are ideal for ultra low power applications. However, they suffer from low operating speeds. By improving the speed of subthreshold circuits their application spectrum can be expanded. In this paper, two existing biasing methods and a new approach to substrate biasing to improve the performance of subthreshold circuits is presented. We derive an approximate expression for the drain current of MOS transistor when substrate biased. We also present a new performance enhancement technique using charge-boosting-buffers to improve the performance of subthreshold circuits. A performance-enhanced standard cell library is built by implementing these techniques on a standard subthreshold cell library. When the enhanced library is applied on ISCAS85 benchmark circuits a 10 times improvement in frequency with an overhead of approximately 2 times in the energy-delay product is observed. Sumanth Amarchinta, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2010 | Variation tolerant 9T SRAM cell designabstractNanoscale SRAM memory design has become increasingly challenging due to the reducing noise margins and increased sensitivity to threshold voltage variations. These issues oppose our ability to achieve stable bitcells and acceptable performance while maintaining density using the standard six-transistor(6T) circuit. To overcome these challenges, researchers have proposed different topologies for SRAMs with single-ended 8T, 9T, 10T bitcell designs. These designs improve the cell stability in the subthreshold regime but suffer from bit-line leakage noise, placing constraints on the number of cells shared by each bitline. In this paper, we propose a novel 9T SRAM cell topology which achieves both cell stability as well as prevents bit-line leakage. With the proposed 9T SRAM circuit, the read static noise margin is nearly twice that of conventional 6T SRAM circuit. Furthermore, the bitline leakage power consumption of the proposed 9T SRAM cell is reduced by up to 79\%, 76\% and 39\% when compared to the previously published 8T, 10T and 9T SRAM cells, respectively. Sreeharsha Tavva, Dhireesha Kudithipudi |
ACM Great Lakes Symposium on VLSI | 2 |
| 2008 | Caches for Multimedia Workloads: Power and Energy TradeoffsabstractOne of the significant workloads in current generation desktop processors and mobile devices is multimedia processing. Large on-chip caches are common in modern processors, but large caches will result in increased power consumption and increased access delays. Regular data access patterns in streaming multimedia applications and video processing applications can provide high hit-rates, but due to issues associated with access time, power and energy, caches cannot be made very large. Characterizing and optimizing the memory system is conducive for designing power and performance efficient multimedia application processors. Performance tradeoffs for multimedia applications have been studied in the past, however, power and energy tradeoffs for caches for multimedia processing have not been adequately studied in the past. In this paper, we characterize multimedia applications for I-cache and D-cache power and energy using a multilevel cache hierarchy. Both dynamic and static power increase with increasing cache sizes, however, the increase in dynamic power is small. The increase in static power is significant, and becomes increasingly relevant for smaller feature sizes. There is significant static power dissipation, ~ 45%, in L1 & L2 caches at 70 nm technology sizes, emphasizing the fact that future multimedia systems must be designed by taking leakage power reduction techniques into account. The energy consumption of on-chip L2 caches is seen to be very sensitive to cache size variations. Sizes larger than 16 k for I-caches and 32 k for D-caches will not be efficient choices to maintain power and performance balance. Since multimedia applications spend significant amounts of time in integer operations, to improve the performance, we propose implementing low power full adders and hybrid multipliers in the data path, which results in 9% to 21% savings in the overall power consumption. Dhireesha Kudithipudi, Stefan Petko, Eugene John |
IEEE Trans. Multim. | 1 |