EDBT 2026 Demo / reviewers in the wild / expert
Nathaniel C. Cady
dblp:52/9430
· DBLP profile ↗
14ranked-venue papers
1as first author
7since 2021 · last 2023
0000-0003-4345-3627ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | RFAM: RESET-Failure-Aware-Model for HfO2-based Memristor to Enhance the Reliability of Neuromorphic DesignabstractMemristors are a suitable candidate to design synapse circuits and neuromorphic systems. Due to device and voltage variability, operating a memristive device with reliability is a big challenge. To enhance the reliability of memristive synapse, RESET failure needs to be considered. In this work, we are focused on RESET failure modeling with RESET voltage variation. Here, the RESET failure is defined as hard failure of the memristive synapse due to a high RESET voltage being applied. The proposed Verilog-A model is derived based on experimental data collected from 1T1R devices, which are fabricated on 65 nm CMOS process. To enhance the reliability of system-level simulation, this device model will provide better guidelines to the designer. In addition, power consumption for a successful RESET operation is 7.065 μW at 1.5 V, which can RESET the memristor resistance from 5 kΩ to 200 kΩ. Hritom Das, Manu Rathore, Rocco D. Febbo, Maximilian Liehr, Nathaniel C. Cady, Garrett S. Rose |
ACM Great Lakes Symposium on VLSI | 5 |
| 2023 | An Efficient and Accurate Memristive Memory for Array-Based Spiking Neural NetworksabstractMemristors provide a tempting solution for weighted synapse connections in neuromorphic computing due to their size and non-volatile nature. However, memristors are unreliable in the commonly used voltage-pulse-based programming approaches and require precisely shaped pulses to avoid programming failure. In this paper, we demonstrate a current-limiting-based solution that provides a more predictable analog memory behavior when reading and writing memristive synapses. With our proposed design READ current can be optimized by ~19x compared to the 1T1R design. Moreover, our proposed design saves ~9x energy compared to the 1T1R design. Our 3T1R design also shows promising write operation which is less affected by the process variation in MOSFETs and the inherent stochastic behavior of memristors. Memristors used for testing are hafnium oxide based and were fabricated in a 65 nm hybrid CMOS-memristor process. The proposed design also shows linear characteristics between the voltage applied and the resulting resistance for the writing operation. The simulation and measured data show similar patterns with respect to voltage pulse based programming and current compliance based programming. We further observed the impact of this behavior on neuromorphic-specific applications such as a spiking neural network. Hritom Das, Rocco D. Febbo, Sree Nirmillo Biswash Tushar, Nishith N. Chakraborty, Maximilian Liehr, Nathaniel C. Cady, Garrett S. Rose |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Work-in-Progress: A Processing-in-Pixel Accelerator based on Multi-level HfOx ReRAMabstractThis work paves the way to realize a processing-in-pixel accelerator based on a multi-level HfOxReRAM as a flexible, energy-efficient, and high-performance solution for real-time and smart image processing at edge devices. The proposed design intrinsically implements and supports a coarse-grained convolution operation in low-bit-width neural networks leveraging a novel compute-pixel with non-volatile weight storage at the sensor side. Our evaluations show that such a design can remarkably reduce the power consumption of data conversion and transmission to an off-chip processor maintaining accuracy compared with the recent in-sensor computing designs. Minhaz Abedin, Arman Roohi, Nathaniel C. Cady, Shaahin Angizi |
CASES | 3 |
| 2022 | Exploring Model Stability of Deep Neural Networks for Reliable RRAM-Based In-Memory AccelerationabstractRRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs). Furthermore, model compression techniques, such as quantization and pruning, are necessary to improve algorithm mapping and hardware performance. However, in the presence of RRAM device variations, low-precision and sparse DNNs suffer from severe post-mapping accuracy loss. To address this, in this work, we investigate a new metric,model stability, from the loss landscape to help shed light on accuracy loss under variations and model compression, which guides an algorithmic solution to maximize model stability and mitigate accuracy loss. Based on statistical data from a CMOS/RRAM 1T1R test chip at 65nm, we characterize wafer-level RRAM variations and develop a cross-layer benchmark tool that incorporates quantization, pruning, device variations, model stability, and IMC architecture parameters to assess post-mapping accuracy and hardware performance. Leveraging this tool, we show that a loss-landscape-based DNN model selection for stability effectively tolerates device variations and achieves a post-mapping accuracy higher than that with 50% lower RRAM variations. Moreover, we quantitatively interpret why model pruning increases the sensitivity to variations, while a lower-precision model has better tolerance to variations. Finally, we propose a novel variation-aware training method to improve model stability, in which there exists the most stable model for the best post-mapping accuracy of compressed DNNs. Experimental evaluation of the method shows up to 19%, 21%, and 11% post-mapping accuracy improvement for our 65nm RRAM device, across various precision and sparsity, on CIFAR-10, CIFAR-100, and SVHN datasets, respectively. Li Yang 0009, Jingbo Sun 0003, Jubin Hazra, Xiaocong Du, Maximilian Liehr, Zheng Li 0020, Karsten Beckmann, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Yu Cao 0001 |
IEEE Trans. Computers | 10 |
| 2022 | Hybrid RRAM/SRAM in-Memory Computing for Robust DNN AccelerationabstractRRAM-based in-memory computing (IMC) effectively accelerates deep neural networks (DNNs) and other machine learning algorithms. On the other hand, in the presence of RRAM device variations and lower precision, the mapping of DNNs to RRAM-based IMC suffers from severe accuracy loss. In this work, we propose a novel hybrid IMC architecture that integrates an RRAM-based IMC macro with a digital SRAM macro using a programmable shifter to compensate for the RRAM variations and recover the accuracy. The digital SRAM macro consists of a small SRAM memory array and an array of multiply-and-accumulate (MAC) units. The nonideal output from the RRAM macro, due to device and circuit nonidealities, is compensated by adding the precise output from the SRAM macro. In addition, the programmable shifter allows for different scales of compensation by shifting the SRAM macro output relative to the RRAM macro output. On the algorithm side, we develop a framework for the training of DNNs to support the hybrid IMC architecture through ensemble learning. The proposed framework performs quantization (weights and activations), pruning, RRAM IMC-aware training, and employs ensemble learning through different compensation scales by utilizing the programmable shifter. Finally, we design a silicon prototype of the proposed hybrid IMC architecture in the 65-nm SUNY process to demonstrate its efficacy. Experimental evaluation of the hybrid IMC architecture shows that the SRAM compensation allows for a realistic IMC architecture with multilevel RRAM cells (MLCs) even though they suffer from high variations. The hybrid IMC architecture achieves up to 21.9%, 12.65%, and 6.52% improvement in post-mapping accuracy over state-of-the-art techniques, at minimal overhead, for ResNet-20 on CIFAR-10, VGG-16 on CIFAR-10, and ResNet-18 on ImageNet, respectively. Zhenyu Wang 0016, Injune Yeo, Li Yang 0009, Jian Meng, Maximilian Liehr, Rajiv V. Joshi, Nathaniel C. Cady, Deliang Fan, Jae-sun Seo, Yu Cao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2021 | Stochasticity and robustness in spiking neural networks
Wilkie Olin-Ammentorp, Karsten Beckmann, Catherine D. Schuman, James S. Plank, Nathaniel C. Cady |
Neurocomputing | 5 |
| 2021 | Investigation of ReRAM Variability on Flow-Based Edge Detection Computing Using HfO2-Based ReRAM ArraysabstractResistive random-access memory (ReRAM) memristors are promising candidates for various compute in memory and flow-based computing approaches. As an alternative to traditional von Neumann computation, flow-based computing avoids serial movement of data between memory and processor. In this paper, we demonstrate arrays of 1 transistor 1 ReRAM (1T1R) to detect edges between 8 bit pixels using flow-based computing, and the effects of stochastic variation of ReRAM on edge detection outputs. Three different tRoff/Ronresistance ratios (1.5:1, 2.5:1 or 28.6:1) were utilized to implement multiple flow-based edge detection computation matrices for 8 bit pixels. Edge detection was distinguishable for all Roff/Ronratios used, for all flow-based computing matrices. However, the binary output resistance ratio of the matrices improved 3-fold when the patterned Roff/Ronratio was increased to 28.6:1. A Gaussian simulation of ReRAM resistance variability validates the experimental data, with a correlation coefficient (r) of 0.9547. These results suggest a trade-off between the flow-based edge detection output ratio and the variability of the ReRAM resistance in Roff/Ronresistance ratio. Sarah Rafiq, Jubin Hazra, Maximilian Liehr, Karsten Beckmann, Minhaz Abedin, Jodh S. Pannu, Sumit Kumar Jha 0001, Nathaniel C. Cady |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2020 | Accurate Inference with Inaccurate RRAM Devices: Statistical Data, Model Transfer, and On-line AdaptationabstractResistive random-access memory (RRAM) is a promising technology for in-memory computing with high storage density, fast inference, and good compatibility with CMOS. However, the mapping of a pre-trained deep neural network (DNN) model on RRAM suffers from realistic device issues, especially the variation and quantization error, resulting in a significant reduction in inference accuracy. In this work, we first extract these statistical properties from 65 nm RRAM data on 300mm wafers. The RRAM data present 10-levels in quantization and 50% variance, resulting in an accuracy drop to 31.76% and 10.49% for MNIST and CIFAR-10 datasets, respectively. Based on the experimental data, we propose a combination of machine learning algorithms and on-line adaptation to recover the accuracy with the minimum overhead. The recipe first applies Knowledge Distillation (KD) to transfer an ideal model into a student model with statistical variations and 10 levels. Furthermore, an on-line sparse adaptation (OSA) method is applied to the DNN model mapped on to the RRAM array. Using importance sampling, OSA adds a small SRAM array that is sparsely connected to the main RRAM array; only this SRAM array is updated to recover the accuracy. As demonstrated on MNIST and CIFAR-10 datasets, a 7.86% area cost is sufficient to achieve baseline accuracy for the 65 nm RRAM devices. Gouranga Charan, Jubin Hazra, Karsten Beckmann, Xiaocong Du, Rajiv V. Joshi, Nathaniel C. Cady, Yu Cao 0001 |
DAC | 7 |
| 2020 | Towards Synaptic Behavior of Nanoscale ReRAM Devices for Neuromorphic Computing ApplicationsabstractResistive Random Access Memory (ReRAM), a form of non-volatile memory, has been proposed as a Flash memory replacement. In addition, novel circuit architectures have been proposed that rely on newly discovered or predicted behavior of ReRAM. One such architecture is the memristive Dynamic Adaptive Neural Network Array, developed to emulate the functionality of a biological neuron system. We demonstrated ReRAM devices that show a synaptic tendency by changing their resistance in an analog fashion. The CMOS compatible nanoscale ReRAM devices shown are based on an HfO2switching layer that sits on a tungsten electrode and is covered by a titanium oxygen scavenger layer and a titanium nitride top electrode. In this work, we showed devices exceeding endurance values of 10B cycles with a discrete Roff/Ronratio of 15. Multi-level states were achieved by using consecutive ultra-short 5/1.5 ns pulses during the reset operation. A neural network simulation was performed in which the synaptic weights were perturbed with the ReRAM variability, which was extracted from two different characterization methods: (1) via direct write, and (2) via a write/read verification approach during the reset operation. A substantial improvement of the neural network fitness was demonstrated when using the write/read verification approach. Karsten Beckmann, Wilkie Olin-Ammentorp, Gangotree Chakma, Sherif Amer, Garrett S. Rose, Chris Hobbs, Joseph Van Nostrand, Martin Rodgers, Nathaniel C. Cady |
ACM J. Emerg. Technol. Comput. Syst. | 9 |
| 2018 | Design Considerations for Memristive Crossbar Physical Unclonable FunctionsabstractHardware security has emerged as a field concerned with issues such as integrated circuit (IC) counterfeiting, cloning, piracy, and reverse engineering. Physical unclonable functions (PUF) are hardware security primitives useful for mitigating such issues by providing hardware-specific fingerprints based on intrinsic process variations within individual IC implementations. As technology scaling progresses further into the nanometer region, emerging nanoelectronic technologies, such as memristors or RRAMs (resistive random-access memory), have become interesting options for emerging computing systems. In this article, using a comprehensive temperature dependent model of an HfO x (hafnium-oxide) memristor, based on experimental measurements, we explore the best region of operation for a memristive crossbar PUF (XbarPUF). The design considered also employs XORing and a column shuffling technique to improve reliability and resilience to machine learning attacks. We present a detailed analysis for the noise margin and discuss the scalability of the XbarPUF structure. Finally, we present results for estimates of area, power, and delay alongside security performance metrics to analyze the strengths and weaknesses of the XbarPUF. Our XbarPUF exhibits nearly ideal (near 50%) uniqueness, bit-aliasing and uniformity, good reliability of 90% and up (with 100% being ideal), a very small footprint, and low average power consumption ≈104μW. Mesbah Uddin, Md. Badruddoja Majumder, Karsten Beckmann, Harika Manem, Zahiruddin Alamgir, Nathaniel C. Cady, Garrett S. Rose |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2017 | A practical hafnium-oxide memristor model suitable for circuit design and simulationabstractThis paper proposes a practical polynomial model for HfO2memristor fabricated in-house at SUNY Polytechnic Institute. Although there is no shortage of memristor models in the literature, most models are not general and assume specific switching and conduction mechanisms. This is often deemed impractical for circuit designers who wish to develop a model for a specific technology of interest. Thus, circuit designers have sought empirical models that are easily fit to their specific device. The model should be simple, intuitive, and most importantly, fast to converge. The proposed model is based on measurable parameters and matches the experimental data well. The convergence of our model is tested against other models in the literature and shows comparable results. It is also shown that the smoothness of the model around the memristor threshold is critical for fast convergence time. Sherif Amer, Sagarvarma Sayyaparaju, Garrett S. Rose, Karsten Beckmann, Nathaniel C. Cady |
ISCAS | 5 |
| 2016 | Flow-based computing on nanoscale crossbars: Design and implementation of full addersabstractWe present the design and implementation of a full adder circuit that exploits the natural flow of current through nanowires and More-than-Moore nano-devices in two dimensional crossbars. We evaluate the speed and energy efficiency of our design and compare it to equivalent one-bit adder designs using CMOS and nanoscale memristors. Our memristive full adder circuit has been shown to be an order of magnitude faster and more energy-efficient than equivalent CMOS designs. Our circuit is an order of magnitude more compact that equivalent CMOS designs. We also argue that our design occupies less area and is faster than competing memristor designs. Zahiruddin Alamgir, Karsten Beckmann, Nathaniel C. Cady, Alvaro Velasquez, Sumit Kumar Jha 0001 |
ISCAS | 3 |
| 2015 | An extendable multi-purpose 3D neuromorphic fabric using nanoscale memristorsabstractNeuromorphic computing offers an attractive means for processing and learning complex real-world data. With the emergence of the memristor, the physical realization of cost-effective artificial neural networks is becoming viable, due to reduced area and increased performance metrics than strictly CMOS implementations. In the work presented here, memristors are utilized as synapses in the realization of a multi-purpose heterogeneous 3D neuromorphic fabric. This paper details our in-house memristor and 3D technologies in the design of a fabric that can perform real-world signal processing (i.e., image/video etc.) as well as everyday Boolean logic applications. The applicability of this fabric is therefore diverse with applications ranging from general-purpose and high performance logic computing to power-conservative image detection for mobile and defense applications. The proposed system is an area-effective heterogeneous 3D integration of memristive neural networks, that consumes significantly less power and allows for high speeds (3D ultra-high bandwidth connectivity) in comparison to a purely CMOS 2D implementation. Images and results provided will illustrate our state of the art 3D and memristor technology capabilities for the realization of the proposed 3D memristive neural fabric. Simulation results also show the results for mapping Boolean logic functions and images onto perceptron based neural networks. Results demonstrate the proof of concept of this system, which is the first step in the physical realization of the multi-purpose heterogeneous 3D memristive neuromorphic fabric. Harika Manem, Karsten Beckmann, Robert Carroll, Robert E. Geer, Nathaniel C. Cady |
CISDA | 6 |
| 2010 | Biologically self-assembled memristive circuit elementsabstractBoth TiO2and HfO2are common materials for semiconductor fabrication that have shown memristive properties. Nanoparticles of these metal oxides have great prospect to provide nanoscale materials with tunable electronic properties for integration in advanced circuits such as neuromorphic networks based on memristive crossbar elements. We seek to take advantage of the unique interaction between the phosphate end-group of DNA and TiO2/HfO2nanoparticles to enable a guided assembly of circuit elements via specific nucleotide sequences in a bottom-up fashion. Nathaniel C. Cady, Magnus Bergkvist, Nicholas M. Fahrenkopf, Phillip Z. Rice, Joseph Van Nostrand |
ISCAS | 1 |