EDBT 2026 Demo / reviewers in the wild / expert
Matthew J. Marinella
dblp:147/4876
· DBLP profile ↗
19ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-6537-1836ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Agnostic One-Shot Training and Retraining for Robust DNN Inferencing on Analog Compute-in-Memory SystemsabstractAnalog Compute-in-Memory (ACiM) architectures are a promising alternatives to traditional von Neumann-based systems for accelerating deep neural networks (DNNs), as they alleviate the memory bottleneck by performing in-situ matrixvector multiplications. However, the analog nature of computation in ACiM makes DNNs highly susceptible to noise and process variations. To mitigate the effects of analog noise, existing approaches rely on variation-aware or noise-aware training, retraining, or fine-tuning. These methods, however, are not scalable, as they require chip-specific retraining and typically involve separate training runs for different levels of noise tolerance. Moreover, they overlook the inherent fault tolerance of analog-to-digital converters (ADCs). To address these limitations, we propose a one-shot training and retraining strategy for robust DNN inferencing on ACiM platforms. Our method is guided by a detailed analysis of error propagation through ADCs, revealing that robustness can be enhanced by strategically reshaping the weight distribution to better align with ADC resilience characteristics. Simulation results and experimental results with fabricated chips show that the proposed method improves inferencing accuracy by $\mathbf{7 0 \%} \boldsymbol{-} \mathbf{9 0 \%}$ for ResNet-18 and DenseNet-121 under $\mathbf{7 0 \%}$ noise injection on CIFAR-10 and SVHN, and by $\mathbf{5 0 \%}$-80% for VGG-16 under $\mathbf{5 0 \%}$ noise. These gains are achieved with only a $5 \%$ energy overhead due to the modified weight distribution. Ashish Reddy Bommana, Ben Feinberg, T. Patrick Xiao, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty |
ASP-DAC | 5 |
| 2025 | ReSpike: A Co-Design Framework for Evaluating SNNs on ReRAM-Based Neuromorphic Processors
Kazi Asifuzzaman, Aaron R. Young, Prasanna Date, Shruti R. Kulkarni, Narasinga Rao Miniskar, Matthew J. Marinella, Jeffrey S. Vetter |
Euro-Par (2) | 6 |
| 2025 | Fault Tolerance in RRAM-based AI Accelerator with Guided Randomized ActivationabstractResistive Random Access Memory (RRAM)-based analog in-memory computing (IMC) AI accelerators offer significant advantages over digital accelerators, including lower power consumption, reduced data movement, and higher computational efficiency. However, their deployment in safety-critical and edge applications is challenging due to their hardware non-idealities, such as programming error, conductance drift, and read noise, which degrade the inferencing accuracy of the implemented neural networks (NNs). Existing methods, including noise injection during training and activation function modifications, provide limited fault-tolerance in realistic scenarios with non-idealities. We propose a fault-tolerant activation function with architectural optimization that enhances robustness against hardware-induced variations with minimal hardware and NN architectural changes. During training, the proposed activation function features a stochastic negative region, which inherently injects noise into the negative region of the activation. During inferencing, the proposed activation function operates deterministically, ensuring compatibility with existing hardware while maintaining computational efficiency. Extensive evaluations with benchmark datasets demonstrate that the proposed approach significantly improves inferencing accuracy by up to 60% under varying noise levels, outperforming conventional activation functions as well as existing fault-tolerant activation functions. By enhancing fault-tolerance to hardware-induced errors, the proposed method enables reliable and energy-efficient RRAM-based analog IMC. Soyed Tuhin Ahmed, Eduardo Ortega, Ryan Depsey, T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty |
ITC | 7 |
| 2024 | A Discovery Platform to Characterize Emerging Nonvolatile Memories for ComputingabstractMemory-centric architectures such as analog in memory computing (IMC) offer the potential for orders of magnitude improvements in energy efficiency and performance beyond state of the art. These architectures perform computations such as multiply-accumulate directly within memory array circuitry. Analog IMC and related architectures create markedly different requirements for memory devices than those of digital systems and a wide array of emerging memory candidate devices have been proposed to best meet these requirements. Accurate assessment of candidate device suitability requires characterizing the behavior in CMOS-integrated arrays, closely representing operation in a real IMC system. To address this, we have developed an analog memory array characterization platform that enables the detailed electrical characterization and optimization of these candidate memory device arrays, allowing accurate modeling and prediction of their behavior in IMC systems. D. Wilson, Nad E. Gilbert, Matthew Spear, J. Short, Christopher H. Bennett, William Wahby, Joshua E. Kim, Robin Jacobs-Gedrim, T. Patrick Xiao, Sapan Agarwal, Matthew J. Marinella |
VTS | 11 |
| 2022 | Purely Spintronic Leaky Integrate-and-Fire NeuronsabstractNeuromorphic computing promises revolutionary improvements over conventional systems for applications that process unstructured information. To fully realize this potential, neuromorphic systems should exploit the biomimetic behavior of emerging nanodevices. In particular, exceptional opportunities are provided by the non-volatility and analog capabilities of spintronic devices. While spintronic devices that emulate neurons have been previously proposed, they require complementary metal-oxide semiconductor (CMOS) technology to function. In turn, this significantly increases the power consumption, fabrication complexity, and device area of a single neuron. This work reviews three previously proposed CMOS-free spintronic neurons designed to resolve this issue. Wesley H. Brigner, Naimul Hassan, Xuan Hu 0002, Christopher H. Bennett, Felipe García-Sánchez, Matthew J. Marinella, Jean Anne C. Incorvia, Joseph S. Friedman |
ISCAS | 6 |
| 2022 | Self-correcting Flip-flops for Triple Modular Redundant Logic in a 12-nm TechnologyabstractArea efficient self-correcting flip-flops for use with triple modular redundant (TMR) soft-error hardened logic are implemented in a 12-nm finFET process technology. The TMR flip-flop slave latches self-correct in the clock low phase using Muller C-elements in the latch feedback. These C-elements are driven by the two redundant stored values and not by the slave latch itself, saving area over a similar implementation using majority gate feedback. These flip-flops are implemented as large shift-register arrays on a test chip and have been experimentally tested for their soft-error mitigation in static and dynamic modes of operation using heavy ions and protons. We show how high clock skew can result in susceptibility to soft-errors in the dynamic mode, and explain the potential failure mechanism. Lawrence T. Clark, Alen Duvnjak, Clifford Young-Sciortino, Matthew Cannon, John S. Brunhaver, Sapan Agarwal, Jereme Neuendank, Donald Wilson, Hugh J. Barnaby, Matthew J. Marinella |
ISCAS | 10 |
| 2022 | Intrinsic Lateral Inhibition Facilitates Winner-Take-All in Domain Wall Racetrack Arrays for Neuromorphic ComputingabstractNeuromorphic computing is a promising candidate for beyond-von Neumann computer architectures, featuring low power consumption and high parallelism. Lateral inhibition and winner-take-all (WTA) features play a crucial role in neuronal competition of the nervous system as well as neuromorphic hardwares. The domain wall - magnetic tunnel junction (DWMTJ) neuron is an emerging spintronic artificial neuron device exhibiting intrinsic lateral inhibition. In this paper we show that lateral inhibition parameters modulate the neuron firing statistics in a DW-MTJ neuron array, thus emulating soft-winner-take-all (WTA) and firing group selection. Can Cui 0020, Otitoaleke G. Akinola, Naimul Hassan, Christopher H. Bennett, Matthew J. Marinella, Joseph S. Friedman, Jean Anne C. Incorvia |
ISCAS | 5 |
| 2022 | Eris: Fault Injection and Tracking Framework for Reliability Analysis of Open-Source HardwareabstractAs transistors have been scaled over the past decade, modern systems have become increasingly susceptible to faults. Increased transistor densities and lower capacitances make a particle strike more likely to cause an upset. At the same time, complex computer systems are increasingly integrated into safety-critical systems such as autonomous vehicles. These two trends make the study of system reliability and fault tolerance essential for modern systems. To analyze and improve system reliability early in the design process, new tools are needed for RTL fault analysis.This paper proposes Eris, a novel framework to identify vulnerable components in hardware designs through fault-injection and fault propagation tracking. Eris builds on ESSENT—a fast C/C++ RTL simulation framework—to provide fault injection, fault tracking, and control-flow deviation detection capabilities for RTL designs. To demonstrate Eris’ capabilities, we analyze the reliability of the open source Rocket Chip SoC by randomly injecting faults during thousands of runs on four microbenchmarks. As part of this analysis we measure the sensitivity of different hardware structures to faults based on the likelihood of a random fault causing silent data corruption, unrecoverable data errors, program crashes, and program hangs. We detect control flow deviations and determine whether or not they are benign. Additionally, using Eris’ novel fault-tracking capabilities we are able to find 78% more vulnerable components in the same number of simulations compared to RTL-based fault injection techniques without these capabilities. We will release Eris as an open-source tool to aid future research into processor reliability and hardening. Shubham Nema, Justin Kirschner, Debpratim Adak, Sapan Agarwal, Ben Feinberg, Arun Rodrigues, Matthew J. Marinella, Amro Awad |
ISPASS | 7 |
| 2022 | An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog MemoryabstractWe demonstrate SONOS (silicon-oxide-nitride-oxide-silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a$> 10\times $gain in energy efficiency over state-of-the-art digital and analog inference accelerators. T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Vineet Agrawal, Prashant Saxena, Venkatraman Prabhakar, Krishnaswamy Ramkumar, Harsha Medu, Ramesh Chettuvetty, Sapan Agarwal, Matthew J. Marinella |
IEEE Trans. Circuits Syst. I Regul. Pap. | 12 |
| 2021 | An Analog Preconditioner for Solving Linear SystemsabstractOver the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively. Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek |
HPCA | 7 |
| 2020 | Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online LearningabstractMachine learning implements backpropagation via abundant training samples. We demonstrate a multi-stage learning system realized by a promising non-volatile memory device, the domain-wall magnetic tunnel junction (DW-MTJ). The system consists of unsupervised (clustering) as well as supervised sub-systems, and generalizes quickly (with few samples). We demonstrate interactions between physical properties of this device and optimal implementation of neuroscience-inspired plasticity learning rules, and highlight performance on a suite of tasks. Our energy analysis confirms the value of the approach, as the learning budget stays below 20μJ even for large tasks used typically in machine learning. Christopher H. Bennett, T. Patrick Xiao, Can Cui 0020, Naimul Hassan, Otitoaleke G. Akinola, Jean Anne C. Incorvia, Alvaro Velasquez, Joseph S. Friedman, Matthew J. Marinella |
ISCAS | 9 |
| 2020 | CMOS-Free Magnetic Domain Wall Leaky Integrate-and-Fire Neurons with Intrinsic Lateral InhibitionabstractSpintronic devices, especially those based on motion of a domain wall (DW) through a ferromagnetic track, have received a significant amount of interest in the field of neuromorphic computing because of their non-volatility and intrinsic current integration capabilities. Many spintronic neurons using this technology have already been proposed, but they also require external circuitry or additional device layers to implement other important neuronal behaviors. Therefore, they result in an increase in fabrication complexity and/or energy consumption. In this work, we discuss three neurons that implement these functions without the use of additional circuitry or material layers. Naimul Hassan, Wesley H. Brigner, Xuan Hu 0002, Otitoaleke G. Akinola, Christopher H. Bennett, Matthew J. Marinella, Felipe García-Sánchez, Jean Anne C. Incorvia, Joseph S. Friedman |
ISCAS | 6 |
| 2020 | Process Variation Model and Analysis for Domain Wall-Magnetic Tunnel Junction LogicabstractThe domain wall-magnetic tunnel junction (DW-MTJ) is a spintronic device that enables efficient logic circuit design because of its low energy consumption, small size, and non-volatility. Furthermore, the DW-MTJ is one of the few spintronic devices for which a direct cascading mechanism is experimentally demonstrated without any extra buffers; this enables potential design and fabrication of a large-scale DW-MTJ logic system. However, DW-MTJ logic relies on the conversion between electrical signals and magnetic states which is sensitive to process imperfection. Therefore, it is important to analyze the robustness of such DW-MTJ devices to anticipate the system reliability before fabrication. Here we propose a new DW-MTJ model that integrates the impacts of process variation to enable the analysis and optimization of DW-MTJ logic. This will allow circuit and device design that enhances the robustness of DW-MTJ logic and advances the development of energy-efficient spintronic computing systems. Xuan Hu 0002, Alexander J. Edwards, T. Patrick Xiao, Christopher H. Bennett, Jean Anne C. Incorvia, Matthew J. Marinella, Joseph S. Friedman |
ISCAS | 6 |
| 2020 | PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAMabstractThe wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively. Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Sapan Agarwal, Matthew J. Marinella, Martin Foltin, John Paul Strachan, Dejan S. Milojicic, Wen-Mei W. Hwu, Kaushik Roy 0001 |
IEEE Trans. Computers | 5 |
| 2020 | Memristor Model Optimization Based on Parameter Extraction From Device Characterization DataabstractThis paper presents a memristive device model capable of accurately matching a wide range of characterization data collected from a tantalum oxide memristor. Memristor models commonly use a set of equations and fitting parameters to match the complex dynamic conductivity pattern observed in these devices. Along with the proposed model, a procedure is also described that can be used to optimize each fitting parameter in the model relative to an I-V curve. Therefore, model parameters are self-updated based on this procedure when a new cyclic I-V sweep is provided for model optimization. This model will automatically provide the best possible match to the characterization data without any additional optimization from the user. In this paper, multiple cyclic I-V characterizations are modeled from ten different tantalum oxide devices (on the same wafer). Additionally, studies were completed to demonstrate the amount of variation present between devices on a wafer, as well as the amount of variation present within a single device. Methods for modeling this variation are then proposed, resulting in an accurate and complete, automated, memristor modeling approach. Chris Yakopcic, Tarek M. Taha, David J. Mountain, Thomas Salter, Matthew J. Marinella, Mark R. McLean |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Ziksa: On-chip learning accelerator with memristor crossbars for multilevel neural networksabstractMemristor crossbars support efficient realizations of spiking and non-spiking neural networks designs. In most of these designs off-chip/ex-situ training is used to set/update the state of the memrisitve devices. However, there is a growing need to design an efficient on-chip/in-situ learning for mobile autonomous systems. In this research, we propose an on-chip learning accelerator, known as Ziksa, that is integrated with the memristor crossbars. We demonstrate how regression and back-propagation in multi-level networks can be realized through Ziksa. The proposed accelerator is evaluated on a fabricated TiN-TaOx-TaTiN memristor crossbar. A 3-layer feedforward network was tested using Ziksa for classification. An accuracy of 95.3% was achieved on Wisconsin breast cancer dataset. The proposed learning accelerator can be envisioned as a core building block in a wide-range of cognitive algorithms that rely on on-chip online learning. Abdullah M. Zyarah, Nicholas Soures, Lydia Hays, Robin Jacobs-Gedrim, Sapan Agarwal, Matthew J. Marinella, Dhireesha Kudithipudi |
ISCAS | 6 |
| 2016 | Resistive memory device requirements for a neural algorithm acceleratorabstractResistive memories enable dramatic energy reductions for neural algorithms. We propose a general purpose neural architecture that can accelerate many different algorithms and determine the device properties that will be needed to run backpropagation on the neural architecture. To maintain high accuracy, the read noise standard deviation should be less than 5% of the weight range. The write noise standard deviation should be less than 0.4% of the weight range and up to 300% of a characteristic update (for the datasets tested). Asymmetric nonlinearities in the change in conductance vs pulse cause weight decay and significantly reduce the accuracy, while moderate symmetric nonlinearities do not have an effect. In order to allow for parallel reads and writes the write current should be less than 100 nA as well. Sapan Agarwal, Steven J. Plimpton, David R. Hughart, Alexander H. Hsia, Isaac Richter, Jonathan A. Cox, Conrad D. James, Matthew J. Marinella |
IJCNN | 8 |
| 2015 | Reconfigurable Memristive Device TechnologiesabstractIn this paper, we present a review of the state of the art in memristor technologies. Along with ionic conducting devices [i.e., conductive bridging random access memory (CBRAM)], we include phase change, and organic/organo-metallic technologies, and we review the most recent advances in oxide-based memristor technologies. We present progress on 3-D integration techniques, and we discuss the behavior of more mature memristive technologies in extreme environments. Arthur H. Edwards, Hugh J. Barnaby, Kristy A. Campbell, Michael N. Kozicki, Wei Liu 0003, Matthew J. Marinella |
Proc. IEEE | 6 |
| 2014 | Emerging resistive switching memory technologies: Overview and current statusabstractResistive memory technologies, in particular redox random access memory (ReRAM), are poised as one of the most prominent emerging memory categories to replace NAND flash and fill the important need for a Storage Class Memory (SCM). This is due to low switching energy, low current switching, high speed, outstanding endurance, scalability below 10 nm, and excellent back-end-of-line CMOS compatibility. Furthermore, the analog aspects of memristors have opened the door for many novel applications such as analog math accelerators and neuromorphic computers. This paper provides an overview of resistive memory technologies and their current status, with a focus on redox RAM (ReRAM). Matthew J. Marinella |
ISCAS | 1 |