Christopher H. Bennett

dblp:166/7012 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-6989-292XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
YearPublicationVenuePosition
2026 Noise-Agnostic One-Shot Training and Retraining for Robust DNN Inferencing on Analog Compute-in-Memory Systems
abstract
Analog Compute-in-Memory (ACiM) architectures are a promising alternatives to traditional von Neumann-based systems for accelerating deep neural networks (DNNs), as they alleviate the memory bottleneck by performing in-situ matrixvector multiplications. However, the analog nature of computation in ACiM makes DNNs highly susceptible to noise and process variations. To mitigate the effects of analog noise, existing approaches rely on variation-aware or noise-aware training, retraining, or fine-tuning. These methods, however, are not scalable, as they require chip-specific retraining and typically involve separate training runs for different levels of noise tolerance. Moreover, they overlook the inherent fault tolerance of analog-to-digital converters (ADCs). To address these limitations, we propose a one-shot training and retraining strategy for robust DNN inferencing on ACiM platforms. Our method is guided by a detailed analysis of error propagation through ADCs, revealing that robustness can be enhanced by strategically reshaping the weight distribution to better align with ADC resilience characteristics. Simulation results and experimental results with fabricated chips show that the proposed method improves inferencing accuracy by $\mathbf{7 0 \%} \boldsymbol{-} \mathbf{9 0 \%}$ for ResNet-18 and DenseNet-121 under $\mathbf{7 0 \%}$ noise injection on CIFAR-10 and SVHN, and by $\mathbf{5 0 \%}$-80% for VGG-16 under $\mathbf{5 0 \%}$ noise. These gains are achieved with only a $5 \%$ energy overhead due to the modified weight distribution.
Ashish Reddy Bommana, Ben Feinberg, T. Patrick Xiao, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty
ASP-DAC4
2025 Fault Tolerance in RRAM-based AI Accelerator with Guided Randomized Activation
abstract
Resistive Random Access Memory (RRAM)-based analog in-memory computing (IMC) AI accelerators offer significant advantages over digital accelerators, including lower power consumption, reduced data movement, and higher computational efficiency. However, their deployment in safety-critical and edge applications is challenging due to their hardware non-idealities, such as programming error, conductance drift, and read noise, which degrade the inferencing accuracy of the implemented neural networks (NNs). Existing methods, including noise injection during training and activation function modifications, provide limited fault-tolerance in realistic scenarios with non-idealities. We propose a fault-tolerant activation function with architectural optimization that enhances robustness against hardware-induced variations with minimal hardware and NN architectural changes. During training, the proposed activation function features a stochastic negative region, which inherently injects noise into the negative region of the activation. During inferencing, the proposed activation function operates deterministically, ensuring compatibility with existing hardware while maintaining computational efficiency. Extensive evaluations with benchmark datasets demonstrate that the proposed approach significantly improves inferencing accuracy by up to 60% under varying noise levels, outperforming conventional activation functions as well as existing fault-tolerant activation functions. By enhancing fault-tolerance to hardware-induced errors, the proposed method enables reliable and energy-efficient RRAM-based analog IMC.
Soyed Tuhin Ahmed, Eduardo Ortega, Ryan Depsey, T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty
ITC6
2024 A Discovery Platform to Characterize Emerging Nonvolatile Memories for Computing
abstract
Memory-centric architectures such as analog in memory computing (IMC) offer the potential for orders of magnitude improvements in energy efficiency and performance beyond state of the art. These architectures perform computations such as multiply-accumulate directly within memory array circuitry. Analog IMC and related architectures create markedly different requirements for memory devices than those of digital systems and a wide array of emerging memory candidate devices have been proposed to best meet these requirements. Accurate assessment of candidate device suitability requires characterizing the behavior in CMOS-integrated arrays, closely representing operation in a real IMC system. To address this, we have developed an analog memory array characterization platform that enables the detailed electrical characterization and optimization of these candidate memory device arrays, allowing accurate modeling and prediction of their behavior in IMC systems.
D. Wilson, Nad E. Gilbert, Matthew Spear, J. Short, Christopher H. Bennett, William Wahby, Joshua E. Kim, Robin Jacobs-Gedrim, T. Patrick Xiao, Sapan Agarwal, Matthew J. Marinella
VTS5
2022 Purely Spintronic Leaky Integrate-and-Fire Neurons
abstract
Neuromorphic computing promises revolutionary improvements over conventional systems for applications that process unstructured information. To fully realize this potential, neuromorphic systems should exploit the biomimetic behavior of emerging nanodevices. In particular, exceptional opportunities are provided by the non-volatility and analog capabilities of spintronic devices. While spintronic devices that emulate neurons have been previously proposed, they require complementary metal-oxide semiconductor (CMOS) technology to function. In turn, this significantly increases the power consumption, fabrication complexity, and device area of a single neuron. This work reviews three previously proposed CMOS-free spintronic neurons designed to resolve this issue.
Wesley H. Brigner, Naimul Hassan, Xuan Hu 0002, Christopher H. Bennett, Felipe García-Sánchez, Matthew J. Marinella, Jean Anne C. Incorvia, Joseph S. Friedman
ISCAS4
2022 Intrinsic Lateral Inhibition Facilitates Winner-Take-All in Domain Wall Racetrack Arrays for Neuromorphic Computing
abstract
Neuromorphic computing is a promising candidate for beyond-von Neumann computer architectures, featuring low power consumption and high parallelism. Lateral inhibition and winner-take-all (WTA) features play a crucial role in neuronal competition of the nervous system as well as neuromorphic hardwares. The domain wall - magnetic tunnel junction (DWMTJ) neuron is an emerging spintronic artificial neuron device exhibiting intrinsic lateral inhibition. In this paper we show that lateral inhibition parameters modulate the neuron firing statistics in a DW-MTJ neuron array, thus emulating soft-winner-take-all (WTA) and firing group selection.
Can Cui 0020, Otitoaleke G. Akinola, Naimul Hassan, Christopher H. Bennett, Matthew J. Marinella, Joseph S. Friedman, Jean Anne C. Incorvia
ISCAS4
2022 An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory
abstract
We demonstrate SONOS (silicon-oxide-nitride-oxide-silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a$> 10\times $gain in energy efficiency over state-of-the-art digital and analog inference accelerators.
T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Vineet Agrawal, Prashant Saxena, Venkatraman Prabhakar, Krishnaswamy Ramkumar, Harsha Medu, Ramesh Chettuvetty, Sapan Agarwal, Matthew J. Marinella
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 An Analog Preconditioner for Solving Linear Systems
abstract
Over the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively.
Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek
HPCA4
2020 Plasticity-Enhanced Domain-Wall MTJ Neural Networks for Energy-Efficient Online Learning
abstract
Machine learning implements backpropagation via abundant training samples. We demonstrate a multi-stage learning system realized by a promising non-volatile memory device, the domain-wall magnetic tunnel junction (DW-MTJ). The system consists of unsupervised (clustering) as well as supervised sub-systems, and generalizes quickly (with few samples). We demonstrate interactions between physical properties of this device and optimal implementation of neuroscience-inspired plasticity learning rules, and highlight performance on a suite of tasks. Our energy analysis confirms the value of the approach, as the learning budget stays below 20μJ even for large tasks used typically in machine learning.
Christopher H. Bennett, T. Patrick Xiao, Can Cui 0020, Naimul Hassan, Otitoaleke G. Akinola, Jean Anne C. Incorvia, Alvaro Velasquez, Joseph S. Friedman, Matthew J. Marinella
ISCAS1
2020 CMOS-Free Magnetic Domain Wall Leaky Integrate-and-Fire Neurons with Intrinsic Lateral Inhibition
abstract
Spintronic devices, especially those based on motion of a domain wall (DW) through a ferromagnetic track, have received a significant amount of interest in the field of neuromorphic computing because of their non-volatility and intrinsic current integration capabilities. Many spintronic neurons using this technology have already been proposed, but they also require external circuitry or additional device layers to implement other important neuronal behaviors. Therefore, they result in an increase in fabrication complexity and/or energy consumption. In this work, we discuss three neurons that implement these functions without the use of additional circuitry or material layers.
Naimul Hassan, Wesley H. Brigner, Xuan Hu 0002, Otitoaleke G. Akinola, Christopher H. Bennett, Matthew J. Marinella, Felipe García-Sánchez, Jean Anne C. Incorvia, Joseph S. Friedman
ISCAS5
2020 Process Variation Model and Analysis for Domain Wall-Magnetic Tunnel Junction Logic
abstract
The domain wall-magnetic tunnel junction (DW-MTJ) is a spintronic device that enables efficient logic circuit design because of its low energy consumption, small size, and non-volatility. Furthermore, the DW-MTJ is one of the few spintronic devices for which a direct cascading mechanism is experimentally demonstrated without any extra buffers; this enables potential design and fabrication of a large-scale DW-MTJ logic system. However, DW-MTJ logic relies on the conversion between electrical signals and magnetic states which is sensitive to process imperfection. Therefore, it is important to analyze the robustness of such DW-MTJ devices to anticipate the system reliability before fabrication. Here we propose a new DW-MTJ model that integrates the impacts of process variation to enable the analysis and optimization of DW-MTJ logic. This will allow circuit and device design that enhances the robustness of DW-MTJ logic and advances the development of energy-efficient spintronic computing systems.
Xuan Hu 0002, Alexander J. Edwards, T. Patrick Xiao, Christopher H. Bennett, Jean Anne C. Incorvia, Matthew J. Marinella, Joseph S. Friedman
ISCAS4
2016 Exploiting the short-term to long-term plasticity transition in memristive nanodevice learning architectures
abstract
Memristive nanodevices offer new frontiers for computing systems that unite arithmetic and memory operations on-chip. Here, we explore the integration of electrochemical metallization cell (ECM) nanodevices with tunable filamentary switching in nanoscale learning systems. Such devices offer a natural transition between short-term plasticity (STP) and long-term plasticity (LTP). In this work, we show that this property can be exploited to efficiently solve noisy classification tasks. A single crossbar learning scheme is first introduced and evaluated. Perfect classification is possible only for simple input patterns, within critical timing parameters, and when device variability is weak. To overcome these limitations, a dual-crossbar learning system partly inspired by the extreme learning machine (ELM) approach is then introduced. This approach outperforms a conventional ELM-inspired system when the first layer is imprinted before training and testing, and especially so when variability in device timing evolution is considered: variability is therefore transformed from an issue to a feature. In attempting to classify the MNIST database under the same conditions, conventional ELM obtains 84% classification, the imprinted, uniform device system obtains 88% classification, and the imprinted, variable device system reaches 92% classification. We discuss benefits and drawbacks of both systems in terms of energy, complexity, area imprint, and speed. All these results highlight that tuning and exploiting intrinsic device timing parameters may be of central interest to future bio-inspired approximate computing systems.
Christopher H. Bennett, Selina La Barbera, Adrien F. Vincent, Jacques-Olivier Klein, Fabien Alibart, Damien Querlioz
IJCNN1