Ben Feinberg

dblp:217/0635 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-0450-0067ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Noise-Agnostic One-Shot Training and Retraining for Robust DNN Inferencing on Analog Compute-in-Memory Systems
abstract
Analog Compute-in-Memory (ACiM) architectures are a promising alternatives to traditional von Neumann-based systems for accelerating deep neural networks (DNNs), as they alleviate the memory bottleneck by performing in-situ matrixvector multiplications. However, the analog nature of computation in ACiM makes DNNs highly susceptible to noise and process variations. To mitigate the effects of analog noise, existing approaches rely on variation-aware or noise-aware training, retraining, or fine-tuning. These methods, however, are not scalable, as they require chip-specific retraining and typically involve separate training runs for different levels of noise tolerance. Moreover, they overlook the inherent fault tolerance of analog-to-digital converters (ADCs). To address these limitations, we propose a one-shot training and retraining strategy for robust DNN inferencing on ACiM platforms. Our method is guided by a detailed analysis of error propagation through ADCs, revealing that robustness can be enhanced by strategically reshaping the weight distribution to better align with ADC resilience characteristics. Simulation results and experimental results with fabricated chips show that the proposed method improves inferencing accuracy by $\mathbf{7 0 \%} \boldsymbol{-} \mathbf{9 0 \%}$ for ResNet-18 and DenseNet-121 under $\mathbf{7 0 \%}$ noise injection on CIFAR-10 and SVHN, and by $\mathbf{5 0 \%}$-80% for VGG-16 under $\mathbf{5 0 \%}$ noise. These gains are achieved with only a $5 \%$ energy overhead due to the modified weight distribution.
Ashish Reddy Bommana, Ben Feinberg, T. Patrick Xiao, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty
ASP-DAC2
2026 DARTH-PUM: A Hybrid Processing-Using-Memory Architecture
abstract
Analog processing-using-memory (PUM; a.k.a. in-memory computing) makes use of electrical interactions inside memory arrays to perform bulk matrix–vector multiplication (MVM) operations. However, many popular matrix-based kernels need to execute non-MVM operations, which analog PUM cannot directly perform. To retain its energy efficiency, analog PUM architectures augment memory arrays with CMOS-based domain-specific fixed-function hardware to provide complete kernel functionality, but the difficulty of integrating such specialized CMOS logic with memory arrays has largely limited analog PUM to being an accelerator for machine learning inference, or for closely related kernels. An opportunity exists to harness analog PUM for general-purpose computation: recent works have shown that memory arrays can also perform Boolean PUM operations, albeit with very different supporting hardware and electrical signals than analog PUM.
Ryan Wong 0001, Ben Feinberg, Saugata Ghose
ASPLOS (2)2
2026 TensorDynamic: Bridging Application- and Instruction-Level Fault Injection for DNN Tensor Core Execution
abstract
Deep neural network (DNN) inference relies heavily on Tensor Core operations, which are vulnerable to transient hardware faults in computation pipelines not protected by errorcorrecting codes (ECC). Prior fault injection work has explored both application-level and instruction-level effects on DNN accuracy. However, existing application-level approaches support only coarse perturbations and do not capture hardware execution details, while instruction-level approaches lack application-level context.To address this gap, we propose TensorDynamic, an application-aware instruction-level dynamic fault injection tool for Tensor Core execution in DNN workloads. TensorDynamic enables fine-grained fault injection into MMA (matrix-multiplyaccumulate) instructions during DNN execution. Across multiple models, we show that, under the same error injection rate and severity, application-level fault injection can produce substantially different inference outcomes from instruction-level fault injection. This result underscores the need for execution-aware fault injection when evaluating DNN resilience on GPU Tensor Cores.
Yuxiao Jia, Euijun Chung, Huanzhi Pu, Ben Feinberg, Hyesoon Kim
ISPASS4
2025 ANVIL: An In-Storage Accelerator for Name-Value Data Stores
abstract
Name-value pairs (NVPs) are a widely-used abstraction to organize data in millions of applications.At a high level, an NVP associates a name (e.g., array index, key, hash) with each value in a collection of data.Specific NVP data store formats can vary widely, ranging from simple arrays/dictionaries and lookup tables to key-value stores and data mining workloads.Despite their importance, existing optimizations for NVPs are limited to only a single data store format, as the broad definition of NVPs allows for significant heterogeneity in encoding and implementation.We propose ANVIL, the first end-to-end system that allows programmers to broadly accelerate most formats of NVPs.With a conventional solid-state drive (SSD), large-scale NVP lookups can saturate both external and internal SSD bandwidth, as every NVP in the data store needs to be sent back to the host CPU to check for a matching name.ANVIL makes use of in-storage processing to avoid reading out any data for names that do not match, by performing name match checks directly inside the SSD's NAND flash chips.We demonstrate that ANVIL can substantially reduce disk I/O, reduce metadata overheads, and provide speedups of 4.0×, 25×, and 14.6% over a conventional SSD, for three different NVP workloads (database transactions, analytics, and graph processing).
Ryan Wong 0001, Nikita Kim, Aniket Das, Kevin Higgs, Engin Ipek, Sapan Agarwal, Saugata Ghose, Ben Feinberg
ISCA8
2025 Fault Tolerance in RRAM-based AI Accelerator with Guided Randomized Activation
abstract
Resistive Random Access Memory (RRAM)-based analog in-memory computing (IMC) AI accelerators offer significant advantages over digital accelerators, including lower power consumption, reduced data movement, and higher computational efficiency. However, their deployment in safety-critical and edge applications is challenging due to their hardware non-idealities, such as programming error, conductance drift, and read noise, which degrade the inferencing accuracy of the implemented neural networks (NNs). Existing methods, including noise injection during training and activation function modifications, provide limited fault-tolerance in realistic scenarios with non-idealities. We propose a fault-tolerant activation function with architectural optimization that enhances robustness against hardware-induced variations with minimal hardware and NN architectural changes. During training, the proposed activation function features a stochastic negative region, which inherently injects noise into the negative region of the activation. During inferencing, the proposed activation function operates deterministically, ensuring compatibility with existing hardware while maintaining computational efficiency. Extensive evaluations with benchmark datasets demonstrate that the proposed approach significantly improves inferencing accuracy by up to 60% under varying noise levels, outperforming conventional activation functions as well as existing fault-tolerant activation functions. By enhancing fault-tolerance to hardware-induced errors, the proposed method enables reliable and energy-efficient RRAM-based analog IMC.
Soyed Tuhin Ahmed, Eduardo Ortega, Ryan Depsey, T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Matthew J. Marinella, Krishnendu Chakrabarty
ITC5
2024 SEFsim: A Statistically-Guided Fast DRAM Simulator
abstract
In academia and industry, computer architects rely heavily on performance models for design space exploration. However, performance models are now experiencing longer simulation times due to the increasing design complexity of modern computing systems. DDR memory, a critical component in a computing system, requires an accurate performance model to properly evaluate the instructions per cycle (IPC). However, a detailed DRAM simulator models each DDR event and, therefore, contributes a considerable simulation time. This paper proposes Satistically-guided Epoch-evolving Fixed-latency Simulator (SEFsim), an approximate and fast DRAM simulation model, to significantly improve the simulation speed. The key design principle of SEFsim is to statistically capture the performance model of DRAM using a large number of patterns, enabling the model to accurately predict the latency and behavior of new workloads. Based on our evaluation using a detailed memory model and 10 workloads, SEFsim captures the original model with 96.16% accuracy while speeding up the simulation by 10.3X and 8.25 % in the standalone and full system evaluations, respectively.
Debpratim Adak, Hyokeun Lee, Ben Feinberg, Gwendolyn Voskuilen, Clay Hughes, Huiyang Zhou, Amro Awad
ISPASS3
2022 Eris: Fault Injection and Tracking Framework for Reliability Analysis of Open-Source Hardware
abstract
As transistors have been scaled over the past decade, modern systems have become increasingly susceptible to faults. Increased transistor densities and lower capacitances make a particle strike more likely to cause an upset. At the same time, complex computer systems are increasingly integrated into safety-critical systems such as autonomous vehicles. These two trends make the study of system reliability and fault tolerance essential for modern systems. To analyze and improve system reliability early in the design process, new tools are needed for RTL fault analysis.This paper proposes Eris, a novel framework to identify vulnerable components in hardware designs through fault-injection and fault propagation tracking. Eris builds on ESSENT—a fast C/C++ RTL simulation framework—to provide fault injection, fault tracking, and control-flow deviation detection capabilities for RTL designs. To demonstrate Eris’ capabilities, we analyze the reliability of the open source Rocket Chip SoC by randomly injecting faults during thousands of runs on four microbenchmarks. As part of this analysis we measure the sensitivity of different hardware structures to faults based on the likelihood of a random fault causing silent data corruption, unrecoverable data errors, program crashes, and program hangs. We detect control flow deviations and determine whether or not they are benign. Additionally, using Eris’ novel fault-tracking capabilities we are able to find 78% more vulnerable components in the same number of simulations compared to RTL-based fault injection techniques without these capabilities. We will release Eris as an open-source tool to aid future research into processor reliability and hardening.
Shubham Nema, Justin Kirschner, Debpratim Adak, Sapan Agarwal, Ben Feinberg, Arun Rodrigues, Matthew J. Marinella, Amro Awad
ISPASS5
2022 An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory
abstract
We demonstrate SONOS (silicon-oxide-nitride-oxide-silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a$> 10\times $gain in energy efficiency over state-of-the-art digital and analog inference accelerators.
T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Vineet Agrawal, Prashant Saxena, Venkatraman Prabhakar, Krishnaswamy Ramkumar, Harsha Medu, Ramesh Chettuvetty, Sapan Agarwal, Matthew J. Marinella
IEEE Trans. Circuits Syst. I Regul. Pap.2
2021 An Analog Preconditioner for Solving Linear Systems
abstract
Over the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively.
Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek
HPCA1
2020 Commutative Data Reordering: A New Technique to Reduce Data Movement Energy on Sparse Inference Workloads
abstract
Data movement is a significant and growing consumer of energy in modern systems, from specialized low-power accelerators to GPUs with power budgets in the hundreds of Watts. Given the importance of the problem, prior work has proposed designing interconnects on which the energy cost of transmitting a 0 is significantly lower than that of transmitting a 1. With such an interconnect, data movement energy is reduced by encoding the transmitted data such that the number of 1s is minimized. Although promising, these data encoding proposals do not take full advantage of application level semantics. As an example of a neglected optimization opportunity, consider the case of a dot product computation as part of a neural network inference task. The order in which the neural network weights are fetched and processed does not affect correctness, and can be optimized to further reduce data movement energy.This paper presents commutative data reordering (CDR), a hardware-software approach that leverages the commutative property in linear algebra to strategically select the order in which weight matrix coefficients are fetched from memory. To find a low-energy transmission order, weight ordering is modeled as an instance of one of two well-studied problems, the Traveling Salesman Problem and the Capacitated Vehicle Routing Problem. This reduction makes it possible to leverage the vast body of work on efficient approximation methods to find a good transmission order. CDR exploits the indirection inherent to sparse matrix formats such that no additional metadata is required to specify the selected order. The hardware modifications required to support CDR are minimal, and incur an area penalty of less than 0.01% when implemented on top of a mobile-class GPU. When applied to 7 neural network inference tasks running on a GPU-based system, CDR respectively reduces average DRAM IO energy by 53.1% and 22.2% over the data bus invert encoding scheme used by LPDDR4, and the recently proposed Base + XOR encoding. These savings are attained with no changes to the mobile system software and no runtime performance penalty.
Ben Feinberg, Benjamin C. Heyman, Darya Mikhailenko, Ryan Wong 0001, An C. Ho, Engin Ipek
ISCA1
2018 Making Memristive Neural Network Accelerators Reliable
abstract
Deep neural networks (DNNs) have attracted substantial interest in recent years due to their superior performance on many classification and regression tasks as compared to other supervised learning models. DNNs often require a large amount of data movement, resulting in performance and energy overheads. One promising way to address this problem is to design an accelerator based on in-situ analog computing that leverages the fundamental electrical properties of memristive circuits to perform matrix-vector multiplication. Recent work on analog neural network accelerators has shown great potential in improving both the system performance and the energy efficiency. However, detecting and correcting the errors that occur during in-memory analog computation remains largely unexplored. The same electrical properties that provide the performance and energy improvements make these systems especially susceptible to errors, which can severely hurt the accuracy of the neural network accelerators. This paper examines a new error correction scheme for analog neural network accelerators based on arithmetic codes. The proposed scheme encodes the data through multiplication by an integer, which preserves addition operations through the distributive property. Error detection and correction are performed through a modulus operation and a correction table lookup. This basic scheme is further improved by data-aware encoding to exploit the state dependence of the errors, and by knowledge of how critical each portion of the computation is to overall system accuracy. By leveraging the observation that a physical row that contains fewer 1s is less susceptible to an error, the proposed scheme increases the effective error correction capability with less than 4.5% area and less than 4.7% energy overheads. When applied to a memristive DNN accelerator performing inference on the MNIST and ILSVRC-2012 datasets, the proposed technique reduces the respective misclassification rates by 1.5x and 1.1x.
Ben Feinberg, Engin Ipek
HPCA1
2018 Enabling Scientific Computing on Memristive Accelerators
abstract
Linear algebra is ubiquitous across virtually every field of science and engineering, from climate modeling to macroeconomics. This ubiquity makes linear algebra a prime candidate for hardware acceleration, which can improve both the run time and the energy efficiency of a wide range of scientific applications. Recent work on memristive hardware accelerators shows significant potential to speed up matrix-vector multiplication (MVM), a critical linear algebra kernel at the heart of neural network inference tasks. Regrettably, the proposed hardware is constrained to a narrow range of workloads: although the eight-to 16-bit computations afforded by memristive MVM accelerators are acceptable for machine learning, they are insufficient for scientific computing where high-precision floating point is the norm. This paper presents the first proposal to enable scientific computing on memristive crossbars. Three techniques are explored — reducing overheads by exploiting exponent range locality, early termination of fixed-point computation, and static operation scheduling — that together enable a fixed-point memristive accelerator to perform high-precision floating point without the exorbitant cost of naïve floating-point emulation on fixed-point hardware. A heterogeneous collection of crossbars with varying sizes is proposed to efficiently handle sparse matrices, and an algorithm for mapping the dense subblocks of a sparse matrix to an appropriate set of crossbars is investigated. The accelerator can be combined with existing GPU-based systems to handle datasets that cannot be efficiently handled by the memristive accelerator alone. The proposed optimizations permit the memristive MVM concept to be applied to a wide range of problem domains, respectively improving the execution time and energy dissipation of sparse linear solvers by 10.3x and 10.9x over a purely GPU-based system.
Ben Feinberg, Uday Kumar Reddy Vengalam, Nathan Whitehair, Engin Ipek
ISCA1