Sapan Agarwal

dblp:187/9763 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-3676-6986ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 ANVIL: An In-Storage Accelerator for Name-Value Data Stores
abstract
Name-value pairs (NVPs) are a widely-used abstraction to organize data in millions of applications.At a high level, an NVP associates a name (e.g., array index, key, hash) with each value in a collection of data.Specific NVP data store formats can vary widely, ranging from simple arrays/dictionaries and lookup tables to key-value stores and data mining workloads.Despite their importance, existing optimizations for NVPs are limited to only a single data store format, as the broad definition of NVPs allows for significant heterogeneity in encoding and implementation.We propose ANVIL, the first end-to-end system that allows programmers to broadly accelerate most formats of NVPs.With a conventional solid-state drive (SSD), large-scale NVP lookups can saturate both external and internal SSD bandwidth, as every NVP in the data store needs to be sent back to the host CPU to check for a matching name.ANVIL makes use of in-storage processing to avoid reading out any data for names that do not match, by performing name match checks directly inside the SSD's NAND flash chips.We demonstrate that ANVIL can substantially reduce disk I/O, reduce metadata overheads, and provide speedups of 4.0×, 25×, and 14.6% over a conventional SSD, for three different NVP workloads (database transactions, analytics, and graph processing).
Ryan Wong 0001, Nikita Kim, Aniket Das, Kevin Higgs, Engin Ipek, Sapan Agarwal, Saugata Ghose, Ben Feinberg
ISCA6
2024 A Discovery Platform to Characterize Emerging Nonvolatile Memories for Computing
abstract
Memory-centric architectures such as analog in memory computing (IMC) offer the potential for orders of magnitude improvements in energy efficiency and performance beyond state of the art. These architectures perform computations such as multiply-accumulate directly within memory array circuitry. Analog IMC and related architectures create markedly different requirements for memory devices than those of digital systems and a wide array of emerging memory candidate devices have been proposed to best meet these requirements. Accurate assessment of candidate device suitability requires characterizing the behavior in CMOS-integrated arrays, closely representing operation in a real IMC system. To address this, we have developed an analog memory array characterization platform that enables the detailed electrical characterization and optimization of these candidate memory device arrays, allowing accurate modeling and prediction of their behavior in IMC systems.
D. Wilson, Nad E. Gilbert, Matthew Spear, J. Short, Christopher H. Bennett, William Wahby, Joshua E. Kim, Robin Jacobs-Gedrim, T. Patrick Xiao, Sapan Agarwal, Matthew J. Marinella
VTS10
2022 Self-correcting Flip-flops for Triple Modular Redundant Logic in a 12-nm Technology
abstract
Area efficient self-correcting flip-flops for use with triple modular redundant (TMR) soft-error hardened logic are implemented in a 12-nm finFET process technology. The TMR flip-flop slave latches self-correct in the clock low phase using Muller C-elements in the latch feedback. These C-elements are driven by the two redundant stored values and not by the slave latch itself, saving area over a similar implementation using majority gate feedback. These flip-flops are implemented as large shift-register arrays on a test chip and have been experimentally tested for their soft-error mitigation in static and dynamic modes of operation using heavy ions and protons. We show how high clock skew can result in susceptibility to soft-errors in the dynamic mode, and explain the potential failure mechanism.
Lawrence T. Clark, Alen Duvnjak, Clifford Young-Sciortino, Matthew Cannon, John S. Brunhaver, Sapan Agarwal, Jereme Neuendank, Donald Wilson, Hugh J. Barnaby, Matthew J. Marinella
ISCAS6
2022 Eris: Fault Injection and Tracking Framework for Reliability Analysis of Open-Source Hardware
abstract
As transistors have been scaled over the past decade, modern systems have become increasingly susceptible to faults. Increased transistor densities and lower capacitances make a particle strike more likely to cause an upset. At the same time, complex computer systems are increasingly integrated into safety-critical systems such as autonomous vehicles. These two trends make the study of system reliability and fault tolerance essential for modern systems. To analyze and improve system reliability early in the design process, new tools are needed for RTL fault analysis.This paper proposes Eris, a novel framework to identify vulnerable components in hardware designs through fault-injection and fault propagation tracking. Eris builds on ESSENT—a fast C/C++ RTL simulation framework—to provide fault injection, fault tracking, and control-flow deviation detection capabilities for RTL designs. To demonstrate Eris’ capabilities, we analyze the reliability of the open source Rocket Chip SoC by randomly injecting faults during thousands of runs on four microbenchmarks. As part of this analysis we measure the sensitivity of different hardware structures to faults based on the likelihood of a random fault causing silent data corruption, unrecoverable data errors, program crashes, and program hangs. We detect control flow deviations and determine whether or not they are benign. Additionally, using Eris’ novel fault-tracking capabilities we are able to find 78% more vulnerable components in the same number of simulations compared to RTL-based fault injection techniques without these capabilities. We will release Eris as an open-source tool to aid future research into processor reliability and hardening.
Shubham Nema, Justin Kirschner, Debpratim Adak, Sapan Agarwal, Ben Feinberg, Arun Rodrigues, Matthew J. Marinella, Amro Awad
ISPASS4
2022 An Accurate, Error-Tolerant, and Energy-Efficient Neural Network Inference Engine Based on SONOS Analog Memory
abstract
We demonstrate SONOS (silicon-oxide-nitride-oxide-silicon) analog memory arrays that are optimized for neural network inference. The devices are fabricated in a 40nm process and operated in the subthreshold regime for in-memory matrix multiplication. Subthreshold operation enables low conductances to be implemented with low error, which matches the typical weight distribution of neural networks, which is heavily skewed toward near-zero values. This leads to high accuracy in the presence of programming errors and process variations. We simulate the end-to-end neural network inference accuracy, accounting for the measured programming error, read noise, and retention loss in a fabricated SONOS array. Evaluated on the ImageNet dataset using ResNet50, the accuracy using a SONOS system is within 2.16% of floating-point accuracy without any retraining. The unique error properties and high On/Off ratio of the SONOS device allow scaling to large arrays without bit slicing, and enable an inference architecture that achieves 20 TOPS/W on ResNet50, a$> 10\times $gain in energy efficiency over state-of-the-art digital and analog inference accelerators.
T. Patrick Xiao, Ben Feinberg, Christopher H. Bennett, Vineet Agrawal, Prashant Saxena, Venkatraman Prabhakar, Krishnaswamy Ramkumar, Harsha Medu, Ramesh Chettuvetty, Sapan Agarwal, Matthew J. Marinella
IEEE Trans. Circuits Syst. I Regul. Pap.11
2021 An Analog Preconditioner for Solving Linear Systems
abstract
Over the past decade as Moore's Law has slowed, the need for new forms of computation that can provide sustainable performance improvements has risen. A new method, called in situ computing, has shown great potential to accelerate matrix vector multiplication (MVM), an important kernel for a diverse range of applications from neural networks to scientific computing. Existing in situ accelerators for scientific computing, however, have a significant limitation: these accelerators provide no acceleration for preconditioning-a key bottleneck in linear solvers and in scientific computing workflows. This paper enables in situ acceleration for state-of-the-art linear solvers by demonstrating how to use a new in situ matrix inversion accelerator for analog preconditioning. As existing techniques that enable high precision and scalability for in situ MVM are inapplicable to in situ matrix inversion, new techniques to compensate for circuit non-idealities are proposed. Additionally, a new approach to bit slicing that enables splitting operands across multiple devices without external digital logic is proposed. For scalability, this paper demonstrates how in situ matrix inversion kernels can work in tandem with existing domain decomposition techniques to accelerate the solutions of arbitrarily large linear systems. The analog kernel can be directly integrated into existing preconditioning workflows, leveraging several well-optimized numerical linear algebra tools to improve the behavior of the circuit. The result is an analog preconditioner that is more effective (up to 50% fewer iterations) than the widely used incomplete LU factorization preconditioner, ILU(0), while also reducing the energy and execution time of each approximate solve operation by 1025x and 105x respectively.
Ben Feinberg, Ryan Wong 0001, T. Patrick Xiao, Christopher H. Bennett, Jacob N. Rohan, Erik G. Boman, Matthew J. Marinella, Sapan Agarwal, Engin Ipek
HPCA8
2020 Tunnel-FET Switching Is Governed by Non-Lorentzian Spectral Line Shape
abstract
In tunnel field-effect transistors (tFETs), the preferred mechanism for switching occurs by alignment (on) or misalignment (off) of two energy levels or band edges. Unfortunately, energy levels are never perfectly sharp. When a quantum dot interacts with a wire, its energy is broadened. Its actual spectral shape controls the current/voltage response of such transistor switches, from on (aligned) to off (misaligned). The most common model of spectral line shape is the Lorentzian, which falls off as reciprocal energy offset squared. Unfortunately, this is too slow a turnoff, algebraically, to be useful as a transistor switch. Electronic switches generally demand an on/off ratio of at least a million. Steep exponentially falling spectral tails would be needed for rapid off-state switching. This requires a new electronic feature, not previously recognized: narrowband, heavy-effective mass, quantum wire electrical contacts, to the tunneling quantum states. These are a necessity for spectrally sharp switching.
Sri Krishna Vadlamani, Sapan Agarwal, David T. Limmer, Steven G. Louie, Felix R. Fischer, Eli Yablonovitch
Proc. IEEE2
2020 PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAM
abstract
The wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively.
Aayush Ankit, Izzat El Hajj, Sai Rahul Chalamalasetti, Sapan Agarwal, Matthew J. Marinella, Martin Foltin, John Paul Strachan, Dejan S. Milojicic, Wen-Mei W. Hwu, Kaushik Roy 0001
IEEE Trans. Computers4
2017 Ziksa: On-chip learning accelerator with memristor crossbars for multilevel neural networks
abstract
Memristor crossbars support efficient realizations of spiking and non-spiking neural networks designs. In most of these designs off-chip/ex-situ training is used to set/update the state of the memrisitve devices. However, there is a growing need to design an efficient on-chip/in-situ learning for mobile autonomous systems. In this research, we propose an on-chip learning accelerator, known as Ziksa, that is integrated with the memristor crossbars. We demonstrate how regression and back-propagation in multi-level networks can be realized through Ziksa. The proposed accelerator is evaluated on a fabricated TiN-TaOx-TaTiN memristor crossbar. A 3-layer feedforward network was tested using Ziksa for classification. An accuracy of 95.3% was achieved on Wisconsin breast cancer dataset. The proposed learning accelerator can be envisioned as a core building block in a wide-range of cognitive algorithms that rely on on-chip online learning.
Abdullah M. Zyarah, Nicholas Soures, Lydia Hays, Robin Jacobs-Gedrim, Sapan Agarwal, Matthew J. Marinella, Dhireesha Kudithipudi
ISCAS5
2016 Resistive memory device requirements for a neural algorithm accelerator
abstract
Resistive memories enable dramatic energy reductions for neural algorithms. We propose a general purpose neural architecture that can accelerate many different algorithms and determine the device properties that will be needed to run backpropagation on the neural architecture. To maintain high accuracy, the read noise standard deviation should be less than 5% of the weight range. The write noise standard deviation should be less than 0.4% of the weight range and up to 300% of a characteristic update (for the datasets tested). Asymmetric nonlinearities in the change in conductance vs pulse cause weight decay and significantly reduce the accuracy, while moderate symmetric nonlinearities do not have an effect. In order to allow for parallel reads and writes the write current should be less than 100 nA as well.
Sapan Agarwal, Steven J. Plimpton, David R. Hughart, Alexander H. Hsia, Isaac Richter, Jonathan A. Cox, Conrad D. James, Matthew J. Marinella
IJCNN1