Tanvi Sharma

dblp:15/8792 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MemRaptor: Magnetoresistive Array as Matrix Vector Multiplication and Transcendental Function Operator for NLP Applications
abstract
Compute-In-Memory (CiM) is emerging as a promising paradigm to design energy-efficient hardware accelerators for AI, addressing the processor-memory data transfer bottleneck. The popularity of CiM can be attributed to their ability to perform massively parallel in-situ matrix vector multiplications (MVMs), the dominant computation in neural networks (NNs). However, NNs used in NLP applications such as Long-Short Term Memory (LSTM) and transformers also frequently perform other operations such as transcendental functions (tanh, sigmoid, and softmax). To that effect, we present MemRaptor, utilizing CiM with magnetoresistive random access memory (MRAM) technology, that can perform both MVM and transcendental functions in the same memory array. MemRaptor overlays a read only memory (ROM) on an MRAM array through hard-wiring the connection of bit-cell with an additional bitline (a bit-cell connected to either of the bitlines but not to both), incurring no array area overhead and a minimal peripheral area overhead. Note, the bitline connection of bit-cell stores the ROM value while the magnetic tunnel junction (MTJ) in the bit-cell stores the RAM data. Particularly, the magnetization state of the 1T-1MTJ bit-cell in the array stores the weight value (RAM data) of the neural network, and the bitline connection of the bit-cell stores the look-up table (ROM data) used for computing transcendental functions. We demonstrate the working of our proposed design through circuit-level simulations for a 64×64 array, using a compact model of CoFeB/MgO PMA MTJ with 120% tunnelling magnetoresistance, 5kΩ RON, in 65nm technology. Further, we showcase the advantage of MemRaptor over standard MVM-based CiM accelerator architecture, PUMA, through comprehensive system-level evaluations for LSTM, BERT, and GPT models. Our results show up to 30% and 5.3% improvements in terms of throughput and energy-efficiency, respectively, on an average across different workloads.
Dong Eun Kim, Tanvi Sharma, Anushka Mukherjee, Mainakh Mukherjee, Kaushik Roy 0001
ISLPED2
2025 Evaluating Compute in Memory Architectures for Matrix Multiplication: A Dataflow-Centric Perspective
abstract
Compute in memory (CIM) is a promising technique to reduce data movement costs in traditional hardware by efficiently performing in-situ matrix multiplication, the dominant computation during deep learning (DL) inference. However, the broader question of how CIM architectures compare to tensor-core-like architectures remains largely unexplored. In this work, we take a dataflow-centric approach utilizing classic parameters such as compute latency, bandwidth, capacity and compute/memory access costs to determine throughput and energy consumption of a given CIM architecture. To that effect, we perform an iso-area comparison of tensorcore-like (or PE array) architecture with different CIM integrated architectures for matrix multiplication kernels. Our results demonstrate that CIM integrated memory can improve energy efficiency by up to$3.5 \times$and throughput by up to$11 \times$compared to tensorcore baseline, considering INT8 precision.
Tanvi Sharma, Indranil Chakraborty, Mustafa Fayez Ali, Kaushik Roy 0001
ISPASS1
2022 Compute-in-Memory Technologies and Architectures for Deep Learning Workloads
abstract
The use of deep learning (DL) to real-world applications, such as computer vision, speech recognition, and robotics, has become ubiquitous. This can be largely attributed to a virtuous cycle between algorithms, data and computing, and storage capacity, which has driven rapid advances in all these dimensions. The ever-increasing demand for computation and memory from DL workloads presents challenges across the entire spectrum of computing platforms, from edge devices to the cloud. Hence, there is a need to explore new hardware paradigms that go well beyond the current mainstays such as graphical processing units (GPUs), tensor processing units (TPUs), and neural processing units (NPUs). A key bottleneck of current platforms is the so-called memory wall, which arises from the need to move large amounts of data between memory and compute units, expending considerable time and energy. One promising solution to this challenge is to move some computations either closer to memory, within the memory subsystem, or even within individual memory arrays. This approach, which is broadly referred to as compute-in-memory (CiM), has the potential to break the memory wall, and thereby greatly improve speed and power consumption. In this article, we provide an overview of CiM techniques used at different levels of the memory hierarchy and based on different memory technologies, including static random access memories (SRAMs), nonvolatile memories (NVMs), and DRAMs. We also discuss architectural approaches to designing CiM-based DL accelerators. Finally, we discuss the challenges associated with adopting CiM in future DL accelerators.
Mustafa Fayez Ali, Sourjya Roy, Utkarsh Saxena, Tanvi Sharma, Anand Raghunathan, Kaushik Roy 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2021 Enabling Robust SOT-MTJ Crossbars for Machine Learning using Sparsity-Aware Device-Circuit Co-design
abstract
Embedded non-volatile memory (eNVM) based crossbars have emerged as energy-efficient building blocks for machine learning accelerators. However, the analog computations in crossbars introduce errors due to several non-idealities. Moreover, since communications between crossbars are usually done in the digital domain, the energy and area costs are dominated by the Analog-to-Digital Converters (ADC). Among the eNVM technologies, Resistive Random-Access-Memory (RRAM) and Phase-Change Memory (PCM) devices suffer from poor endurance, Write variability and conductance drift. Whereas magneto-resistive technologies provide superior endurance, write stability and reliability. To that effect, we propose sparsity-aware device/circuit co-design of robust crossbars using Spin-Orbit-Torque Magnetic Tunnel Junctions (SOT-MTJs). Note, standard MTJs have low $\mathrm{R}_{\mathrm{O}\mathrm{F}\mathrm{F}}/\mathrm{R}_{\mathrm{O}\mathrm{N}}$ and low $\mathrm{R}_{\mathrm{O}\mathrm{N}}$, making them unsuitable for crossbars. In this work, we first demonstrate SOT-MTJs as crossbar elements With high $\mathrm{R}_{\mathrm{O}\mathrm{N}}$ and high $\mathrm{R}_{\mathrm{O}\mathrm{F}\mathrm{F}}/\mathrm{R}_{\mathrm{O}\mathrm{N}}$ by allowing the read-path to have thicker tunneling-barrier, leaving the write path undisturbed. Second, through extensive simulations, we quantitatively assess the impact of various device-circuit parameters such as $\mathrm{R}_{\mathrm{O}\mathrm{N}}, \mathrm{R}_{\mathrm{O}\mathrm{F}\mathrm{F}}/\mathrm{R}_{\mathrm{O}\mathrm{N}}$ ratio, crossbar size, along With input and weight sparsity, on both circuit and application level accuracy and energy consumption. We evaluate system accuracy for Resnet-20 inference on CIFAR-10 dataset and show that leveraging sparsity allows reduced ADC precision, Without degrading accuracy. Our results show that an SOT-MTJ $(\mathrm{R}_{\mathrm{O}\mathrm{N}}=200\mathrm{k}\Omega$ and $\mathrm{R}_{\mathrm{O}\mathrm{F}\mathrm{F}}/\mathrm{R}_{\mathrm{O}\mathrm{N}}=7)$ crossbar array of size 32×32 could achieve near-software accuracy. The 64×64 and 128×128 crossbars show an accuracy degradation of 2% and 9.8%, respectively, from the software accuracy and an energy improvement of upto 3.8× and 6.3× compared to a 32×32 array with 4bit-ADC.
Tanvi Sharma, Cheng Wang 0036, Amogh Agrawal, Kaushik Roy 0001
ISLPED1
2021 Magnetoresistive Circuits and Systems: Embedded Non-Volatile Memory to Crossbar Arrays
abstract
This overview article describes Magnetoresistive Random Access Memory (MRAM) from a circuits and systems perspective. We discuss various tradeoffs and design challenges of MRAM in three broad application areas: 1) embedded non-volatile memory (eNVMs), 2) crossbar-based analog in-memory computing, and 3) stochastic computing. Certain MRAM characteristics, such as high retention, high endurance and fast read and write operations, make them ideal for replacing the standard CMOS memories for last-level cache applications with future scaling. However, various tradeoffs in power, performance and area pose conflicting requirements on MRAM design. We explore these challenges and various circuit techniques that have been developed to mitigate them. Further, we present various requirements of memristive crossbar arrays for accelerating matrix-vector-multiplication (MVM) operations in light of MRAM devices, and highlight various challenges, design considerations, and applicability of MRAM as crossbar arrays. Finally, we will elaborate on how inherent stochasticity of MRAM devices can be leveraged for implementing energy-efficient true random number generators (TRNGs) and stochastic units for performing certain tasks, such as developing fast solvers for combinatorial optimization, and stochastic neural networks.
Amogh Agrawal, Cheng Wang 0036, Tanvi Sharma, Kaushik Roy 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3