EDBT 2026 Demo / reviewers in the wild / expert
Sumeet Kumar Gupta
dblp:97/7449 · also Sumeet Gupta 0001
· DBLP profile ↗
42ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-5609-9722ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 2 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCION: A Comprehensive Simulation Framework for Charge-based In-Memory Computing for Rapid Evaluation of Hardware Non-Idealities and DNNabstractCharge-based in-memory computing (IMC) has shown great potential in achieving higher computational robustness compared to current-based IMC. However, it suffers from its own non-idealities such as parasitic capacitive coupling. Accurately evaluating these effects requires time-intensive SPICE simulations, making it challenging to conduct cross-layer exploration. To overcome these limitations, we propose SCION, a PyTorch-based framework that rigorously models hardware non-idealities in charge-based IMC and integrates them directly with DNN inference for rapid accuracy evaluation. We show that SCION predicts the IMC output with more than 99% accuracy with respect to SPICE while offering four orders of magnitude speedup. We demonstrate the capability of our framework by analyzing an SRAM-based charge-IMC accelerator deploying ResNet-50 and ViT-small DNNs with CIFAR-100 dataset. We show how inter-column capacitive coupling leads to data-dependent non-idealities, severely impairing the inference accuracy. We also explore techniques to mitigate non-idealities using SCION. To that end, we propose alternate column enablement (ACE) to eliminate inter-column coupling. Our results show that compared to the baseline design and another non-ideality mitigation approach based on ground shielding, ACE achieves significant improvement in sense margin and near-software accuracy under nominal conditions. Further, compared to ground shielding, ACE exhibits a higher tolerance to analog-to-digital converter (ADC) noise and superior voltage scalability. Doug Hyun Kim, Akul Malhotra, Sumeet Kumar Gupta |
ASP-DAC | 3 |
| 2026 | X-CODA: Xbar-Output-Dependent Compensation of In-Memory-Computing Errors for DNN Accuracy Enhancement
Adrija Chakraborty, Akul Malhotra, Sumeet Kumar Gupta |
ISLPED | 3 |
| 2026 | InterAxNN: Reconfigurable and Approximate in-Memory Processing Accelerator for Ultra-Low-Power Binary Neural Network Inference in Intermittently Powered SystemsabstractIn this work, we propose InterAxNN , an energy-aware approximate hardware architecture to perform vector-matrix multiplications in the binary precision regime for energy-constrained intermittently powered systems (IPS). In contrast to existing XNOR multiply-and-accumulate (MAC) operations implemented widely for binary neural networks (BNNs), we design a novel reconfigurable XNOR-MAC and AND-MAC memory macro to perform approximate binary precision operations, targeted for systems with extreme energy constraints. The proposed macro design is integrated with the ability to modify the MAC mode during run-time depending on instantaneous energy and power transients. We utilize the unique attributes of ferroelectric transistors (FeFETs) to implement the proposed ultra-low power BNN engine performing in-memory computing for artificial intelligence (AI) workloads. Subsequently, we leverage the quality configurable compute-in-memory-based hardware accelerator to implement InterAxNN based on a TI MSP430-based microcontroller. We evaluate the proposed InterAxNN concerning two baselines: (a) standard von Neumann computing architecture-based-microcontroller platform (MCU), and (b) MCU with a state-of-the-art low energy accelerator (MCU+LEA), and observe significant performance and energy benefits. Experimental results performed using a TI MSP430FR5379 IPS system show 448×–581× uplift in forward progress for 2%–8% accuracy loss for MNIST, 4%–5% accuracy loss for EMNIST, and 1%–4% reduction in accuracy for QMNIST, respectively, using MLP on a MCU+LEA platform with Unified NVM architecture. The AND-MAC mode in InterAxNN results in 91×–127× amount of additional forward progress over XNOR-MAC for 1%–2%, 4%–19%, and 1%–5% higher quality degradation for MNIST, EMNIST, and QMNIST, respectively. Arnab Raha, Sandeep Krishna Thirumala, Sumeet Kumar Gupta, Vijay Raghunathan |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | CREST-CiM: Cross-Coupling-Enhanced Differential STT-MRAM for Robust Computing-in-Memory in Binary Neural NetworksabstractWe propose CREST-CiM, an STT-MRAM-based Computing-in-Memory (CiM) technique, targeted for binary neural networks. To circumvent the low-distinguishability issue in standard MRAM-based CiM, CREST-CiM utilizes two magnetic tunnel junctions (MTJs) to store +1 and -1 weights in a bitcell and cross-couples the MTJs, achieving a high-to-low current ratio of up to 8100 for a bit-cell. Our analysis for $64 \times 64$ arrays shows up to 3.4 x higher CiM sense-margin, 27.6% higher read-disturb-margin, and resilience to process variations, and other hardware non-idealities, albeit at the cost of just 7.9% overall-area overhead, and $\lt1 \%$ energy and latency overhead compared to a 2T-2MTJ-CiM design. Our system-level analysis for ResNet-18 trained on CIFAR-10 shows near-sofware inference accuracy with CREST-CiM, with 10.7% improvement over 2T2MTJ baseline. Akul Malhotra, Sumeet Kumar Gupta |
DAC | 3 |
| 2025 | Harnessing Unipolar Threshold Switches for Enhanced RectificationabstractPhase transition materials (PTMs) have drawn significant attention in recent years due to their abrupt threshold switching characteristics and hysteretic behavior. Augmentation of the PTM with a transistor has been shown to provide enhanced selectivity (as high as ~107 for Ag/HfO2/Pt) leading to unique circuit-level advantages. Previously, a unipolar PTM, Ag-HfO2-Pt, was reported as a replacement for diodes due to its polarity-dependent high selectivity and hysteretic properties. It was shown to achieve ~50% higher-DC output compared to a diode-based design in a Cockcroft-Walton multiplier circuit. In this article, we take a deeper dive into this design. We augment two different PTMs (unipolar Ag-HfO2-Pt and bipolar VO2) with diode-connected MOSFETs to retain the benefits of hysteretic rectification. Our proposed hysteretic diodes (Hyperdiodes) exhibit a low-forward voltage drop owing to their volatile hysteretic characteristics. However, augmenting a hysteretic PTM with a transistor brings an additional stability concern due to their complex interplay. Hence, we perform a comprehensive stability analysis for a range of threshold voltages (−0.2 V$V_{\mathrm { th}}$$3 {\sigma }$Monte-Carlo variation analysis for a Cockcroft-Walton multiplier considering the nonidealities in the host transistor and the PTM. We observe that, hyperdiode-based design achieves ~20% higher-output voltage compared with the conventional designs within a fixed timeframe ($200~\boldsymbol {\mu }$s). Md. Mazharul Islam 0006, Shamiul Alam, Garrett S. Rose, Aly E. Fathy, Sumeet Kumar Gupta, Ahmedullah Aziz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Memory Technologies for Crossbar Array Design: A Comparative Evaluation of Their Impact on DNN AccuracyabstractIn-memory computing (IMC) using synaptic crossbar arrays offers a promising pathway toward energy-efficient deep neural network (DNN) accelerators. The potential of IMC is accentuated by virtue of IMC-compatible CMOS and non-volatile memory technologies, which offer various appealing features. However, each technology suffers from its own issues, which along with the crossbar non-idealities, significantly impair the DNN accuracy, especially in scaled technologies. In this context, a comprehensive cross-layer optimization to address the hardware non-idealities, coupled with a comparative evaluation of optimized technologies, remains largely unexplored. To address this need, we conduct a design space exploration and comparative evaluation of four prominent technologies—8T SRAM, ferroelectric transistors (FeFETs), resistive RAM (ReRAM), and spin-orbit torque magnetic RAM (SOT-MRAM) – at the 7nm node. We focus on the computational robustness in crossbar arrays and the influence of the technologies on DNN inference accuracy. We conduct IMC-driven optimization of each technology, accounting for both device- and circuit-level non-idealities. Using a cross-layer simulation framework that integrates physical models of synaptic devices and interconnects, we compare the inference accuracy of these technologies. Further, we throw light on the unique device attributes and device-circuit interactions that impact the DNN accuracy. Following this, we analyze the response of each technology to array size, bit-slice, and dataset/network complexity and how such design knob impact DNN accuracy. Our results for ResNet-20 with CIFAR-10 show that optimized FeFETs, due to their compact bit-cell layout and high distinguishability, lead to the highest accuracy, especially for large arrays. For more complex datasets (such as ResNet-50 with CIFAR-100) and larger bit-slices, ReRAM achieves performance comparable to that of FeFET. Lastly, we evaluate Partial Wordline Activation (PWA) and propose custom ADC reference levels as non-ideality-mitigating solutions and compare the response of each technology to these techniques. Victor Jeffry Louis, Sumeet Kumar Gupta |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | TWINN: Training-Free Weight-Input Flipping for Mitigating Crossbar Non-Idealities in Binary Neural Network AcceleratorsabstractCompute-in-memory (CiM)-based binary neural network (CiM-BNN) accelerators marry the benefits of CiM and ultra-low precision quantization, making them highly suitable for edge computing. However, CiM-enabled crossbar (Xbar) arrays are plagued with hardware non-idealities like parasitic resistances and device non-linearities that impair inference accuracy, especially in scaled technologies. In this work, we first analyze the impact of Xbar non-idealities on the inference accuracy of various CiM-BNNs, establishing that the unique properties of CiM-BNNs make them more prone to hardware non-idealities compared to higher precision deep neural networks (DNNs). To address this issue, we propose TWINN, a training-free technique that mitigates non-idealities in CiM-BNNs. TWINN utilizes the distinct attributes of BNNs to reduce the average current generated during the CiM operations in Xbar arrays. This is achieved by statically and dynamically flipping the BNN weights and activations, respectively. This minimizes the IR drops across the parasitic resistances, drastically mitigating their impact on inference accuracy. To evaluate our technique, we conduct experiments on ResNet-18 and VGG-small CiM-BNNs designed at the 7nm technology node using 8T-SRAM and 1T-1ReRAM. Our results show that TWINN is highly effective in alleviating the impact of non-idealities, recouping the inference accuracy to near-ideal (software) levels in some cases and providing accuracy boost of up to 77.25%. These benefits are accompanied by energy reduction, albeit at the cost of mild latency/area increase. Akul Malhotra, Sumeet Kumar Gupta |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | WAGONN: Weight Bit Agglomeration in Crossbar Arrays for Reduced Impact of Interconnect Resistance on DNN Inference AccuracyabstractDeep neural network (DNN) accelerators employing crossbar arrays capable of in-memory computing (IMC) are highly promising for neural computing platforms. However, in deeply scaled technologies, interconnect resistance severely impairs IMC robustness, leading to a drop in the system accuracy. To address this problem, we propose WAGONN - a technique based on agglomerating weight bits in crossbar arrays which alleviates the detrimental effect of wire resistance. For 8T-SRAM-based$128\times 128$crossbar arrays in 7nm technology, WAGONN enhances the accuracy from 47.78% to 83.5% for ResNet-20/CIFAR-10. We also show that WAGONN can be used synergistically with Partial-Word-Line-Activation, further boosting the accuracy. Further, we evaluate the implications of WAGONN for compact ferroelectric transistor-based crossbar arrays and show accuracy enhancement. WAGONN incurs minimal hardware overhead, with less than a 1% increase in energy consumption. Additionally, the latency and area overheads of WAGONN are ~1% and ~16%, respectively when 1 ADC is utilized per crossbar array. Jeffry Victor, Dong Eun Kim, Kaushik Roy 0001, Sumeet Kumar Gupta |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory-Based Ternary LLMsabstractTernary large language models (LLMs), which use ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision LLMs. The energy efficiency and performance of ternary LLMs can be further improved by deploying them on ternary computing-in-memory (TCiM) accelerators, thereby alleviating the von-Neumann bottleneck. However, TCiM accelerators are prone to memory stuck-at faults (SAFs) leading to degradation in model accuracy. This is particularly severe for LLMs due to their low weight sparsity. To boost SAF tolerance of TCiM accelerators, we propose ReTern that is based on 1) fault-aware sign transformations (FASTs) and 2) TCiM bitcell reprogramming exploiting their natural redundancy. The key idea is to use FAST to minimize computation errors due to SAFs in +1/−1 weights, while the natural bitcell redundancy is exploited to target SAFs in 0 weights (zero-fix). Our experiments on BitNet b1.58 700M and 3B ternary LLMs show that our technique furnishes significant fault tolerance, notably ~35% reduction in perplexity on the Wikitext dataset in the presence of faults. These benefits come at the cost of <3%, <7%, and <1% energy, latency, and area overheads, respectively. Akul Malhotra, Sumeet Kumar Gupta |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | BNN-Flip: Enhancing the Fault Tolerance and Security of Compute-in-Memory Enabled Binary Neural Network AcceleratorsabstractCompute-in-memory based binary neural networks or CiM-BNNs offer high energy/area efficiency for the design of edge deep neural network (DNN) accelerators, with only a mild accuracy reduction. However, for successful deployment, the design of CiM-BNNs must consider challenges such as memory faults and data security that plague existing DNN accelerators. In this work, we aim to mitigate both these problems simultaneously by proposing BNN-Flip, a training-free weight transformation algorithm that not only enhances the fault tolerance of CiM-BNNs but also protects them from weight theft attacks. BNN-Flip inverts the rows and columns of the BNN weight matrix in a way that reduces the impact of memory faults on the CiM-BNN’s inference accuracy, while preserving the correctness of the CiM operation. Concurrently, our technique encodes the CiM-BNN weights, securing them from weight theft. Our experiments on various CiM-BNNs show that BNN-Flip achieves an inference accuracy increase of up to 10.55% over the baseline (i.e. CiM-BNNs not employing BNN-Flip) in the presence of memory faults. Additionally, we show that the encoded weights generated by BNN-Flip furnish extremely low (near ‘random guess’) inference accuracy for the adversary attempting weight theft. The benefits of BNN-Flip come with an energy overhead of < 3%. Akul Malhotra, Sumeet Kumar Gupta |
ASPDAC | 3 |
| 2024 | Fault Tolerant In-Memory Computing based on Emerging Technologies for Ultra-Low Precision Edge AI AcceleratorsabstractEdge Artificial Intelligence (AI) demands ultra-low power data processing on highly resource-constrained platforms, mandating a departure from conventional computing architectures. In this context, in-memory computing (IMC) coupled with ultra-low precision (ULP) neural architectures have gained traction. Furthermore, non-volatile memories such as Resistive RAMs (ReRAMs), ferroelectric-transistors (FeFETs) and others have shown an immense promise to further enhance the efficiency of deep neural network (DNN) accelerators by enabling compact and low-leakage solutions. However, the design of ULP IMC platforms must counter the impairment of inference accuracy due to manufacturing defects such as stuck-at faults (SAFs), especially those based on relatively immature emerging technologies. To that end, we present two training-free and IMC-compatible fault tolerant techniques that utilize the unique properties of ULP (binary and ternary) DNNs to mitigate the impact of SAFs. For binary neural networks (BNNs), we present BNN-Flip, a weight transformation technique that inverts rows and columns of the weight matrices to convert harmful unmasked faults to innocuous masked faults, while preserving the correctness of matrix-vector multiplication operation. Our experiments show that BNN-Flip recovers the inference accuracy of binary-precision edge devices by up to 10.55% with an energy overhead of < 3%. For ternary neural networks (TNNs), we propose TFix, a technique that exploits the natural redundancy of ternary memory arrays and high weight sparsity of TNNs to enhance fault tolerance. Our experiments show that TFix reprograms a majority of the faulty weights to their correct values, regaining the inference accuracy by up to 11.07%, with an energy overhead of < 6%. Akul Malhotra, Sumeet Kumar Gupta |
ICCAD | 2 |
| 2024 | Design Space Exploration for Phase Transition Material-Augmented MRAMs With Separate Read-Write PathsabstractThis report presents a design space analysis for the phase transition material (PTM)-augmented magnetic random-access memories (MRAMs) with separate read–write paths. PTM is augmented in parallel with the magnetic tunnel junction (MTJ), improving the read performance along with providing separate read–write paths. Compared to the standard MRAM, PTM-augmented design achieves up to$1.7 \times $boost in cell tunnel magnetoresistance (CTMR),$1.2 \times $increase in read disturb margin (RDM), and${\sim }3.75 \times $increase in sense margin (SM) at the cost of${\sim }4.75 \times $more power consumption. Here, we first discuss the operating region and biasing requirements to achieve performance improvement. Then, we thoroughly explore the design space to put more options on the table for choosing the material and device structure. Finally, we perform the variation analysis where we address the performance and variation immunity tradeoffs. We demonstrate a 1000-point Monte-Carlo analysis to illustrate the effects of process variations on the performance. With lower distinguishability and read stability, the variation tolerance of the design can be improved manifold employing device-circuit co-design methodology and vice versa. Shamiul Alam, William Mitchell Hunter, Nazmul Amin, Md. Mazharul Islam 0006, Sumeet Kumar Gupta, Ahmedullah Aziz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | TFix: Exploiting the Natural Redundancy of Ternary Neural Networks for Fault Tolerant In-Memory Vector Matrix MultiplicationabstractIn-memory computing (IMC) and quantization have emerged as promising techniques for edge-based deep neural network (DNN) accelerators by reducing their energy, latency and storage requirements. In pursuit of ultra-low precision, ternary precision DNNs (TDNNs) offer high efficiency without sacrificing much inference accuracy. In this work, we explore the impact of hard faults on IMC based TDNNs and propose TFix to enhance their fault tolerance. TFix exploits the natural redundancy present in most ternary IMC bitcells as well as the high weight sparsity in TDNNs to provide up to 40.68% accuracy increase over the baseline with < 6% energy overhead. Akul Malhotra, Sumeet Kumar Gupta |
DAC | 3 |
| 2022 | Energy Efficient Cache Design with Piezoelectric FETsabstractPiezoelectric FETs (PeFETs) are a promising class of ferroelectric devices that use the piezoelectric effect to modulate strain in the channel. They present several desirable properties for on-chip memory, such as non-volatility, high-density, and low-power write capability. In this work, we present the first effort to design and evaluate cache architectures using PeFETs. Reena Elangovan, Ashish Ranjan 0001, Niharika Thakuria, Sumeet Kumar Gupta, Anand Raghunathan |
ISLPED | 4 |
| 2022 | RIBoNN: Designing Robust In-Memory Binary Neural Network AcceleratorsabstractRRAM crossbar-based accelerators show promise to execute compute intensive Deep Learning applications at the edge. For highly energy-constrained systems, Binary Neural Networks (BNNs) have gained momentum in recent times as the reduced precision alleviates the costs associated with storage, compute and communication. However, faults manifested in a unit bitcell of a RRAM crossbar-based accelerator may lead to drastic degradation in accuracy of the BNN, resulting in unintended system behavior. In this paper, we propose RIBoNN, a robust RRAM-based in-memory BNN accelerator, that consists of a 2T2R differential bitcell as the basic element of the crossbar. By leveraging the inherent characteristics of the proposed bitcell, RIBoNN is capable of achieving in-situ fault tolerance, thus circumventing the need to stall the deployed application for detection or diagnosis at the edge. RIBoNN, when evaluated on image-based datasets yields up to 96.57 % improvement in BNN classification accuracy, at a fault rate of 5 %; thereby demonstrating significant fault-tolerance over the state-of-the-art XNOR-RRAM BNN accelerator. Even though RiBoNN furnishes a negligible energy overhead of 2.62% over XNOR-RRAM, our proposed accelerator significantly reduces the inference latency by performing 24.4 % faster MAC operations with identical area footprint, while providing immense fault tolerance at the edge. Shamik Kundu, Akul Malhotra, Arnab Raha, Sumeet Kumar Gupta, Kanad Basu |
ITC | 4 |
| 2022 | Exploring the Design of Energy-Efficient Intermittently Powered Systems Using Reconfigurable Ferroelectric TransistorsabstractIn this article, we explore the design of energy-efficient intermittently powered systems (IPSs) using reconfigurable-ferroelectric transistors (R-FEFETs). Utilizing the dynamic tunability between volatile and nonvolatile modes of operation in R-FEFETs, we design nonvolatile flip-flops (NVFFs) and memory (NVM) suitable for IPS. We present two variants of R-FEFET-based NVFFs (RNVFFs): 1) with automatic backup and 2) with need-based backup. While the former offers high backup energy efficiency, the latter offers low normal operation energy. We also present an IPS-specific R-FEFET-based NVM (3T-R) with high energy efficiency compared with FEFET-based 2T NVM. Leveraging these nonvolatile circuits, we map the microcontroller unit (MCU) core registers of an IPS to RNVFFs and its on-chip memory to 3T-R. Subsequently, we analyze system-level implications of improving NVFFs and NVM individually by using R-FEFETs compared with existing FEFET-based designs. Our system-level simulations demonstrate that although we improve the register energy by 55%–67%, the total memory and system-level energy savings obtained from just improving the NVFFs (registers) in the microcontroller core are only 0.60%–5.78% and 0.31%–3.18%, respectively. However, improving the NVM by using 3T-R results in a much larger total memory and system-level energy savings in the range of 37%–40% and 20%–22%, respectively, in the context of a state-of-the-art IPS. Sandeep Krishna Thirumala, Arnab Raha, Sumeet Kumar Gupta, Vijay Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | Monte Carlo Variation Analysis of NCFET-based 6-T SRAM: Design Opportunities and Trade-offsabstractNegative Capacitance FET (NCFET) is one of the most promising variants of the emerging steep-slope transistors, able to overcome the ?Boltzmann limit'. The ferroelectric layer in the gate stack brings in new dynamics to the transistor operation by amplifying the surface potential. Steeper subthreshold slope, higher ON/OFF ratio, and the possibility to attain negative output conductance provide unique opportunities for NCFET-based circuit design. However, NCFETs inherently possess additional sources of variation, and hence, the promise of performance benefits in the nominal designs must be examined through extensive variation analysis. The non-volatile ferroelectric FETs (FEFETs) are promising candidates for storage-class memory, whereas the volatile NCFETs are suitable for high-speed SRAM design. In this work, we first draw a contrast between the modeling approaches ideal for the non-volatile FEFETs and volatile NCFETs. We then utilize a compact model for NCFET to analyze the design possibilities in an NCFET-based 6-T SRAM cell compared with its conventional counterpart ? both implemented in the 10 nm technology node. We examine the read, write, and hold performance of the SRAM cells through Monte Carlo variation analysis. We show that, even with additional variation induced spread in the device characteristics, NCFET-based SRAM cell can achieve better Static Noise Margin (SNM) during read/hold modes and allows more aggressive supply voltage scaling. The increased hold stability imposes a penalty in the write performance ? forcing design trade-offs. Shamiul Alam, Nazmul Amin, Sumeet Kumar Gupta, Ahmedullah Aziz |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Ternary Compute-Enabled Memory using Ferroelectric Transistors for Accelerating Deep Neural NetworksabstractTernary Deep Neural Networks (DNNs), which employ ternary precision for weights and activations, have recently been shown to attain accuracies close to full-precision DNNs, raising interest in their efficient hardware realization. In this work we propose a Non-Volatile Ternary Compute-Enabled memory cell (TeC-Cell) based on ferroelectric transistors (FEFETs) for inmemory computing in the signed ternary regime. In particular, the proposed cell enables storage of ternary weights and employs multi-word-line assertion to perform massively parallel signed dot-product computations between ternary weights and ternary inputs. We evaluate the proposed design at the array level and show 72% and 74% higher energy efficiency for multiply-andaccumulate (MAC) operations compared to standard nearmemory computing designs based on SRAM and FEFET, respectively. Furthermore, we evaluate the proposed TeC-Cell in an existing ternary in-memory DNN accelerator. Our results show 3.3X-3.4X reduction in system energy and 4.3X-7X improvement in system performance over SRAM and FEFET based nearmemory accelerators, across a wide range of DNN benchmarks including both deep convolutional and recurrent neural networks. Sandeep Krishna Thirumala, Shubham Jain 0004, Sumeet Kumar Gupta, Anand Raghunathan |
DATE | 3 |
| 2020 | IPS-CiM: Enhancing Energy Efficiency of Intermittently-Powered Systems with Compute-in-MemoryabstractIntermittently Powered Systems (IPS) have an ability to sustain computation progress across multiple power cycles in the presence of unreliable and sporadic harvested energy. However, with the emergence of data-intensive applications to be processed on energy-constrained IPS, it becomes challenging to handle large amounts of data with standard IPS architectures due to the von-Neumann bottleneck. To address this issue, we propose a compute-in-memory (CiM) engine which alleviates the memory-processor bottleneck and enhances energy-efficiency for transient computing workloads in IPS. We present a ferroelectric transistor (FEFET) based memory architecture which supports (a) nonvolatile memory (NVM) storage, (b) standard Boolean and arithmetic operations, (c) cyclic redundancy check for error detection and (d) edge-sensing for wireless sensory networks. Using the proposed CiM engine as a unified NVM, we construct an integrated IPS-CiM architecture based on the TI MSP430 microcontroller system with supply capacitances in the range of 10 nF -1 μF. We evaluate the proposed design with two baselines: hybrid SRAM+NVM and unified NVM architectures, both of which perform standard out-of-memory computing. We observe that for 1μF supply capacitance, IPS-CiM results in energy and performance benefits in the range of 35X-450X and 32X-400X, respectively over conventional microcontroller-based systems. Sandeep Krishna Thirumala, Arnab Raha, Vijay Raghunathan, Sumeet Kumar Gupta |
ICCD | 4 |
| 2020 | TiM-DNN: Ternary In-Memory Accelerator for Deep Neural NetworksabstractThe use of lower precision has emerged as a popular technique to optimize the compute and storage requirements of complex deep neural networks (DNNs). In the quest for lower precision, recent studies have shown that ternary DNNs (which represent weights and activations by signed ternary values) represent a promising sweet spot, achieving accuracy close to full-precision networks on complex tasks. We propose TiM-DNN, a programmable in-memory accelerator that is specifically designed to execute ternary DNNs. TiM-DNN supports various ternary representations including unweighted {-1, 0, 1}, symmetric weighted {-a, 0, a}, and asymmetric weighted {-a, 0, b} ternary systems. The building blocks of TiM-DNN are TiM tiles- specialized memory arrays that perform massively parallel signed ternary vector-matrix multiplications with a single access. TiM tiles are in turn composed of ternary processing cells (TPCs), bit-cells that function as both ternary storage units and signed ternary multiplication units. We evaluate an implementation of TiM-DNN in 32-nm technology using an architectural simulator calibrated with SPICE simulations and RTL synthesis. We evaluate TiM-DNN across a suite of state-of-the-art DNN benchmarks including both deep convolutional and recurrent neural networks. A 32-tile instance of TiM-DNN achieves a peak performance of 114 TOPs/s, consumes 0.9-W power, and occupies 1.96 mm2 chip area, representing a 300× and 388× improvement in TOPS/W and TOPS/mm2, respectively, compared to an NVIDIA Tesla V100 GPU. In comparison to specialized DNN accelerators, TiM-DNN achieves 55×-240× and 160×-291× improvement in TOPS/W and TOPS/mm2, respectively. Finally, when compared to a well-optimized near-memory accelerator for ternary DNNs, TiM-DNN demonstrates 3.9×-4.7× improvement in system-level energy and 3.2×-4.2× speedup, underscoring the potential of in-memory computing for ternary DNNs. Shubham Jain 0004, Sumeet Kumar Gupta, Anand Raghunathan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Non-Volatile Memory utilizing Reconfigurable Ferroelectric Transistors to enable Differential Read and Energy-Efficient In-Memory ComputationabstractWe propose a non-volatile memory based on cross-coupled reconfigurable ferroelectric transistors (R-FEFETs) which features differential read along with low power computation-in-memory (CiM). Exploiting the dynamic modulation of hysteresis in R-FEFETs, we achieve the aforementioned functionalities with just 2 access transistors (in addition to 2 R-FEFETs). The differential access of the proposed memory not only enhances the sense margin during read, but also enables natural computation of AND and NOR logic functions between two bits stored in the array, with the assertion of two word-lines. Using this feature, we propose a CiM architecture involving the use of a compact compute module integrated to a sense amplifier which performs Boolean logic as well as arithmetic operations such as addition of two words with a single array access. Unlike existing non-volatile CiM designs, our work features: (i) a self-referenced read operation due to differential access and (ii) a single universal voltage reference for all compute operations. At the array-level, the proposed design (R-FEFET-CiM) achieves 33%, 27% and 12% lower write, read and compute energies respectively, at iso-access time compared to FEFET based CiM (FEFET-CiM). System analysis performed by integrating our R-FEFET-CiM in the Nios II processor shows total system energy savings of 24% and 14% across various benchmarks, compared to near-memory computing and FEFET-CiM, respectively. Sandeep Krishna Thirumala, Shubham Jain 0004, Anand Raghunathan, Sumeet Kumar Gupta |
ISLPED | 4 |
| 2019 | Utilization of Negative-Capacitance FETs to Boost Analog Circuit PerformancesabstractNegative-capacitance FETs (NCFETs) are a promising candidate for low-power circuits with intrinsic features, e.g., the steep switching slope. Prior works have shown potential for enabling low-power digital logic and memory design with NCFETs. Yet, it is still not quite clear how to harness these new features of NCFETs for analog functionalities. This article provides more insights into the circuit design space with new device characteristics and investigates its deployment in analog circuits, specifically, time-domain analog-to-digital converters (ADCs) and phase-locked loops (PLLs). We propose and optimize a novel digital-based clocked comparator and a capacitor-based voltage-to-time converter (VTC), which are essential building blocks in ADCs and PLLs. Evaluation results show beyond-FinFET comparison speed and enhanced linearity for the proposed NCFET-based clocked comparator and VTC, respectively. Such improvement is achieved by exploiting the steeper slope and increased output impedance of NCFETs. More details on design details and a discussion are provided in this article. Yuhua Liang, Zhangming Zhu, Xueqing Li 0002, Sumeet Kumar Gupta, Suman Datta, Narayanan Vijaykrishnan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Computing with ferroelectric FETs: Devices, models, systems, and applicationsabstractIn this paper, we consider devices, circuits, and systems comprised of transistors with integrated ferroelectrics. Said structures are actively being considered by various semiconductor manufacturers as they can address a large and unique design space. Transistors with integrated ferroelectrics could (i) enable a better switch (i.e., offer steeper subthreshold swings), (ii) are CMOS compatible, (iii) have multiple operating modes (i.e., I-V characteristics can also enable compact, 1-transistor, non-volatile storage elements, as well as analog synaptic behavior), and (iv) have been experimentally demonstrated (i.e., with respect to all of the aforementioned operating modes). These device-level characteristics offer unique opportunities at the circuit, architectural, and system-level, and are considered here from device, circuit/architecture, and foundry-level perspectives. Ahmedullah Aziz, Evelyn T. Breyer, Xiaoming Chen 0003, Suman Datta, Sumeet Kumar Gupta, Michael Hoffmann 0008, Xiaobo Sharon Hu, Adrian M. Ionescu, Matthew Jerry, Thomas Mikolajick, Halid Mulaosmanovic, Kai Ni 0004, Michael T. Niemier, Ian O'Connor, Atanu Saha, Stefan Slesazeck, Sandeep Krishna Thirumala, Xunzhao Yin |
DATE | 6 |
| 2018 | A Monolithic-3D SRAM Design with Enhanced Robustness and In-Memory Computation SupportabstractWe present a novel 3D-SRAM cell using a Monolithic 3D integration (M3D-IC) technology for realizing both robustness and In-memory Boolean logic compute support. The proposed two-layer design makes use of additional transistors over the SRAM layer to enable assist techniques as well as provide logic functions (such as AND/NAND, OR/NOR, XNOR/XOR) without degrading cell density. Through analysis, we provide insights into the benefits provided by three memory assist and two logic modes and evaluate the energy efficiency of our proposed design. Assist techniques improve SRAM read stability by 2.2x and increase the write margin by 17.6%, while staying within the SRAM footprint. By virtue of increased robustness, the cell enables seamless operation at lower supply voltages and thereby ensures energy efficiency. Energy Delay Product (EDP) reduces by 1.6x over standard 6T SRAM with a faster data access. Transistor placement and their biasing technique in layer-2 enables In-memory bitwise Boolean computation. When computing bulk In-memory operations, 6.5x energy savings is achieved as compared to computing outside the memory system. Srivatsa Rangachar Srinivasa, Akshay Krishna Ramanathan, Xueqing Li 0002, Wei-Hao Chen, Fu-Kuo Hsueh, Chih-Chao Yang, Chang-Hong Shen, Jia-Min Shieh, Sumeet Kumar Gupta, Meng-Fan Chang, Swaroop Ghosh, Jack Sampson, Narayanan Vijaykrishnan |
ISLPED | 9 |
| 2018 | Dual Mode Ferroelectric Transistor based Non-Volatile Flip-Flops for Intermittently-Powered SystemsabstractIn this work, we propose dual mode ferroelectric transistors (D-FEFETs) that exhibit dynamic tuning of operation between volatile and non-volatile modes with the help of a control signal. We utilize the unique features of D-FEFET to design two variants of non-volatile flip-flops (NVFFs). In both designs, D-FEFETs are operated in the volatile mode for normal operations and in the non-volatile mode to backup the state of the flip-flop during a power outage. The first design comprises of a truly embedded non-volatile element (D-FEFET) which enables a fully automatic backup operation. In the second design, we introduce need-based backup, which lowers energy during normal operation at the cost of area with respect to the first design. Compared to a previously proposed FEFET based NVFF, the first design achieves 19% area reduction along with 96% lower backup energy and 9% lower restore energy, but at 14%-35% larger operation energy. The second design shows 11% lower area, 21% lower backup energy, 16% decrease in backup delay and similar operation energy but with a penalty of 17% and 19% in the restore energy and delay, respectively. System-level analysis of the proposed NVFFs in context of a state-of-the-art intermittently-powered system using real benchmarks yielded 5%-33% energy savings. Sandeep Krishna Thirumala, Arnab Raha, Hrishikesh Jayakumar, Kaisheng Ma, Narayanan Vijaykrishnan, Vijay Raghunathan, Sumeet Kumar Gupta |
ISLPED | 7 |
| 2018 | Symmetric 2-D-Memory Access to Multidimensional Data
Sumitha George, Xueqing Li 0002, Minli Julie Liao, Kaisheng Ma, Srivatsa Rangachar Srinivasa, Karthik Mohan, Ahmedullah Aziz, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2018 | Compact 3-D-SRAM Memory With Concurrent Row and Column Data Access Capability Using Sequential Monolithic 3-D IntegrationabstractThis paper proposes the use of monolithic 3-D integration technology in designing a novel two-layer 3-D-static random access memory (3-D-SRAM) cell in standard 6T-SRAM footprint. The proposed 3-D-SRAM cell is capable of data access from both the layers. The cell is designed to retrieve row-wise and column-wise data concurrently from the memory array. This memory design can cater to applications and workloads requiring multidimensional data access for enhancing system performance. The novel 3-D layout technique ensures the same footprint as a 6T-SRAM cell despite enhancing the functionality. The design ensures no degradation in the cell stability and performance. Voltage reduction in layer-2 provides 5.4× power savings during column-wise data access. We analyze the implications of employing the proposed SRAM to achieve efficient data access for integral image algorithm. We obtain 2.15× savings in access time and 7.81% access energy savings while accessing data from a 32-kB memory array to compute integral image for a region of 32 rows and 16 columns. Srivatsa Rangachar Srinivasa, Xueqing Li 0002, Meng-Fan Chang, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2017 | In Quest of the Next Information Processing Substrate: Extended Abstract: InvitedabstractConventional CMOS scaling and the Moore's law have been the cornerstone of progress in computing hardware technology. However, with dimensional scaling expected to end soon, there is a pressing need to find the next information processing hardware that can continue to support the technology revolution. Will this hardware solution be an enhanced or an augmented version of MOSFET or a switch based on a radically new switching mechanism. Ultimately, do we require a complete deviation from the Boolean paradigm itself? In this invited paper, we will review some of the actively pursued future logic, merged logic-memory and related concepts. Suman Datta, Alan C. Seabaugh, Michael T. Niemier, Arijit Raychowdhury, Darrell Schlom, Debdeep Jena, Huili Grace Xing, H.-S. Philip Wong, Eric Pop, Sayeef S. Salahuddin, Sumeet Kumar Gupta, Supratik Guha |
DAC | 11 |
| 2016 | Nonvolatile memory design based on ferroelectric FETsabstractFerroelectric FETs (FEFETs) offer intriguing possibilities for the design of low power nonvolatile memories by virtue of their three-terminal structure coupled with the ability of the ferroelectric (FE) material to retain its polarization in the absence of an electric field. Utilizing the distinct features of FEFETs, we propose a 2-transistor (2T) FEFET-based nonvolatile memory with separate read and write paths. With proper co-design at the device, cell and array levels, the proposed design achieves non-destructive read and lower write power at iso-write speed compared to standard FERAM. In addition, the FEFET-based memory exhibits high distinguishability with six orders of magnitude difference in the read currents corresponding to the two states. Comparative analysis based on experimentally calibrated models shows significant improvement of access energy-delay. For example, at a fixed write time of 550ps, the write voltage and energy are 58.5% and 67.7% lower than FERAM, respectively. These benefits are achieved with 2.4 times the area overhead. Further exploration of the proposed FEFET memory in energy harvesting nonvolatile processors shows an average improvement of 27% in forward progress over FERAM. Sumitha George, Kaisheng Ma, Ahmedullah Aziz, Xueqing Li 0002, Asif Islam Khan, Sayeef S. Salahuddin, Meng-Fan Chang, Suman Datta, Jack Sampson, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
DAC | 10 |
| 2016 | Exploiting ferroelectric FETs for low-power non-volatile logic-in-memory circuitsabstractNumerous research efforts are targeting new devices that could continue performance scaling trends associated with Moore's Law and/or accomplish computational tasks with less energy. One such device is the ferroelectric FET (FeFET), which offers the potential to be scaled beyond the end of the silicon roadmap as predicted by ITRS. Furthermore, the Ids vs. Vgs characteristics of FeFETs may allow a device to function as both a switch and a non-volatile storage element. We exploit this FeFET property to enable fine-grained logic-in-memory (LiM). We consider three different circuit design styles for FeFET-based LiM: complementary (differential), dynamic current mode, and dynamic logic. Our designs are compared with existing approaches for LiM (i.e., based on magnetic tunnel junctions (MTJs), CMOS, etc.) that afford the same circuit-level functionality. Assuming similar feature sizes, non-volatile FeFET-based LiM circuits are more efficient than functional equivalents based on MTJs when considering metrics such as propagation delay (2.9×, 6.8×) and dyanmic power (3.7×, 2.3×) (for 45 nm, 22 nm technology respectively). Compared to CMOS functional equivalents, FeFET designs still exhibit modest improvements in the aforementioned metrics while also offering non-volatility and reduced device count. Xunzhao Yin, Ahmedullah Aziz, Joseph Nahas, Suman Datta, Sumeet Kumar Gupta, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 5 |
| 2016 | On the potential of correlated materials in the design of spin-based cross-point memories (Invited)abstractCross-point architectures are promising for designing dense memory arrays. However, sneak current paths in a cross-point array necessitates the use of non-linear selectors. In this paper, we analyze the potential of employing correlated materials exhibiting abrupt insulator-metal transitions as selectors to design cross-point memories based on magnetic tunnel junctions (MTJs). We analyze the properties of the correlated materials and co-design MTJs and the selector to optimize the energy efficiency and robustness of the memory array. Our analysis points to the need of a correlated material with a large ratio of insulator and metal resistivities along with appropriate critical currents for the phase transitions (the values of which depend on the absolute value of the resistivities). We discuss that the design constraints lead to a restriction on the range of the selector length, which is closely related to the oxide thickness of the MTJ. Comparison of the cross-point architecture with standard architecture shows the benefits in the former in terms of 7% larger sense margin and 5X higher integration density at iso-read stability. However, this comes at the cost of 2X lower write speed (due to two-cycle write) and 11%-19% increase in the read/write power (due to sneak current in the cross-point array). Sumeet Kumar Gupta, Ahmedullah Aziz, Nikhil Shukla, Suman Datta |
ISCAS | 1 |
| 2016 | Ferroelectric Transistor based Non-Volatile Flip-FlopabstractWe present a non-volatile flip-flop with a feature to back-up the state in a ferroelectric transistor (FEFET) during power failure or supply gating. The data is stored in the form of polarization of the ferroelectric (FE) layer in the gate stack of the FEFET. The proposed flip-flop utilizes the non-volatility of the three-terminal FEFET to optimize the data backup and restore operations. We perform an extensive device-circuit analysis to provide insights into the design of the proposed flip-flop. We discuss the optimization of the FE thickness in the gate stack of the FEFET to introduce suitable non-volatility and present the implications at the circuit level. Our analysis shows that by virtue of the three terminal structure of the FEFET and the order of magnitude difference in the current for the two polarization states, the design of the backup/restore module is considerably simplified. Compared to a FE capacitor based non-volatile flip-flop, the proposed flip-flop achieves 40%--50% smaller backup delay, 27%--40% lower backup energy, comparable restore delay and up to an order of magnitude lower restore energy. While the FE capacitor based design leads to 76% area penalty compared to a conventional (volatile) flip-flop, the proposed design incurs only 35% area overhead. Danni Wang, Sumitha George, Ahmedullah Aziz, Suman Datta, Narayanan Vijaykrishnan, Sumeet Kumar Gupta |
ISLPED | 6 |
| 2016 | Comparative Area and Parasitics Analysis in FinFET and Heterojunction Vertical TFET Standard CellsabstractVertical tunnel field-effect transistors (VTFETs) have been extensively explored to overcome the scaling limits and to improve on-current ( I ON ) compared to standard lateral device structures for the future technologies. The benefits in terms of reduced footprint, high I ON and feasibility of fabrication have been demonstrated in several works. Among various VTFETs, the asymmetric heterojunction vertical tunnel FETs (HVTFETs) have emerged as one of the promising alternatives to standard transistors for low-voltage applications. However, while such device-level benefits without parasitics have been widely investigated, logic-gate design with parasitics and layout implications are not clear. In this article, we investigate and compare the layouts and parasitic capacitances and resistances of HVTFETs with FinFETs. Due to the vertical device structure of HVTFETs, a smaller footprint is observed compared to FinFETs in cells with small fan-in. However, for high fan-in cells, HVTFETs exhibit area overheads due to infeasibility of contact sharing in parallel and series transistors. These area overheads also lead to approximately 48% higher parasitic capacitance and resistance compared to FinFETs when the number of parallel and series connections increases. Further, in order to analyze the impact of parasitics, we modeled the analytical parasitics in SPICE. The models for both HVTFETs and FinFETs with parasitics were used to simulate a 15-stage inverter-based ring oscillator (RO) in order to compare the delay and energy. Our simulation results clearly show that HVTFETs exhibit less delay at a V DD < 0.45 V and higher energy efficiency for V DDs in the range of 0.3V--0.7V, albeit at the cost of 8% performance degradation. Moon Seok Kim, William Cane-Wissing, Xueqing Li 0002, Jack Sampson, Suman Datta, Sumeet Kumar Gupta, Narayanan Vijaykrishnan |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2015 | COAST: Correlated material assisted STT MRAMs for optimized read operationabstractWe present a novel technique for optimizing the read operation of spin-transfer torque (STT) MRAMs by employing a correlated material in conjunction with a magnetic tunnel junction (MTJ). The design of the proposed memory cell is based on exploiting the orders-of-magnitude difference in the resistance of the two phases of the correlated material (CM) and triggering operation-driven phase transitions in the CM by judiciously co-optimizing devices and the memory cell. During read, the CM operates in the metallic and insulating phases when the MTJ is in the low resistance and high resistance states, respectively. This leads to superior distinguishability, read efficiency and stability. During write, the CM operates in the metallic phase, which minimizes the impact of the CM resistance on the write speed. Our analysis shows that CM amplifies the cell tunneling magneto-resistance from 107% (for the standard STT MRAM) to 1878% (for the proposed cell) leading to 68% higher sense margin. In addition, 45% enhancement in the read disturb margin and 36% reduction in the cell read power is achieved. At the same time, the write asymmetry associated with different state transitions is mildly mitigated, leading to 9% reduction in the write power. This comes at a negligible cost of 4% larger write time. We also discuss the layout implications of our technique and propose the sharing of the CM amongst multiple cells. As a result of the sharing, the proposed technique incurs no area penalty. Ahmedullah Aziz, Nikhil Shukla, Suman Datta, Sumeet Kumar Gupta |
ISLPED | 4 |
| 2013 | Dual pillar spin-transfer torque MRAMs for low power applicationsabstractElectron-spin based data storage for on-chip memories has the potential for ultra-high density, low power consumption, very high endurance, and reasonably low read/write latency. In this article, we discuss the design challenges associated with spin-transfer torque (STT) MRAM in its state-of-the-art configuration. We propose an alternative bit cell configuration and three new genres of magnetic tunnel junction (MTJ) structures to improve STT-MRAM bit cell stabilities, write endurance, and reduce write energy consumption. The proposed multi-port, multi-pillar MTJ structures offer the unique possibility of electrical and spatial isolation of memory read and write. In order to realize ultralow power under process variations, we propose device, bit-cell and architecture level design techniques. Such design alternatives at multiple levels of design abstraction has been found to achieve substantially enhanced robustness, density, reliability and low power as compared to their charge-based counterparts for future embedded applications. Niladri Narayan Mojumder, Xuanyao Fong, Charles Augustine, Sumeet Kumar Gupta, Sri Harsha Choday, Kaushik Roy 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2012 | Future cache design using STT MRAMs for improved energy efficiency: devices, circuits and architectureabstractSpin-transfer torque magnetic RAM (STT MRAM) has emerged as a promising candidate for on-chip memory in future computing platforms. We present a cross-layer (device-circuit-architecture) approach to energy-efficient cache design using STT MRAM. At the device and circuit levels, we consider different genres of MTJs and bitcells, and evaluate their impact on the area, energy and performance of caches. In addition, we propose micro-architectural techniques viz. sequential cache read and partial cache line update, which exploit the non-volatility of STT MRAM to further improve energy efficiency of STT MRAM caches. A detailed comparison of STT MRAM caches with SRAM-based caches is also presented. Our results indicate that the proposed optimizations significantly enhance the efficiency of STT MRAM for designing lower level caches. Sang Phill Park, Sumeet Kumar Gupta, Niladri Narayan Mojumder, Anand Raghunathan, Kaushik Roy 0001 |
DAC | 2 |
| 2012 | Layout-aware optimization of stt mramsabstractWe present a layout-aware optimization methodology for spin-transfer torque (STT) MRAMs, considering the dependence of cell area on the access transistor width (WFET), number of fingers in the access transistor and the metal pitch of bit- and source-lines. It is shown that for WFETless than a critical value (~7 times the minimum feature length), one-finger transistor yields minimum cell area. For large WFET, minimum cell area is achieved with a two-finger transistor. We also show that for a range of WFET, the cell area is limited by the metal pitch of bit- and source-lines. As a result, in the metal pitch limited (MPL) region, WFETcan be increased with no change in the cell area. We analyze the impact of increase in WFETin the MPL region on the write margin and cell tunneling magneto-resistance (CTMR) of different genres of STT MRAMs. We consider conventional STT MRAM cells in the standard and reverse-connected configurations and STT MRAMs with tilted magnetic anisotropy for the analysis. By increasing WFETfrom the minimum to the maximum value in the MPL region (at iso-cell area) and reducing read voltage to achieve iso-read disturb margin, 2X improvement in write margin and 27% improvement in CTMR is achieved for the reverse-connected STT MRAM. Similar trends are observed for other STT MRAM cells. Sumeet Kumar Gupta, Sang Phill Park, Niladri Narayan Mojumder, Kaushik Roy 0001 |
DATE | 1 |
| 2012 | Write-optimized reliable design of STT MRAMabstractSpin transfer torque magnetic random access memory (STT MRAM) is a promising non-volatile memory due to its outstanding potential for high integration density and excellent scalability. Despite the attractive features, high write current and power is still a major challenge. As a result, the optimization of the memory for write is critical. Yusung Kim 0002, Sumeet Kumar Gupta, Sang Phill Park, Georgios Panagopoulos, Kaushik Roy 0001 |
ISLPED | 2 |
| 2012 | High-performance low-energy STT MRAM based on balanced write schemeabstractIt is well known that high write time/energy in STT MRAM are aggravated by the asymmetry in write currents for '0'→'1' and '1'→'0' transitions. This asymmetry is primarily due to the source degeneration of the access transistor during write. In this work, we propose a design methodology which avoids the source degeneration of the access transistor, leading to balanced switching times for '0'→'1' and '1'→'0' transitions. Dongsoo Lee, Sumeet Kumar Gupta, Kaushik Roy 0001 |
ISLPED | 2 |
| 2012 | Low-Power Architecture for Epileptic Seizure Detection Based on Reduced Complexity DWTabstractIn this article, we present a low-power, user-programmable architecture for discrete wavelet transform (DWT) based epileptic seizure detection algorithm. A simplified, low-pass filter (LPF)-only-DWT technique is employed in which energy contents of different frequency bands are obtained by subtracting quasi-averaged, consecutive LPF outputs. Training phase is used to identify the range of critical DWT coefficients that are in turn used to set patient-specific system level parameters for minimizing power consumption. The proposed optimizations allow the design to work at significantly lower power in the normal operation mode. The system has been tested on neural data obtained from kainate-treated rats. The design was implemented in TSMC-65nm technology and consumes less than 550-nW power at 250-mV supply. Mrigank Sharad, Sumeet Kumar Gupta, Raghunathan Shriram, Pedro P. Irazoqui, Kaushik Roy 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2010 | Digital Computation in Subthreshold Region for Ultralow-Power Operation: A Device-Circuit-Architecture Codesign PerspectiveabstractUltralow-power dissipation can be achieved by operating digital circuits with scaled supply voltages, albeit with degradation in speed and increased susceptibility to parameter variations. However, operating digital logic and memory circuits in the subthreshold region (supply voltage less than the transistor threshold voltage) for ultralow-power operations requires device, circuit as well as architectural design optimizations, different from the conventional superthreshold design. This paper analyzes such optimizations from energy dissipation point of view and shows that it is feasible to achieve robust operation of ultralow-voltage systems. Operation with power supply as low as 60 mV is demonstrated. Techniques to reduce the impact of process variations on subthreshold circuits are also discussed. In addition, it is shown that subthreshold leakage current can be useful for other applications like thermal sensors. Sumeet Kumar Gupta, Arijit Raychowdhury, Kaushik Roy 0001 |
Proc. IEEE | 1 |
| 2009 | Device/circuit interactions at 22nm technology nodeabstractAs transition is being made into 22nm node, technology considerations and device architectures suitable for such scaled technologies are being explored. To design circuits and systems at scaled nodes, we believe there is a need for technology aware circuit and system design methodology that considers device architecture, and technology challenges to achieve design optimality. In this paper, we discuss the challenges of device-circuit-system design at the 22 nm node and present techniques at different levels of design abstraction to meet these challenges. In particular, we discuss different device options for multi-gate FETs. Logic and memory design using multi-gate FETs is also considered. Finally, we briefly discuss process variation tolerant system design methodologies for such scaled technologies. Kaushik Roy 0001, Jaydeep P. Kulkarni, Sumeet Kumar Gupta |
DAC | 3 |