EDBT 2026 Demo / reviewers in the wild / expert
Mohit Gupta 0004
dblp:181/2677-4 · also Mohit Kumar Gupta 0001
· DBLP profile ↗
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-1924-1264ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benchmarking of Scaled Majority-Logic-Synthesized Spintronic Circuits Based on Magnetic Tunnel Junction TransducersabstractIt is envisaged that spintronic logic devices will ultimately be utilized in hybrid CMOS-spintronic systems where signal interconversion between magnetic and electrical domains via transducers takes place. This underscores the vital role of transducers in influencing the overall performance of such hybrid systems. This paper addresses the question: Can spintronic circuits based on Magnetic Tunnel Junction (MTJ) transducers outperform their state-of-the-art CMOS counterparts? To this end, we use the EPFL (École Polytechnique Fédérale de Lausanne) combinational benchmark sets, synthesize them in 7 nm CMOS and in MTJ transducer based spintronic technologies, and compare the two implementation methods in terms of Energy-Delay-Product (EDP). To fully utilize the technologies’ potential, CMOS and spintronic implementations are built upon standard Boolean and Majority Gates, respectively. For the spintronic circuits, we assumed that domain conversion (electric/magnetic to magnetic/electric) is performed by means of MTJs and the computation is accomplished by domain wall (DW)-based majority gates, and considered two EDP estimation scenarios: (i) Uniform Benchmarking, which ignores the circuit’s internal structure and only includes domain transducers’ power and delay contributions into the calculations, and (ii) Majority-Inverter-Graph Benchmarking, which also embeds the circuit structure, the associated critical path delay and energy consumption by DW propagation. Our results indicate that, for the uniform case, the spintronic route is better suited for the implementation of complex circuits with few inputs and outputs. On the other hand, when the circuit structure is also considered via majority and inverter synthesis, our analysis clearly indicates that in order to match and eventually outperform CMOS performance, MTJ transducers’ efficiency has to be improved by 3-4 orders of magnitude. While it is clear that for the time being the MTJ-based-spintronic way cannot compete with CMOS, further technological transducer developments may tip the balance, which, when combined with information non-volatility, may make spintronic implementation for certain applications that require a large number of calculations and have a rather limited amount of interaction with the environment. Fanfan Meng, Siang-Yun Lee, Odysseas Zografos, Mohit Gupta 0004, Van D. Nguyen, Giovanni De Micheli, Sorin Cotofana, Inge Asselberghs, Christoph Adelmann, Gouri Sankar Kar, Sebastien Couet, Florin Ciubotaru |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | An Energy Efficient Soft SIMD Microarchitecture and Its Application on Quantized CNNsabstractThe ever-increasing computational complexity and energy consumption of today’s applications, such as machine learning (ML) algorithms, not only strain the capabilities of the underlying hardware but also significantly restrict their wide deployment at the edge. Addressing these challenges, novel architecture solutions are required by leveraging opportunities exposed by algorithms, e.g., robustness to small-bitwidth operand quantization and high intrinsic data-level parallelism. However, traditional hardware single instruction multiple data (Hard SIMD) architectures only support a small set of operand bitwidths, limiting performance improvement. To fill the gap, this manuscript introduces a novel pipelined processor microarchitecture for arithmetic computing based on the software-defined SIMD (Soft SIMD) paradigm that can define arbitrary SIMD modes through control instructions at run-time. This microarchitecture is optimized for parallel fine-grained fixed-point arithmetic, such as shift/add. It can also efficiently execute sequential shift-add-based multiplication over SIMD subwords, thanks to zero-skipping and canonical signed digit (CSD) coding. A lightweight repacking unit allows changing subword bitwidth dynamically. These features are implemented within a tight energy and area budget. An energy consumption model is established through post-synthesis for performance assessment. We select heterogeneously quantized (HQ) convolutional neural networks (CNNs) from the ML domain as the benchmark and map it onto our microarchitecture. Experimental results showcase that our approach dramatically outperforms traditional Hard SIMD Multiplier-Adder regarding area and energy requirements. In particular, our microarchitecture occupies up to 59.9% less area than a Hard SIMD that supports fewer SIMD bitwidths, while consuming up to 50.1% less energy on average to execute HQ CNNs. Pengbo Yu, Flavio Ponzina, Alexandre Levisse, Mohit Gupta 0004, Dwaipayan Biswas, Giovanni Ansaloni, David Atienza 0001, Francky Catthoor |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Design Technology co-optimization of 1D-1VCMA to improve read performance for SCM applicationsabstract1-diode 1-Voltage controlled magnetic anisotropy (1D-1VCMA) can be an option for Storage Class Memory (SCM) to bridge the latency gap between DRAM and flash memory. It has low sneak current, high non-linearity and low IR drop. This paper presents the Design Technology Co-optimization (DTCO) study of 1D-1VCMA stack to improve the performance and energy. Thanks to precessional switching of VCMA, the write operation is very fast, but the read determines overall latency as read before write is needed to ensure reliable write operations. The read performance of 1D-1VCMA is penalized due to high VCMA MTJ resistance, hence impacting the overall performance. To improve the read performance, this paper explores two solutions: 1) reducing the VCMA RA product, and 2) improving the read circuit. These solutions improve the read performance by 36% and 260%, respectively. Mohit Gupta 0004, Manu Perumkunnil Komalan, Dwaipayan Biswas, Saeideh Alinezhad Chamazcoti, Gouri Sankar Kar, Arnaud Furnémont, Julien Ryckaert |
ISCAS | 1 |
| 2023 | Impact of interconnects enhancement on SRAM design beyond 5nm technology nodeabstractThis paper presents an extensive study of 6T-SRAM based on FinFET for advanced technology nodes beyond 5nm. We deduce that parasitic resistance becomes the main bottleneck for SRAM design at these nodes. SRAM's writing margin and read speed are impacted due to the increased Bit-Line (BL) and Word-Line (WL) resistance. This work primarily explores two possible solutions to improve the parasitic resistance at advanced process technology nodes: 1) strapping of BL and WL to higher metal, and 2) adopting the resistance optimized BEOL. Strapping BL and WL to higher metal layer improves the Write Trip Point (WTP) by ~100mV and the critical path delay by 24% at the cost of 50% higher energy. Resistance optimized BEOL can improve WTP by ~50mV more and delay by 25% more, at the cost of increased energy consumption (8%). Mohit Gupta 0004, Pieter Weckx, Manu Perumkunnil Komalan, Julien Ryckaert |
ISCAS | 1 |
| 2023 | Exploring Pareto-Optimal Hybrid Main Memory Configurations Using Different Emerging MemoriesabstractMain memory system design and corresponding technology requirements have become increasingly challenging for data-dominated high-performance applications. To address the leakage and scalability issues of the conventional DRAM-based memory, new memory technologies with ultra-low leakage and potential for high scalability have been explored extensively over the last decade. However, none of them are mature enough to serve as a drop-in replacement for DRAM. In this paper, we propose a hybrid main memory system solution for utilizing new memory technologies with specific features, based on the target application characteristics and system configurations. To this end, we examine two new memories, 1S-1VCMA and IGZO-based DRAM, along with conventional DRAM in the context of hybrid main memory solutions for high-capacity and low-power Pareto-optimizations, respectively. To better evaluate the power and performance, we consider the page-fault modeling in our evaluations. The results of the simulation show that different combinations of memory technologies in the hybrid memory system, different memory capacities, and different storage systems could provide a promising solution in the system regarding the characteristics of running applications and the requirements of the system. Saeideh Alinezhad Chamazcoti, Mohit Gupta 0004, Hyungrock Oh, Timon Evenblij, Francky Catthoor, Manu Perumkunnil Komalan, Gouri Sankar Kar, Arnaud Furnémont |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Microarchitectural Exploration of STT-MRAM Last-level Cache Parameters for Energy-efficient DevicesabstractAs the technology scaling advances, limitations of traditional memories in terms of density and energy become more evident. Modern caches occupy a large part of a CPU physical size and high static leakage poses a limit to the overall efficiency of the systems, including IoT/edge devices. Several alternatives to CMOS SRAM memories have been studied during the past few decades, some of which already represent a viable replacement for different levels of the cache hierarchy. One of the most promising technologies is the spin-transfer torque magnetic RAM (STT-MRAM), due to its small basic cell design, almost absent static current and non-volatility as an added value. However, nothing comes for free, and designers will have to deal with other limitations, such as the higher latencies and dynamic energy consumption for write operations compared to reads. The goal of this work is to explore several microarchitectural parameters that may overcome some of those drawbacks when using STT-MRAM as last-level cache (LLC) in embedded devices. Such parameters include: number of cache banks, number of miss status handling registers (MSHRs) and write buffer entries, presence of hardware prefetchers. We show that an effective tuning of those parameters may virtually remove any performance loss while saving more than 60% of the LLC energy on average. The analysis is then extended comparing the energy results from calibrated technology models with data obtained with freely available tools, highlighting the importance of using accurate models for architectural exploration. Tommaso Marinelli, José Ignacio Gómez, Christian Tenllado, Manu Perumkunnil Komalan, Mohit Gupta 0004, Francky Catthoor |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Single-Ended Schmitt-Trigger-Based Robust Low-Power SRAM CellabstractThis paper presents a Schmitt-trigger-based single-ended 11T SRAM cell, which significantly improves read and write static noise margin (SNM) and consumes low power. Simulation results show that the cell also achieves the lowest leakage power dissipation among the cells considered for comparison. We also investigate the impact of process, voltage, and temperature variations on various performance parameters, such as hold SNM, read SNM, write margin, immunity to half-select issue, ION/IOFFratio of read path, and leakage power of the cell; Monte Carlo simulation results confirm the robustness of the proposed cell toward these issues. Layout drawn in a 45-nm technology rule shows that the proposed cell occupies 2.02× greater area as compared with 6T SRAM cells. However, 6.9× higher ION/IOFFratio of the read path of the proposed cell as compared with 6T cell holds potential to significantly subside the area overhead. A new figure of merit that comprehensively captures stability, delay, power dissipation, and area of an SRAM cell is also proposed. Based on the proposed metric, we observe that the proposed cell outperforms all, but one of the SRAM cells considered in this paper. Sayeed Ahmad, Mohit Gupta 0004, Naushad Alam, Mohd. Hasan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | A Low-Power Robust Easily Cascaded PentaMTJ-Based Combinational and Sequential CircuitsabstractAdvanced computing systems embed spintronic devices to improve the leakage performance of conventional CMOS systems. High speed, low power, and infinite endurance are important properties of magnetic tunnel junction (MTJ), a spintronic device, which assures its use in memories and logic circuits. This paper presents a PentaMTJ-based logic gate, which provides easy cascading, self-referencing, less voltage headroom problem in precharge sense amplifier and low area overhead contrary to existing MTJ-based gates. PentaMTJ is used here because it provides guaranteed disturbance free reading and increased tolerance to process variations along with compatibility with CMOS process. The logic gate is validated by simulation at the 45-nm technology node using a VerilogA model of the PentaMTJ. Mohit Gupta 0004, Mohd. Hasan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |