EDBT 2026 Demo / reviewers in the wild / expert
Shubham Sahay
dblp:210/0107
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-9992-3240ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Ferroelectric FET-Based Bayesian Inference Engine for Disease DiagnosisabstractProbabilistic/stochastic computations form the backbone of autonomous systems and classifiers. Recently, biomedical applications of probabilistic computing such as Bayesian networks for disease diagnosis, DNA sequencing, etc. have attracted significant attention owing to their high energy-efficiency. Bayesian inference is widely used for decision making based on independent (often conflicting) sources of information/evidence. A cascaded chain or tree structure of asynchronous circuit elements known as Muller C-elements can effectively implement Bayesian inference. Such circuits utilize stochastic bit streams to encode input probabilities which enhances their robustness and fault-tolerance. However, the CMOS implementations of Muller C-element are bulky and energy hungry which restricts their widespread application in resource constrained IoT and mobile devices such as UAVs, robots, space rovers, etc. In this work, for the first time, we propose a compact and energy-efficient implementation of Muller C-element utilizing a single Ferroelectric FET and use it for cancer diagnosis task by performing Bayesian inference with high accuracy on Wisconsin data set. The proposed implementation exploits the unique drain-erase, program inhibit and drain-erase inhibit characteristics of FeFETs to yield the output as the polarization-state of the ferroelectric layer. Our extensive investigation utilizing an in-house developed experimentally calibrated compact model of FeFET reveals that the proposed C-element consumes (worst-case) energy of 4.1 fJ and an area$0.07~\mu m^{2}$and outperforms the prior implementations in terms of energy-efficiency and footprint while exhibiting a comparable delay. We also propose a novel read circuitry for realising a Bayesian inference engine by cascading a network of proposed FeFET-based C-elements for practical applications. Furthermore, for the first time, we analyze the impact of cross-correlation between the stochastic input bit streams on the accuracy of the C-element based Bayesian inference implementation. Arka Chakraborty, Musaib Rafiq, Yawar Hayat Zarkob, Yogesh Singh Chauhan, Shubham Sahay |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Satisfiability Attack-Resilient Camouflaged Multiple Multivariable Logic-in-Memory Exploiting 3D NAND Flash ArrayabstractLogic-in-memory implementations have attracted significant attention recently for energy efficient in-situ processing of big data in this era of IoT. However, the emerging memory technologies such as RRAMs, PCMs, STT-MRAMs, etc. are still immature and exhibit significant spatial and temporal variations limiting the yield and the size of crossbar arrays available for implementing logic functions. Considering the technological maturity, ultra-high density and ultra-low cost of 3D NAND flash memory, in this work, we have proposed a novel methodology to exploit 3D NAND flash memory for realizing any logic function in sum-of-product form (SOP) with ≤177 literals/inputs and$\le 2^{14}$minterms parallelly. Moreover, all the logic functions realized using the proposed technique appear same at the layout level rendering the logic-in-memory implementation utilizing the 3D NAND flash memory an innate camouflaging property and an inherent immunity against security vulnerabilities in the semiconductor supply chain. We have also evaluated the resiliency of the proposed technique against reverse engineering attacks such as SAT attacks, ATPG attacks and brute force attacks on ISCAS’85 and ISCAS’89 benchmark circuits. Our results indicate that the proposed logic-in-memory implementation facilitates complete obfuscation of the logic function without introducing any area overhead and exhibits a strong resiliency against reverse engineering. Bhogi Satya Swaroop, Ayush Saxena, Shubham Sahay |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | An Automatic Leakage Compensation Technique for Capacitively Coupled Class-AB Operational AmplifiersabstractLow voltage class-AB operational amplifiers which need to drive large loads, for low-frequency applications require signal coupling between the gates of the push-pull transistors. When realized through a coupling capacitor, this scheme requires a large bias setting resistor to ensure extremely low corner frequency of operation. In the presence of gate leakage, the voltage drop across the large resistor often makes this topology unusable in the modern leakage-prone technology nodes. In this work, we introduce an automatic leakage compensation mechanism that makes the use of capacitively-coupled class-AB stage feasible even in the presence of non-negligible gate-leakage. Shubham Sahay, Imon Mondal |
ISCAS | 2 |
| 2023 | A Computationally Efficient Compact Model for Ferroelectric Switching With Asymmetric Nonperiodic Input SignalsabstractIn this article, we develop a Verilog-A implementable compact model for the dynamic switching of ferroelectric FinFETs (Fe-FinFETs) for asymmetric nonperiodic input signals. We use the multidomain Preisach Model to capture the saturated$P$–$E $loop of the ferroelectric capacitors. In addition to the saturation loop, we model the history-dependent minor loop paths in the$P$–$E $by tracing input signals’ turning points. To capture the input signals’ turning points, we propose an RC circuit-based approach in this work. We calibrate our proposed model with the experimental data, and it accurately captures the history effect and minor loop paths of the ferroelectric capacitor. Furthermore, the elimination of storage of each turning point makes the proposed model computationally efficient compared with the previous implementations. We also demonstrate the unique electrical characteristics of Fe-FinFETs by integrating the developed compact model of Fe-Cap with the BSIM-CMG model of the 7-nm FinFET. Amol D. Gaidhane, Raghvendra Dangi, Shubham Sahay, Amit Verma 0006, Yogesh Singh Chauhan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | A Flash-Based Multi-Bit Content-Addressable Memory with Euclidean Squared DistanceabstractContent-addressable memories (CAMs) can perform fast and energy-efficient search operations. Recently, ternary CAMs (TCAMs) have been utilized to measure Hamming distance for machine learning applications, where they offer significant energy savings and speed-ups. However, the binary precision of the Hamming distance can lead to severe degradation in application-level accuracies, thus mitigating the impact of gains with respect to other figures of merit. To enhance accuracy, multi-bit CAMs (MCAMs) have been proposed that offer higher density and energy savings than TCAMs by storing multiple bits in each cell. However, existing MCAMs are based on emerging nonvolatile memory technologies that are yet to be established. To this end, we propose a fast and extremely energy-efficient MCAM based on mature and widely used flash cells, called $\mathrm{E}^{2} -$MCAM. $\mathrm{E}^{2} -$MCAM can measure the Euclidean squared distance between search queries and data stored in the MCAM “in-memory”, and in a single cycle. We evaluate $\mathrm{E}^{2} -$MCAM using an experimentally calibrated flash model in HSPICE with 3-bit precision for proof of concept demonstration. $\mathrm{A}64 \times 32 \mathrm{E}^{2} -$MCAM array achieves a 0.34 fJ energy per bit per search and a 2.7 ns latency while operating at a $770 \mu \mathrm{W}$ power. Fast and efficient hardware support for Euclidean squared distance is highly valuable as it is widely used in a plethora of machine learning applications. As an example, we show that $\mathrm{E}^{2} -$MCAM achieves accuracies comparable to floating-point GPU implementations with only 3-bit precision for few-shot learning tasks with the ImageNet dataset while offering improvements in energy and latency. Arman Kazemi, Shubham Sahay, Ayush Saxena, Mohammad Mehdi Sharifi, Michael T. Niemier, Xiaobo Sharon Hu |
ISLPED | 2 |
| 2020 | Mixed-Signal Vector-by-Matrix Multiplier Circuits Based on 3D-NAND Memories for NeurocomputingabstractWe propose an extremely dense, energy-efficient mixed-signal vector-by-matrix-multiplication (VMM) circuits based on the existing 3D-NAND flash memory blocks, without any need for their modification. Such compatibility is achieved using time-domain-encoded VMM design. We have performed rigorous simulations of such a circuit, taking into account non-idealities such as drain-induced barrier lowering, capacitive coupling, charge injection, parasitics, process variations, and noise. Our results, for example, show that the 4-bit VMM of 200-element vectors, using the commercially available 64-layer gate-all-around macaroni-type 3D-NAND memory blocks designed in the 55-nm technology node, may provide an unprecedented area efficiency of 0.14 pm2/byte and energy efficiency of ~11 fJ/Op, including the input/output and other peripheral circuitry overheads. Mohammad Bavandpour, Shubham Sahay, Mohammad Reza Mahmoodi, Dmitri B. Strukov |
DATE | 2 |
| 2020 | Efficient Mixed-Signal Neurocomputing Via Successive Integration and RescalingabstractThe widespread and ever-increasing demand for performing in situ inference, signal processing, and other computationally intensive applications in mobile Internet-of-Things (IoT) devices requires fast, compact, and energy-efficient vector-by-matrix multipliers (VMMs). The time-domain VMMs based on emerging nonvolatile memory devices exhibit significantly higher circuit density and energy efficiency than their current-mode counterparts. However, the load capacitors used to accumulate the weighted summation of the inputs in the time-domain-based circuits dominate their energy dissipation and footprint area. The true potential of the time-domain-based VMMs may be realized only when this overhead is minimized. To this end, in this brief, we propose a novel successive integration and rescaling (SIR) approach for implementing a highly efficient mixed-signal time-domain VMM for low-to-medium-precision computing. For a proof of concept, we quantitatively evaluated the performance of the proposed SIR VMM and compared it with the results of the conventional time-domain VMM, using a similar 1T-1R array. Preliminary simulation results for the 4-bit $200\, \times \, 200$ VMM, implemented using a 55-nm technology node, show area and energy efficiencies of 1.33 bits/m2and ~1.3 POp/J-the numbers, respectively, $\sim 2.5\times $ and $\sim 2.65\times $ higher than those for the prior-work time-domain VMM. Furthermore, we analyze the system-level performance of the proposed SIR VMM engine in the neuromorphic accelerator architectures and provide the preliminary estimates for various deep/recurrent neural network (DNN/RNN) applications. Mohammad Bavandpour, Shubham Sahay, Mohammad Reza Mahmoodi, Dmitri B. Strukov |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |