EDBT 2026 Demo / reviewers in the wild / expert
Yogesh Singh Chauhan
dblp:95/1810
· DBLP profile ↗
31ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0002-3356-8917ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 30 · 25 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance-Reconfigurable Multi-Mode Doherty Power Amplifier Using Bias-Space Control
Siddharth Thakur, Mir Mohammad Shayoub, Yogesh Singh Chauhan, Nagaditya Poluri |
ISCAS | 3 |
| 2026 | Evaluation of Radiation Resilience, Performance, and Vmin of Sub-3 nm FSFET Based SRAM ArraysabstractIn this work, we present single-event upset (SEU) analysis for Forksheet FET (FSFET) based CMOS circuits. Next, we present an array-level power and performance analysis along with the Vminevaluation for the FSFET-based SRAM. Physics based TCAD and industry-standard BSIM-CMG compact models are calibrated for accurate circuit analysis in SPICE. The impact of varying Heavy-Ion Radiation (HIR) doses and strike orientations is investigated for the FSFETs. The robustness of CMOS inverter against HIR is also reported in terms of failure time (tfail) and output voltage swing ($Δ$VDrop). For the SRAM, we determine the critical Linear Energy Transfer (LET). For FSFET, the individual n-/p-FETs are more vulnerable to the irradiation incident on nearby devices. At the circuit level, in comparison to perpendicular strikes, the$Δ$VDropincreases by 1.25V and 2.75V respectively, for oblique and transverse incidences, at a dose of 2.0MeVcm2/mg. The tfail also increases by 43% and 60% and the SRAM critical LET also decreases by 85% and 57.5%, respectively. The array level SRAM evaluation shows that the FSFET enables reliable operation with low-power consumption, impressive noise margins, and low minimum operating voltage (Vmin) values. FSFET SRAM power dissipation during the read and write operations is as low as 7.02$μ$W, and 3.00$μ$W respectively. At VDD=0.70V, the noise margins for hold, read, and write operations are 289.27mV, 122.89mV, and 297.79mV. The Vminfor read and write operations are 0.30V and 0.35V respectively. Hafeez Raza, Mahdi Benkhelifa, Koshal Kumar, Shivendra Singh Parihar, Yogesh Singh Chauhan, Hussam Amrouch, Avinash Lahgere |
IEEE Trans. Computers | 5 |
| 2026 | Ferroelectric Digital In-Memory Computing for Scalable, Reliable, and Efficient Similarity ComputationabstractClassification-based learning in deep neural networks, particularly few-shot learning, demands efficient similarity metrics such as Hamming distance. Conventional architectures suffer from high energy overheads due to frequent data movement between memory and processing units, hindering scalability. In-memory computing addresses this by integrating computation within memory, yet analog-based systems rely on power-hungry analog-to-digital converters (ADCs) and face scalability challenges due to device variability, especially in emerging memories. This work presents a fully digital Ferroelectric FET (FeFET)-based Logic-in-Memory (LiM) XOR cell, designed using GlobalFoundries’ 28 nm technology, eliminating ADCs and ensuring robust, energy-efficient, and scalable operation. Our 2T FeFET XOR cell, applied to 4096-bit Hamming distance calculations, achieves$23\times $lower energy,$3\times $faster latency, and$14\times $area reduction over state-of-the-art designs. Delivering 2337 Gsamples/(s$\cdot $W$\cdot $mm2) — a$300\times $improvement — this architecture offers a compelling solution for energy-efficient, reliable, and scalable AI hardware, driving sustainable computing. Anirban Kar, Albi Mema, Thorgund Nemec, Stefan Dünkel, Halid Mulaosmanovic, Sven Beyer, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | Pushing the Boundaries of AI Chips: From Monolithic 3D CMOS to Cryogenic ComputingabstractAs CMOS scaling approaches its fundamental limits, the explosive rise of AI and LLMs has unveiled profound bottlenecks in computing architectures. This paper presents two groundbreaking paradigms poised to reshape the landscape of high-performance computing and meet the surging demands of AI-driven workloads. The first paradigm is 3D monolithic integration, a revolutionary approach that achieves unprecedented logic density through Complementary FETs (CFETs), where pMOS and nMOS transistors are vertically stacked, and a dramatic expansion of on-chip memory capacity by integrating memory layers atop logic transistors. The second paradigm leverages the transformative potential of operating chips at cryogenic temperatures where transistors exhibit enhanced performance, and parasitic resistances are substantially minimized. These advancements hold the promise of redefining computing efficiency and performance for the AI era. Mahdi Benkhelifa, Shivendra Singh Parihar, Anirban Kar, Girish Pahwa, Yogesh Singh Chauhan, Hussam Amrouch |
DATE | 5 |
| 2025 | Self-Aware Silicon: Enhancing Lifecycle Management with Intelligent Testing and Data Insights
Fabian Vargas 0001, Marko S. Andjelkovic, Milos Krstic, Anirban Kar, Swati Deshwal, Yogesh Singh Chauhan, Hussam Amrouch, Daniel Tille, Sebastian Huhn 0001 |
ETS | 6 |
| 2025 | Benchmarking Cryogenic Circuits using 5 nm FinFETs for Quantum ProcessingabstractQuantum computing offers the potential to solve problems that are intractable for classical computers. A major challenge in scaling quantum computers lies in bridging the gap between cryogenic qubits, operating at millikelvin to few kelvin temperatures, and the classical CMOS-based system-on-chip (SoC) typically located at room temperature (300K). This connection introduces heat leakage, which can destabilize the qubit states. A promising solution is to relocate the control circuits and processors to the cryogenic environment, but this imposes strict constraints on power consumption due to limited cooling capacity. Additionally, the SoC must meet stringent timing requirements for qubit measurement classification. In this work, we investigate the performance of CMOS-based circuits for cryogenic operations using 5 nm FinFET technology. We begin by measuring the electrical characteristics of advanced 5 nm FinFETs at both 10K and 300K. Using these measured data, we calibrate the industry-standard compact model (BSIM-CMG) and develop two standard cell libraries for each temperature. Through the logic synthesis of six circuits from the EPFL benchmark suite, we analyze their behavior at cryogenic temperatures. Our results show that circuits at 10K achieve a 41% increase in speed compared to 300K. Further, they operate efficiently at lower supply voltages, which enables reduced power consumption while maintaining high-speed performance in cryogenic environments. Anirban Kar, Shivendra Singh Parihar, Florian Klemme, Yogesh Singh Chauhan, Hussam Amrouch |
ISCAS | 4 |
| 2025 | Transistor-to-GDS Reliability Analysis in Sub-3nm: Impact of Self-Heating and Aging on TimingabstractAs transistor scaling advances into the sub-3 nm regime, self-heating effects (SHE) and aging-induced degradation emerge as profound challenges that threaten timing closure, signal integrity, guardbands, and long-term reliability. This work presents a comprehensive transistor-to-GDS reliability analysis that captures the impact of SHE and aging in nanosheet field-effect transistors (NSFETs) and propagates it through the entire design stack to full-chip signoff. We evaluate a 64-bit RISC-V processor core and an AI accelerator containing 4096 Multiply-and-Accumulate (MAC) units, both implemented using gate-all-around (GAA) NSFET technology. TCAD simulations, carefully calibrated against measurement data, reveal local temperature rises up to 124 K in multi-stack sheet structures, which exacerbate aging and result in a threshold voltage shift of up to 42.3 mV. Incorporating these effects into standard cell characterization and commercial signoff timing analysis uncovers substantial End-of-Life (EOL) timing degradation—37.7 % for the RISC-V core and 61.7 % for the AI accelerator—highlighting the urgent need for SHE- and aging-aware methodologies, as well as reliability-optimized standard cell libraries for advanced nodes. Swati Deshwal, Hadi Nour Eddine, Mahdi Benkhelifa, Albi Mema, Yogesh Singh Chauhan, Hussam Amrouch |
ISLPED | 5 |
| 2025 | Cryo-CACTI: Cryogenic-Aware CACTI for Cache Modeling Down to 10K in Advanced 7nm FinFETsabstractCryogenic circuits are currently employed in fields such as quantum computing, particle detectors, magnetic resonance imaging, and space applications. While cryogenic circuits are being researched, there is limited work on designing cryogenic caches at temperatures below 77K. Moreover, there is no tool to estimate the delay, power, and area of cryogenic caches at advanced technology nodes. Our research focuses on the development of cryogenic caches tailored for the 7nm technology node, operating at 10K. However, a key challenge is the lack of cryogenic measurement data, especially in recent technologies. Consequently, through conducting our own FinFET transistor measurements, we calibrate cryogenic transistor models at 10K. With the 7nm cryogenic transistor data, we modelCryo-CACTIfor cryogenic caches (due to cache’s vital role in improving performance and their considerable share in area and power of the processor). Using Cryo-CACTI, our evaluation reveals considerable improvements in the energy efficiency (up to 99%) of cryogenic caches of larger sizes compared to the caches at room temperature (300K). Additionally, we explore alternative cache configurations at circuit-level to optimize cryogenic operation. Furthermore, we use Cryo-CACTI to explore the performance/energy consumption of cryogenic caches while simulating workloads such as SPEC CPU2017 and machine learning via neural networks.Cryo-CACTI is available for download athttps://github.com/marg-tools/Cryo-CACTI Divya Praneetha Ravipati, Victor M. van Santen, Shivendra Singh Parihar, Yogesh Singh Chauhan, Preeti Ranjan Panda, Hussam Amrouch |
IEEE Trans. Computers | 4 |
| 2025 | Ferroelectric FET-Based Bayesian Inference Engine for Disease DiagnosisabstractProbabilistic/stochastic computations form the backbone of autonomous systems and classifiers. Recently, biomedical applications of probabilistic computing such as Bayesian networks for disease diagnosis, DNA sequencing, etc. have attracted significant attention owing to their high energy-efficiency. Bayesian inference is widely used for decision making based on independent (often conflicting) sources of information/evidence. A cascaded chain or tree structure of asynchronous circuit elements known as Muller C-elements can effectively implement Bayesian inference. Such circuits utilize stochastic bit streams to encode input probabilities which enhances their robustness and fault-tolerance. However, the CMOS implementations of Muller C-element are bulky and energy hungry which restricts their widespread application in resource constrained IoT and mobile devices such as UAVs, robots, space rovers, etc. In this work, for the first time, we propose a compact and energy-efficient implementation of Muller C-element utilizing a single Ferroelectric FET and use it for cancer diagnosis task by performing Bayesian inference with high accuracy on Wisconsin data set. The proposed implementation exploits the unique drain-erase, program inhibit and drain-erase inhibit characteristics of FeFETs to yield the output as the polarization-state of the ferroelectric layer. Our extensive investigation utilizing an in-house developed experimentally calibrated compact model of FeFET reveals that the proposed C-element consumes (worst-case) energy of 4.1 fJ and an area$0.07~\mu m^{2}$and outperforms the prior implementations in terms of energy-efficiency and footprint while exhibiting a comparable delay. We also propose a novel read circuitry for realising a Bayesian inference engine by cascading a network of proposed FeFET-based C-elements for practical applications. Furthermore, for the first time, we analyze the impact of cross-correlation between the stochastic input bit streams on the accuracy of the C-element based Bayesian inference implementation. Arka Chakraborty, Musaib Rafiq, Yawar Hayat Zarkob, Yogesh Singh Chauhan, Shubham Sahay |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | A Lightweight PUF-Based Weights Obfuscation Technique for Secure In-Memory AI InferenceabstractIn-Memory Computing (IMC) has introduced a novel computational approach that substantially improves emerging embedded AI accelerators’ latency and power consumption efficiency. Despite the numerous advantages, IMC architectures also introduce new security vulnerabilities that may compromise the confidentiality of the deployed Neural Network (NN) algorithms. In this work, following an analysis of the potential threats, we present a novel lightweight security countermeasure for IMC accelerators. This methodology can be employed to de-obfuscate the pre-trained weights of NN architectures whose bits’ significance has been reordered prior to the deployment phase onto the IMC crossbar. The proposed solution is based on the coordinated action of a Ferroelectric Field-Effect Transistor (FeFET) based Physical Unclonable Function (PUF) design and shifting registers. These components perform custom arithmetic shift operations on the values calculated by the IMC device at runtime to obtain a coherent inference computation. Furthermore, a design-space exploration method is proposed to investigate the trade-off between area overhead and the level of security provided by the implementation. The results show that with less than 3% of area overhead our design is robust against all the tested attack strategies. Luca Parrini, Anirban Kar, Benjamin Hettwer, Taha Soliman, Yogesh Singh Chauhan, Hussam Amrouch, Norbert Wehn |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Design Automation for Cryogenic CMOS CircuitsabstractCryogenic CMOS circuits operate at temperatures close to absolute zero and are essential in many applications such as controllers for quantum computing but also medical engineering, space technology, or physical instruments. However, operating circuits at cryogenic temperatures fundamentally changes the underlying semiconductor physics that governs the CMOS transistor—rendering existing design automation approaches infeasible. In this work, we propose and implement the first end-to-end approach that enables design automation for cryogenic CMOS circuits. To this end, we (1) perform the first-of-its-kind measurements of commercial 5nm FinFET transistors from 300K down to 10K, (2) use the results to validate and calibrate the first cryogenic-aware industrial-standard compact model for FinFET technology, (3) create cryogenic-aware standard cell libraries that are compatible with the existing EDA tool flows, and (4) propose an initial cryogenic-aware logic synthesis approach that re-uses established design automation expertise but optimizes it for cryogenic purposes. Evaluations, comparisons, and discussions of all these novel contributions confirm the applicability and validity of the resulting cryogenic-aware design automation flow. Victor M. van Santen, Marcel Walter, Florian Klemme, Shivendra Singh Parihar, Girish Pahwa, Yogesh Singh Chauhan, Robert Wille, Hussam Amrouch |
DAC | 6 |
| 2023 | Invited Paper: Ultra-Efficient Edge AI Using FeFET-based Monolithic 3D IntegrationabstractMonolithic three-dimensional (M3D) integration signifies a notable technological leap by providing solutions of high density and energy efficiency, particularly for ultra-efficient edge AI solutions. Among emerging technologies, ferroelectric thin-film transistors (FeTFTs) have attracted substantial interest due to their potential in neuromorphic computing and their compatibility with back-end-of-the-line (BEOL) fabrication processes. Nevertheless, the challenge of M3D integrated circuits lies in elevated temperatures resulting from limited heat dissipation across various stacked tiers, subsequently exerting adverse effects on system reliability. In this work, we explore how compute-in-memory (CIM) architectures can be realized using BEOL FeTFT to accelerate deep learning applications. We demonstrate how generated temperature impacts the reliability and ultimately degrade the inference accuracy. To achieve that, we have modified the open-source “3D+NeuroSim” framework to introduce FeTFT and then employed it to estimate the key figures of merit, such as throughput, power density, area, etc. Then, we perform a comprehensive thermal analysis for the simulated M3D architecture to explore how the DNN accuracy will be impacted due to the generated heat. We also demonstrate how FeTFT devices can be accurately modeled using TCAD simulations and how run-time variability (due to temperature effects) and design-time variability (due to process variation) impact the reliability of FeTFT transistors. Yogesh Singh Chauhan, Hussam Amrouch |
ICCAD | 2 |
| 2023 | Frontiers in AI Acceleration: From Approximate Computing to FeFET Monolithic 3D IntegrationabstractWith the rapidly expanding applications of artificial intelligence (AI), the quest for hardware acceleration to foster high-speed and energy-efficient AI computation has become ever more important. In this work, we first explore the performance and energy advantages of employing classical AI acceleration with conventional systolic multiply-accumulate (MAC) arrays. We then highlight the growing importance of monolithic 3D integration as a transformative hardware acceleration strategy, moving beyond the constraints of classical von Neumann architectures. We also discuss how brain-inspired hyperdimensional computing (HDC) offers an exciting avenue for overcoming the power-hungry requirements often associated with MAC arrays, which are inevitable in deep learning hardware. Addressing the limitations of von Neumann architectures, we present the potential of monolithic 3D integration to enable ultra-dense Processing-in-Memory (PiM) layers stacked on top of high-performance CMOS logic. This novel approach offers to enhance computational performance. Recognizing the need for compatibility with low thermal budgets, we identify ferroelectric thin-film transistors (FeTFT) as a promising candidate for back-end-ofline (BEOL) fabrication. We highlight recent advances in BEOL FeTFT technology and demonstrate how technology/algorithm co-optimization plays a crucial role in the successful realization of reliable brain-inspired HDC on potentially unreliable FeTFT-based PiM layers. Our results showcase the potential of these innovations for the development of next-generation, energy-efficient AI hardware. Paul R. Genssler, Somaya Mansour, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 4 |
| 2023 | A Computationally Efficient Compact Model for Ferroelectric Switching With Asymmetric Nonperiodic Input SignalsabstractIn this article, we develop a Verilog-A implementable compact model for the dynamic switching of ferroelectric FinFETs (Fe-FinFETs) for asymmetric nonperiodic input signals. We use the multidomain Preisach Model to capture the saturated$P$–$E $loop of the ferroelectric capacitors. In addition to the saturation loop, we model the history-dependent minor loop paths in the$P$–$E $by tracing input signals’ turning points. To capture the input signals’ turning points, we propose an RC circuit-based approach in this work. We calibrate our proposed model with the experimental data, and it accurately captures the history effect and minor loop paths of the ferroelectric capacitor. Furthermore, the elimination of storage of each turning point makes the proposed model computationally efficient compared with the previous implementations. We also demonstrate the unique electrical characteristics of Fe-FinFETs by integrating the developed compact model of Fe-Cap with the BSIM-CMG model of the 7-nm FinFET. Amol D. Gaidhane, Raghvendra Dangi, Shubham Sahay, Amit Verma 0006, Yogesh Singh Chauhan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Robust Compact Model of High-Voltage MOSFET's Drift RegionabstractThis brief presents a compact model to capture the major difference between high-voltage (HV) and low-voltage MOSFETs, i.e., the carrier velocity saturation effect in the drift region of HV MOSFETs. We discuss the numerical and behavioral issues that can arise in SPICE simulations with the existing current-dependent formulation in Berkeley-Short-Channel-IGFET model (BSIM) for HV transistors. We then demonstrate how a voltage-dependent formulation can mitigate them without losing simplicity and accuracy. We also validate the proposed model against experimental data of HV transistors. Girish Pahwa, Ravi Goel, Garima Gill, Harshit Agarwal, Yogesh Singh Chauhan, Chenming Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | An Ultra-Low Noise Figure and Multi-Band Re-Configurable Low Noise AmplifierabstractIn this article, we design and fabricate an ultra-low noise figure, multi-band low noise amplifier (LNA) monolithic microwave integrated circuit (MMIC) using${0.25~\mu \text {m}}$GaAs pHEMT process. The proposed LNA’s performance results from a simultaneous match of optimal impedance for minimum noise figure and maximum power gain at the input. It also includes theoretical analysis of key factors in which simultaneous match of noise figure and input power gain depend. The presented LNA design can be re-configured anywhere in the frequency range from 1.8 to 5.0 GHz with the bandwidth of 200 MHz to 600 MHz. The experimentally reported performance of the proposed LNA includes a maximum 22 dB small signal gain, 18.6 dBm output power at 1-dB gain compression (OP1dB), 39.4 dBm output power at 3$^{\mathrm{ rd}}$-order intercept point (OIP3), and 0.35 dB minimum noise figure in the band 1.8 – 2.1 GHz while consuming only 0.165 W of dc power. Moreover, our LNA MMIC has an inbuilt input and output Electro-Static-Discharge (ESD) limiter, which can handle a 250 V charged device model (CDM) and 650 V human body model (HBM) according to JEDEC standards while occupying only${0.32~\text {mm}^{2}}$of area. Neha Bajpai, Paramita Maity, Manish Shah, Yogesh Singh Chauhan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Cross-Layer Reliability Modeling of Dual-Port FeFET: Device-Algorithm InteractionabstractThe Ferroelectric Field-Effect Transistor (FeFET) is an emerging Non-Volatile Memory (NVM) technology enabling novel data-centric architectures that go far beyond von Neumann principles. Nevertheless, FeFET devices exhibit significant variations that can severely restrict their applicability. Temperature further exacerbates variation effects because it degrades ferroelectric parameters. Hence, it is indispensable to investigate and model design-time variations, run-time variations, and stochastic variations due to spatial fluctuation of ferroelectric domains under different temperatures. Dual-port FeFET has been recently proposed and demonstrated as a novel structure that offers for the first time disturb-free read operation along with$>\,\,\mathrm {10\,\,\times }$larger memory window (MW) compared to conventional FeFETs. However, all the before-mentioned variations are amplified in such a new structure. This work analyses the impact of temperature variation for dual-port FeFETs for the first time in a cross-layer manner starting from the device level to the circuit/system levels, and compared to conventional FeFET. Through our cross-layer framework, we demonstrate the severe impact of variation on FeFET reliability despite the significant increase in the MW that dual-port FeFET offers. Even Hyperdimensional Computing is affected, despite its remarkable robustness against errors. All in all, our work reveals that a larger MW at the device level does not necessarily translate to benefits at the application level.Hence, investigating and modeling variability effects in a cross-layermanner is indispensable. Swetaki Chatterjee, Simon Thomann, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Cryogenic CMOS for Quantum Processing: 5-nm FinFET-Based SRAM Arrays at 10 KabstractIn this work, we are the first to investigate and model the characteristics of a commercial 5nm FinFET technology from room temperature (300K) all the way down to cryogenic temperature (10K). We focus on SRAM circuits demonstrating how cryogenic temperatures impact their power, delay, and reliability. SRAM memories are key components in quantum read-out and control circuits, and therefore characterizing their key figure of merits when building cryogenic-CMOS circuits is essential. To achieve that, we first measure the electrical characteristics of nFinFET and pFinFET devices from 300K down to 10K. Then, we carefully calibrate the cryogenic-aware BSIM-CMG, which is the first industry-standard compact model for FinFET technologies designed for cryogenic temperatures. This enables us to reproduce the experimental data in which SPICE simulations come with an excellent agreement with the measurements. Using our well-calibrated transistor models, we simulate a complete 32-bit SRAM memory array, including a write driver, sense amplifier, pre-charger, and output latch. Then, we investigate how cryogenic temperatures impact the SRAM read and write delays at several stages during the operation, as well as the power and energy. For a more comprehensive analysis, we perform our studies for different SRAM types covering high-density, high-performance, and low-voltage cells. All transistor and SRAM analyses are performed at both room temperature and cryogenic temperature to obtain detailed comparisons revealing the exact role that cryogenic temperature plays in SRAMs. All in all, we demonstrate that commercial 5nm FinFET is indeed suitable for cryogenic-CMOS circuits required in quantum processors, revealing that the performance of SRAMs at 10K does improve while power and energy consumption are reduced. Nevertheless, SRAM reliability is more challenging in which noise margins need to be carefully engineered to remain sufficient at 10K. Shivendra Singh Parihar, Victor M. van Santen, Simon Thomann, Girish Pahwa, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | SCANet: Securing the Weights With Superparamagnetic-MTJ Crossbar Array NetworksabstractDeep neural networks (DNNs) form a critical infrastructure supporting various systems, spanning from the iPhone neural engine to imaging satellites and drones. The design of these neural cores is often proprietary or a military secret. Nevertheless, they remain vulnerable to model replication attacks that seek to reverse engineer the network's synaptic weights. In this article, we propose SCANet (Superparamagnetic-MTJ Crossbar Array Networks), a novel defense mechanism against such model stealing attacks by utilizing the innate stochasticity in superparamagnets. When used as the synapse in DNNs, superparamagnetic magnetic tunnel junctions (s-MTJs) are shown to be significantly more secure than prior memristor-based solutions. The thermally induced telegraphic switching in the s-MTJs is robust and uncontrollable, thus thwarting the attackers from obtaining sensitive data from the network. Using a mixture of both superparamagnetic and conventional MTJs in the neural network (NN), the designer can optimize the time period between the weight updation and the power consumed by the system. Furthermore, we propose a modified NN architecture that can prevent replication attacks while minimizing power consumption. We investigate the effect of the number of layers in the deep network and the number of neurons in each layer on the sharpness of accuracy degradation when the network is under attack. We also explore the efficacy of SCANet in real-time scenarios, using a case study on object detection. Dinesh Rajasekharan, Nikhil Rangarajan, Satwik Patnaik, Ozgur Sinanoglu, Yogesh Singh Chauhan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Novel FDSOI-based Dynamic XNOR Logic for Ultra-Dense Highly-Efficient ComputingabstractFor the first time, we propose a novel circuit for dynamic 2-input XNOR gate that merely employs two n-type Fully-Depleted Silicon on Insulator (nFDSOI) FETs along with one additional precharging pFDSOI FET. Our design exploits the threshold voltage (Vt) tuning feature (i.e., 1ow-Vtand high-Vtstates) of FDSOI FET using the back bias as one input. The front gate bias is used as a second input. The proposed novel XNOR design reduces the number of transistors and significantly reduces power, delay, and energy compared to state-of-the-art dynamic XNOR gates. To accurately evaluate the Figure of merits, the industrial transistor compact model has been carefully calibrated against industrial measurements. The analysis demonstrates that our novel XNOR gates exhibits $8\times$ improvement in the propagation delay and $17\times$ improvement in the power consumption compared to the state-of-the-art dynamic XNOR design. Additionally, we explore the critical role of the buried oxide (BOX) thickness on the performance of proposed XNOR design. Swetaki Chatterjee, Chetan K. Dabhi, Hussam Amrouch, Yogesh Singh Chauhan |
ISCAS | 5 |
| 2022 | Cross-layer FeFET Reliability Modeling for Robust Hyperdimensional ComputingabstractHyperdimensional computing (HDC) is an emerging learning paradigm that has gained a lot of attention due to its ability to train with fewer data, lightweight implementation, and resiliency against errors. Similar to the brain, HDC can learn patterns in one iteration from small training data by computing a similarity metric such as Hamming distance. Ferroelectric Field-Effect-Transistor (FeFET) based Ternary Content Addressable Memory (TCAM) has been demonstrated as an excellent candi-date for computing this similarity metric. However, variations in the underlying ferroelectric transistor does impact the reliable HDC operation. In this paper, we demonstrate an end-to-end cross-layer FeFET reliability modeling to obtain robust HDC across the computing stack starting from transistor physics all the way to circuits and systems. The effect of random spatial fluctuation of ferroelectric (FE) domains and other variability sources on electrical characteristics of FeFET is computed through detailed physics-based TCAD simulations. Then, the entire TCAM array is simulated in SPICE using a carefully designed and calibrated compact model to capture the effect of transistor variability on the error probability for individual Hamming distances. Finally, the error probability is employed to compute the loss of inference accuracy of HDC with a language recognition task. We observe very little loss in accuracy even with a high degree of variation. Swetaki Chatterjee, Simon Thomann, Paul R. Genssler, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 5 |
| 2022 | Impact of NCFET Technology on Eliminating the Cooling Cost and Boosting the Efficiency of Google TPUabstractRecent breakthroughs in Neural Networks (NNs) led to significant accuracy improvements of several machine learning applications such as image classification and voice recognition. However, this accuracy improvement comes at the cost of an immense increase in computation demands. NNs became one of the most common and computationally intensive workloads in today's datacenters. To address these computational demands, Google announced in 2016 the Tensor Processing Unit (TPU), an advanced custom ASIC accelerator for NN inference. Two new TPU versions (v2 and v3) followed in 2017 and 2018 that support also training. Google TPUv3 packs an immense processing power ($\mathrm{90TFLOPS}$per chip) in a tiny and condensed area, leading to very high on-chip power densities and thus excessive temperature. In this article, superlattice thermoelectric cooling, which is one of the emerging on-chip cooling, is considered as an advanced cooling example for Google TPU and we investigate the impact of Negative Capacitance FET (NCFET), which is one of the recent emerging technologies, on the cooling and efficiency of TPU. Through full-chip design, of the computational core of the TPU, based on$14\mathrm{nm}$Intel FinFET technology and multiphysics temperature simulations, we demonstrate that NCFET can significantly minimize the required cooling-cost. More than 4000 NCFET configurations are evaluated in order to traverse the entire design space defined by the thickness of the ferroelectric layer of NCFET, the operating voltage, cooling, and the operating frequency, in addition to all possible FinFET's configurations. Moreover, our experimental evaluation shows that by eliminating the cooling cost, NCFET delivers 2.8x higher efficiency compared to the conventional FinFET baseline. Sami Salamin, Georgios Zervakis 0001, Florian Klemme, Hammam Kattan, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 5 |
| 2022 | Ferroelectric FET-Based Implementation of FitzHugh-Nagumo Neuron ModelabstractFerroelectric field-effect transistor (FeFET)-based circuit implementation mimicking FitzHugh-Nagumo neuron is proposed in this work. The proposed circuit is shown to mimic biological neuron properties, such as excitation block and anodal break excitation which are not mimicked by an integrate and fire neuron model. We also show a winner-take-all circuit that can be used with this proposed neuron implementation. The neuron implementation requires just one FeFET, three baseline field-effect transistors, and one capacitor, making it area and energy-efficient. The neuron circuit, with minimum sized transistors, consumes approximately 10 pJ per spike. The neuron’s energy consumption per spike can be reduced to as low as 100 fJ by designing some of the transistors with aspect ratio less than one. Dinesh Rajasekharan, Amol D. Gaidhane, Amit Ranjan Trivedi, Yogesh Singh Chauhan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Transistor Self-Heating: The Rising Challenge for Semiconductor TestingabstractQuantum confinement in 3-D device structure together with the newly employed materials like silicon-germanium (SiGe) in advanced technologies (e.g., FinFET, nanowire, nanosheets, etc.) makes transistors seriously suffer from localized self-heating effects in which generated heat within the transistor's channel is trapped inside. This is mainly due to the much lower channel and surrounding material thermal conductivity and hence lower ability for heat dissipation along with the firm isolation needed for better gate control. Self-heating effects strongly accelerate transistor aging and all the underlying defect generation mechanisms leading to serious reliability problems during the early life of chips. The key challenge in transistor self-heating when it comes to semiconductor testing is the profound difficulty in measuring self-heating directly as generated heat is trapped inside the transistor. Failing in capturing self-heating phenomenon during IC testing would later lead to chips malfunctions at run-time and hence early life failures because of reliability degradations and failure mechanisms will be unexpectedly accelerated akin to excessive internal temperatures. In this paper, we investigate the impact of self-heating effects on n-type and p-type FinFET transistors calibrated with Intel 14 nm measurement data using mature Technology CAD (TCAD) simulations. Then, the industry standard compact model for FinFET technologies (BSIM-CMG) is carefully calibrated to accurately model and reproduce all measurements. This enables circuit's designers, for the first time, to accurately investigate how emerging self-heating effects in transistors impacts the performance and power of large circuits. This opens new doors for developing novel Design-for-Testing methods that effectively reveal self-heating effects and increase the yield of chips. Om Prakash 0007, Chetan K. Dabhi, Yogesh Singh Chauhan, Hussam Amrouch |
VTS | 3 |
| 2021 | On the Resiliency of NCFET Circuits Against Voltage Over-ScalingabstractApproximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds. Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | PROTON: Post-Synthesis Ferroelectric Thickness Optimization for NCFET CircuitsabstractFor the first time, we demonstrate an optimization technique to synthesize circuits in the Negative Capacitance FET (NCFET) technology. NCFET is a rapidly emerging technology to replace the currently employed CMOS technology due to its profound ability to overcome the fundamental limit in scaling along with its full compatibility with the existing fabrication process. This is achieved by replacing the traditional transistor gate dielectric with a ferroelectric layer that manifests itself as a Negative Capacitance (NC), which magnifies the electric field. As a result, NCFET-based circuits can operate at a higher clock frequency without the need to increase the operating voltage. NC breaks one of the fundamental laws in physics in which the total capacitance of two capacitors connected in series becomes larger–instead of smaller in ordinary capacitors– than each of them. This could lead to sub-optimal netlists, suffering from significant increase in dynamic power and IR-drops. To suppress that, we employ the relation between delay decrease and capacitance increase of gates w.r.t ferroelectric thickness. Our technique takes an optimized netlist, obtained from commercial EDA tools, and then selectively determines the optimal ferroelectric thickness for each gate in the netlist, so that the maximum performance provided by NCFET is still achieved while the dynamic power is considerably decreased (45% on average),i.e., no trade-offs. Particularly, our technique enables the full exploitation of the performance benefits originating by NCFET, at a significantly lower (power) cost. Compared to state of the art, our technique decreases the energy-delay-product of circuits by 25% on average and reduces the deleterious effects of IR-drop by 56%. Hence, efficiency and reliability of circuits are improved without any loss in the obtained performance from NCFET. Sami Salamin, Georgios Zervakis 0001, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | NCFET to Rescue Technology Scaling: Opportunities and ChallengesabstractNegative Capacitance Field Effect Transistor (NCFET) is one of the promising emerging technologies that may overcome the fundamental limits of conventional CMOS technology. NCFET features a ferroelectric (FE) layer within the transistor's gate, which internally amplifies the voltage, allowing NCFET to operate at a lower voltage while sustaining performance at considerable energy savings. In this work, we raise awareness that n- and p-NCFET transistors are asymmetrically affected by the FE layer and show, for the first time, how this asymmetry results in unbalanced circuit performance (e.g., longer fall than rise propagation delay, reduced noise margins). As NCFET are meant to maintain performance while reducing power, we present a solution by scaling the number of fins in n-NCFET to regain symmetry. We optimize iteratively in conjunction with supply voltage scaling to find the minimal energy consumption while maintaining performance. In our first case study, we achieve at least 34% lower power consumption and thus 34% higher energy efficiency as the circuit exhibits identical propagation delay. However, our second case study reveals that NCFETs can consume 3× more power and energy than the FinFET design. In summary, not considering the asymmetry and replacing FinFET with current-matched NCFET results in unreliable circuits (timing violations). This work exemplifies how the power and energy consumption of a NCFET circuit might surpass that of a FinFET, if circuits are designed considering asymmetry and circuit metric matching. Hussam Amrouch, Victor M. van Santen, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
ASP-DAC | 4 |
| 2020 | Cell Library Characterization using Machine Learning for Design Technology Co-OptimizationabstractTo explore the full potential of any circuit and ensure its functionality at run-time, cell libraries beyond the typical PVT corners are needed. This holds even more for emerging technologies like Negative Capacitance (NC)-FinFET, where research in finding the optimal set of transistor parameters is still in its infancy. Design Technology Co-Optimization (DTCO) tackles bridging the large existing gap between device physics and the figures of merit of circuits. In this paper, we propose a Machine Learning (ML) approach to rapidly generate full cell libraries on demand. This enables the designer to perform extensive design space exploration and fully automated Design Technology Co-Optimization while lowering the barrier of accessibility. We demonstrate library prediction with an R2 score of around 98% for individual values and Static Timing Analysis (STA) reports. Experimental results show that our DTCO approach overestimates the achievable improvement by around 5%, nevertheless improving upon the baseline configuration. Florian Klemme, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
ICCAD | 2 |
| 2019 | Performance, Power and Cooling Trade-Offs with NCFET-based Many-CoresabstractNegative Capacitance Field-Effect Transistor (NCFET) is an emerging technology that incorporates a ferroelectric layer within the transistor gate stack to overcome the fundamental limit of sub-threshold swing in transistors. Even though physics-based NCFET models have been recently proposed, system-level NCFET models do not exist and research is still in its infancy. In this work, we are the first to investigate the impact of NCFET on performance, energy and cooling costs in many-core processors. Our proposed methodology starts from accurate physics models all the way up to the system level, where the performance and power of a many-core are widely affected. Our new methodology and system-level models allow, for the first time, the exploration of the novel trade-offs between performance gains and power losses that NCFET now offers to system-level designers. We demonstrate that an optimal ferroelectric thickness does exist. In addition, we reveal that current state-of-the-art power management techniques fail when NCFET (with a thick ferroelectric layer) comes into play. Martin Rapp, Sami Salamin, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
DAC | 5 |
| 2019 | NCFET-Aware Voltage ScalingabstractNegative Capacitance Field-Effect Transistor (NCFET) has recently attracted significant attention. In the NCFET technology with a thick ferroelectric layer, voltage reduction increases the leakage power, rather than decreases, due to the negative Drain-Induced Barrier Lowering (DIBL) effect. This work is the first to demonstrate the far-reaching consequences of such an inverse dependency w.r.t. the existing power management techniques. Moreover, this work is the first to demonstrate that state-of-the-art Dynamic Voltage Scaling (DVS) techniques are sub-optimal for NCFET. Our investigation revealed that the optimal voltage at which the total power is minimized is not necessarily at the point of the minimum voltage required to fulfill the performance constraint (as in traditional DVS). Hence, an NCFET-aware DVS is key for high energy efficiency. In this work, we therefore propose the first NCFET-aware DVS technique that selects the optimal voltage to minimize the power following the dynamics of workloads. Our experimental results of a multi-core system demonstrate that NCFET-aware DVS results in 20% on average, and up to 27% energy saving while still fulfilling the same performance constraint (i.e., no trade-offs) compared to traditional NCFET-unaware DVS techniques. Sami Salamin, Martin Rapp, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
ISLPED | 5 |
| 2015 | Modeling STI Edge Parasitic Current for Accurate Circuit SimulationsabstractWe enhance the capability of industry standard compact model BSIM6 to model the parasitic current$ I_{{\text {edge}}}$at the shallow trench isolation edge. Accurate, efficient, and scalable model for$ I_{{\text {edge}}}$is developed by finding the key differences between$ I_{{\text {edge}}}$and main device drain current ($ I_{{\text {main}}}$). It is found that$ I_{{\text {edge}}}$has a different sub-threshold slope, body-bias coefficient, and short-channel behavior as compared to$ I_{{\text {main}}}$. These important effects along with their dependencies on device geometry, bias conditions, and temperature are accounted for in the model. The model is in excellent agreement with experimental data verifying its scalability and readiness for production level usage. Sourabh Khandelwal, Harshit Agarwal, Juan Pablo Duarte, Kaiman Chan, Sagnik Dey, Yogesh Singh Chauhan, Chenming Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |