VLDB 2026 Research / reviewers in the wild / expert
Hussam Amrouch
dblp:94/10663
· DBLP profile ↗
176ranked-venue papers
28as first author
125since 2021 · last 2026
0000-0002-5649-3102ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 171 · 28 first-author · 121 since 2021Software engineering, systems software and programming languages · 27 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: Advancing European Semiconductor and Chiplet Innovation Through the Bavarian Chip Design CenterabstractEurope’s semiconductor industry relies heavily on Asian and US manufacturers. The EU Chips Act seeks to strengthen Europe’s capabilities across the semiconductor value chain. Aligned with this goal, the Bavarian Chip Design Center (BCDC) supports local chip design, manufacturing, and talent development, with a focus on RISC-V computing and heterogeneous integration. Within BCDC, the Technical University of Munich and Fraunhofer are developing a chiplet-based architecture optimized for low-power edge AI. The system integrates two chiplets, combining a security-enhanced RISC-V core and AI accelerators, connected via a chiplet-optimized serial interface that supports encrypted data. The chiplets are mounted on a custom interposer with low-capacitance wires for efficient data transmission. System-and component-level development is currently ongoing, with a tapeout in 22 nm FD-SOI planned for 2027. The overall goal is to deliver a proof of concept for a small-scale energy-efficient chiplet system that demonstrates Bavaria’s and Europe’s capability to drive innovation in novel chip design fields. Hussam Amrouch, Jehaan Joseph, Michael Schirmer, Johannes Geier, Ulf Schlichtmann, Michael Meidinger, Thomas Wild, Andreas Herkersdorf, Jens Nöpel, Georg Sigl, Carsten Trinitis, Aswathy Nedumpalli Sankaranarayanan, Martin Schulz 0001, Andreas Korb, Konrad Hohentanner |
DATE | 1 |
| 2026 | Focus Session: Advanced CMOS and 5.5D Packaging: Perspectives and Challenges for Design, Reliability and SecurityabstractAI workloads are driving exceptional demand for performance and energy efficiency, forcing semiconductor innovation to advance along two major directions simultaneously. On the device roadmap, the transition from FinFETs to gate-all-around nanosheet FETs and, subsequently, monolithic 3D Complementary FETs (CFETs) is enabling scaling toward the 2 nm era and beyond while targeting aggressive logic density. In parallel, advanced packaging, spanning 2.5D integration on silicon interposers, true 3D stacking, and hybrid 5.5D assemblies, is becoming essential to deliver ultra-high bandwidth, low-energy die-to-die connectivity required by rapidly growing AI model sizes and the resulting memory-wall bottlenecks. This focus session discusses the opportunities and challenges of this co-evolution, with emphasis on system-technology co-optimization and the inevitable need for multiphysics analysis across electrical, thermal, mechanical, and reliability domains. We highlight how reliability and security concerns are increasingly shaping architectural and packaging choices, and we discuss the role of deep learning as a practical enabler for faster simulation and design-space exploration under rising complexity. Hussam Amrouch, Dragomir Milojevic, Giorgio Di Natale, Jérôme Toublanc |
DATE | 1 |
| 2026 | Securing Hyper-Dimensional Computing: A Locking Mechanism with FPGA Implementation
Rupesh Raj Karn, Paul R. Genssler, Hussam Amrouch, Ozgur Sinanoglu |
ICISSP (2) | 3 |
| 2026 | Power Side-Channel Attacks in Nanosheet Circuits
Mohammed Nabeel Thari Moopan, Hadi Nour Eddine, Mahdi Benkhelifa, Ozgur Sinanoglu, Michail Maniatakos, Johann Knechtel, Hussam Amrouch |
ISCAS | 7 |
| 2026 | Robust Adaptive DLBIST for Delay Fault Testing: Minimizing PVT Variability with Zero Temperature Coefficient (ZTC) VoltageabstractAbstract Safety-critical automotive systems require on-chip testing methods to ensure high fault coverage and reliable operation. Periodic Deterministic Logic Built-In Self-Test (DLBIST) is often used to meet these demands. For automotive applications, DLBIST must operate reliably despite temperature variations, including those caused by ambient changes and self-heating in FinFET transistors. A test set effective for all temperatures typically requires a large volume, which can make DLBIST impractical. This paper proposes a robust DLBIST scheme which applies multiple voltages during power-on and power-off tests and the optimal or adapted voltage during periodic tests in system operation. If distributed sensors for on-chip temperature are available for DVFS control, they can be exploited for an adaptive DLBIST scheme. During the periodic test phase, the BIST Control Unit (BCU) dynamically selects and applies the pre-generated test set corresponding to the current operating voltage and measured temperature. This adaptive selection ensures that testing conditions precisely match the real operating points. If temperature sensors are not available, testing at the so-called Zero Temperature Coefficient (ZTC) voltage is one alternative, which is the voltage where the temperature-induced variability is minimized. This makes periodic DLBIST a feasible solution for in-field self-testing, even in cases where on-chip temperature sensors are not available. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Electron. Test. | 3 |
| 2026 | CNN-Based Transistor Modeling Under Self-Heating EffectsabstractDue to the steady shrinking of technology node sizes, Self-Heating Effect (SHE) has become at the forefront of reliability and longevity concerns. Without careful consideration of the increased temperatures, reliability effects such as aging will be underestimated, putting the transistor and circuit at risk. SHE challenges and concerns are all amplified at cryogenic temperatures, as at low temperatures, multiple physical changes take effect, like the significant decrease of thermal conductivity, carriers’ mobility increase, and lower heat dissipation due to the low ambient temperature. The study of SHE at cryogenic temperatures is especially important in applications like quantum computing and space electronics. To accurately anticipate the SHE on the transistor, physics-based transistor and semiconductor simulation tools, Technology CAD (TCAD), are used. TCAD tools offer accurate, physics-based transistor SHE results, which designers can use to make informed decisions. However, TCAD tools can only deliver accurate SHE results after a manual, time-intensive calibration process. This calibration can only be done for operating temperatures and transistor configurations with existing experimental data. The reliance on experimental data for calibration and the constant need for recalibration restrict the scope of the design process to only a limited subset of configurations. In this work, we propose the first data-driven surrogate Convolutional Neural Network (CNN)-based TCAD thermal model for SHE thermal profile and channel hotspot prediction. Our approach enables designers to sweep a huge amount of different operating temperatures without the dependency on available experimental data or the constant need for model recalibration, while achieving a thermal profile prediction accuracy of 99.91%, a thermal profile extrapolation accuracy of 96.88%, hotspot prediction average error of 0.127%, and hotspot extrapolation average error of 1.65%. All results are compared to calibrated TCAD thermal model simulation results. Our model also alleviates the high computational cost and low throughput associated with TCAD tools by offering a time speedup of more than 13 000 x for both prediction and extrapolation tasks. Tarek Mohamed, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2026 | Evaluation of Radiation Resilience, Performance, and Vmin of Sub-3 nm FSFET Based SRAM ArraysabstractIn this work, we present single-event upset (SEU) analysis for Forksheet FET (FSFET) based CMOS circuits. Next, we present an array-level power and performance analysis along with the Vminevaluation for the FSFET-based SRAM. Physics based TCAD and industry-standard BSIM-CMG compact models are calibrated for accurate circuit analysis in SPICE. The impact of varying Heavy-Ion Radiation (HIR) doses and strike orientations is investigated for the FSFETs. The robustness of CMOS inverter against HIR is also reported in terms of failure time (tfail) and output voltage swing ($Δ$VDrop). For the SRAM, we determine the critical Linear Energy Transfer (LET). For FSFET, the individual n-/p-FETs are more vulnerable to the irradiation incident on nearby devices. At the circuit level, in comparison to perpendicular strikes, the$Δ$VDropincreases by 1.25V and 2.75V respectively, for oblique and transverse incidences, at a dose of 2.0MeVcm2/mg. The tfail also increases by 43% and 60% and the SRAM critical LET also decreases by 85% and 57.5%, respectively. The array level SRAM evaluation shows that the FSFET enables reliable operation with low-power consumption, impressive noise margins, and low minimum operating voltage (Vmin) values. FSFET SRAM power dissipation during the read and write operations is as low as 7.02$μ$W, and 3.00$μ$W respectively. At VDD=0.70V, the noise margins for hold, read, and write operations are 289.27mV, 122.89mV, and 297.79mV. The Vminfor read and write operations are 0.30V and 0.35V respectively. Hafeez Raza, Mahdi Benkhelifa, Koshal Kumar, Shivendra Singh Parihar, Yogesh Singh Chauhan, Hussam Amrouch, Avinash Lahgere |
IEEE Trans. Computers | 6 |
| 2026 | Ferroelectric Digital In-Memory Computing for Scalable, Reliable, and Efficient Similarity ComputationabstractClassification-based learning in deep neural networks, particularly few-shot learning, demands efficient similarity metrics such as Hamming distance. Conventional architectures suffer from high energy overheads due to frequent data movement between memory and processing units, hindering scalability. In-memory computing addresses this by integrating computation within memory, yet analog-based systems rely on power-hungry analog-to-digital converters (ADCs) and face scalability challenges due to device variability, especially in emerging memories. This work presents a fully digital Ferroelectric FET (FeFET)-based Logic-in-Memory (LiM) XOR cell, designed using GlobalFoundries’ 28 nm technology, eliminating ADCs and ensuring robust, energy-efficient, and scalable operation. Our 2T FeFET XOR cell, applied to 4096-bit Hamming distance calculations, achieves$23\times $lower energy,$3\times $faster latency, and$14\times $area reduction over state-of-the-art designs. Delivering 2337 Gsamples/(s$\cdot $W$\cdot $mm2) — a$300\times $improvement — this architecture offers a compelling solution for energy-efficient, reliable, and scalable AI hardware, driving sustainable computing. Anirban Kar, Albi Mema, Thorgund Nemec, Stefan Dünkel, Halid Mulaosmanovic, Sven Beyer, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2026 | Introduction to the Special Issue on Machine Learning for CAD, Part I
Yibo Lin, Siddharth Garg, Hussam Amrouch, Cong Hao |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | Pushing the Boundaries of AI Chips: From Monolithic 3D CMOS to Cryogenic ComputingabstractAs CMOS scaling approaches its fundamental limits, the explosive rise of AI and LLMs has unveiled profound bottlenecks in computing architectures. This paper presents two groundbreaking paradigms poised to reshape the landscape of high-performance computing and meet the surging demands of AI-driven workloads. The first paradigm is 3D monolithic integration, a revolutionary approach that achieves unprecedented logic density through Complementary FETs (CFETs), where pMOS and nMOS transistors are vertically stacked, and a dramatic expansion of on-chip memory capacity by integrating memory layers atop logic transistors. The second paradigm leverages the transformative potential of operating chips at cryogenic temperatures where transistors exhibit enhanced performance, and parasitic resistances are substantially minimized. These advancements hold the promise of redefining computing efficiency and performance for the AI era. Mahdi Benkhelifa, Shivendra Singh Parihar, Anirban Kar, Girish Pahwa, Yogesh Singh Chauhan, Hussam Amrouch |
DATE | 6 |
| 2025 | Late Breaking Results: Leveraging Approximate Computing for Carbon-Aware DNN AcceleratorsabstractThe rapid growth of Machine Learning (ML) has increased demand for DNN hardware accelerators, but their embodied carbon footprint poses significant environmental challenges. This paper leverages approximate computing to design sustainable accelerators by minimizing the Carbon Delay Product (CDP). Using gate-level pruning and precision scaling, we generate area-aware approximate multipliers and optimize the accelerator design with a genetic algorithm. Results demonstrate reduced embodied carbon while meeting performance and accuracy requirements. Aikaterini Maria Panteleaki, Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Iraklis Anagnostopoulos |
DATE | 4 |
| 2025 | Self-Aware Silicon: Enhancing Lifecycle Management with Intelligent Testing and Data Insights
Fabian Vargas 0001, Marko S. Andjelkovic, Milos Krstic, Anirban Kar, Swati Deshwal, Yogesh Singh Chauhan, Hussam Amrouch, Daniel Tille, Sebastian Huhn 0001 |
ETS | 7 |
| 2025 | Heterogeneous Integration of Advanced CMOS and Emerging Devices: Challenges and Solutions
Letícia Maria Veiras Bolzani, André Lucas Chinazzo, Mahdi Benkhelifa, Anirban Kar, Hussam Amrouch, Milos Krstic |
ETS | 5 |
| 2025 | Benchmarking Cryogenic Circuits using 5 nm FinFETs for Quantum ProcessingabstractQuantum computing offers the potential to solve problems that are intractable for classical computers. A major challenge in scaling quantum computers lies in bridging the gap between cryogenic qubits, operating at millikelvin to few kelvin temperatures, and the classical CMOS-based system-on-chip (SoC) typically located at room temperature (300K). This connection introduces heat leakage, which can destabilize the qubit states. A promising solution is to relocate the control circuits and processors to the cryogenic environment, but this imposes strict constraints on power consumption due to limited cooling capacity. Additionally, the SoC must meet stringent timing requirements for qubit measurement classification. In this work, we investigate the performance of CMOS-based circuits for cryogenic operations using 5 nm FinFET technology. We begin by measuring the electrical characteristics of advanced 5 nm FinFETs at both 10K and 300K. Using these measured data, we calibrate the industry-standard compact model (BSIM-CMG) and develop two standard cell libraries for each temperature. Through the logic synthesis of six circuits from the EPFL benchmark suite, we analyze their behavior at cryogenic temperatures. Our results show that circuits at 10K achieve a 41% increase in speed compared to 300K. Further, they operate efficiently at lower supply voltages, which enables reduced power consumption while maintaining high-speed performance in cryogenic environments. Anirban Kar, Shivendra Singh Parihar, Florian Klemme, Yogesh Singh Chauhan, Hussam Amrouch |
ISCAS | 5 |
| 2025 | Transistor-to-GDS Reliability Analysis in Sub-3nm: Impact of Self-Heating and Aging on TimingabstractAs transistor scaling advances into the sub-3 nm regime, self-heating effects (SHE) and aging-induced degradation emerge as profound challenges that threaten timing closure, signal integrity, guardbands, and long-term reliability. This work presents a comprehensive transistor-to-GDS reliability analysis that captures the impact of SHE and aging in nanosheet field-effect transistors (NSFETs) and propagates it through the entire design stack to full-chip signoff. We evaluate a 64-bit RISC-V processor core and an AI accelerator containing 4096 Multiply-and-Accumulate (MAC) units, both implemented using gate-all-around (GAA) NSFET technology. TCAD simulations, carefully calibrated against measurement data, reveal local temperature rises up to 124 K in multi-stack sheet structures, which exacerbate aging and result in a threshold voltage shift of up to 42.3 mV. Incorporating these effects into standard cell characterization and commercial signoff timing analysis uncovers substantial End-of-Life (EOL) timing degradation—37.7 % for the RISC-V core and 61.7 % for the AI accelerator—highlighting the urgent need for SHE- and aging-aware methodologies, as well as reliability-optimized standard cell libraries for advanced nodes. Swati Deshwal, Hadi Nour Eddine, Mahdi Benkhelifa, Albi Mema, Yogesh Singh Chauhan, Hussam Amrouch |
ISLPED | 6 |
| 2025 | Small Delay Fault Testing with Multiple Voltages under Variations: Defect vs. Fault CoverageabstractAbstract It has been known and explored for many years that low voltage testing amplifies the effect of a defect, increasing the size of a Small Delay Fault (SDF) and, in the best case, turning SDFs into easily detectable stuck-at-faults. It is often overlooked that $$V_{\textrm{min}}$$ V min testing poses an additional challenge to the test pattern generation method under process variations. The standard deviation of gate delays under $$V_{\textrm{min}}$$ V min is a multiple of that under nominal voltage. The increased variation will invalidate the efficiency of test patterns generated under nominal voltage and significantly reduce fault coverage. This paper presents the first algorithm for test pattern generation specifically tuned for $$V_{\textrm{min}}$$ V min testing which obtains higher fault coverage by smaller test sets than those generated for nominal voltage. The patterns applicable to other voltage levels can be derived from the pattern set generated under extreme variations at low supply voltage. Experimental results demonstrate that the proposed method produces test patterns that outperform N-detection test sets in terms of test set volume and fault efficiency across different voltage levels. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Electron. Test. | 3 |
| 2025 | Cryo-CACTI: Cryogenic-Aware CACTI for Cache Modeling Down to 10K in Advanced 7nm FinFETsabstractCryogenic circuits are currently employed in fields such as quantum computing, particle detectors, magnetic resonance imaging, and space applications. While cryogenic circuits are being researched, there is limited work on designing cryogenic caches at temperatures below 77K. Moreover, there is no tool to estimate the delay, power, and area of cryogenic caches at advanced technology nodes. Our research focuses on the development of cryogenic caches tailored for the 7nm technology node, operating at 10K. However, a key challenge is the lack of cryogenic measurement data, especially in recent technologies. Consequently, through conducting our own FinFET transistor measurements, we calibrate cryogenic transistor models at 10K. With the 7nm cryogenic transistor data, we modelCryo-CACTIfor cryogenic caches (due to cache’s vital role in improving performance and their considerable share in area and power of the processor). Using Cryo-CACTI, our evaluation reveals considerable improvements in the energy efficiency (up to 99%) of cryogenic caches of larger sizes compared to the caches at room temperature (300K). Additionally, we explore alternative cache configurations at circuit-level to optimize cryogenic operation. Furthermore, we use Cryo-CACTI to explore the performance/energy consumption of cryogenic caches while simulating workloads such as SPEC CPU2017 and machine learning via neural networks.Cryo-CACTI is available for download athttps://github.com/marg-tools/Cryo-CACTI Divya Praneetha Ravipati, Victor M. van Santen, Shivendra Singh Parihar, Yogesh Singh Chauhan, Preeti Ranjan Panda, Hussam Amrouch |
IEEE Trans. Computers | 6 |
| 2025 | On the Efficacy and Vulnerabilities of Logic Locking in Tree-Based Machine LearningabstractThe popularity and widespread usage of machine learning (ML) hardware have created challenges for its intellectual property (IP) protection. Logic locking is a widely used technique for IP protection but has received little attention in error-resilient applications such as ML hardware modules. This work investigates the effectiveness of logic locking when applied to tree-based ML circuits and reveals a critical vulnerability that undermines its effectiveness for single-label ML classifiers. We propose a logic locking scheme to eliminate the vulnerabilities in decision trees (DTs) and random forests (RFs) circuits. In our extensive simulation involving 16 DTs and 16 RFs, our solution consistently thwarts the vulnerability. We further evaluated the security of our approach by considering different obfuscation percentages and launching state-of-the-art oracle-less attacks on logic locking. Our method proves resilient, indicating that by fixing the identified vulnerability, we did not introduce new attack vectors. Further, our investigation indicates that DT/RF accelerators are significantly less vulnerable to oracle-less attacks compared to exact circuits. Overall, our work lays the foundation for future investigations into the effectiveness of logic locking for ML circuits. Brunno Abreu, Guilherme Paim, Lilas Alrahis, Paulo F. Flores, Ozgur Sinanoglu, Sergio Bampi, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2025 | A Lightweight PUF-Based Weights Obfuscation Technique for Secure In-Memory AI InferenceabstractIn-Memory Computing (IMC) has introduced a novel computational approach that substantially improves emerging embedded AI accelerators’ latency and power consumption efficiency. Despite the numerous advantages, IMC architectures also introduce new security vulnerabilities that may compromise the confidentiality of the deployed Neural Network (NN) algorithms. In this work, following an analysis of the potential threats, we present a novel lightweight security countermeasure for IMC accelerators. This methodology can be employed to de-obfuscate the pre-trained weights of NN architectures whose bits’ significance has been reordered prior to the deployment phase onto the IMC crossbar. The proposed solution is based on the coordinated action of a Ferroelectric Field-Effect Transistor (FeFET) based Physical Unclonable Function (PUF) design and shifting registers. These components perform custom arithmetic shift operations on the values calculated by the IMC device at runtime to obtain a coherent inference computation. Furthermore, a design-space exploration method is proposed to investigate the trade-off between area overhead and the level of security provided by the implementation. The results show that with less than 3% of area overhead our design is robust against all the tested attack strategies. Luca Parrini, Anirban Kar, Benjamin Hettwer, Taha Soliman, Yogesh Singh Chauhan, Hussam Amrouch, Norbert Wehn |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | Workload Compression Techniques to Scale Defect-Centric BTI Models to the Circuit LevelabstractBias Temperature Instability (BTI) poses a significant challenge in ensuring the reliability of digital systems, affecting the delay of digital logic gates, which ultimately can lead into timing failures. Sophisticated defect-centric models have been developed and successfully calibrated against empirical data to forecast the impacts of BTI at the device level. However, their application to large-scale digital circuits operating under realistic workloads over typical system lifetimes is limited because of the computational complexity of defect-centric models. To make the application of aging models in that context feasible, a useful technique is to compress the transistor workloads into simplified and hence manageable representative workloads. While fast in terms of execution speed, previous techniques struggle with accuracy when predicting aging degradation, and can reach a very high average error in threshold voltage increase prediction. In this work, we review the compression techniques described in the literature and propose two novel approaches that surpass existing ones in terms of accuracy, which is demonstrated for a complex digital design used as benchmark. Specifically, our best compression technique matches the predictions obtained through the reference uncompressed workloads, introducing negligible error, and maintains low execution times to efficiently and accurately scale defect-centric models to the circuit level. Andrés Santana-Andreo, Victor M. van Santen, Rafael Castro-López, Elisenda Roca, Hussam Amrouch, Francisco V. Fernández 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Domain-Specific Hyperdimensional RISC-V Processor for Edge-AI TrainingabstractEdge AI has become the cornerstone of many applications. Yet, progress is limited by the large complexity of training a deep neural network (a DNN). hyperdimensional computing (HDC) is positioned as an alternative approach for Edge AI that is compact enough to enable training. The main challenge for an HDC model is to maintain its key features while balancing high inference accuracy with efficiency. A simple binary HDC model lacks accuracy, while the computational complexity of a floating-point model is too high. This work presents FixedHD, a novel 16-bit fixed-point HDC model enabling training at the Edge. FixedHD achieves an accuracy similar to floating-point model while lowering computational complexity. The model is supported by a customized RISC-V processor tailored to speedup both training and inference. The processor is extended with advanced HDC-specific instructions, a vector unit to utilize HDC’s parallel nature, and, for the first time, approximate computing to exploit its robustness. Further, memory requirements are reduced by quantizing mathematical functions and reducing the large HDC encoding matrix by up to 390 x. Compared to the baseline processor, inference and training are accelerated on average by 6.9 x and 3 x, respectively. The energy consumption is reduced by 4.6 x and 1.9 x at the cost of an increase in area by 45 %. The inference accuracy remains at the high level of floating-point models despite the heavy quantization and approximation. Sandy A. Wasif, Miran Wael, Paul R. Genssler, Eman Azab, Maggie Mashaly, Mohamed Abdelghany, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | Optimized Detection of Marginal Defects in Standard Cells Using Unsupervised LearningabstractMarginal defects, such as high-resistance short or low-resistance open defects, are hard to detect by conventional pass-fail test methods because their manifestations are practically indistinguishable from the effects of regular variations. However, their coverage is essential for circuits with high-quality requirements and/or when early-life failures are a concern. In this paper, we propose an alternative detection concept based on evaluating several parametric responses of a circuit against a machine learning (ML) model. We use a 14nm FinFET transistor model validated against industrial measurements. We show that high detection performance is possible even when unsupervised learning that does not consider defective behavior is used; to this end, the procedure is generic. Moreover, an AUC score of over 0.96 is achieved when only measurements from a single voltage level are utilized, in contrast to earlier work. We also present a procedure to select a reduced set of test sequences, achieving an improvement of 50% reduction with a limited impact on detection performance. Karthik Pandaram, Hussam Amrouch, Ilia Polian |
ATS | 2 |
| 2024 | HDCircuit: Brain-Inspired HyperDimensional Computing for Circuit RecognitionabstractCircuits possess a non-Euclidean representation, necessitating the encoding of their data structure (e.g., gate-level netlists) into fixed formats like vectors. This work is the first to propose brain-inspired hyperdimensional computing (HDC) for optimized circuit encoding. HDC does not require extensive training to encode a gate-level netlist into a hypervector and simplifies the similarity check between circuits from graph-based to the similarity between their hypervectors. We introduce a versatile HDC-based encoding method for circuit encoding. We demonstrate its effectiveness with the application of circuit recognition using ITC-99 and ISCAS-85 benchmarks. We maintain a 98.2% accuracy, even when the designs are obfuscated using logic locking. Paul R. Genssler, Lilas Alrahis, Ozgur Sinanoglu, Hussam Amrouch |
DATE | 4 |
| 2024 | DropHD: Technology/Algorithm Co-Design for Reliable Energy-Efficient NVM-Based Hyper-Dimensional Computing Under Voltage ScalingabstractBrain-inspired hyperdimensional computing (HDC) offers much more efficient computing compared to other classical deep learning and related machine learning algorithms. Unlike classical CMOS, emerging non-volatile memories (NVMs) used in the realization of HDC are susceptible to failures under voltage scaling, which is essential for energy saving. Although HDC is inherently robust against errors, this is only possible when hypervectors with a large dimension (e.g., 10,000 bits) are being used, resulting in significant energy consumption. This work demonstrates, for the first time, that different NVM technologies exhibit different error characteristics under voltage scaling. In contrast to conventional CMOS-based SRAM, we demonstrate that the error behavior is data-dependent and not captured by simple bit flips in emerging NVMs. We employ our cross-layer framework that starts from the underlying technology all the way up to the algorithm to develop the novel HDC training approach DropHD. DropHD considerably shrinks the size of hypervectors (e.g., from 10,000 bits down to merely 3000 bits), while maintaining a high inference accuracy. The use of aggressive voltage scaling reduces energy consumption by 1.6 x. DropHD further reduces it to up to 9.5 × while fully recovering the induced accuracy drop, i.e. without a tradeoff. Paul R. Genssler, Mahta Mayahinia, Simon Thomann, Mehdi Baradaran Tahoori, Hussam Amrouch |
DATE | 5 |
| 2024 | Frontiers in Edge AI with RISC-V: Hyperdimensional Computing vs. Quantized Neural NetworksabstractHyperdimensional Computing (HDC) is an emerging paradigm that stands as a compelling alternative to conventional Deep Learning algorithms. HDC holds four key promises. First, the ability to learn from little data. Second, to be robust against noise in this data. HDC also promises to be resilient against errors in the underlying hardware. This includes the memory on which the model is stored and errors in the computations of the operations, which is attributed to the encoding of information across an expansive dimensional space. Fourth, HDC can be implemented efficiently in hardware due to its lightweight and embarrassingly parallel computations. In this work, those four key promises are evaluated in a holistic way. A fixed-point and a binary HDC implementation are compared against neural network implementations. The models are executed on a RISC-V processor to ensure a fair comparison. While the results confirm the ability to learn from little data and the resiliency against errors, the higher inference accuracy of neural networks favors them in most experiments. Based on these insights, we formulate challenges and opportunities for HDC. Our implementations for QNN, binary and fixed-point HDC are available online: https://github.com/TUM-AIPro/HDC-vs-QNN Paul R. Genssler, Sandy A. Wasif, Miran Wael, Rodion Novkin, Hussam Amrouch |
DATE | 5 |
| 2024 | Algorithm to Technology Co-Optimization for CiM-Based Hyperdimensional ComputingabstractHyperdimensional computing (HDC) has been recognized as an efficient machine learning algorithm in recent years. Robustness against noise and simple computational operations, while being limited by the memory bandwidth, make it a perfect fit for the concept of computation in memory (CiM) with emerging nonvolatile memory (NVM) technologies. For an HDC accelerator based on NVM-CiM, there are different parameters from the algorithm all the way down to the technology that interact with each other and affect the overall inference accuracy as well as the energy efficiency of the accelerator. Therefore, in this paper, we propose, for the first time, a full-stack co-optimization method and use it to design an HDC accelerator based on NVM-based content addressable memory (CAM). By incorporating the device manufacturing variability and co-optimizing the algorithm and hardware design, HDC inference on our proposed NVM-based CiM accelerator can reduce the energy consumption by 3.27x, while compared to the purely software-based implementation, the inference accuracy loss is merely 0.125%. Mahta Mayahinia, Simon Thomann, Paul R. Genssler, Christopher Münch, Hussam Amrouch, Mehdi Baradaran Tahoori |
DATE | 5 |
| 2024 | Low Power and Temperature- Resilient Compute-In-Memory Based on Subthreshold-FeFETabstractCompute-in-memory (CiM) is a promising solution for addressing the challenges of artificial intelligence (AI) and the Internet of Things (IoT) hardware such as “memory wall” issue. Specifically, CiM employing nonvolatile memory (NVM) devices in a crossbar structure can efficiently accelerate multiply-accumulation (MAC) computation, a crucial operator in neural networks among various AI models. Low power CiM designs are thus highly desired for further energy efficiency optimization on AI models. Ferroelectric FET (FeFET), an emerging device, is attractive for building ultra-low power CiM array due to CMOS compatibility, high ION /$I$O F F ratio, etc. Recent studies have explored FeFET based CiM designs that achieve low power consumption. Nevertheless, subthreshold-operated FeFETs, where the operating voltages are scaled down to the subthreshold region to reduce array power consumption, are particularly vulnerable to temperature drift, leading to accuracy degradation. To address this challenge, we propose a temperature-resilient 2T-1FeFET CiM design that performs MAC operations reliably at subthreahold region from 0°C to 85°C, while consuming ultra-low power. Benchmarked against the VGG neural network architecture running the CIFAR-10 dataset, the proposed 2T1FeFET CiM design achieves 89.45% CIFAR-10 test accuracy. Compared to previous FeFET based CiM designs, it exhibits immunity to temperature drift at an 8-bit wordlength scale, and achieves better energy efficiency with 2866 TOPS/W. Xuchu Huang, Jianyi Yang 0003, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
DATE | 5 |
| 2024 | Time and Space Optimized Storage-based BIST under Multiple Voltages and VariationsabstractLogic Built-In Self-Test (LBIST) with stored deterministic patterns is supported by the major CAD vendors and is gaining increasing attention, especially for safety-critical applications such as automotive. It is used for both manufacturing and periodic in-field testing. An unresolved challenge so far stems from the inevitable process variations. This paper presents the first approach for storage-based BIST addressing delay faults under process variations and multiple voltages. A unified solution for pattern generation, test set compaction and BIST hardware is presented that is compatible with commercial schemes. The solution significantly outperforms traditional N-detect for transition faults in terms of test set size, test application time and fault efficiency. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ETS | 3 |
| 2024 | TReCiM: Lower Power and Temperature-Resilient Multibit 2FeFET-1T Compute-in-Memory DesignabstractCompute-in-memory (CiM) emerges as a promising solution to solve hardware challenges in artificial intelligence (AI) and the Internet of Things (IoT), particularly addressing the "memory wall" issue. By utilizing nonvolatile memory (NVM) devices in a crossbar structure, CiM efficiently accelerates multiplyaccumulate (MAC) computations, the crucial operations in neural networks and other AI models. Among various NVM devices, Ferroelectric FET (FeFET) is particularly appealing for ultra-low-power CiM arrays due to its CMOS compatibility, voltage-driven write/read mechanisms and high ION/IOFF ratio. Moreover, subthreshold-operated FeFETs, which operate at scaling voltages in the subthreshold region, can further minimize the power consumption of CiM array. However, subthreshold-FeFETs are susceptible to temperature drift, resulting in computation accuracy degradation. Existing solutions exhibit weak temperature resilience at larger array size and only support 1-bit. In this paper, we propose TReCiM, an ultra-low-power temperature-resilient multibit 2FeFET-1T CiM design that reliably performs MAC operations in the subthreshold-FeFET region with temperature ranging from 0°C to 85°C at scale. We benchmark our design using NeuroSim framework in the context of VGG-8 neural network architecture running the CIFAR-10 dataset. Benchmarking results suggest that when considering temperature drift impact, our proposed TReCiM array achieves 91.31% accuracy, with 1.86% accuracy improvement compared to existing 1-bit 2T-1FeFET CiM array. Furthermore, our proposed design achieves 48.03 TOPS/W energy efficiency at system level, comparable to existing designs with smaller technology feature sizes. Thomas Kämpfe, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
ICCAD | 4 |
| 2024 | Minimizing PVT-Variability by Exploiting the Zero Temperature Coefficient (ZTC) for Robust Delay Fault TestingabstractProcess, Voltage, Temperature (PVT) variations impede the test generation for Small Delay Faults (SDFs) significantly as test patterns effective for one circuit instance may not be valid for a different one. Temperature-induced timing variations in FinFET and Gate-All-Around (GAA) technologies are especially severe due to temperature fluctuations and self-heating. Depending on the supply voltage, they show the Temperature Effect Inversion (TEI) which describes the increase of the circuit speed with increasing temperature. The Zero Temperature Coefficient (ZTC) specifies a supply voltage where TEI approaches 0, and the optimal voltage is determined, such that the effects of temperature-induced variability are minimized. Simulation results are reported, which demonstrate that test generation at the ZTC voltage leads to higher fault coverage of SDFs while using significantly less test patterns. Hanieh Jafarzadeh, Florian Klemme, Jan Dennis Reimer, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ITC | 4 |
| 2024 | Exploring BTI aging effects on spatial power density and temperature profiles of VLSI chips
Sachin Sachdeva, Jincong Lu, Hussam Amrouch, Sheldon X.-D. Tan |
Integr. | 3 |
| 2024 | Approximation- and Quantization-Aware Training for Graph Neural NetworksabstractGraph Neural Networks (GNNs) are one of the best-performing models for processing graph data. They are known to have considerable computational complexity, despite the smaller number of parameters compared to traditional Deep Neural Networks (DNNs). Operations-to-parameters ratio for GNNs can be tens and hundreds of times higher than for DNNs, depending on the input graph size. This complexity indicates the importance of arithmetic operation optimization within GNNs through model quantization and approximation. In this work, for the first time, we combine both approaches and implementquantization-andapproximation-aware trainingfor GNNs to sustain their accuracy under the errors induced by inexact multiplications. We employ matrix multiplication CUDA kernel to speed up the simulation of approximate multiplication within GNNs. Further, we demonstrate the execution speed, accuracy, and energy efficiency of GNNs with approximate multipliers in comparison with quantized low-bit GNNs. We evaluate the performance of state-of-the-art GNN architectures (i.e., GIN, SAGE, GCN, and GAT) on various datasets and tasks (i.e., Reddit-Binary, Collab for graph classification, Cora and PubMed for node classification) with a wide range of approximate multipliers. Our framework is available online:https://github.com/TUM-AIPro/AxC-GNN. Rodion Novkin, Florian Klemme, Hussam Amrouch |
IEEE Trans. Computers | 3 |
| 2024 | CAPE: Criticality-Aware Performance and Energy Optimization Policy for NCFET-Based CachesabstractCaches are crucial yet power-hungry components in present-day computing systems. With the Negative Capacitance Fin Field-Effect Transistor (NCFET) gaining significant attention due to its internal voltage amplification, allowing for better operation at lower voltages (stronger ON-current and reduced leakage current), the introduction of NCFET technology in caches can reduce power consumption without loss in performance. Apart from the benefits offered by the technology, we leverage the unique characteristics offered by NCFETs and propose a dynamic voltage scaling based criticality-aware performance and energy optimization policy (CAPE) for on-chip caches. We present the first work towards optimizing energy in NCFET-based caches with minimal impact on performance. Compared to operating at a nominal voltage of 0.7 V, CAPE shows improvement in Last-Level Cache (LLC) energy savings by up to 19.2%, while the baseline policies devised for traditional CMOS- (/FinFET-) based caches are ineffective in improving NCFET-based LLC energy savings. Compared to the considered baseline policies, our CAPE policy also demonstrates better LLC energy-delay product (EDP) and throughput savings. Divya Praneetha Ravipati, Ramanuj Goel, Victor M. van Santen, Hussam Amrouch, Preeti Ranjan Panda |
IEEE Trans. Computers | 4 |
| 2024 | In-Memory Acceleration of Hyperdimensional Genome Matching on Unreliable Emerging TechnologiesabstractNovel computer architectures like Compute-in-Memory (CiM) merge the memory and processing units, mimicking the human brain. Simultaneously, Hyperdimensional Computing (HDC) is emerging as a brain-inspired machine learning (ML) approach. Both developments hold promise for the realm of AI and computing, especially for genome-matching tasks, where large data movements overwhelm traditional von Neumann architectures. FeFET is one of the up-and-coming emerging technologies that promises to enable ultra-efficient and compact CiM architectures. However, the adoption of FeFETs is hindered by their 10 nm-thick Ferroelectric (FE) layer and process variation. Thus, calculations with FeFETs have errors (noise) that traditional ML genome-matching models cannot tolerate. To overcome these challenges, this work is the first one to i) present a reliable HDC framework (HDGIM) for highly-scaled (down to merely 3nm), multi-bit FeFET technology, ii) introduce temperature-thickness modeled noise from FeFET to the HDC system, and iii) extensively define the memorization capacity of HDC hyperparameters in order to evaluate the performance before deployment theoretically. Our novel HDC learning framework iteratively uses two models: a full-precision 32-bit HDC model, an ideal model for training, and a reduced bit-precision by a novel quantization method for validation and inference. Our results demonstrate that highly-scaled FeFET, realizing 3-bit and even 4-bit, can withstand any modeled noise given high dimensionality during inference. Considering the noise during model adjustment improves the inherent robustness by almost 9% on the 4-bit case. Hamza Errahmouni Barkam, Sanggeon Yun, Paul R. Genssler, Che-Kai Liu, Zhuowen Zou, Hussam Amrouch, Mohsen Imani |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | Graph Attention Networks to Identify the Impact of Transistor Degradation on Circuit ReliabilityabstractReliability is one of the key concerns in circuit design. The circuit must be able to tolerate transistor degradation to sustain reliability against timing failure. Whether a transistor is degraded due to noise, aging, or poor manufacturing, a circuit must uphold a timing error-free functionality over its entire projected lifetime. Transistors are hardened (designed stronger than necessary) to tolerate these degradations. However, hardening (e.g., widening the transistors) comes at the cost of additional area and power. Hence, it is necessary to identify and selectively harden specific transistors within a circuit. In this work, transistors that prolong a circuit’s delay when they are degraded are termed “susceptible”, and thus, are to be hardened. Identifying the susceptible transistors within a circuit is a complex task, for example, Monte Carlo circuit simulations require days to identify susceptible transistors in a single circuit. Consequently, current solutions are costly in terms of time and limited in their application. Instead, machine learning (ML) can offer a fast (inference in seconds) and universal (applicable to unseen circuits) alternative. However, traditional ML techniques struggle with inference on topology-based problems, while recent graph neural networks (GNNs) excel in these applications. Therefore, this work presents the first ML to classify susceptible transistors with GNNs. We use GNNs, specifically Graph Attention Networks (GAT), because the topology of the cell strongly affects how each transistor degradation affects performance. For instance, series-connected transistors amplify their impact, while parallel-connected ones can offset each other’s influence. Our GAT-based approach employs a heterogeneous graph in combination with GAT’s attention mechanism to capture the circuit’s topology and its impact on the analysis. Our evaluation demonstrates the capability of our approach to classifying transistors according to their impact within the ASAP7 standard cell library’s standard cells in mere 0.04 s (compared to days of Monte Carlo simulation time) while achieving 80.4% accuracy on unseen circuits. Tarek Mohamed, Victor M. van Santen, Lilas Alrahis, Ozgur Sinanoglu, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | WaSSaBi: Wafer Selection With Self-Supervised Representations and Brain-Inspired Active LearningabstractLarge datasets are often available for machine learning tasks. However, only very few contain labels for all the samples because labeling is a very labor-intensive process. Hence, large unlabeled datasets are available but inaccessible to traditional supervised learning methods. In this work, we combine two approaches to reduce the number of required labels. First, self-supervised learning (SSL) to utilize the large unlabeled dataset. SSL creates an encoder from those unlabeled samples that transforms the input into intermediate feature representations. Second, active learning is employed for the classification where labels are required. Active learning intelligently selects the most informative samples for manual labeling. Thus, it reduces the amount the labels required to achieve a high classification accuracy. The selected samples are used to train a brain-inspired hyperdimensional computing and random forest classifier. We demonstrate the outstanding performance of our approach with the example of wafer map defect pattern classification. It is a crucial diagnostic task helping to identify systematic problems in the manufacturing and improving yield. With our proposed method, a high 97% classification accuracy is achieved with only 6% of the labeled dataset for the first time. Our approach demonstrates the potential for training a machine learning model from less labeled samples by combining SSL with active learning. Karthik Pandaram, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Introduction to Special Issue on In/Near Memory and Storage Computing for Embedded Systems
Liang Shi 0001, Jingtong Shi, Hussam Amrouch, Kuan-Hsun Chen, Mengying Zhao, Weichen Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Introduction to the Special Issue on Design for Testability and Reliability of Security-aware HardwareabstractThe research on design for testability and reliability of security-aware hardware has been important in both academia and industry. With ever-growing globalization, commercial hardware design, manufacturing, transportation, and supply now involve many different countries, resulting in aggravated vulnerability from hardware design to manufacturing. Hardware with malicious purposes implanted from the third-party manufacturing process may control the operation of a circuit and tamper its functions, causing serious security issues. However, hardware includes not only devices and circuits but also systems. An important fact is that testability, reliability, and security technologies come from different design layers, but the impact evaluation is conducted at the system level. In other words, the testability, reliability, and security design of different layers can be carried out in a holistic manner to achieve optimization for the whole system. In addition, the testability, reliability, and security design technologies of each design layer can be collaboratively conducted to achieve better performance. The testability, reliability, and security tradeoff has garnered attention from academia and industry, particularly in the Post-Moore Era, due to the complexities and opportunities arising from new architectures and technologies. Tianming Ni, Xiaoqing Wen, Hussam Amrouch, Cheng Zhuo, Peilin Song |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Beyond von Neumann Era: Brain-Inspired Hyperdimensional Computing to the RescueabstractBreakthroughs in deep learning (DL) continuously fuel innovations that profoundly improve our daily life. However, DNNs overwhelm conventional computing architectures by their massive data movements between processing and memory units. As a result, novel computer architectures are indispensable to improve or even replace the decades-old von Neumann architecture. Nevertheless, going far beyond the existing von Neumann principles comes with profound reliability challenges for the performed computations. This is due to analog computing together with emerging beyond-CMOS technologies being inherently noisy and inevitably leading to unreliable computing. Hence, novel robust algorithms become a key to go beyond the boundaries of the von Neumann era. Hyper-dimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional DL and ML algorithms. Unlike conventional DL and ML algorithms, HDC is inherently robust against errors along a much more efficient hardware implementation. In addition to these advantages at hardware level, HDC's promise to learn from little data and the underlying algebra enable new possibilities at the application level. In this work, the robustness of HDC algorithms against errors and beyond von Neumann architectures are discussed. Further, the benefits of HDC as a machine learning algorithm are demonstrated with the example of outlier detection and reinforcement learning. Hussam Amrouch, Paul R. Genssler, Mohsen Imani, Mariam Issa, Xun Jiao 0002, Wegdan Mohammad, Gloria Sepanta |
ASP-DAC | 1 |
| 2023 | ML to the Rescue: Reliability Estimation from Self-Heating and Aging in Transistors All the Way up ProcessorsabstractWith increasingly confined 3D structures and newly-adopted materials of higher thermal resistance, transistor self-heating has risen to a critical reliability threat in state-of-the-art and emerging process nodes. One of the challenges of transistor self-heating is accelerated transistor aging, which leads to earlier failure of the chip if not considered appropriately. Nevertheless, adequate consideration of accelerated aging effects, induced by self-heating, throughout a large circuit design is profoundly challenging due to the large gap between where self-heating does originate (i.e., at the transistor level) and where its ultimate effect occurs (i.e., at the circuit and system levels). In this work, we demonstrate an end-to-end workflow starting from self-heating and aging effects in individual transistors all the way up to large circuits and processor designs. We demonstrate that with our accurately estimated degradations, the required timing guardband to ensure reliable operation of circuits is considerably reduced by up to 96% compared to otherwise worst-case estimations that are conventionally employed. Hussam Amrouch, Florian Klemme |
ASP-DAC | 1 |
| 2023 | Tutorial: The Synergy of Hyperdimensional and In-Memory Computing
Paul R. Genssler, Simon Thomann, Hussam Amrouch |
CODES+ISSS | 3 |
| 2023 | Compact and High-Performance TCAM Based on Scaled Double-Gate FeFETsabstractTernary content addressable memory (TCAM), widely used in network routers and high-associativity caches, is gaining popularity in machine learning and data-analytic applications. Ferroelectric FETs (FeFETs) are a promising candidate for implementing TCAM owing to their high ON/OFF ratio, non-volatility, and CMOS compatibility. However, conventional single-gate FeFETs (SG-FeFETs) suffer from relatively high write voltage, low endurance, potential read disturbance, and face scaling challenges. Recently, a double-gate FeFET (DG-FeFET) has been proposed and outperforms SG-FeFETs in many aspects. This paper investigates TCAM design challenges specific to DG-FeFETs and introduces a novel 1.5T1Fe TCAM design based on DG-FeFETs. A 2-step search with early termination is employed to reduce the cell area and improve energy efficiency. A shared driver design is proposed to reduce the peripherals area. Detailed analysis and SPICE simulation show that the 1.5T1Fe DGTCAM leads to superior search speed and energy efficiency. The 1.5T1Fe TCAM design can also be built with SG-FeFETs, which achieve search latency and energy improvement compared with 2FeFET TCAM. Liu Liu 0023, Simon Thomann, Hussam Amrouch, Xiaobo Sharon Hu |
DAC | 4 |
| 2023 | Design Automation for Cryogenic CMOS CircuitsabstractCryogenic CMOS circuits operate at temperatures close to absolute zero and are essential in many applications such as controllers for quantum computing but also medical engineering, space technology, or physical instruments. However, operating circuits at cryogenic temperatures fundamentally changes the underlying semiconductor physics that governs the CMOS transistor—rendering existing design automation approaches infeasible. In this work, we propose and implement the first end-to-end approach that enables design automation for cryogenic CMOS circuits. To this end, we (1) perform the first-of-its-kind measurements of commercial 5nm FinFET transistors from 300K down to 10K, (2) use the results to validate and calibrate the first cryogenic-aware industrial-standard compact model for FinFET technology, (3) create cryogenic-aware standard cell libraries that are compatible with the existing EDA tool flows, and (4) propose an initial cryogenic-aware logic synthesis approach that re-uses established design automation expertise but optimizes it for cryogenic purposes. Evaluations, comparisons, and discussions of all these novel contributions confirm the applicability and validity of the resulting cryogenic-aware design automation flow. Victor M. van Santen, Marcel Walter, Florian Klemme, Shivendra Singh Parihar, Girish Pahwa, Yogesh Singh Chauhan, Robert Wille, Hussam Amrouch |
DAC | 8 |
| 2023 | HDGIM: Hyperdimensional Genome Sequence Matching on Unreliable highly scaled FeFETabstractThis is the first work to present a reliable application for highly scaled (down to merely 3nm), multi-bit Ferroelectric FET (FeFET) technology. FeFET is one of the up-and-coming emerging technologies that is not only fully compatible with the existing CMOS but does hold the promise to realize ultra-efficient and compact Compute-in-Memory (CiM) architectures. Nevertheless, FeFETs struggle with the 10nm thickness of the Ferroelectric (FE) layer. This makes scaling profoundly challenging if not impossible because thinner FE significantly shrinks the memory window leading to large error probabilities that cannot be tolerated. To overcome these challenges, we propose HDGIM, a hyperdimensional computing framework catered to FeFET in the context of genome sequence matching. Genome Sequence Matching is known to have high computational costs, primarily due to huge data movement that substantially overwhelms von-Neuman architectures. On the one hand, our cross-layer FeFET reliability modeling (starting from device physics to circuits) accurately captures the impact of FE scaling on errors induced by process variation and inherent stochasticity in multi-bit FeFETs. On the other hand, our HDC learning framework iteratively adapts by using two models, a full-precision, ideal model for training and a quantized, noisy version for validation and inference. Our results demonstrate that highly scaled FeFET realizing 3-bit and even 4-bit can withstand any noise given high dimensionality during inference. If we consider the noise during model adjustment, we can improve the inherent robustness compared to adding noise during the matching process. Hamza Errahmouni Barkam, Sanggeon Yun, Paul R. Genssler, Zhuowen Zou, Che-Kai Liu, Hussam Amrouch, Mohsen Imani |
DATE | 6 |
| 2023 | Upheaving Self-Heating Effects from Transistor to Circuit Level using Conventional EDA Tool FlowsabstractIn this work, we are the first to demonstrate how well-established EDA tool flows can be employed to upheave Self- Heating Effects (SHE) from individual devices at the transistor level all the way up to complete large circuits at the final layout (i.e., GDS-II) level. Transistor SHE imposes an ever-growing reliability challenge due to the continuous shrinking of geometries alongside the non-ideal voltage scaling in advanced technology nodes. The challenge is largely exacerbated when more confined 3D structures are adopted to build transistors such as upcoming Nanosheet FETs and Ribbon FETs. By employing increasingly-confined structures and materials of poorer thermal conductance, heat arising within the transistor's channel is trapped inside and cannot escape. This leads to accelerated defect generation and, if not considered carefully, a profound risk to IC reliability. Due to the lack of EDA tool flows that can consider SHE, circuit designers are forced to take pessimistic worst-case assumptions (obtained at the transistor level) to ensure reliability of the complete chip for the entire projected lifetime - at the cost of sub-optimal circuit designs and considerable efficiency losses. Our work paves the way for designers to estimate less pessimistic (i.e., small yet sufficient) safety margins for their circuits leading to higher efficiency without compromising reliability. Further, it provides new perspectives and opens new doors to estimate and optimize reliability correctly in the presence of emerging SHE challenge through identifying early the weak spots and failure sources across the design. Florian Klemme, Sami Salamin, Hussam Amrouch |
DATE | 3 |
| 2023 | Robust Resistive Open Defect Identification Using Machine Learning with Efficient Feature SelectionabstractResistive open defects in FinFET circuits are reliability threats and should be ruled out before deployment. The performance variations due to these defects are similar to the effect of process variations which are mostly benign. In order not to sacrifice yield for reliability the effect of defects should be distinguished from process variations. It has been shown that machine learning (ML) schemes are able to classify defective circuits with high accuracy based on the maximum frequencies$F_{max}$obtained under multiple supply voltages$V_{dd} \in V_{op}$. The paper at hand presents a method to minimize the number of required measurements. Each supply voltage$V_{dd}$defines a feature$F_{max}(V_{dd})$. A feature selection technique is presented, which uses also the already available$F_{max}$measurements. It is shown that ML-based techniques can work efficiently and accurately with this reduced number of$F_{max}(V_{dd})$measurements. Zahra Paria Najafi-Haghi, Florian Klemme, Hanieh Jafarzadeh, Hussam Amrouch, Hans-Joachim Wunderlich |
DATE | 4 |
| 2023 | Learning-Oriented Reliability Improvement of Computing Systems From Transistor to Application LevelabstractDue to technology scaling in modern computing platforms, the safety and reliability issues have increased tremendously, which often accelerate aging, lead to permanent faults, and cause unreliable execution of applications. Failure in some computing systems like avionics may cause catastrophic consequences. Therefore, managing reliability under all circumstances of stress and environmental changes is crucial in all abstraction layers, from application to transistor levels. Machine learning techniques are recently being employed for dynamic reliability estimation and optimization. They can adapt to varying workloads and system conditions. This paper presents reliability improvement approaches from multiple perspectives-from transistor-level to application-level-and discusses their effectiveness and limitations as well as open challenges. Behnaz Ranjbar, Florian Klemme, Paul R. Genssler, Hussam Amrouch, Jinhyo Jung, Shail Dave, Hwisoo So, Kyongwoo Lee, Aviral Shrivastava, Ji-Yung Lin, Pieter Weckx, Subrat Mishra, Francky Catthoor, Dwaipayan Biswas, Akash Kumar 0001 |
DATE | 4 |
| 2023 | Stress-Resiliency of AI Implementations on FPGAsabstractFPGAs have become a popular choice for machine learning acceleration for both cloud and edge devices. While traditional neural networks show impressive performance in classification tasks, Hyperdimensional Computing (HDC) is rapidly emerging as a promising novel machine learning approach for its hardware-friendly inference. In HDC, classes are embedded into high-dimensional vectors during training, and inputs can be classified by computing similarity metrics between class-vectors during inference. HDC inference is especially promoted in terms of its resiliency against errors, attributed to the large inherent redundancy. In this work, we perform a thorough experimental investigation of the fault resiliency of various FPGA-based machine learning implementations under different aspects of stress, comparing HDC with classical neural network approaches. We explore both the amount of faulty classifications as well as system crashes while subjecting the designs to timing stress using overclocking, voltage stress with excessive switching activity, and thermal stress. Jonas Krautter, Paul R. Genssler, Gloria Sepanta, Hussam Amrouch, Mehdi Baradaran Tahoori |
FPL | 4 |
| 2023 | Reliable Hyperdimensional Reasoning on Unreliable Emerging TechnologiesabstractWhile Graph Neural Networks (GNNs) have demonstrated remarkable achievements in knowledge graph reasoning, their computational efficiency on conventional computing platforms is impeded by the memory wall problem. To overcome these challenges, we introduce an innovative algorithm-hardware solution that harnesses the potential of hyperdimensional computing (HDC) for robust and memory-centric computation on computing in-memory (CiM) platforms. Departing from traditional graph neural networks, the proposed HDC reasoning model employs a symbolic approach to effectively encode graph entities and their relationships as high-dimensional neural activity. Complementing this approach is a customized Computing-in-Memory (CiM) architecture based on advanced Ferroelectric Field-Effect Transistor (FeFET) technology, which incorporates a precise characterization of non-idealities. This modeling enables the generation of an HDC-tailored model that faithfully represents the hardware architecture. Despite the non-idealities inherent in emerging CiM technologies, our platform demonstrates performance on par with traditional von Neumann architectures for substantial combinations of FeFET device parameters. Our solution overcomes FeFET CiM the increased non-idealities from down-scaled 3nm, operating effectively under all possible configurations when 50 graph edges are considered. Scenarios with less than 4-bit precision per FeFET device cannot handle graphs with more than 200 edges, whereas the 4-bit case can achieve a 90.3% graph reconstruction rate on the worst-case scenario of 80% of noise. Hamza Errahmouni Barkam, Sanggeon Yun, Hanning Chen, Paul Gensler, Albi Mema, Andrew Ding, George Michelogiannakis, Hussam Amrouch, Mohsen Imani |
ICCAD | 8 |
| 2023 | Invited Paper: Ultra-Efficient Edge AI Using FeFET-based Monolithic 3D IntegrationabstractMonolithic three-dimensional (M3D) integration signifies a notable technological leap by providing solutions of high density and energy efficiency, particularly for ultra-efficient edge AI solutions. Among emerging technologies, ferroelectric thin-film transistors (FeTFTs) have attracted substantial interest due to their potential in neuromorphic computing and their compatibility with back-end-of-the-line (BEOL) fabrication processes. Nevertheless, the challenge of M3D integrated circuits lies in elevated temperatures resulting from limited heat dissipation across various stacked tiers, subsequently exerting adverse effects on system reliability. In this work, we explore how compute-in-memory (CIM) architectures can be realized using BEOL FeTFT to accelerate deep learning applications. We demonstrate how generated temperature impacts the reliability and ultimately degrade the inference accuracy. To achieve that, we have modified the open-source “3D+NeuroSim” framework to introduce FeTFT and then employed it to estimate the key figures of merit, such as throughput, power density, area, etc. Then, we perform a comprehensive thermal analysis for the simulated M3D architecture to explore how the DNN accuracy will be impacted due to the generated heat. We also demonstrate how FeTFT devices can be accurately modeled using TCAD simulations and how run-time variability (due to temperature effects) and design-time variability (due to process variation) impact the reliability of FeTFT transistors. Yogesh Singh Chauhan, Hussam Amrouch |
ICCAD | 3 |
| 2023 | Comprehensive Reliability Analysis of 22nm FDSOI SRAM from Device Physics to Deep LearningabstractThis work investigates the joint impact of device variability and transistor aging on the data integrity of SRAM cells implemented using 22 FDSOI. Our analysis is based on well-calibrated TCAD simulations that reproduce measurements from a commercial 22nm FDSOI technology node. The calibrations are done against measurement data for both I-V characteristics and variability data. We perform error analysis for SRAMs during hold and read operations under three different scenarios: (i) Fresh: time-zero variation (PV) alone caused by manufacturing variability, (ii) Aged: combined impact of PV and aging-induced increase in the transistor threshold voltage ($V_{TH}$) at the room temperature, (iii) Aged@85°C: combined impact of PV and transistor aging but at an elevated temperature of 85°C. Further, we explore how SRAM errors are exacerbated when the voltage is scaled down due to the reductions in noise margins. All error analyses were accurately performed in TCAD mixed-mode simulations for a complete 6-T SRAM cell. Finally, to investigate further how such errors impact the system level, we explore the corresponding induced accuracy drop in Deep Neural Networks (DNNs). Different quantized NNs are studied, and their sensitivity to errors in weights and activations is also explored. We demonstrate that short-term aging (i.e., when aging effects are combined with voltage scaling) results in a noticeable accuracy drop when ResNet20 and ResNet18 DNN models are examined on the CIFAR100 and Imagenet datasets, respectively. Om Prakash 0007, Rodion Novkin, Virinchi Roy Surabhi, Prashanth Krishnamurthy, Ramesh Karri, Farshad Khorrami, Hussam Amrouch |
ISCAS | 7 |
| 2023 | Temperature-Aware Memory Mapping and Active Cooling of Neural Processing UnitsabstractNeural processing units (NPUs) have become indispensable for meeting the high computational demands of deep neural networks (DNNs). They provide a very efficient solution, thanks to having a huge MAC array that enables massive parallelism. Nevertheless, such an architecture exhibits excessive on-chip power densities leading to a localized hot-spot that seriously heats its surroundings. This work demonstrates how the on-chip temperatures induced by the MAC array create a spatial thermal gradient through the on-chip SRAM memory. This makes the memory regions sensitive to different error probabilities (Perror), leading to significant accuracy drops when DNNs are being executed. To surmount this challenge, we employ on-chip superlattice thermoelectric (TEC) cooling devices that effectively reduce the memory temperature. Although scaling the memory voltage makes SRAM cells more sensitive to errors, it significantly decreases the leakage power, which compensates for the power consumed by the incorporated TEC devices. Furthermore, operating the SRAM at a lower voltage and temperature substantially increases its lifetime because voltage and temperature are key stimuli of transistor aging. By running multi-physics simulations using commercial finite-element tools and SPICE simulations for the 14nm FinFET technology, we accurately derive the relation between the Perror in different memory regions and the corresponding cooling cost. We then propose a three-stage temperature-aware layer-wise memory mapping that exploits different degrees of the sensitivity of NN layers to errors towards maximizing the DNN accuracy while minimizing the cooling cost. Experimental results reveal that our method notably improves the DNN accuracy compared to existing temperature-oblivious memory mapping. Vahidreza Moghaddas, Hammam Kattan, Tim Bücher, Mikail Yayla, Jian-Jia Chen, Hussam Amrouch |
ISLPED | 6 |
| 2023 | Robust Pattern Generation for Small Delay Faults Under Process VariationsabstractSmall Delay Faults (SDFs) introduce additional delays smaller than the capture time and require timing-aware test pattern generation. Since process variations can invalidate the effectiveness of such patterns, different circuit instances may show a different fault coverage for the same test pattern set. This paper presents a method to generate test pattern sets for SDFs which are valid for all circuit timings. The method overcomes the limitations of known timing-aware Automatic Test Pattern Generation (ATPG) which has to use fault sampling under process variations due to the computational complexity. A statistical learning scheme maximises the coverage of SDFs in circuits following the variation parameters of a calibrated industrial FinFET transistor model. The method combines efficient ATPG for Transition Faults (TFs) with fast timing-aware fault simulation on GPUs. Simulation experiments show that the size of the pattern set is significantly reduced in comparison to standard N-detection while the fault coverage even increases. Hanieh Jafarzadeh, Florian Klemme, Jan Dennis Reimer, Zahra Paria Najafi-Haghi, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ITC | 5 |
| 2023 | Analysis and Characterization of Defects in FeFETsabstractEmerging devices are susceptible to manufacturing defects due to immature fabrication processes. Ferroelectric field-effect transistors, referred to as FeFETs, are promising emerging devices, but the impact of manufacturing imperfections on these devices has yet to be studied. Thus, we combine a technology CAD (TCAD) model with a fault-injection technique to represent fabrication defects in a FeFET. The TCAD model is calibrated against a fabricated metal-ferroelectric-metal capacitor and uses a multi-domain ferroelectric-layer structure. We address two classes of defects in the ferroelectric layer and map them to stuck-at-fault models referred to as neutral faults (SAP°) and stuck-at-plus and stuck-at-minus (SAP+and SAP−) faults. We also develop a machine-learning (ML) framework to characterize these fault-injected FeFET devices. The ML framework provides a significant speedup in predicting the health of the FE layer as compared to computationally heavy TCAD simulations. Our study of defects in ferroelectric FET (FeFET), which is done for the first time, and the insights gained thereof can provide valuable feedback for the fabrication and yield learning of FeFET-based circuits. Dhruv Thapar, Simon Thomann, Arjun Chaudhuri, Hussam Amrouch, Krishnendu Chakrabarty |
ITC | 4 |
| 2023 | SyncTREE: Fast Timing Analysis for Integrated Circuit Design through a Physics-informed Tree-based Graph Neural NetworkabstractNowadays integrated circuits (ICs) are underpinning all major information technology innovations including the current trends of artificial intelligence (AI). Modern IC designs often involve analyses of complex phenomena (such as timing, noise, and power etc.) for tens of billions of electronic components, like resistance (R), capacitance (C), transistors and gates, interconnected in various complex structures. Those analyses often need to strike a balance between accuracy and speed as those analyses need to be carried out many times throughout the entire IC design cycles. With the advancement of AI, researchers also start to explore news ways in leveraging AI to improve those analyses. This paper focuses on one of the most important analyses, timing analysis for interconnects. Since IC interconnects can be represented as an RC-tree, a specialized graph as tree, we design a novel tree-based graph neural network, SyncTREE, to speed up the timing analysis by incorporating both the structural and physical properties of electronic circuits. Our major innovations include (1) a two-pass message-passing (bottom-up and top-down) for graph embedding, (2) a tree contrastive loss to guide learning, and (3) a closed formular-based approach to conduct fast timing. Our experiments show that, compared to conventional GNN models, SyncTREE achieves the best timing prediction in terms of both delays and slews, all in reference to the industry golden numerical analyses results on real IC design data. Jiajie Li 0002, Florian Klemme, Gi-Joon Nam, Tengfei Ma 0001, Hussam Amrouch, Jinjun Xiong |
NeurIPS | 6 |
| 2023 | Frontiers in AI Acceleration: From Approximate Computing to FeFET Monolithic 3D IntegrationabstractWith the rapidly expanding applications of artificial intelligence (AI), the quest for hardware acceleration to foster high-speed and energy-efficient AI computation has become ever more important. In this work, we first explore the performance and energy advantages of employing classical AI acceleration with conventional systolic multiply-accumulate (MAC) arrays. We then highlight the growing importance of monolithic 3D integration as a transformative hardware acceleration strategy, moving beyond the constraints of classical von Neumann architectures. We also discuss how brain-inspired hyperdimensional computing (HDC) offers an exciting avenue for overcoming the power-hungry requirements often associated with MAC arrays, which are inevitable in deep learning hardware. Addressing the limitations of von Neumann architectures, we present the potential of monolithic 3D integration to enable ultra-dense Processing-in-Memory (PiM) layers stacked on top of high-performance CMOS logic. This novel approach offers to enhance computational performance. Recognizing the need for compatibility with low thermal budgets, we identify ferroelectric thin-film transistors (FeTFT) as a promising candidate for back-end-ofline (BEOL) fabrication. We highlight recent advances in BEOL FeTFT technology and demonstrate how technology/algorithm co-optimization plays a crucial role in the successful realization of reliable brain-inspired HDC on potentially unreliable FeTFT-based PiM layers. Our results showcase the potential of these innovations for the development of next-generation, energy-efficient AI hardware. Paul R. Genssler, Somaya Mansour, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 5 |
| 2023 | Reliable Brain-inspired AI Accelerators using Classical and Emerging MemoriesabstractBy taking inspiration from the operation of biological brains, emerging brain-inspired hardware has the potential to revolutionize the way computations are performed. Brain-inspired computing can be realized using both classical CMOS and emerging beyond-CMOS technologies, whereas the latter holds the promise to provide substantial energy savings akin to the employment of non-volatile memories. One way to implement highly efficient brain-inspired AI applications is through analog computing schemes, such as Integrate-and-Fire (IF) Spiking Neural Networks (SNNs), which can be implemented using both CMOS and beyond-CMOS technologies as synaptic storage. However, managing the inherent degradation of computing accuracy in analog circuits and mitigating their effects on the predictive accuracy of AI systems remains a key challenge due to the inherent nature of analog computing.In this paper, we discuss how the aforementioned challenges can be addressed. In the first part, we present our SPICE-Torch, a framework that connects low-level SPICE simulations of circuits and memories performing analog computations with high-level accuracy evaluations of NN models based on PyTorch. Furthermore, we present an example of neuromorphic optimization using classical CMOS technology. In the second part, we introduce memristors as an emerging beyond-CMOS technology that can retain their state without any outside influence and are well-suited for brain-inspired neuromorphic hardware. We demonstrate that brain-inspired hardware, realized using classical CMOS or beyond-CMOS technologies, has the potential to revolutionize the way we process information and solve complex computation problems. Nevertheless, to harness its full potential, reliability issues have to be managed carefully and HW/SW codesign is key. Our presented framework SPICE-Torch, which connects low-level SPICE simulations of circuits performing analog computations with high-level accuracy evaluations of NN models based on PyTorch is available as open-source in https://github.com/myay/SPICE-Torch. Mikail Yayla, Simon Thomann, Md. Mazharul Islam 0006, Ming-Liang Wei, Shu-Yin Ho, Ahmedullah Aziz, Chia-Lin Yang, Jian-Jia Chen, Hussam Amrouch |
VTS | 9 |
| 2023 | Hot-spot aware thermoelectric array based cooling for multicore processors
Sheriff Sadiqbatcha, Liang Chen 0025, Cuong Thi, Sachin Sachdeva, Hussam Amrouch, Sheldon X.-D. Tan |
Integr. | 6 |
| 2023 | Machine Learning-Based Microarchitecture- Level Power Modeling of CPUsabstractEnergy efficiency has emerged as a key concern for modern processor design, especially when it comes to embedded and mobile devices. It is vital to accurately quantify the power consumption of different micro-architectural components in a CPU. Traditional RTL or gate-level power estimation is too slow for early design-space exploration studies. By contrast, existing architecture-level power models suffer from large inaccuracies. Recently, advanced machine learning techniques have been proposed for accurate power modeling. However, existing approaches still require slow RTL simulations, have large training overheads or have only been demonstrated for fixed-function accelerators and simple in-order cores with predictable behavior. In this work, we present a novel machine learning-based approach for microarchitecture-level power modeling of complex CPUs. Our approach requires only high-level activity traces obtained from microarchitecture simulations. We extract representative features and develop low-complexity learning formulations for different types of CPU-internal structures. Cycle-accurate models at the sub-component level are trained from a small number of gate-level simulations and hierarchically composed to build power models for complete CPUs. We apply our approach to both in-order and out-of-order RISC-V cores. Cross-validation results show that our models predict cycle-by-cycle power consumption to within 3% of a gate-level power estimation on average. In addition, our power model for the Berkeley Out-of-Order (BOOM) core trained on micro-benchmarks can predict the cycle-by-cycle power of real-world applications with less than 3.6% mean absolute error. Ajay Krishna Ananda Kumar, Sami Salamin, Hussam Amrouch, Andreas Gerstlauer |
IEEE Trans. Computers | 3 |
| 2023 | Massively Parallel Circuit Setup in GPU-SPICEabstractSPICE simulations are the industry standard to analyze circuits for decades. However, they are computationally complex as each circuit is simulated at the transistor-level where individual transistor is modeled with dozens of sophisticated equations. This limits the practicality of SPICE simulations to relatively small circuits. However, this is in a direct conflict with the ever-increasing demands of circuit designers in which SPICE simulations for large circuits (e.g., DSPs, AES, etc.) at full accuracy are inevitably required to fulfill new industrial standards like automotive safety ISO 26262 with tool confidence level 1. To accelerate SPICE simulation without sacrificing accuracy, state-of-the-art approaches have started to employ GPUs to parallelize the LU-factorization and device linearization phases. Instead of focusing on these phases, this article demonstrates for the first time that when large circuits come into play, a new and equally important performance bottleneck emerges at the circuit setup phase. Speeding up the circuit setup phase in SPICE is our key focus in this paper. Our two implementations demonstrate that our GPU-based circuit setup reduces the analysis time from 4.5 days to merely 89 seconds for a 256-bit multiplier, which consists of more than 1M transistors. Our achieved speedup is 4396x compared to the baseline (open-source NGSPICE) and more than 2x compared to commercial (HSPICE and Spectre) SPICE circuit setup. Victor M. van Santen, Fu Lam Florian Diep, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 4 |
| 2023 | HW/SW Co-Design for Reliable TCAM- Based In-Memory Brain-Inspired Hyperdimensional ComputingabstractBrain-inspired hyperdimensional computing (HDC) is continuously gaining remarkable attention. It is a promising alternative to traditional machine-learning approaches due to its ability to learn from little data, lightweight implementation, and resiliency against errors. However, HDC is overwhelmingly data-centric similar to traditional machine-learning algorithms. In-memory computing is rapidly emerging to overcome the von Neumann bottleneck by eliminating data movements between compute and storage units. In this work, we investigate and model the impact of imprecise in-memory computing hardware, namely TCAM cells, on the inference accuracy of HDC. Our modeling is based on 14nm FinFET technology fully calibrated with Intel measurement data. We accurately model, for the first time, the voltage-dependent error probability in SRAM-based and FeFET-based in-memory computing. Thanks to HDC's resiliency against errors, the complexity of the underlying hardware can be reduced, providing large energy savings of up to 6x. Experimental results for SRAM reveal that variability-induced errors have a probability of up to 39%. Despite such a high error probability, the inference accuracy is only marginally impacted. This opens doors to explore new tradeoffs. We also demonstrate that the resiliency against errors is application-dependent. In addition, we investigate the robustness of HDC against errors with emerging non-volatile FeFET devices instead of mature CMOS-based SRAMs. We demonstrate that inference accuracy does remain high despite the larger error probability, while large area and power savings can be obtained.All in all, HW/SW co-design is the key for efficient yet reliable in-memory HDC for both conventional CMOS technology and upcoming emerging technologies. Simon Thomann, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Computers | 3 |
| 2023 | Golden-Free Robust Age Estimation to Triage Recycled ICsabstractNondestructive golden-free detection of recycled/counterfeit integrated circuits (ICs) is the focus of this article. This is achieved by estimating the functional/operational age of the IC. The age estimation method is based on exploiting short-term aging effects in advanced transistor technologies to induce bit errors at the IC’s output. Gate-level simulations are used to capture the impact of workload on short-term aging. In advanced technology nodes, including bulk CMOS at 45 nm or below and FinFET, combining transistor aging with ultrafast voltage scaling magnifies the effects of aging-induced degradation at high voltage when voltage scales to a lower level, causing short-term aging-based timing violations. These timing violations create bit errors at IC outputs. We employ the bit error patterns to build a machine learning (ML)-based nonlinear regression model to estimate the IC’s age. Our study confirms that short-term aging-induced output bit error patterns can be used to estimate long-term age of an IC. If the IC’s age is beyond a predefined threshold, it can be marked as recycled. Although this article considers the FinFET technology, the method applies to bulk CMOS advanced nodes at 45 nm or below. We model IC-to-IC variations taking into account the voltage scaling. We demonstrate the approach on two cryptographic ICs and the method accurately estimates the long-term age of an IC, facilitating recycled IC detection. Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Modeling and Predicting Transistor Aging Under Workload Dependency Using Machine LearningabstractThe pivotal issue of reliability is one of the major concerns for circuit designers. The driving force is transistor aging, dependent on operating voltage and workload. At the design time, it is difficult to estimate close-to-the-edge guardbands that keep aging effects during the lifetime at bay. This is because the foundry does not share its calibrated physics-based models, comprised of highly confidential technology and material parameters. However, the unmonitored yet necessary overestimation of degradation amounts to a performance decline, which could be preventable. Furthermore, these physics-based models are computationally complex. The costs of modeling millions of individual transistors at design time can be exorbitant. We propose the use of a machine learning model trained to replicate the physics-based model, such that no confidential parameters are disclosed. This effectual workaround is fully accessible to circuit designers for the purposes of design optimization. We demonstrate the model’s ability to generalize by training on data from one circuit and applying it successfully to a benchmark circuit. The mean relative error is as low as 1.7%, with a speedup of up to$20\times $. Circuit designers, for the first time ever, will have ease of access to a high-precision aging model, which is paramount for efficient designs. In contrast to existing work, our approach takes the full switching activity into account to model recovery effects. This work is a promising step in the direction of bridging the gap between the foundry and circuit designers. Paul R. Genssler, Hamza Errahmouni Barkam, Karthik Pandaram, Mohsen Imani, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Transistor Self-Heating-Aware Synthesis for Reliable Digital Circuit DesignsabstractWith the continuous scaling in technology nodes, the transistor self-heating effect (SHE) emerges as a growing threat to circuit reliability. Increasingly confined transistor structures and advanced materials exacerbate thermal insulation, concealing temperature hotspots in the transistor’s channel. Without the consideration of these increased temperatures, reliability effects such as aging will be underestimated, putting the circuit at risk. In this work, we propose a novel design flow that enables designers to extract accurate SHE temperatures at the circuit level and harden their design with SHE-aware synthesis. Our approach employs customized standard cell libraries to convey SHE information and guide logic synthesis toward SHE-resilient designs. Using our approach, we demonstrate effective suppression of SHE in circuits by up to 50%, trading off improved SHE resilience against timing, power, and area goals. In addition, our SHE analysis allows for accurate estimation of timing guardbands, reducing unnecessary pessimism of conventional approaches by over 50%. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Cross-Layer Reliability Modeling of Dual-Port FeFET: Device-Algorithm InteractionabstractThe Ferroelectric Field-Effect Transistor (FeFET) is an emerging Non-Volatile Memory (NVM) technology enabling novel data-centric architectures that go far beyond von Neumann principles. Nevertheless, FeFET devices exhibit significant variations that can severely restrict their applicability. Temperature further exacerbates variation effects because it degrades ferroelectric parameters. Hence, it is indispensable to investigate and model design-time variations, run-time variations, and stochastic variations due to spatial fluctuation of ferroelectric domains under different temperatures. Dual-port FeFET has been recently proposed and demonstrated as a novel structure that offers for the first time disturb-free read operation along with$>\,\,\mathrm {10\,\,\times }$larger memory window (MW) compared to conventional FeFETs. However, all the before-mentioned variations are amplified in such a new structure. This work analyses the impact of temperature variation for dual-port FeFETs for the first time in a cross-layer manner starting from the device level to the circuit/system levels, and compared to conventional FeFET. Through our cross-layer framework, we demonstrate the severe impact of variation on FeFET reliability despite the significant increase in the MW that dual-port FeFET offers. Even Hyperdimensional Computing is affected, despite its remarkable robustness against errors. All in all, our work reveals that a larger MW at the device level does not necessarily translate to benefits at the application level.Hence, investigating and modeling variability effects in a cross-layermanner is indispensable. Swetaki Chatterjee, Simon Thomann, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | FDSOI-Based Analog Computing for Ultra-Efficient Hamming Distance Similarity CalculationabstractComputing the similarity between two binary strings is a frequently used operation in cryptography, machine learning, and other areas. The Hamming distance is a simple yet costly to compute similarity metric. A common way is to XOR both binary input strings and then count the number of 1s. Especially the latter popcount part is inefficient with purely digital circuits. In this paper, a novel analog circuit is proposed to compute the Hamming distance in an ultra-efficient way. Contrary to the major trend in the state of the art, no emerging technology is required. Instead, the unique feature of the mature FDSOI transistor technology is exploited for the first time to perform analog-based similarity calculation. Thanks to the additional back gate available in this technology, the transistor’s threshold voltage can be modulated by more than 1 V. Through this key feature, an ultra-efficient analog computing is realized, replacing the inefficient digital popcount traditionally built from expensive adder tree structures. The design is evaluated with an FDSOI transistor model calibrated with industrial measurements. The energy-delay product is at least 24$\times $smaller than purely digital implementations and the transistor count is reduced by over 2.6$\times $. Albi Mema, Simon Thomann, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Cryogenic CMOS for Quantum Processing: 5-nm FinFET-Based SRAM Arrays at 10 KabstractIn this work, we are the first to investigate and model the characteristics of a commercial 5nm FinFET technology from room temperature (300K) all the way down to cryogenic temperature (10K). We focus on SRAM circuits demonstrating how cryogenic temperatures impact their power, delay, and reliability. SRAM memories are key components in quantum read-out and control circuits, and therefore characterizing their key figure of merits when building cryogenic-CMOS circuits is essential. To achieve that, we first measure the electrical characteristics of nFinFET and pFinFET devices from 300K down to 10K. Then, we carefully calibrate the cryogenic-aware BSIM-CMG, which is the first industry-standard compact model for FinFET technologies designed for cryogenic temperatures. This enables us to reproduce the experimental data in which SPICE simulations come with an excellent agreement with the measurements. Using our well-calibrated transistor models, we simulate a complete 32-bit SRAM memory array, including a write driver, sense amplifier, pre-charger, and output latch. Then, we investigate how cryogenic temperatures impact the SRAM read and write delays at several stages during the operation, as well as the power and energy. For a more comprehensive analysis, we perform our studies for different SRAM types covering high-density, high-performance, and low-voltage cells. All transistor and SRAM analyses are performed at both room temperature and cryogenic temperature to obtain detailed comparisons revealing the exact role that cryogenic temperature plays in SRAMs. All in all, we demonstrate that commercial 5nm FinFET is indeed suitable for cryogenic-CMOS circuits required in quantum processors, revealing that the performance of SRAMs at 10K does improve while power and energy consumption are reduced. Nevertheless, SRAM reliability is more challenging in which noise margins need to be carefully engineered to remain sufficient at 10K. Shivendra Singh Parihar, Victor M. van Santen, Simon Thomann, Girish Pahwa, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | Impact of Non-Volatile Memory Cells on Spiking Neural Network Annealing Machine With In-Situ Synapse ProcessingabstractSolving constraint satisfaction problems (CSPs) is in high demand for various applications. SNN serves as a competitive annealing machine that can solve the CSP more efficiently than well-known Metropolis sampling and Hopfield networks. NVM-based crossbars with analog Integrate and Fire (IF) neurons can evolve the state of SNN to solve CSP more efficiently. However, analog computations inherently suffer from imprecisions in NVM cells, e.g., current variation, OFF-state leakage, and temperature-induced drift. We are the first to analyze the impacts of various memory technologies, including 2T-NOR, FeFET, WOx ReRAM, and HfOx ReRAM, on solving the Ising model, Sudoku, and Traveling-salesman-problem (TSP). The results show that both 2T-NOR Flash and FeFET with normalized standard deviation( ${\sigma}/{u}$ ) $<$ $5\%$ and ON-OFF ratio $>$ $1000$ are both ideal candidates as synapse devices at room temperature, while other devices suffer from the effects of current variation and OFF-state leakage, which would require the neuron circuits to have infeasible membrane capacitance size. However, the drift of cell current and the reduction of the ON-OFF ratio drops the success rate as the temperature increases. The success rate of solving TSP drops by 60 $\%$ and 90 $\%$ while the temperature increases from 300K to 358K for 2T-NOR and FeFET, respectively. Throughout the simulation, we show that the transistor-based memory is suggested to be a synapse device. Yet, we also find that the tolerance of temperature is inevitable under limited capacitance. Exploration of temperature-tolerated design of circuit and memory design is still in demand for future works. Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Hussam Amrouch, Chia-Lin Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Performance and Energy Studies on NC-FinFET Cache-Based Systems With FN-McPATabstractTo understand performance and energy tradeoffs in CPU–memory systems at lower geometries and new technologies, there is a need to update the processor and cache models used by instruction-level simulators. We improve the existing McPAT tool to support the 14-nm FinFET commercial technology, while respecting McPAT’s overall modeling methodology. We also include the results from the BOOM CPU core, synthesized with FinFET technology, into the McPAT tool to model the core components. For the first time, we extend McPAT to support the negative capacitance fin field-effect transistor (NC-FinFET), an emerging transistor technology with subthreshold swing (SS) below 60 mV/decade and unique leakage characteristics. Experiments using our FN-McPAT tool indicate that the NC-FinFET-based system is more energy-efficient relative to the FinFET-based system for memory-intensive workloads and vice versa for the compute-intensive workloads while operating at the highest voltage and frequency. In addition, we analyze the performance and energy consumption of last-level caches (LLCs) operating at various voltages and report novel insights into the energy consumption behavior for the NC-FinFET-based LLC. FN-McPAT is available for download athttps://github.com/marg-tools/FN-McPAT. Divya Praneetha Ravipati, Victor M. van Santen, Sami Salamin, Hussam Amrouch, Preeti Ranjan Panda |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | Design Close to the Edge for Advanced Technology using Machine Learning and Brain-Inspired AlgorithmsabstractIn advanced technology nodes, transistor performance is increasingly impacted by different types of design-time and run-time degradation. First, variation is inherent to the manufacturing process and is constant over the lifetime. Second, aging effects degrade the transistor over its whole life and can cause failures later on. Both effects impact the underlying electrical properties of which the threshold voltage is the most important. To estimate the degradation-induced changes in the transistor performance for a whole circuit, extensive SPICE simulations have to be performed. However, for large circuits, the computational effort of such simulations can become infeasible very quickly. Furthermore, the SPICE simulations cannot be delegated to circuit designers, since the required underlying transistor models cannot be shared due to their high confidentiality for the foundry. In this paper, we tackle these challenges at multiple levels, ranging from transistor to memory to circuit level. We employ machine learning and brain-inspired algorithms to overcome computational infeasibility and confidentiality problems, paving the way towards design close to the edge. Hussam Amrouch, Florian Klemme, Paul R. Genssler |
ASP-DAC | 1 |
| 2022 | Brain-Inspired Hyperdimensional Computing for Ultra-Efficient Edge AIabstractHyperdimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional deep learning algorithms. Despite the profound success of Deep Neural Networks (DNNs) in many domains, the amount of computational power and storage that they demand during training makes deploying them in edge devices very challenging if not infeasible. This, in turn, inevitably necessitates streaming the data from the edge to the cloud which raises serious concerns when it comes to availability, scalability, security, and privacy. Further, the nature of data that edge devices often receive from sensors is inherently noisy. However, DNN algorithms are very sensitive to noise, which makes accomplishing the required learning tasks with high accuracy immensely difficult. In this paper, we aim at providing a comprehensive overview of the latest advances in HDC. HDC aims at realizing real-time performance and robustness through using strategies that more closely model the human brain. HDC is, in fact, motivated by the observation that the human brain operates on high-dimensional data representations. In HDC, objects are thereby encoded with high-dimensional vectors which have thousands of elements. In this paper, we will discuss the promising robustness of HDC algorithms against noise along with the ability to learn from little data. Further, we will present the outstanding synergy between HDC and beyond von Neumann architectures and how HDC opens doors for efficient learning at the edge due to the ultra-lightweight implementation that it needs, contrary to traditional DNNs. Hussam Amrouch, Mohsen Imani, Xun Jiao 0002, Yiannis Aloimonos, Cornelia Fermüller, Dehao Yuan, Dongning Ma, Hamza Errahmouni Barkam, Paul R. Genssler, Peter Sutor Jr. |
CODES+ISSS | 1 |
| 2022 | Intelligent Methods for Test and ReliabilityabstractTest methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects. Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009 |
DATE | 1 |
| 2022 | Machine Learning for Test, Diagnosis, Post-Silicon Validation and Yield OptimizationabstractRecent breakthroughs in machine learning (ML) technology are shifting the boundaries of what is technologically possible in several areas of Computer Science and Engineering. This paper discusses ML in the context of test-related activities, including fault diagnosis, post-silicon validation and yield optimization. ML is by now an established scientific discipline, and a large number of successful ML techniques have been developed over the years. This paper focuses on how to adapt ML approaches that were originally developed with other applications in mind to test-related problems. We consider two specific applications of learning in more depth: delay fault diagnosis in three-dimensional integrated circuits and tuning performed during post-silicon validation. Moreover, we examine the emerging concept of brain-inspired hyperdimensional computing (HDC) and its potential for addressing test and reliability questions. Finally, we show how to integrate ML into actual industrial test and yield-optimization flows. Hussam Amrouch, Krishnendu Chakrabarty, Dirk Pflüger, Ilia Polian, Matthias Sauer 0002, Matteo Sonza Reorda |
ETS | 1 |
| 2022 | On Extracting Reliability Information from Speed BinningabstractAdaptive Voltage Frequency Scaling (AVFS) is an important means to overcome process-induced variability challenges for advanced high-performance circuits. AVFS requires and allows determining the maximum speed Fmax(Vdd) reachable under a set of certain operation voltages Vdd. In this paper, it is shown that the Fmax(Vdd) measurements contain relevant data to identify some hidden defects in a chip which are reliability threats and can cause device failures, but pass the speed binning procedure within the given specifications.Static Timing Analysis (STA) is applied to a circuit designed by using standard cell libraries in which the underlying transistors along with process variations have been carefully calibrated against industrial 14nm FinFET measurement data, and in-stances with and without injected small resistive open defects are generated. From the slope of the function Fmax(Vdd), a machine learning procedure can identify some defects with high precision and few false positives. These chips can be then discarded without any further need and cost for testing. It has to be noted that this reliability information comes for free from the data which is already generated, and does not need any additional measurements. Zahra Paria Najafi-Haghi, Florian Klemme, Hussam Amrouch, Hans-Joachim Wunderlich |
ETS | 3 |
| 2022 | AppGNN: Approximation-Aware Functional Reverse Engineering Using Graph Neural NetworksabstractThe globalization of the Integrated Circuit (IC) market is attracting an ever-growing number of partners, while remarkably lengthening the supply chain. Thereby, security concerns, such as those imposed by functional Reverse Engineering (RE), have become quintessential. RE leads to disclosure of confidential information to competitors, potentially enabling the theft of intellectual property. Traditional functional RE methods analyze a given gate-level netlist through employing pattern matching towards reconstructing the underlying basic blocks, and hence, reverse engineer the circuit's function. Tim Bücher, Lilas Alrahis, Guilherme Paim, Sergio Bampi, Ozgur Sinanoglu, Hussam Amrouch |
ICCAD | 6 |
| 2022 | Novel FDSOI-based Dynamic XNOR Logic for Ultra-Dense Highly-Efficient ComputingabstractFor the first time, we propose a novel circuit for dynamic 2-input XNOR gate that merely employs two n-type Fully-Depleted Silicon on Insulator (nFDSOI) FETs along with one additional precharging pFDSOI FET. Our design exploits the threshold voltage (Vt) tuning feature (i.e., 1ow-Vtand high-Vtstates) of FDSOI FET using the back bias as one input. The front gate bias is used as a second input. The proposed novel XNOR design reduces the number of transistors and significantly reduces power, delay, and energy compared to state-of-the-art dynamic XNOR gates. To accurately evaluate the Figure of merits, the industrial transistor compact model has been carefully calibrated against industrial measurements. The analysis demonstrates that our novel XNOR gates exhibits $8\times$ improvement in the propagation delay and $17\times$ improvement in the power consumption compared to the state-of-the-art dynamic XNOR design. Additionally, we explore the critical role of the buried oxide (BOX) thickness on the performance of proposed XNOR design. Swetaki Chatterjee, Chetan K. Dabhi, Hussam Amrouch, Yogesh Singh Chauhan |
ISCAS | 4 |
| 2022 | Wafer Map Defect Classification Based on the Fusion of Pattern and Pixel InformationabstractWith the dramatically increasing requirements on semiconductor products, improving the yield is one of the major tasks for semiconductor manufacturers. To minimize losses, automatic and efficient wafer testing tools are required to quickly notify the engineers of potential problems. One such technique is wafer map defect pattern classification, which has inspired and motivated extensive research over the last decades. Many popular studies often design novel wafer map defect identification algorithms based on manual feature extraction, statistical learning and deep neural networks, having achieved significant advancement and success. However, these methods often face challenges of training large-scale networks and few of them have noticed the full usage of the information within each wafer map. Based on the concerns above, this paper proposes a multi-task learning framework based on neural networks that fuses the information of the entire wafer map as well as the state of each individual die to enhance the defect pattern classification capability. Extensive experiments on a public real-world dataset have been conducted to justify the effectiveness of our method. Specifically, our method achieved an classification accuracy of 96.3%, which was better or comparable to other state-of-the-art approaches that required notably larger network sizes and heavy data augmentation. Yiwen Liao, Raphaël Latty, Paul R. Genssler, Hussam Amrouch, Bin Yang 0009 |
ITC | 4 |
| 2022 | Cross-layer FeFET Reliability Modeling for Robust Hyperdimensional ComputingabstractHyperdimensional computing (HDC) is an emerging learning paradigm that has gained a lot of attention due to its ability to train with fewer data, lightweight implementation, and resiliency against errors. Similar to the brain, HDC can learn patterns in one iteration from small training data by computing a similarity metric such as Hamming distance. Ferroelectric Field-Effect-Transistor (FeFET) based Ternary Content Addressable Memory (TCAM) has been demonstrated as an excellent candi-date for computing this similarity metric. However, variations in the underlying ferroelectric transistor does impact the reliable HDC operation. In this paper, we demonstrate an end-to-end cross-layer FeFET reliability modeling to obtain robust HDC across the computing stack starting from transistor physics all the way to circuits and systems. The effect of random spatial fluctuation of ferroelectric (FE) domains and other variability sources on electrical characteristics of FeFET is computed through detailed physics-based TCAD simulations. Then, the entire TCAM array is simulated in SPICE using a carefully designed and calibrated compact model to capture the effect of transistor variability on the error probability for individual Hamming distances. Finally, the error probability is employed to compute the loss of inference accuracy of HDC with a language recognition task. We observe very little loss in accuracy even with a high degree of variation. Swetaki Chatterjee, Simon Thomann, Paul R. Genssler, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 6 |
| 2022 | Design-time exploration of voltage switching against power analysis attacks in 14 nm FinFET technology
Johann Knechtel, Tarek Ashraf, Natascha Fernengel, Satwik Patnaik, Mohammed Nabeel Thari Moopan, Mohammed Ashraf, Ozgur Sinanoglu, Hussam Amrouch |
Integr. | 8 |
| 2022 | Brain-Inspired Computing for Circuit Reliability CharacterizationabstractTransistor scaling steadily approaches fundamental limits. Sustaining circuit reliability becomes an overwhelming challenge for foundries. Therefore, early and rapid characterization of degradation effects impacting the circuits transistors becomes essential. Such degradation effects are caused by design-time variation due to manufacturing variability and/or run-time variation due to transistor aging. In this work, we are the first to employ Brain-Inspired Hyperdimensional computing (HDC) for circuit reliability. HDC is an emerging light-weight machine-learning solution. Nowadays, it is mainly applied to bio-signal processing. We bring the research of HDC to the next level by demonstrating how it can be applied to address the challenges in circuit reliability. This has far-reaching consequences due to the large savings achieved by 1) reducing the amount of training data, 2) removing the need to send the data to the Cloud for model training, and 3) significantly speeding up the characterization and classification tasks. We demonstrate the viability of HDC, using SRAM and other circuits as an example. HDC outperforms traditional machine learning methods, such as support vector machine, in accuracy and requires up to 20x fewer training samples. Our implementation and analysis are based on industrial 14 nm FinFET fully calibrated with Intel measurements. Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2022 | On the Reliability of FeFET On-Chip MemoryabstractFerroelectric Field-Effect Transistor (FeFET) is a promising future technology for non-volatile on-chip memories. It is rapidly attracting an ever-increasing attention from industry. The key advantage of FeFETs is full compatibility with the existing CMOS fabrication process beside their very low power consumption. To enable ultra-dense memories, 1-FeFET AND Arrays were proposed in which a memory cell is formed from merely a single FeFET. All access transistors, which are traditionally needed to operate memory cells, are removed. However, this imposes a new challenge ofindirect write disturbances. Neighboring memory cells are indirectly degraded whenever adirect write operationoccurs to a particular FeFET cell. Only recently the impact of such indirect disturbances on the FeFET reliability was experimentally investigated at device (i.e., transistor) level. However, to explore and properly judge the feasibility of 1-FeFET AND Arrays for on-chip memories, investigating only the reliability of individual cells is indeed insufficient. Bridging the gap between the device level and system (i.e., chip) level is inevitable. In the presence of indirect disturbances, the position of a write access within the array plays a key role, which is governed by the running workloads. In addition, whether the write operation flips the previously stored value or not also plays an important role with regards to reliability. Hence, running workloads, which determine not only the position of the memory cells to be written but also the values written to them, plays an essential role in determining 1-FeFET AND Array reliability over time. Therefore, studying the reliability of FeFETs only at the device level (as done in state of the art) is insufficient. In this work, we investigate, for the first time, the reliability of FeFET memories from device to system level. To achieve that, we develop a unified model capturing the impact of bothindirectdisturbances anddirectwrites on the reliability of FeFET cells. Our study at system level then employs the unified model in the context of application workloads. We investigate different array sizes, write voltages, write methods and a wide range of workloads using the example of CPU caches as an example of on-chip memory. We demonstrate that indirect write disturbances are the dominate effect degrading the reliability of FeFET memories. For most cells, it contributes over 90 percent to the overall induced degradation. This provides guidelines for researchers at both device and circuit level to optimize the FeFET reliability further while considering thehiddenimpact of indirect write disturbances. Paul R. Genssler, Victor M. van Santen, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 4 |
| 2022 | A Framework for Crossing Temperature-Induced Timing Errors Underlying Hardware Accelerators to the Algorithm and Application LayersabstractTemperature rising is an unavoidable effect on VLSI and has always been a critical issue in any system-on-chip – especially when targeting compute-intensive applications. This effect increases the delay in hardware accelerators, resulting in timing errors due to unsustainable clock frequency, whose impact must be carefully evaluated on design time to measure the performance degradation of the hardware accelerator. Further, a hardware operating at a higher temperature accelerates device aging, which incurs in more timing errors. This issue is usually addressed with the inclusion of timing guardbands that compensate for the deleterious effects of temperature, ensuring the hardware accelerator works within a reliable zone, i.e., without any timing errors caused by temperature effects at runtime. However, guardbands directly result in considerable performance and efficiency losses because the circuit will be clocked at a frequency lower than its full potential. Accelerators on edge devices often dismiss such guardbands to explore the full potential of the designed circuits, posing an enormous design challenge as this approach requires a careful evaluation of the impact of timing errors on the quality of the target applications. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating hardware errors. Yet, these algorithms have a dynamic behavior (i.e., closed-loop) where a timing error can be propagated, affecting subsequent steps. Measuring the degradation-induced errors in these applications is very challenging given that an accurate gate-level simulation to investigate degradation-induced timing errors needs to be coupled dynamically with a system-level simulator to unveil how induced errors in the underlying hardware ultimately impact the algorithm execution in the hardware accelerator.This is the first work to achieve this goal. State-of-the-art works have studied accelerators under timing-errors when removing (or narrowing) guardbands. However, their approach was suitableonly for open-loop hardware accelerators which are entirely agnostic of complex interactions of the algorithms. Unlike prior work, this paper investigates temperature- and aging-induced timing-errors in the joint accelerator-algorithm interactions and their runtime impacts. Our framework investigates aging effects across the different layers starting from transistor physics all the way up to the algorithm layer. The hardware accelerator employed as a case study in this work is the sum of absolute differences (SAD), which is the most compute-intensive accelerator on commercial video encoder for mobile applications. Our results demonstrate the runtime behavior impacts of three advanced block-matching algorithms of the video encoder in a joint operation by a SAD accelerator under timing-errors induced by temperature and aging effects considering a 14nm FinFET technology. Guilherme Paim, Hussam Amrouch, Leandro M. G. Rocha, Brunno Abreu, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Computers | 2 |
| 2022 | Real-Time Full-Chip Thermal Tracking: A Post-Silicon, Machine Learning PerspectiveabstractThis article presents a novel approach to real-time tracking of full-chip heatmaps for off-the-shelf microprocessors based on machine-learning. The proposed post-silicon approach, named RealMaps, only uses the existing temperature sensors and workload-independent utilization information. RealMaps does not require any knowledge of the proprietary design or manufacturing process-specific details of the chip. Consequently, the methods presented in this work can be implemented by either the original chip manufacturer or a third party alike. The approach involves offline acquisition of spatial heatmaps using a thermal imaging setup. To build the dynamic thermal model, a temporal-aware long-short-term-memory neutral network is trained with system-level features as inputs. 2D discrete cosine transformation (DCT) is performed on the heatmaps so that they can be expressed with just a few dominant DCT coefficients. This allows the model to be built to estimate just the dominant spatial features of the heatmaps, rather than the entire heatmap images, making it significantly more efficient. Experimental results from two commercial chips show that RealMaps can estimate the full-chip heatmaps with 0.9C and 1.2C root-mean-square-error respectively and take only 0.4ms for each inference. Compared to the state of the art pre-silicon approach, RealMaps shows similar accuracy, but with much less computational cost. Sheriff Sadiqbatcha, Hussam Amrouch, Sheldon X.-D. Tan |
IEEE Trans. Computers | 3 |
| 2022 | Impact of NCFET Technology on Eliminating the Cooling Cost and Boosting the Efficiency of Google TPUabstractRecent breakthroughs in Neural Networks (NNs) led to significant accuracy improvements of several machine learning applications such as image classification and voice recognition. However, this accuracy improvement comes at the cost of an immense increase in computation demands. NNs became one of the most common and computationally intensive workloads in today's datacenters. To address these computational demands, Google announced in 2016 the Tensor Processing Unit (TPU), an advanced custom ASIC accelerator for NN inference. Two new TPU versions (v2 and v3) followed in 2017 and 2018 that support also training. Google TPUv3 packs an immense processing power ($\mathrm{90TFLOPS}$per chip) in a tiny and condensed area, leading to very high on-chip power densities and thus excessive temperature. In this article, superlattice thermoelectric cooling, which is one of the emerging on-chip cooling, is considered as an advanced cooling example for Google TPU and we investigate the impact of Negative Capacitance FET (NCFET), which is one of the recent emerging technologies, on the cooling and efficiency of TPU. Through full-chip design, of the computational core of the TPU, based on$14\mathrm{nm}$Intel FinFET technology and multiphysics temperature simulations, we demonstrate that NCFET can significantly minimize the required cooling-cost. More than 4000 NCFET configurations are evaluated in order to traverse the entire design space defined by the thickness of the ferroelectric layer of NCFET, the operating voltage, cooling, and the operating frequency, in addition to all possible FinFET's configurations. Moreover, our experimental evaluation shows that by eliminating the cooling cost, NCFET delivers 2.8x higher efficiency compared to the conventional FinFET baseline. Sami Salamin, Georgios Zervakis 0001, Florian Klemme, Hammam Kattan, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 7 |
| 2022 | Trojan Detection in Embedded Systems With FinFET TechnologyabstractThis study considers detecting Trojans in circuits using FinFET technology non-destructively, when a golden Integrated Circuit (IC) is unavailable. The method employs short-term aging effects in FinFET transistors and circuit overclocking to induce bit errors at the circuit outputs in conjunction with Machine Learning (ML) tools learning Trojan-free behavior. Short-term aging causes delays along multiple paths in the IC to vary dynamically, causing bit errors at circuit outputs. Overclocking enhances this in FinFET but is not necessary for bulk CMOS technology. We use bit error patterns at the output of the circuit to detect Trojans using an ML classifier trained on simulations of the Trojan-free circuit. The study shows efficacy of the method by using dynamic short-term aging-aware standard cell libraries with FinFET technology that are modeled by considering the dynamic short-term aging of each cell. Trojan detection is robust to chip-to-chip variations. We apply the technique on fourteen Trust-Hub Trojans. Our method detects Trojans with$>$95% accuracy. Trojan detection in FinFET technology is more challenging than in bulk CMOS because the voltage range for switching from a high to low value is smaller. Therefore we use overclocking. Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami |
IEEE Trans. Computers | 3 |
| 2022 | FeFET-Based Binarized Neural Networks Under Temperature-Dependent Bit ErrorsabstractFerroelectric FET (FeFET) is a highly promising emerging non-volatile memory (NVM) technology, especially for binarized neural network (BNN) inference on the low-power edge. The reliability of such devices, however, inherently depends on temperature. Hence, changes in temperature during run time manifest themselves as changes in bit error rates. In this work, we reveal the temperature-dependent bit error model of FeFET memories, evaluate its effect on BNN accuracy, and propose countermeasures. We begin on the transistor level and accurately model the impact of temperature on bit error rates of FeFET. This analysis reveals temperature-dependent asymmetric bit error rates. Afterwards, on the application level, we evaluate the impact of the temperature-dependent bit errors on the accuracy of BNNs. Under such bit errors, the BNN accuracy drops to unacceptable levels when no countermeasures are employed. We propose two countermeasures: (1) Training BNNs for bit error tolerance by injecting bit flips into the BNN data, and (2) applying a bit error rate assignment algorithm (BERA) which operates in a layer-wise manner and does not inject bit flips during training. In experiments, the BNNs, to which the countermeasures are applied to, effectively tolerate temperature-dependent bit errors for the entire range of operating temperature. Mikail Yayla, Sebastian Buschjäger, Aniket Gupta, Jian-Jia Chen, Jörg Henkel, Katharina Morik, Kuan-Hsun Chen, Hussam Amrouch |
IEEE Trans. Computers | 8 |
| 2022 | Thermal-Aware Design for Approximate DNN AcceleratorsabstractRecent breakthroughs in Neural Networks (NNs) have made DNN accelerators ubiquitous and led to an ever-increasing quest on adopting them from Cloud to edge computing. However, state-of-the-art DNN accelerators pack immense computational power in a relatively confined area, inducing significant on-chip power densities that lead to intolerable thermal bottlenecks. Existing state of the art focuses on using approximate multipliers only to trade-off efficiency with inference accuracy. In this work, we present a thermal-aware approximate DNN accelerator design in which we additionally trade-off approximation with temperature effects towards designing DNN accelerators that satisfy tight temperature constraints. Using commercial multi-physics tool flows for heat simulations, we demonstrate how our thermal-aware approximate design reduces the temperature from 139$^{\circ }$C, in an accurate circuit, down to 79$^{\circ }$C. This enables DNN accelerators to fulfill tight thermal constraints, while still maximizing the performance and reducing the energy by around 75% with a negligible accuracy loss of merely 0.44% on average for a wide range of NN models. Furthermore, using physics-based transistor aging models, we demonstrate how reductions in voltage and temperature obtained by our approximate design considerably improve the circuit’s reliability. Our approximate design exhibits around 40% less aging-induced degradation compared to the baseline design. Georgios Zervakis 0001, Iraklis Anagnostopoulos, Sami Salamin, Ourania Spantidi, Isai Roman-Ballesteros, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 7 |
| 2022 | GNN4REL: Graph Neural Networks for Predicting Circuit Reliability DegradationabstractProcess variations and device aging impose profound challenges for circuit designers. Without a precise understanding of the impact of variations on the delay of circuit paths, guardbands, which keep timing violations at bay, cannot be correctly estimated. This problem is exacerbated for advanced technology nodes, where transistor dimensions reach atomic levels and established margins are severely constrained. Hence, traditional worst-case analysis becomes impractical, resulting in intolerable performance overheads. Contrarily, process-variation/aging-aware static timing analysis (STA) equips designers with accurate statistical delay distributions. Timing guardbands that are small, yet sufficient, can then be effectively estimated. However, such analysis is costly as it requires intensive Monte-Carlo simulations. Further, it necessitates access to confidential physics-based aging models to generate the standard-cell libraries required for STA. In this work, we employ graph neural networks (GNNs) to accurately estimate the impact of process variations and device aging on the delay of any path within a circuit. Our proposed GNN4REL framework empowers designers to perform rapid and accurate reliability estimations without accessing transistor models, standard-cell libraries, or even STA; these components are all incorporated into the GNN model via training by the foundry. Specifically, GNN4REL is trained on a FinFET technology model that is calibrated against industrial 14-nm measurement data. Through our extensive experiments on EPFL and ITC-99 benchmarks, as well as RISC-V processors, we successfully estimate delay degradations of all paths—notably within seconds—with a mean absolute error down to 0.01 percentage points. Lilas Alrahis, Johann Knechtel, Florian Klemme, Hussam Amrouch, Ozgur Sinanoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Electrothermal Simulation and Optimal Design of Thermoelectric Cooler Using Analytical ApproachabstractIn this article, electrothermal modeling and simulation of thermoelectric cooling (TEC) in the package design of VLSI systems are performed by solving coupled heat conduction and current continuity equations. We propose a new analytical solution to the coupled partial differential equations (PDEs) which describe temperature and voltage with the reduction from 3-D to 1-D. In addition to this, we derive new analytic expressions for two key performance metrics for TEC devices: 1) the maximum temperature difference and 2) the maximum heat-flux pumping capability, which can be guided for the optimal design of thermoelectric cooler to achieve the maximum cooling performance. Furthermore, for the first time, we observe that when the dimensionless figure of merit$ZT_{0}$value is larger than 1, there is no maximum heat-flux value, which means the heat dissipation due to the Peltier and Fourier transfer effects is larger than the heat generation caused by the Joule heating effect, which can lead to more efficient TEC cooling design. The accuracy of the proposed 1-D formulas is verified by a 3-D finite element method using COMSOL software. The compact model delivers many orders of magnitude speedup and memory saving compared to COMSOL with marginal accuracy loss. Compared with the conventional simplified 1-D energy equilibrium model, the proposed analytical coupled multiphysics model is more robust and accurate. Liang Chen 0025, Sheriff Sadiqbatcha, Hussam Amrouch, Sheldon X.-D. Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A Novel Attack Mode on Advanced Technology Nodes Exploiting Transistor Self-HeatingabstractSelf-heating (SH) is a phenomenon that can induce excessive heat inside the transistor channel. SH represents an emerging and serious concern, especially in advanced technology nodes, where excessive heat acting on elevated channel geometries will notably shift the critical transistor parameters (e.g., threshold-voltage$V_{\text {th}}$and carrier mobility$\mu $). The underlying 3-D device structures (e.g., FinFET, nanowire, or nanosheet structures), along with newly employed materials such as silicon-germanium (SiGe), which show worse thermal conductivity than traditional materials, can considerably exacerbate SH. On top of that, quantum confinement, a phenomenon that becomes dominant at sub-10nm, further increases the intensity of SH. In this article, we are the first to explore SH effects from the perspective of hardware security, rather than the performance, reliability standpoints covered in state-of-the-art (SOTA) work. As proof of concept, we devise an SH-based hardware trojan (HT) that exploits the SH-induced$V_{\text {th}}$change in 7-nm FinFET circuits. Leveraging$V_{\text {th}}$-dependent reconfigurable logic, we design a reconfigurable HT payload that maliciously changes its functional behavior once the SH-induced$V_{\text {th}}$change takes effect. Following SOTA work, we present a comprehensive modeling and analysis of SH effects at the device level and highlight its impact on transistor$V_{\text {th}}$. Next, we study how fabrication-time changes in the transistor doping and geometry can promote the SH-assisted degradation. We then describe various payload configurations for the proposed HT, quantify its overheads, and discuss its resilience against standard HT detection techniques. Finally, we demonstrate two case studies using the proposed HT, one to leak the secret key from a pipelined design of an advanced encryption standard (AES) circuit, and another to showcase denial-of-service for a Gaussian-blur filter circuit. Our work utilizes industry-standard models with parameters extracted from measurements and calibrated with experiments. Our results are obtained from meticulous study and optimization across the device-, circuit-, and system-levels. Nikhil Rangarajan, Johann Knechtel, Nimisha Limaye, Ozgur Sinanoglu, Hussam Amrouch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | MLCAD: A Survey of Research in Machine Learning for CAD Keynote PaperabstractDue to the increasing size of integrated circuits (ICs), their design and optimization phases (i.e., computer-aided design, CAD) grow increasingly complex. At design time, a large design space needs to be explored to find an implementation that fulfills all specifications and then optimizes metrics like energy, area, delay, reliability, etc. At run time, a large configuration space needs to be searched to find the best set of parameters (e.g., voltage/frequency) to further optimize the system. Both spaces are infeasible for exhaustive search typically leading to heuristic optimization algorithms that find some tradeoff between design quality and computational overhead. Machine learning (ML) can build powerful models that have successfully been employed in related domains. In this survey, we categorize how ML may be used and is used for design-time and run-time optimization and exploration strategies of ICs. A metastudy of published techniques unveils areas in CAD that are well explored and underexplored with ML, as well as trends in the employed ML algorithms. We present a comprehensive categorization and summary of the state of the art on ML for CAD. Finally, we summarize the remaining challenges and promising open research directions. Martin Rapp, Hussam Amrouch, Yibo Lin, Bei Yu 0001, David Z. Pan, Marilyn Wolf, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Full-Chip Power Density and Thermal Map Characterization for Commercial Microprocessors Under Heat Sink CoolingabstractIn this article, we address the problem of accurate full-chip power and thermal map estimation for commercial off-the-shelf multicore processors. Processors operating with heat sink cooling remains a challenging problem due to the difficulty in direct measurement. We first propose an accurate full-chip steady-state power density map estimation method for commercial multicore microprocessors. The new method consists of a few steps. First, 2-D spatial Laplace operation is performed on the measured thermal maps (images) without heat sink to obtain the so-calledraw power maps. Then, a novel scheme is developed to generate the true power density maps from the raw power density maps. The new approach is based on thermal measurements of the processor with back-side cooling using an advanced infrared (IR) thermal imaging system. FEM thermal model constructed in COMSOL Multiphysics is used to validate the estimated power density maps and thermal conductivity. Later, this work creates a high-fidelity FEM thermal model with heat sink and reconstructs the full-chip thermal maps while the heat sink is on. Ensuring that power maps are similar under back cooling and heat sink cooling settings, the reconstructed thermal maps are verified by the matching between the on-chip thermal sensor readings and the corresponding elements of thermal maps. Experiments on an Intel i7-8650U 4-core processor with back cooling shows 96% similarity (2-D correlation) between the measured thermal maps and the thermal maps reconstructed from the estimated power maps, with 1.3 °C average absolute error. Under heat sink cooling, the average absolute error is 2.2 °C over a 56 °C temperature range and about 3.9% error between the computed and the real thermal maps at the sensor locations. Furthermore, the proposed power map estimation method achieves higher resolution and at least$100\times $speedup than a recently proposed state-of-art Blind Power Identification method. Sheriff Sadiqbatcha, Michael O'Dea, Hussam Amrouch, Sheldon X.-D. Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Variability-Aware Approximate Circuit Synthesis via Genetic OptimizationabstractOne of the major barriers that CMOS devices face at nanometer scale is increasing parameter variation due to manufacturing imperfections. Process variations severely inhibit the reliable operation of circuits, as the operational frequency at the nominal process corner is insufficient to suppress timing violations across the entire variability spectrum. To avoid variability-induced timing errors, previous efforts impose pessimistic and performance-degrading timing guardbands atop the operating frequency. In this work, we employ approximate computing principles and propose a circuit-agnostic automated framework for generating variability-aware approximate circuits that eliminate process-induced timing guardbands. Variability effects are accurately portrayed with the creation of variation-aware standard cell libraries, fully compatible with standard EDA tools. The underlying transistors are fully calibrated against industrial measurements from Intel 14nm FinFET in which both electrical characteristics of transistors and variability effects are accurately captured. In this work, we explore the design space of approximate variability-aware designs to automatically generate circuits of reduced variability and increased performance without the need for timing guardbands. Experimental results show that by introducing negligible functional error of merely$\boldsymbol {5.3 \times 10^{-3}}$, our variability-aware approximate circuits can be reliably operated under process variations without sacrificing the application performance. Konstantinos Balaskas, Florian Klemme, Georgios Zervakis 0001, Kostas Siozios, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2022 | Scalable Machine Learning to Estimate the Impact of Aging on Circuits Under Workload DependencyabstractTo ensure the correct functionality of a chip throughout its entire lifetime, preliminary circuit analysis with respect to aging-induced degradation is indispensable. However, state-of-the-art techniques only allow for the consideration of uniformly applied degradations, despite the fact that different workloads will lead to different degradations due to their distinct induced activities. This imposes over-pessimism when estimating the required timing guardbands, resulting in an unnecessary loss of performance and efficiency. In this work, we propose an approach that takes real-world workload dependencies into account and generates workload-specific aging-aware standard cell libraries, allowing for accurate analysis of aging-induced degradations. We employ machine learning techniques to overcome infeasible simulation times for individual transistor aging while sustaining high prediction accuracy. We also demonstrate scalability to previously unknown workloads and discuss multiple approaches to estimate the machine learning accuracy by employing coverage metrics. In our evaluation, we achieve predictions of workload-dependent aging-aware standard cells with an average accuracy (R2 score) of 95.28%. Using predicted cell libraries in static timing analysis, timing guardbands for multiple circuits are reported with an error of less than 0.1% on average. We demonstrate that timing guardband requirements can be reduced by up to 30% when considering specific workloads over worst-case estimations as performed in state-of-the-art tool flows. Even for unknown workloads of different circuits, accurate prediction with relative errors below 1% can be achieved. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Efficient Learning Strategies for Machine Learning-Based Characterization of Aging-Aware Cell LibrariesabstractMachine learning (ML)-driven standard cell library characterization enables rapid, on-the-fly generation of cell libraries, opening the door for extensive design-space exploration and other, previously infeasible approaches. However, the benefits of ML-based cell library characterization are strongly limited by its high demand in training data and the costly SPICE simulation required to generate the training samples. Therefore, efficient learning strategies are needed to minimize the required training data for ML models while still sustaining high prediction accuracy. In this work, we explore multiple active and passive learning strategies for ML-based cell library characterization with focus on aging-induced degradation. While random sampling and greedy sampling strategies operate with low computational overhead, active learning considers the performance of ML models to find the most valuable samples for training. We also introduce a hybrid approach of active learning and greedy sampling to optimize the trade-off between reduction in training samples and computational overhead. Our experiments demonstrate an achievable training data reduction of up to 77% compared to the state of the art, depending on the targeted accuracy of the ML models. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Characterizing Approximate Adders and Multipliers for Mitigating Aging and Temperature DegradationsabstractThe performance of nanoscale semiconductor technologies has become susceptible to high temperatures and aging phenomena. While guard-bands have conventionally been used to combat degradation-induced timing violations, approximations have recently been leveraged to compensate for degradations in lieu of adding timing guard-bands, without a loss in performance. However, only simple approximation techniques such as truncation have been considered in prior work. In this paper, a wide range of approximate arithmetic circuits including adders and multipliers using various sophisticated approximation techniques are investigated to cope with aging- and temperature-induced degradations. To this end, approximate circuits are first characterized for their delay increase under degradations. With this, we then determine the approximation level required to compensate for guard-bands under different degradations. Degradation-aware logic synthesis results show that the simple use of truncated arithmetic circuits leads to a higher quality loss compared to using other approximate circuits. However, a truncated multiplier has the lowest error distance towards a reliable operation in 10 years. The approximate multipliers with configurable error recovery are most suitable when the level of degradation is higher, e.g., at a temperature of 70 °C. The characterization of degradation at the circuit level is then used for design exploration at the architecture level without the need for further gate-level simulations. For three different image processing applications, experimental results show that guard-bands can be mitigated while maintaining an output result with a high visual quality. Francisco J. H. Santiago, Honglan Jiang, Hussam Amrouch, Andreas Gerstlauer, Leibo Liu, Jie Han 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Reliable Binarized Neural Networks on Unreliable Beyond Von-Neumann ArchitectureabstractSpecialized hardware accelerators beyond von-Neumann, that offer processing capability in where the data resides without moving it, become inevitable in data-centric computing. Emerging non-volatile memories, like Ferroelectric Field-Effect Transistor (FeFET), are able to build compact Logic-in-Memory (LiM). In this work, we investigate the probability of error (Perror) in FeFET-based XNOR LiM, demonstrating the new trade-off between the speed and reliability. Using our reliability model, we present how Binarized Neural Networks (BNNs) can be proactively trained in the presence of XNOR-induced errors towards obtaining robust BNNs at the design time. Furthermore, leveraging the trade-off between Perror and speed, we present a run-time adaptation technique, that selectively trades-off Perror and XNOR speed for every BNN layer. Our results demonstrate that when a small loss (e.g., 1%) in inference accuracy could be accepted, our design-time and run-time techniques provide error-resilient BNNs that exhibit 75% and 50% (FashionMNIST) and 38% and 24% (CIFAR10) XNOR speedups, respectively. Mikail Yayla, Simon Thomann, Sebastian Buschjäger, Katharina Morik, Jian-Jia Chen, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Bridging the Gap Between Voltage Over-Scaling and Joint Hardware Accelerator-Algorithm Closed-LoopabstractVoltage over-scaling (VOS) optimizes energy while causing timing errors due to an unsustainable clock frequency. Many algorithms, such as in multimedia and machine learning applications, are capable of tolerating such errors. VOS has never been investigated in hardware accelerators running closed-loop algorithms. As the errors impact most decisions and actions in the subsequent steps, closed-loops dynamically change the execution flow. Timing errors should be evaluated by an accurate gate-level simulation, but a large gap still remains: how these timing errors propagate from the underlying hardware all the way up to the entire algorithm run, where they just may degrade the performance and quality of service of the application at stake? This paper tackles this issue showing a framework for VOS investigation, embracing any kind of application. Our framework simulates the VOS-induced timing errors at gate-level, dynamically linking the hardware result with the algorithm and vice versa during the evolution of the runtime of the application. The state-of-the-art VOS literature for video encoding application fails to assess the ultimate impacts of VOS-induced timing errors, as current works open the encoding loops. Unlike those, our work investigates the ultimate impact of a hardware accelerator dynamically carrying through to the video encoder all VOS-induced timing errors and preserving the full compliance to the standard. We employ a parallel sum of absolute differences (SAD) hardware accelerator as a case study. We assess the performance of the overall encoder under varying timing guardbands. Next, it is demonstrated that, under VOS, the ultimate impact in compression efficiency is related to the video’s motion intensity. Additionally, the advantages of timing guardband controlled reduction are clearly quantified in our results by virtue of the framework. Reducing at maximum 9.5% the clock frequency, energy savings (up to 16.5% in energy/operation) are achieved in SAD for video compression. Guilherme Paim, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Towards a New Thermal Monitoring Based Framework for Embedded CPS Device SecurityabstractThis article introduces a thermal side channel as a proxy for the behavior of embedded processors to detect changes in the behavior in a cyber-physical system. Such changes may be due to software/hardware attacks and altered processors. Since control system processes are periodic computations, the thermal side channels exhibit a temporal pattern. This enables the detection of altered code and changed device characteristics. We present a machine learning approach to estimate the activity of the embedded device from the time sequence of thermal images and show that deviations from expected behavior can be detected. The approach is validated on a multi-core processor running a periodic computational code. The infrared imager collects thermal imagery from the processor, which is cooled from the backside. Instead of an external imager, one can deploy a finite number of on-chip temperature sensors. This article shows that integrating on-chip temperature sensors allows robust real-time monitoring of the processor behavior. Finally, we offer a machine learning approach to optimally place the on-chip sensors to aid detection. Naman Patel, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Michael Shamouilian, Ramesh Karri, Farshad Khorrami |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Software-Managed Read and Write Wear-Leveling for Non-Volatile Main MemoryabstractIn-memory wear-leveling has become an important research field for emerging non-volatile main memories over the past years. Many approaches in the literature perform wear-leveling by making use of special hardware. Since most non-volatile memories only wear out from write accesses, the proposed approaches in the literature also usually try to spread write accesses widely over the entire memory space. Some non-volatile memories, however, also wear out from read accesses, because every read causes a consecutive write access. Software-based solutions only operate from the application or kernel level, where read and write accesses are realized with different instructions and semantics. Therefore different mechanisms are required to handle reads and writes on the software level. First, we design a method to approximate read and write accesses to the memory to allow aging aware coarse-grained wear-leveling in the absence of special hardware, providing the age information. Second, we provide specific solutions to resolve access hot-spots within the compiled program code (text segment) and on the application stack. In our evaluation, we estimate the cell age by counting the total amount of accesses per cell. The results show that employing all our methods improves the memory lifetime by up to a factor of 955×. Christian Hakert, Kuan-Hsun Chen, Horst Schirmeier, Lars Bauer, Paul R. Genssler, Georg von der Brüggen, Hussam Amrouch, Jörg Henkel, Jian-Jia Chen |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2022 | FN-CACTI: Advanced CACTI for FinFET and NC-FinFET TechnologiesabstractCache memories are an indispensable component of many processor-based systems and contribute significantly to the overall area, power consumption, and delay. This leads to an important role played by modeling tools for estimating the area, power consumption, and access time of cache memories. However, existing modeling tools such as CACTI and its various extensions have been primarily designed using data from various projections. For the first time, we propose an entire flow for obtaining/calibrating the transistor characteristics from a commercial technology and use these characteristics within CACTI. We also improve the modeling approach to make them more fine-grained and follow recent manufacturing trends suitable for FinFET technology. Further, for the first time, we extend CACTI to support negative capacitance fin field effect transistor (NC-FinFET), an emerging technology depicting negative capacitance whose current and capacitive characteristics are very different compared to those of the FinFET. We use the proposed tool (FN-CACTI) to identify NC-FinFET-based caches to be significantly more energy-efficient than corresponding FinFET-based caches. We also study an application of FN-CACTI to determine optimal voltages corresponding to the lowest energy consumption for NC-FinFET and FinFET-based caches of various sizes. Divya Praneetha Ravipati, Rajesh Kedia, Victor M. van Santen, Jörg Henkel, Preeti Ranjan Panda, Hussam Amrouch |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | Cross-layer Design for Computing-in-Memory: From Devices, Circuits, to Architectures and ApplicationsabstractThe era of Big Data, Artificial Intelligence (AI) and Internet of Things (IoT) is approaching, but our underlying computing infrastructures are not sufficiently ready. The end of Moore's law and process scaling as well as the memory wall associated with von Neumann architectures have throttled the rapid development of conventional architectures based on CMOS technology, and cross-layer efforts that involve the interactions from low-end devices to high-end applications have been prominently studied to overcome the aforementioned challenges. On one hand, various emerging devices, e.g., Ferroelectric FET, have been proposed to either sustain the scaling trends or enable novel circuit and architecture innovations. On the other hand, novel computing architectures/algorithms, e.g., computing-in-memory (CiM), have been proposed to address the challenges faced by conventional von Neumann architectures. Naturally, integrated approaches across the emerging devices and computing architectures/algorithms for data-intensive applications are of great interests. This paper uses the FeFET as a representative device, and discuss about the challenges, opportunities and contributions for the emerging trends of cross-layer co-design for CiM. Hussam Amrouch, Xiaobo Sharon Hu, Mohsen Imani, Ann Franchesca Laguna, Michael T. Niemier, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ASP-DAC | 1 |
| 2021 | Approximate Computing for ML: State-of-the-art, Challenges and VisionsabstractIn this paper, we present our state-of-the-art approximate techniques that cover the main pillars of approximate computing research. Our analysis considers both static and reconfigurable approximation techniques as well as operation-specific approximate components (e.g., multipliers) and generalized approximate highlevel synthesis approaches. As our application target, we discuss the improvements that such techniques bring on machine learning and neural networks. In addition to the conventionally analyzed performance and energy gains, we also evaluate the improvements that approximate computing brings in the operating temperature. Georgios Zervakis 0001, Hassaan Saadat, Hussam Amrouch, Andreas Gerstlauer, Sri Parameswaran, Jörg Henkel |
ASP-DAC | 3 |
| 2021 | Control Variate Approximation for DNN AcceleratorsabstractIn this work, we introduce a control variate approximation technique for low error approximate Deep Neural Network (DNN) accelerators. The control variate technique is used in Monte Carlo methods to achieve variance reduction. Our approach significantly decreases the induced error due to approximate multiplications in DNN inference, without requiring time-exhaustive retraining compared to state-of-the-art. Leveraging our control variate method, we use highly approximated multipliers to generate power-optimized DNN accelerators. Our experimental evaluation on six DNNs, for Cifar-10 and Cifar100 datasets, demonstrates that, compared to the accurate design, our control variate approximation achieves same performance and 24% power reduction for a merely 0.16% accuracy loss. Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel |
DAC | 4 |
| 2021 | Reliability-Aware Quantization for Anti-Aging NPUs
Sami Salamin, Georgios Zervakis 0001, Ourania Spantidi, Iraklis Anagnostopoulos, Jörg Henkel, Hussam Amrouch |
DATE | 6 |
| 2021 | FeFET and NCFET for Future Neural Networks: Visions and OpportunitiesabstractThe goal of this special session paper is to introduce and discuss different emerging technologies for logic circuitry and memory as well as new lightweight architectures for neural networks. We demonstrate how the ever-increasing complexity in Artificial Intelligent (AI) applications, resulting in an immense increase in the computational power, necessitates inevitably employing innovations starting from the underlying devices all the way up to the architectures. Two different promising emerging technologies will be presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new beyond-CMOS technology with advantages for offering low power and/or higher accuracy for neural network inference. (ii) Ferroelectric FET (FeFET) as a novel non-volatile, area-efficient and ultra-low power memory device. In addition, we demonstrate how Binarized Neural Networks (BNNs) offer a promising alternative for traditional Deep Neural Networks (DNNs) due to its lightweight hardware implementation. Finally, we present the challenges from combining FeFET-based NVM with NNs and summarize our perspectives for future NNs and the vital role that emerging technologies may play. Mikail Yayla, Kuan-Hsun Chen, Georgios Zervakis 0001, Jörg Henkel, Jian-Jia Chen, Hussam Amrouch |
DATE | 6 |
| 2021 | Brain-Inspired Computing: Adventure from Beyond CMOS Technologies to Beyond von Neumann Architectures ICCAD Special Session PaperabstractThe goal of this special session paper is to introduce and discuss different breakthrough technologies as well as novel architectures and how they together may reshape the future of Artificial Intelligent. Our aim is to provide a comprehensive overview on the latest advances in brain-inspired computing and how the latter can be realized when emerging technologies, using beyond-CMOS devices, are coupled with novel computing paradigms that go beyond von Neumann architectures. Different emerging technologies like Ferroelectric Field-Effect Transistor (FeFET), Phase Change Memory (PCM), and Resistive RAM (ReRAM) are discussed, demonstrating their promising capability in building neuromorphic computing architectures that are inspired by nature. In addition, this special session paper discusses various novel concepts such as Logic-in-Memory (LIM), Processing-in-Memory (PIM), and Spiking Neural Networks (SNNs) towards exploring the far-reaching consequences of beyond von Neumann computing on accelerating deep learning. Finally, the latest trends in brain-inspired computing are summarized into algorithm, technology, and application-driven innovations towards comparing different PIM architectures. Hussam Amrouch, Jian-Jia Chen, Kaushik Roy 0001, Yuan Xie 0001, Indranil Chakraborty, Wenqin Huangfu, Ling Liang 0003, Fengbin Tu, Cheng Wang 0036, Mikail Yayla |
ICCAD | 1 |
| 2021 | ICCAD Tutorial Session Paper Ferroelectric FET Technology and Applications: From Devices to SystemsabstractThe rapidly increasing volume and complexity of data is demanding the relentless scaling of computing power. With transistor feature size approaching physical limits, the benefits that CMOS technology can provide is diminishing. For future energy efficient computing systems, researchers aim to exploit various emerging nanotechnologies to replace conventional CMOS technology. In particular, ferroelectric FETs (FeFETs) appear to be a promising candidate to continue improving energy efficiency for data-intensive applications. Advances in FeFET scalability and FeFET compatibility with CMOS have sparked growing interest in device, circuit, and system communities. While FeFET is still evolving, many researchers and developers are already cautiously optimistic about its future. This paper provides a review on FeFET's recent technology advances, challenges, and opportunities, with a particular emphasis upon device modeling and circuit design of FeFET content addressable memory, as well as their applications in machine learning. Hussam Amrouch, Xiaobo Sharon Hu, Arman Kazemi, Ann Franchesca Laguna, Kai Ni 0004, Michael T. Niemier, Mohammad Mehdi Sharifi, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ICCAD | 1 |
| 2021 | Security Closure of Physical Layouts ICCAD Special Session PaperabstractComputer-aided design (CAD) tools traditionally optimize for power, performance, and area (PPA). However, given a vast number of hardware security threats, we call for secure-by-design CAD flows, to adopt principles of secure hardware design and streamline security closure throughout the flow. The stakes are high for integrated circuit (IC) vendors and design companies, as security risks that are not addressed during design will inevitably be exploited in the field, where vulnerabilities are almost impossible to fix. This paper highlights the need for security closure of physical layouts because efforts taken toward securing ICs at higher abstraction layers may be futile without support for securing the tape-out ready layouts. Johann Knechtel, Jayanth Gopinath, Jitendra Bhandari, Mohammed Ashraf, Hussam Amrouch, Shekhar Borkar, Sung Kyu Lim, Ozgur Sinanoglu, Ramesh Karri |
ICCAD | 5 |
| 2021 | Toward Security Closure in the Face of Reliability Effects ICCAD Special Session PaperabstractThe reliable operation of ICs is subject to physical effects like electromigration, thermal and stress migration, negative bias temperature instability, hot-carrier injection, etc. While these effects have been studied thoroughly for IC design, threats of their subtle exploitation are not captured well yet. In this paper, we open up a path for security closure of physical layouts in the face of reliability effects. Toward that end, we first review migration effects in interconnects and aging effects in transistors, along with established and emerging means for handling these effects during IC design. Next, we study security threats arising from these effects; in particular, we cover migration effects-based, disruptive Trojans and aging-exacerbated side-channel leakage. Finally, we outline corresponding strategies for security closure of physical layouts, along with an outline for CAD frameworks. Jens Lienig, Susann Rothe, Matthias Thiele, Nikhil Rangarajan, Mohammed Ashraf, Mohammed Nabeel Thari Moopan, Hussam Amrouch, Ozgur Sinanoglu, Johann Knechtel |
ICCAD | 7 |
| 2021 | Positive/Negative Approximate Multipliers for DNN AcceleratorsabstractRecent Deep Neural Networks (DNNs) manage to deliver superhuman accuracy levels on many AI tasks. DNN accelerators are becoming integral components of modern systems-on-chips. DNNs perform millions of arithmetic operations per inference and DNN accelerators integrate thousands of multiply-accumulate units leading to increased energy requirements. To lower the energy consumption of DNN accelerators, approximate computing principles are employed. However, complex DNNs can be increasingly sensitive to approximation. In this work, we present a dynamically configurable approximate multiplier that supports three operation modes, i.e., exact, positive error, and negative error. In addition, we propose a filter-oriented approximation method to map the weights to the appropriate modes of the approximate multiplier. Our mapping algorithm balances the positive with the negative errors due to the approximate multiplications, aiming at maximizing the energy reduction while minimizing the overall convolution error. We evaluate our approach on multiple DNNs and datasets against state-of-the-art approaches, where our method achieves 18.33% energy gains on average across 7 NNs on 4 different datasets for a maximum accuracy drop of only 1%. Ourania Spantidi, Georgios Zervakis 0001, Iraklis Anagnostopoulos, Hussam Amrouch, Jörg Henkel |
ICCAD | 4 |
| 2021 | Binarized SNNs: Efficient and Error-Resilient Spiking Neural Networks through BinarizationabstractSpiking Neural Networks (SNNs) are considered the third generation of NNs and can reach similar accuracy as conventional deep NNs, but with a considerable improvement in efficiency. However, to achieve high accuracy, state-of-the-art SNNs employ stochastic spike coding of the inputs, requiring multiple cycles of computation. Because of this and due to the nature of analog computing, it is required to accumulate and hold the charges of multiple cycles, necessitating a large membrane capacitor. This results in high energy, long latency, and expensive area costs, constituting one of the major bottlenecks in analog SNN implementations. Membrane capacitor size determines the precision of the firing time. Hence reducing the capacitor size considerably degrades the inference accuracy. To alleviate this, we focus on bridging the gap between binarized NNs (BNNs) and SNNs. BNNs are rapidly emerging as an attractive alternative for NNs due to their high efficiency and error tolerance. In this work, we evaluate the impact of deploying error-resilient BNNs, i.e. BNNs that have been proactively trained in the presence of errors, on analog implementation of SNNs. We show that for BNNs, the capacitor size and latency can be reduced significantly compared to state-of-the-art SNNs, which employ multi-bit models. Our experiments demonstrate that when error-resilient BNNs are deployed on analog-based SNN accelerator, the size of the membrane capacitor is reduced by 50%, the inference latency is decreased by two orders of magnitude, and energy is reduced by 57% compared to the baseline 4-bit SNN implementation, under minimal accuracy cost. Ming-Liang Wei, Mikail Yayla, Shu-Yin Ho, Jian-Jia Chen, Chia-Lin Yang, Hussam Amrouch |
ICCAD | 6 |
| 2021 | Brain-Inspired Computing for Wafer Map Defect Pattern ClassificationabstractBrain-Inspired hyperdimensional computing is a quickly emerging alternative machine-learning concept. Hypervectors with thousands of dimensions represent real-world data. Thanks to this redundancy, the system becomes robust against noise in the input data, but also resilient against faults, similar to the human brain. The light-weight operations with hypervectors are fully parallelizable enabling fast learning and inference at the edge. A classifier achieving high accuracies can be created through one-shot learning from few examples. Such a feature is particularly valuable in the area of semiconductor testing, where the number of training samples, especially for cutting-edge technology, is limited. In this work, we explore the applicability of brain-inspired hyperdimensional computing to the field of testing for the first time. With the example of wafer map defect pattern classification, we investigate the challenges and opportunities of this emerging concept. Paul R. Genssler, Hussam Amrouch |
ITC | 2 |
| 2021 | Machine Learning for Circuit Aging Estimation under Workload DependencyabstractCircuit analysis with respect to aging-induced degradation is critical to ensure correct operation throughout the entire lifetime of a chip. However, state-of-the-art techniques only allow for the consideration of uniformly applied degradation, despite the fact that different workloads will lead to different degradations due to the different induced activities. This imposes over-pessimism in estimating the required timing guardbands, resulting in unnecessary losses of performance and efficiency. In this work, we propose an approach that takes real-world workload dependencies into account and generates workload-specific aging-aware standard cell libraries. This allows for accurate analysis of circuits under the actual effect of aging-induced degradation. We make use of machine learning techniques to overcome infeasible simulation times for individual transistor aging while sustaining high accuracy. In our evaluation on the PULP microprocessor, we achieve predictions of workload-dependent aging-aware standard cells with an average accuracy (R2score) of 94.7 %. Using the predicted cell libraries in Static Timing Analysis, timing guardbands are reported with an error of less than 0.1 %. We demonstrate that timing guardband requirements can be reduced by up to 21 % by considering specific workloads over worst-case analysis as performed in state-of-the-art tool flows. Florian Klemme, Hussam Amrouch |
ITC | 2 |
| 2021 | Towards Reliable In-Memory Computing: From Emerging Devices to Post-von-Neumann ArchitecturesabstractBreakthroughs in Deep neural networks (DNNs) steadily bring new innovations that substantially improve our daily life. However, DNNs overwhelm our existing computer architectures because the latter is largely bottlenecked by the data movement between memory and processing units. As a matter of fact, in the current von-Neumann architecture, which has remained unchanged since the beginning, data repeatedly moves back and forth between the physically-separated processing units (e.g., CPU, accelerator, etc.) and memory. This, in turn, inevitably leads to large latency and efficiency losses. In DNNs such a bottleneck becomes more and more prominent due to the massive amount of data that must be frequently transferred. This paper provides a cross-layer overview on how post-von-Neumann in-memory computing (IMC) architectures can be realized using three different emerging technologies: Charge-based ferroelectric transistors for logic-in-memory computations; memristive devices for unconventional brain-inspired computing; and ultra-low-power memristors especially suitable for Edge AI. Various levels of abstraction will be covered starting from semiconductor device physics to circuit and microarchitecture levels all the way up to the system level, but special attention will be put on reliability aspects. Hussam Amrouch, Nan Du 0004, Anteneh Gebregiorgis, Said Hamdioui, Ilia Polian |
VLSI-SoC | 1 |
| 2021 | Transistor Self-Heating: The Rising Challenge for Semiconductor TestingabstractQuantum confinement in 3-D device structure together with the newly employed materials like silicon-germanium (SiGe) in advanced technologies (e.g., FinFET, nanowire, nanosheets, etc.) makes transistors seriously suffer from localized self-heating effects in which generated heat within the transistor's channel is trapped inside. This is mainly due to the much lower channel and surrounding material thermal conductivity and hence lower ability for heat dissipation along with the firm isolation needed for better gate control. Self-heating effects strongly accelerate transistor aging and all the underlying defect generation mechanisms leading to serious reliability problems during the early life of chips. The key challenge in transistor self-heating when it comes to semiconductor testing is the profound difficulty in measuring self-heating directly as generated heat is trapped inside the transistor. Failing in capturing self-heating phenomenon during IC testing would later lead to chips malfunctions at run-time and hence early life failures because of reliability degradations and failure mechanisms will be unexpectedly accelerated akin to excessive internal temperatures. In this paper, we investigate the impact of self-heating effects on n-type and p-type FinFET transistors calibrated with Intel 14 nm measurement data using mature Technology CAD (TCAD) simulations. Then, the industry standard compact model for FinFET technologies (BSIM-CMG) is carefully calibrated to accurately model and reproduce all measurements. This enables circuit's designers, for the first time, to accurately investigate how emerging self-heating effects in transistors impacts the performance and power of large circuits. This opens new doors for developing novel Design-for-Testing methods that effectively reveal self-heating effects and increase the yield of chips. Om Prakash 0007, Chetan K. Dabhi, Yogesh Singh Chauhan, Hussam Amrouch |
VTS | 4 |
| 2021 | Special Session: Machine Learning for Semiconductor Test and ReliabilityabstractWith technology scaling approaching atomic levels, IC test and diagnosis of complex System-on-Chips (SoCs) become overwhelming challenging. In addition, sustaining the reliability of transistors as well as circuits at such extreme feature sizes, for the entire projected lifetime, also become profoundly difficult. This holds even more when it comes to emerging technologies that go beyond convectional CMOS in which the underlying physics are not yet fully understood. In this special session paper, we describe the usage of machine learning in several test and reliability related areas. First, we demonstrate the vital role that machine learning can play in IC test showing the importance of explainability as a frontier for machine learning in IC test. Afterwards, we discuss how novel physics-informed neural networks can be employed to model electrostatic problems in VLSI designs. This is essential to mitigate the deleterious effects of of time dependent dielectric breakdown, which is the key source of reliability degradations. Finally, we discuss the major sources of reliability degradations at the transistor level in advanced technology nodes such as transistor aging phenomena and self-heating effects as well as we demonstrate how machine learning approaches can further help in developing reliable emerging technologies. Hussam Amrouch, Animesh Basak Chowdhury, Wentian Jin, Ramesh Karri, Farshad Khorrami, Prashanth Krishnamurthy, Ilia Polian, Victor M. van Santen, Benjamin Tan 0001, Sheldon X.-D. Tan |
VTS | 1 |
| 2021 | Reliability-Driven Voltage Optimization for NCFET-based SRAM Memory BanksabstractNegative Capacitance Field-Effect Transistors (NCFET) are promising significant power reductions while maintaining performance due to their internal voltage amplification. However, the addition of the ferroelectric layer also introduces a higher gate capacitance, which has to be charged and discharged resulting in higher power consumption. This results in trade-offs when employing NC-FinFET with respect to the thickness of the ferroelectric layer and their operating voltage on power, performance and reliability in circuits. This design-space is currently not explored, as existing research focused on a transistor-to-transistor comparison to show the superiority of NC-FinFET at the same voltage. In this work, we evaluate NC-FinFET employment in a full SRAM memory array (including write driver, sense amplifier, pre-charging, etc.) to obtain circuit delay, read and hold power and reliability metrics. This work shows, that solely evaluating SRAM cells results in inaccurate delay and power estimations compared to a full SRAM array. We explore iso-voltage and iso-performance NC-FinFET operation. Additionally, we explore two new operation modes: operating NC-FinFET within the same overall power consumption (iso-power) and operating at the same noise margins (iso-reliability). This exploration shows, for the first time, how ferroelectric layer thickness plays a role on reliability as a 4 nm layer features a 47% loss compared to FinFET. Lastly, we obtain the activity of a register file in a processor simulator to obtain the ultimate impact on power and energy consumption of employing NC-FinFET in a microprocessor. Victor M. van Santen, Simon Thomann, Yogesh S. Chauchan, Jörg Henkel, Hussam Amrouch |
VTS | 5 |
| 2021 | On the Reliability of In-Memory Computing: Impact of Temperature on Ferroelectric TCAMabstractWith the rapid development of emerging technologies, especially the ferroelectric field-effect transistors (FeFETs), the density and energy efficiency of ternary content addressable memory (TCAM) have been increasingly improved. TCAM plays a major role in realizing In-Memory Computing and other brain-inspired computing concepts. Recently, the parallel search functionality of a FeFET based ultra-dense TCAM design is also enhanced with a Hamming distance-based approximate search scheme. However, in order to realize the highly-promising TCAM design, in which the approximate search function based on Hamming distance is implemented, it is inevitable to investigate the impact of temperature on the reliability of FeFET-based TCAM cells as well as all involved peripheral circuits. In this paper, the temperature impact on the FeFET at the device level and the approximate TCAM design at the circuit level is investigated for the first time. The demonstrated example of a FeFET-based TCAM array shows that the unique temperature dependency of a FeFET device can help mitigate the temperature impact on the FeFET TCAM array. Based on the observation, we showcase, evaluate, and discuss in detail one strategy to eliminate the temperature impact on the approximate TCAM design. Understanding and mitigating the deleterious impact of temperature on the reliability of FeFET-based TCAM circuits is essential to ensure reliable In-Memory Computing. Simon Thomann, Chao Li 0065, Cheng Zhuo, Om Prakash 0007, Xunzhao Yin, Xiaobo Sharon Hu, Hussam Amrouch |
VTS | 7 |
| 2021 | Power-Efficient Heterogeneous Many-Core Design With NCFET TechnologyabstractMulti-/many-core, homogeneous or heterogeneous architectures, using the existing CMOS technology are inevitably approaching the limit of attainable power efficiency due to the fundamental limits in scaling. Negative Capacitance Field-Effect Transistor (NCFET) is rapidly emerging as an alternative technology that promises a multi-fold increase in the power efficiency of transistors, yet is compatible with the existing CMOS fabrication process. NCFET incorporates a ferroelectric (FE) layer within the transistor's gate stack, which exhibits a negative capacitance effect amplifying the internal voltage. NCFET has been in detail studied in both physics and devices/circuits communities where its superiority has been demonstrated in semiconductor measurements. However, the full promise of NCFET remains unmodeled and unquantified unless the research is further continued to the microarchitecture and system levels. This article, for the first time, explores system- and application-level benefits of NCFET-based multi-/many-core designs in terms of performance and power-efficiency compared to state-of-the-art FinFET-based designs. This exploration is done first through analytical modeling in which we extend Amdahl's law for NCFET multi-/many-cores, and then through quantitative modeling. The latter is achieved through RTL- and system-level simulations of NCFET-based multi-cores. The analytical modeling shows that a novel type of technology-based heterogeneity in which cores with the same microarchitecture but different FE thickness are combined is highly beneficial. Our exploration shows that this novel heterogeneity increases the power-efficiency by up to 3.5× over homogeneous systems and even achieves 8.3% better performance and 20% higher power-efficiency than conventional heterogeneity in the microarchitecture without having to cope with the complexity of managing different microarchitectures. Sami Salamin, Martin Rapp, Anuj Pathania, Arka Maity, Jörg Henkel, Tulika Mitra, Hussam Amrouch |
IEEE Trans. Computers | 7 |
| 2021 | Post-Silicon Heat-Source Identification and Machine-Learning-Based Thermal Modeling Using Infrared Thermal ImagingabstractIn this article, we present a novel post-silicon approach to locating the dominant heat sources on commercial multicore processors using heatmaps measured via an infrared (IR) thermal imaging setup. To locate the heat sources, 2-D spatial Laplacian transformation is performed on the heatmaps followed by K-means clustering to find the dominant power/heat-source clusters. This is an exclusively post-silicon approach that does not require any knowledge of the underlying design of the commercial chips other than the information that is publicly available. Since the identified clusters are the thermally vulnerable areas on the die, we then propose a machine-learning-based framework to deriving a thermal model capable of estimating their temperatures during online use. Our approach involves collecting transient temperature data of the aforementioned heat sources and synchronized high-level performance metrics from the chip, and training a long-short-term-memory (LSTM) neural network (NN) that uses the performance metrics as inputs to estimate the temperatures of the identified heat sources in real time. Since the model is meant for real-time use, we explore methods of reducing the performance overhead and inference time of the model. This includes a novel power correlation-based approach to identifying the thermally irrelevant performance metrics and eliminating them in order to reduce the input dimensionality of the model, and an analysis on network sizing to determine the ideal NN configuration for the problem at hand. The model is trained and tested exclusively using measured thermal data from commercial multicore processors. The experimental results from two Intel multicore processors (i5-3337U and i7-8650U) show that the proposed approach achieves very high accuracy (root-mean-square error: 0.55 °C-0.93 °C) in estimating the temperatures of all the identified heat sources on the chip. Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Automated Design Approximation to Overcome Circuit AgingabstractTransistor aging phenomena manifest themselves as degradations in the main electrical characteristics of transistors. Over time, they result in a significant increase of cell propagation delay, leading to errors due to timing violations, since the operating frequency becomes unsustainable as the circuit ages. Conventional techniques employ timing guardbands to mitigate aging-induced delay increase, which leads to considerable performance losses from the beginning of the circuit’s lifetime. Leveraging the inherent error resilience of a vast number of application domains, approximate computing was recently introduced as an aging mitigation mechanism. In this work, we present the first automated framework for generatingaging-aware approximate circuits. Our framework, by applying directed gate-level netlist approximation, induces a small functional error and recovers the delay degradation due to aging. As a result, our optimized circuits eliminate aging-induced timing errors. Experimental evaluation over a variety of arithmetic circuits and image processing benchmarks demonstrates that for an average error of merely$5\times 10^{-3}$, our framework completely eliminates aging-induced timing guardbands. Compared to the respective baseline circuits without timing guardbands (i.e., iso-performance evaluation), the error of the circuits generated by our framework is$1208\times $smaller. Konstantinos Balaskas, Georgios Zervakis 0001, Hussam Amrouch, Jörg Henkel, Kostas Siozios |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | Machine Learning for On-the-Fly Reliability-Aware Cell Library CharacterizationabstractAging-induced degradation imposes a major challenge to the designer when estimating timing guardbands. This problem increases as traditional worst-case corners bring over-pessimism to designers, exacerbating competitive and close-to-the-edge designs. In this work, we present an accurate machine learning approach for aging-aware cell library characterization, enabling the designer to evaluate their circuit under the impact of precisely selected degradation. Unlike state of the art, we bring cell library characterization to the designer, empowering their capability in exploring the impact of aging while protecting confidential information from the foundry at the same time. Furthermore, the fast inference of cell libraries makes it feasible, for the first time, to examine aging-induced variability analysis in a Monte-Carlo fashion. Finally, we show that the designer is able to select a less pessimistic timing guardband by choosing adequate delta threshold voltage ( ΔVth) for their design and their needs. Our machine learning approach reaches an R2score of $>99\%$ for almost all data stored in the cell library. Only timing constraints show slightly less accuracy with an R2score around 95%. When using ML-characterized libraries in static timing analysis, we achieve errors smaller than ±0.5% and ±0.1% for path delay and dynamic power, respectively. Errors in leakage power are negligible and even smaller by orders of magnitude. Our machine learning implementation for standard cell library characterization is publicly available. Download: https://opensource.mlcad.org Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | On the Resiliency of NCFET Circuits Against Voltage Over-ScalingabstractApproximate computing is established as a design alternative to improve the energy requirements of a vast number of applications, leveraging their intrinsic error tolerance. Voltage over-scaling (VOS) is one of the most energy-efficient approximation techniques, but its exploitation is still limited due to the large errors it induces. In this work, we investigate, for the first time, the resiliency of negative capacitance transistor (NCFET) technology to VOS in comparison to conventional CMOS technology. Our work reveals that circuits implemented using the NCFET technology exhibit much less timing errors under VOS due to the inherent voltage amplification provided by the ferroelectric layer. NCFET is one of the very promising emerging technologies that is rapidly evolving for low-power circuit as it enables the transistors to switch faster without the need to increase the voltage. We demonstrate how NCFET technology allows circuit designers to effectively employ VOS to boost the efficiency of their approximate circuits, while still keeping the induced errors marginal. Our analysis shows that the VOS-resilience of NCFET circuits enables maximizing the voltage decrease and thus, NCFET based VOS approximate circuits achieve from 1.83× up to 2.78× higher energy reduction compared to the corresponding FinFET circuits for the same error bounds. Guilherme Paim, Georgios Zervakis 0001, Girish Pahwa, Yogesh Singh Chauhan, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | PROTON: Post-Synthesis Ferroelectric Thickness Optimization for NCFET CircuitsabstractFor the first time, we demonstrate an optimization technique to synthesize circuits in the Negative Capacitance FET (NCFET) technology. NCFET is a rapidly emerging technology to replace the currently employed CMOS technology due to its profound ability to overcome the fundamental limit in scaling along with its full compatibility with the existing fabrication process. This is achieved by replacing the traditional transistor gate dielectric with a ferroelectric layer that manifests itself as a Negative Capacitance (NC), which magnifies the electric field. As a result, NCFET-based circuits can operate at a higher clock frequency without the need to increase the operating voltage. NC breaks one of the fundamental laws in physics in which the total capacitance of two capacitors connected in series becomes larger–instead of smaller in ordinary capacitors– than each of them. This could lead to sub-optimal netlists, suffering from significant increase in dynamic power and IR-drops. To suppress that, we employ the relation between delay decrease and capacitance increase of gates w.r.t ferroelectric thickness. Our technique takes an optimized netlist, obtained from commercial EDA tools, and then selectively determines the optimal ferroelectric thickness for each gate in the netlist, so that the maximum performance provided by NCFET is still achieved while the dynamic power is considerably decreased (45% on average),i.e., no trade-offs. Particularly, our technique enables the full exploitation of the performance benefits originating by NCFET, at a significantly lower (power) cost. Compared to state of the art, our technique decreases the energy-delay-product of circuits by 25% on average and reduces the deleterious effects of IR-drop by 56%. Hence, efficiency and reliability of circuits are improved without any loss in the obtained performance from NCFET. Sami Salamin, Georgios Zervakis 0001, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | NCFET to Rescue Technology Scaling: Opportunities and ChallengesabstractNegative Capacitance Field Effect Transistor (NCFET) is one of the promising emerging technologies that may overcome the fundamental limits of conventional CMOS technology. NCFET features a ferroelectric (FE) layer within the transistor's gate, which internally amplifies the voltage, allowing NCFET to operate at a lower voltage while sustaining performance at considerable energy savings. In this work, we raise awareness that n- and p-NCFET transistors are asymmetrically affected by the FE layer and show, for the first time, how this asymmetry results in unbalanced circuit performance (e.g., longer fall than rise propagation delay, reduced noise margins). As NCFET are meant to maintain performance while reducing power, we present a solution by scaling the number of fins in n-NCFET to regain symmetry. We optimize iteratively in conjunction with supply voltage scaling to find the minimal energy consumption while maintaining performance. In our first case study, we achieve at least 34% lower power consumption and thus 34% higher energy efficiency as the circuit exhibits identical propagation delay. However, our second case study reveals that NCFETs can consume 3× more power and energy than the FinFET design. In summary, not considering the asymmetry and replacing FinFET with current-matched NCFET results in unreliable circuits (timing violations). This work exemplifies how the power and energy consumption of a NCFET circuit might surpass that of a FinFET, if circuits are designed considering asymmetry and circuit metric matching. Hussam Amrouch, Victor M. van Santen, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
ASP-DAC | 1 |
| 2020 | Machine Learning Based Online Full-Chip Heatmap EstimationabstractRuntime power and thermal control is crucial in any modern processor. However, these control schemes require accurate real-time temperature information, ideally of the entire die area, in order to be effective. On-chip temperature sensors alone cannot provide the full-chip temperature information since the number of sensors that are typically available is very limited due to their high area and power overheads. Furthermore, as we will demonstrate, the peak locations within hot-spots are not stationary and are very workload dependent, making it difficult to rely on fixed temperature sensors alone. Therefore, we propose a novel approach to real-time estimation of fullchip transient heatmaps for commercial processors based on machine learning. The model derived in this work supplements the temperature data sensed from the existing on-chip sensors, allowing for the development of more robust runtime power and thermal control schemes that can take advantage of the additional thermal information that is otherwise not available. The new approach involves offline acquisition of accurate spatial and temporal heatmaps using an infrared thermal imaging setup while nominal working conditions are maintained on the chip. To build the dynamic thermal model, we apply LongShort-Term-Memory (LSTM) neutral networks with system-level variables such as chip frequency, instruction counts, and other performance metrics as inputs. To reduce the dimensionality of the model, 2D spatial discrete cosine transformation (DCT) is first performed on the heatmaps so that they can be expressed with just their dominant DCT frequencies. Our study shows that only 6×6 DCT coefficients are required to maintain sufficient accuracy across a variety of workloads. Experimental results show that the proposed approach can estimate the full-chip heatmaps with less than 1.4°C root-mean-square-error and take only ~19ms for each inference which suits well for real-time use. Sheriff Sadiqbatcha, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan |
ASP-DAC | 4 |
| 2020 | Impact of Self-Heating on Performance, Power and Reliability in FinFET TechnologyabstractSelf-heating is one of the biggest threats to reliability in current and advanced CMOS technologies like FinFET and Nanowire, respectively. Encapsulating the channel with the gate dielectric improved electrostatics, but also thermally insulates the channel resulting in elevated channel temperatures as the generated heat is trapped within the channel. Elevated channel temperatures lowers the performance, increases leakage power and degrades the reliability of circuits. Self-heating becomes worse in each new transistor structure (from planar transistor to FinFET to Nanowire) due to the ever-increasing thermal resistance of the transistor. This leads to elevated temperatures, which must be carefully considered while designing circuits. Otherwise, reliability cannot be ensured. This work presents a self-heating study to illustrate how self-heating matters in digital circuits. It also explores the impact of running workloads in SRAM arrays, such as register files in CPUs, and how self-heating effects in SRAM cells can be mitigated. Victor M. van Santen, Paul R. Genssler, Om Prakash 0007, Simon Thomann, Jörg Henkel, Hussam Amrouch |
ASP-DAC | 6 |
| 2020 | Run-Time Accuracy Reconfigurable Stochastic Computing for Dynamic Reliability and Power Management: Work-in-ProgressabstractIn this paper, we propose a novel accuracy-reconfigurable stochastic computing (ARSC) framework for dynamic reliability and power management. Different than the existing stochastic computing works, where the accuracy versus power/energy trade-off is carried out in the design time, the new ARSC design can change accuracy or bit-width of the data in the run-time so that it can accommodate the long-term aging effects by slowing the system clock frequency at the cost of accuracy while maintaining the throughput of the computing. We validate the ARSC concept on a discrete cosine transformation (DCT) and inverse DCT designs for image compressing/decompressing applications, which are implemented on Xilinx Spartan-6 family XC6SLX45 platform. Experimental results show that the new design can easily mitigate the long-term aging induced effects by accuracy trade-off while maintaining the throughput of the whole computing process using simple frequency scaling. We further show that one-bit precision loss for the input data, which translated to 3.44dB of the accuracy loss in term of Peak Signal to Noise Ratio (PSNR) for images, we can sufficiently compensate the NBTI induced aging effects in 10 years while maintaining the pre-aging computing throughput of 7.19 frames per second. At the same time, we can save 74% power consumption by 10.67dB of accuracy loss. The proposed ARSC computing framework also allows much aggressive frequency scaling, which can lead to order of magnitude power savings compared to the traditional dynamic voltage and frequency scaling (DVFS) techniques. Shuyuan Yu, Han Zhou 0002, Shaoyi Peng, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan |
CASES | 4 |
| 2020 | Impact of NBTI Aging on Self-Heating in Nanowire FETabstractThis is the first work that investigates the impact of Negative Bias Temperature Instability (NBTI) on the Self-Heating (SH) phenomenon in Silicon Nanowire Field-Effect Transistors (SiNW-FETs). We investigate the individual as well as joint impact of NBTI and SH on pSiNW-FETs and demonstrate that NBTI-induced traps mitigate SH effects due to reduced current densities. Our Technology CAD (TCAD)-based SiNW-FET device is calibrated against experimental data. It accounts for thermodynamic and hydrodynamic effects in 3-D nano structures for accurate modeling of carrier transport mechanisms. Our analysis focuses on how lattice temperature, thermal resistance and thermal capacitance of pSiNW-FETs are affected due to NBTI, demonstrating that accurate self-heating modeling necessitates considering the effects that NBTI aging has over time. Hence, NBTI and SH effects need to be jointly and not individually modeled. Our evaluation shows that an individual modeling of NBTI and SH effects leads to a noticeable overestimation of the overall induced delay increase in circuits due to the impact of NBTI traps on SH mitigation. Hence, it is necessary to model NBTI and SH effects jointly in order to estimate efficient (i.e. small, yet sufficient) timing guardbands that protect circuits against timing violations, which will occur at runtime due to delay increases induced by aging and self-heating. Om Prakash 0007, Hussam Amrouch, Sanjeev Manhas 0001, Jörg Henkel |
DATE | 2 |
| 2020 | Energy Optimization in NCFET-based ProcessorsabstractEnergy consumption is a key optimization goal for all modern processors. Negative Capacitance Field-Effect Transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. Thickness of the additional ferroelectric layer, frequency, and voltage are the key parameters in NCFET technology that impact the power and frequency of processors. However, their joint impact on energy optimization has not been investigated yet.In this work, we are the first to demonstrate that conventional (i.e., NCFET-unaware) dynamic voltage/frequency scaling (DVFS) techniques to minimize energy are sub-optimal when applied to NCFET-based processors. We further demonstrate that state-of-the-art NCFET-aware voltage scaling for power minimization is also sub-optimal when it comes to energy. This work provides the first NCFET-aware DVFS technique that optimizes the processor's energy through optimal runtime frequency/voltage selection. In NCFETs, energy-optimal frequency and voltage are dependent on the workload and technology parameters. Our NCFET-aware DVFS technique considers these effects to perform optimal voltage/frequency selection at runtime depending on workload characteristics. Results show up to 90 % energy savings compared to conventional DVFS techniques. Compared to state-of-the-art NCFET-aware power management, our technique provides up to 72 % energy savings along with 3.7x higher performance. Sami Salamin, Martin Rapp, Hussam Amrouch, Andreas Gerstlauer, Jörg Henkel |
DATE | 3 |
| 2020 | Cell Library Characterization using Machine Learning for Design Technology Co-OptimizationabstractTo explore the full potential of any circuit and ensure its functionality at run-time, cell libraries beyond the typical PVT corners are needed. This holds even more for emerging technologies like Negative Capacitance (NC)-FinFET, where research in finding the optimal set of transistor parameters is still in its infancy. Design Technology Co-Optimization (DTCO) tackles bridging the large existing gap between device physics and the figures of merit of circuits. In this paper, we propose a Machine Learning (ML) approach to rapidly generate full cell libraries on demand. This enables the designer to perform extensive design space exploration and fully automated Design Technology Co-Optimization while lowering the barrier of accessibility. We demonstrate library prediction with an R2 score of around 98% for individual values and Static Timing Analysis (STA) reports. Experimental results show that our DTCO approach overestimates the achievable improvement by around 5%, nevertheless improving upon the baseline configuration. Florian Klemme, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
ICCAD | 4 |
| 2020 | Modeling Emerging Technologies using Machine Learning: Challenges and OpportunitiesabstractCompact models of transistors act as the link between semiconductor technology and circuit design via circuit simulations. Unfortunately, compact model development and calibration is a challenging and time-intensive task, hindering rapid prototyping of a circuit (via circuit simulations) in emerging technologies. Moreover, foundries want to protect their confidential technology details to prevent reverse engineering. Hence, they limit access to compact transistor models of commercial technologies (e.g., with Non-Disclosure-Agreements). In this work, we propose Machine Learning (ML) to bridge the gap between early device measurements and later occurring compact model development. Our approach employs a Neural Network (NN) that captures the electrical response of a conventional FinFET transistor without knowledge of semiconductor physics. Additionally, our approach can be applied to emerging technologies, using Negative Capacitance FinFET (NC-FinFET) as an example for a (challenging to model) emerging technology. Inherently, the black-box nature of ML approaches keeps technology manufacturing details confidential. Furthermore, we show how using solely R2 score as our fitness function is insufficient and instead propose fitness based on key electrical characteristics or transistors like threshold voltage. Our NN-based transistor modeling can infer FinFET and NC-FinFET with an R2 score larger than 0.99 and transistor characteristics within 5% of experimental data. Florian Klemme, Jannik Prinz, Victor M. van Santen, Jörg Henkel, Hussam Amrouch |
ICCAD | 5 |
| 2020 | NPU Thermal ManagementabstractNeural processing units (NPUs) are becoming an integral part in all modern computing systems due to their substantial role in accelerating neural networks (NNs). The significant improvements in cost-energy-performance stem from the massive array of multiply accumulate (MAC) units that remarkably boosts the throughput of NN inference. In this work, we are the first to investigate the thermal challenges that NPUs bring, revealing how MAC arrays, which form the heart of any NPU, impose serious thermal bottlenecks to on-chip systems due to their excessive power densities. For the first time, we explore: 1) the effectiveness of precision scaling and frequency scaling (FS) in temperature reductions and 2) how advanced on-chip cooling using superlattice thin-film thermoelectric (TE) open doors for new tradeoffs between temperature, throughput, cooling cost, and inference accuracy in NPU chips. Our work unveils that hybrid thermal management, which composes different means to reduce the NPU temperature, is a key. To achieve that, we propose and implement PFS-TE technique that couples precision and FS together with superlattice TE cooling for effective NPU thermal management. Using commercial signoff tools, we obtain accurate power and timing analysis of MAC arrays after a full-chip design is performed based on 14-nm Intel FinFET technology. Then, multiphysics simulations using finite-element methods are carried out for accurate heat simulations in the presence and absence of on-chip cooling. Afterward, comprehensive design-space exploration is presented to demonstrate the Pareto frontier and the existing tradeoffs between temperature reductions, power overheads due to cooling, throughput, and inference accuracy. Using a wide range of NNs trained for image classification, experimental results demonstrate that our novel NPU thermal management increases the inference efficiency (TOPS/Joule) by 1.33×, 1.87×, and 2× under different temperature constraints; 105 °C, 85 °C, and 70 °C, respectively, while the average accuracy drops merely from 89.0% to 85.5%. Hussam Amrouch, Georgios Zervakis 0001, Sami Salamin, Hammam Kattan, Iraklis Anagnostopoulos, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2020 | Dynamic Power and Energy Management for NCFET-Based ProcessorsabstractPower and energy consumption are the key optimization goals in all modern processors. Negative capacitance field-effect transistors (NCFETs) are a leading emerging technology that promises outstanding performance in addition to better energy efficiency. The thickness of the added ferroelectric layer as well as frequency and voltage are the key parameters that impact the power and energy of NCFET-based processors in addition to the characteristics of runtime workloads. Unlike existing CMOS technologies, operating NCFET-based processors at a higher frequency than the required minimum can result in power/energy minimization. The optimal operating point, however, strongly depends on dynamic workload characteristics and technology parameters. In this work, we propose and implement the first NCFET-aware power and energy management approach that minimizes the processor's power and energy through optimal voltage/frequency selection under different runtime scenarios. Such an NCFET-aware approach does not result in any tradeoff between power/energy and performance. Instead, it can achieve higher performance while minimizing energy. A comprehensive, simulation-based evaluation of our runtime management under realistic workloads demonstrates up to 58% energy saving with 2.1× higher performance, and 46% power saving compared to conventional NCFET-unaware management techniques, over the total execution of a benchmark. Compared to state-of-the-art NCFET-aware management techniques, our technique provides up to 49% energy saving and 32% power saving. Sami Salamin, Martin Rapp, Jörg Henkel, Andreas Gerstlauer, Hussam Amrouch |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Exposing Hardware Trojans in Embedded Platforms via Short-Term AgingabstractWe demonstrate a novel technique that employs transistor short-term aging effects in integrated circuits (ICs) to detect hardware Trojans in embedded systems. In advanced technology nodes (≤ 45 nm), voltage scaling in combination with short-term aging opens doors for short-term degradations. The induced short-term degradations result in dynamic variation of delays along various paths within the IC. Aging degradation generated under fast voltage switching from high to low results in bit errors at the circuit output. Our experiments use short-term aging-aware standard cell libraries to show the effectiveness of short-term aging to detect hardware Trojans. We extract a rich set of features that capture bit error patterns at the outputs of the IC. We use a one class SVM-based classifier that uses these features to learn the distribution of bit errors at the outputs of a clean IC. We discern the deviation in the pattern of bit errors due to a Trojan in the IC from the baseline distribution. To reiterate, the method uses the model of a clean IC. Furthermore, it is robust against chip-to-chip variations. We illustrate the technique on six Trojans from Trust-Hub spanning two cryptographic chips and an embedded PIC microcontroller. Our approach detects Trojans with an accuracy ≥ 95%. It is easier to detect Trojans in an optimized-netlist circuit as more paths are close to the critical path. Even when the circuit is not optimized (i.e., when very few paths are close to the critical path), short-term aging plus mild overclocking can detect Trojans with high accuracy. Virinchi Roy Surabhi, Prashanth Krishnamurthy, Hussam Amrouch, Jörg Henkel, Ramesh Karri, Farshad Khorrami |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | A Cross-Layer Gate-Level-to-Application Co-Simulation for Design Space Exploration of Approximate Circuits in HEVC Video EncodersabstractA cross-layer design space exploration (DSE) method based on a proposed co-simulation technique is presented herein. The proposed method is demonstrated evaluating the impacts on both coding efficiency and power dissipation of applying distinct approximate logic operators in a sum of absolute differences (SAD) kernel that accelerates an H.265/HEVC (high-efficiency video coding) encoder. The proposed method simulates the gate-level circuit dynamically inside the application, with realistic results of the impact of the adder-tree approximate logic implementation on both quality and encoder bit-rate results. A comprehensive DSE is shown herein, with 13 types of 6 classes of approximate adders in the SAD accelerator hardware blocks. Over 3,000 logic variants of approximations at gate-level were developed. Actual video sequences as inputs to the x265 software encoder are co-simulated, to dynamically capture the video motion-estimation (ME) behavior in the presence of logic approximations. While the prior art that only estimates the impact of the approximate logic on power, area, and quality on static designs with statistical assumptions, which are agnostic to the actual algorithm data-dependent behavior in the application, our method explores accurately the trade-off between power dissipation and coding efficiency dynamically over the entire HEVC encoding. Our approach shows that the lower-part-or and error-tolerant adder I approximate adders, as well as truncation-to-zero deliver better compression-power trade-offs, with substantial differences from the static analysis. Guilherme Paim, Leandro M. G. Rocha, Hussam Amrouch, Eduardo A. C. da Costa, Sergio Bampi, Jörg Henkel |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Introduction to the Special Issue on Machine Learning for CADabstractNo abstract available. Jörg Henkel, Hussam Amrouch, Marilyn Wolf |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2019 | Performance, Power and Cooling Trade-Offs with NCFET-based Many-CoresabstractNegative Capacitance Field-Effect Transistor (NCFET) is an emerging technology that incorporates a ferroelectric layer within the transistor gate stack to overcome the fundamental limit of sub-threshold swing in transistors. Even though physics-based NCFET models have been recently proposed, system-level NCFET models do not exist and research is still in its infancy. In this work, we are the first to investigate the impact of NCFET on performance, energy and cooling costs in many-core processors. Our proposed methodology starts from accurate physics models all the way up to the system level, where the performance and power of a many-core are widely affected. Our new methodology and system-level models allow, for the first time, the exploration of the novel trade-offs between performance gains and power losses that NCFET now offers to system-level designers. We demonstrate that an optimal ferroelectric thickness does exist. In addition, we reveal that current state-of-the-art power management techniques fail when NCFET (with a thick ferroelectric layer) comes into play. Martin Rapp, Sami Salamin, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
DAC | 3 |
| 2019 | Hot Spot Identification and System Parameterized Thermal Modeling for Multi-Core Processors Through Infrared Thermal ImagingabstractAccurate thermal models suitable for system level dynamic thermal, power and reliability regulation and management are vital for many commercial multi-core processors. However, developing such accurate thermal models and identifying the related thermal-power relevant spatial locations for commercial processors is a challenging task due to the lack of information and available tools. Existing tools such as HotSpot-like thermal models may suffer from inaccuracy or inefficiency for online applications, primarily because most rely on parameters that cannot be precisely quantified, such as power-traces, while others are numerical methods not suitable for runtime use. In this work, we propose a novel approach to automatically detecting the major heat-sources on a commercial multi-core microprocessor using an infrared thermal imaging setup. Our approach involves a number of steps including 2D discrete cosine transformation filter for noise reduction on the measured thermal maps, and Laplacian transformation followed by K-mean clustering for heat-source identification. Since the identified heat-sources are the thermally vulnerable areas of the die, we propose a novel approach to deriving a thermal model capable of predicting their temperatures during runtime. We apply Long-Short-Term-Memory (LSTM) networks to build a dynamic thermal model which uses system-level variables such as chip frequency, voltage and instruction count as inputs. The model is trained and tested exclusively using measured thermal data from a commercial multi-core processor. Experimental results show that the proposed thermal model achieves very high accuracy (root-mean-square-error: 2.04°C to 2.57° C) in predicting the temperature of all the identified heat-sources on the chip. Sheriff Sadiqbatcha, Hengyang Zhao, Hussam Amrouch, Jörg Henkel, Sheldon X.-D. Tan |
DATE | 3 |
| 2019 | Selecting the Optimal Energy Point in Near-Threshold ComputingabstractNear-Threshold Computing (NTC) has recently emerged as an attractive paradigm as it allows devices to operate close to their optimal energy point (OEP). This work demonstrates, for the first time, that determining where the OEP of a processor exists is challenging because standard cells, forming the processor's netlist, unevenly profit w.r.t power and also unevenly degrade w.r.t delay when the voltage approaches the near-threshold region. To precisely explore, at design time, where OEP is, we create voltage-aware cell libraries that enable designers to seamlessly employ the standard tool flows, even they were not designed for that purpose, to perform voltage-aware timing and power analysis. Besides determining where the OEP is, we also demonstrate how providing logic synthesis tool flows with voltage-aware cell libraries results in a 35% higher performance at NTC. In addition, we investigate how the performance loss at NTC can be compensated through parallelized computing demonstrating, for the first time, that the OEP moves far from NTC as the number of cores increases. Our proposed methodology enables designers to select the maximum number of cores along with the optimal operating voltage jointly in which a specific power budget is fulfilled. Finally, we show how voltage-aware design for parallelized NTC provides [40%-50%] performance increase compared to traditional (i.e., voltage-unaware design) parallelized NTC. Sami Salamin, Hussam Amrouch, Jörg Henkel |
DATE | 2 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 2 |
| 2019 | Reliability Challenges with Self-Heating and Aging in FinFET TechnologyabstractThe introduction of FinFET technology as an effective solution to continue technology scaling has pushed self-heating effects to the forefront of reliability challenges, especially at the 14nm technology node and below. Due to limited silicon volume for heat dissipation, elevated temperatures across the transistors channel can be generated during operation. This results in a considerable degradation of the key properties of transistors like decreased drain and increased leakage current. In addition, excessive temperatures considerably accelerate aging phenomena in transistors such as Bias Temperature Instability (BTI) and Hot Carrier Injection (HCI), which shorten the lifetime of circuits. In this work, we discuss how self-heating effects in FinFET transistors can prolong the delay of circuits leading to reliability problems. We evaluate self-heating in an entire SRAM block consisting of SRAM cells, pre-charging circuit, sense amplifiers and an output latch. When it comes to reliability and lifetime, we demonstrate how self-heating effects can result in larger aging-induced degradations which, in turn, enforce designers to include wider and wider safety margins to sustain reliability. Lastly, we provide an outlook of self-heating and reliability concerns in Negative Capacitance Field Effect Transistors (NCFET). Hussam Amrouch, Victor M. van Santen, Om Prakash 0007, Hammam Kattan, Sami Salamin, Simon Thomann, Jörg Henkel |
IOLTS | 1 |
| 2019 | Aging Gracefully with ApproximationabstractThis paper presents a design methodology to turn aging-induced chip slowdown into approximation without adding reliability guardband or increasing supply voltage. It guarantees always-best quality while the system is under aging. It is based on run-time monitoring of critical path delay. If the delay increases due to aging, the proposed approach curtails the critical path at the cost of precision reduction. We evaluate our approach at the component level as well as microarchitecture level. The evaluation results show that the approach reduces the dynamic and static power consumptions by 19.8% and 10.2%, respectively, with minimal area overhead and quality degradation. Heesu Kim, Hussam Amrouch, Jörg Henkel, Andreas Gerstlauer, Kiyoung Choi |
ISCAS | 3 |
| 2019 | NCFET-Aware Voltage ScalingabstractNegative Capacitance Field-Effect Transistor (NCFET) has recently attracted significant attention. In the NCFET technology with a thick ferroelectric layer, voltage reduction increases the leakage power, rather than decreases, due to the negative Drain-Induced Barrier Lowering (DIBL) effect. This work is the first to demonstrate the far-reaching consequences of such an inverse dependency w.r.t. the existing power management techniques. Moreover, this work is the first to demonstrate that state-of-the-art Dynamic Voltage Scaling (DVS) techniques are sub-optimal for NCFET. Our investigation revealed that the optimal voltage at which the total power is minimized is not necessarily at the point of the minimum voltage required to fulfill the performance constraint (as in traditional DVS). Hence, an NCFET-aware DVS is key for high energy efficiency. In this work, we therefore propose the first NCFET-aware DVS technique that selects the optimal voltage to minimize the power following the dynamics of workloads. Our experimental results of a multi-core system demonstrate that NCFET-aware DVS results in 20% on average, and up to 27% energy saving while still fulfilling the same performance constraint (i.e., no trade-offs) compared to traditional NCFET-unaware DVS techniques. Sami Salamin, Martin Rapp, Hussam Amrouch, Girish Pahwa, Yogesh Singh Chauhan, Jörg Henkel |
ISLPED | 3 |
| 2019 | On the Efficiency of Voltage Overscaling under Temperature and Aging EffectsabstractVoltage overscaling has received extensive attention in the last decade as an attractive paradigm for systems in which resulting timing errors and thus a loss in accuracy can be accepted in exchange for an increase in energy efficiency. At the same time, the delay of a circuit is, in turn, and in addition to voltage, also subject to temperature and aging. Existing work has largely studied voltage overscaling in isolation. This ignores interdependencies with temperature and aging, which can lead to wrong or misleading conclusions. In this work, we are the first to model the combined impact of voltage, temperature and aging on the delay of circuits towards investigating the actual existing trade-offs between efficiency and accuracy provided by voltage overscaling. We show that analyzing voltage in isolation overestimates timing errors and thus underestimates the voltage scaling potential. We further develop an approach that leverages interdependencies to optimize energy, delay and accuracy trade-offs. We precisely translate the individual and combined impact of voltage-, temperature-, and aging-induced delay increase into corresponding probability of error (Perror). This reveals that the same amount of timing increase results in different error probabilities depending on the origin (i.e., voltage, temperature or aging). For the same timing increase, voltage reductions result in the smallest-Perror compared to temperature or aging, while also reducing temperature- and aging-induced delay increases themselves. This allows voltage reduction to be employed as an effective means to minimize delay, reduce energy and thus maximize efficiency under a given upper bound on error probability. We apply our approach to multipliers in GPUs exploring the trade-off between efficiency and accuracy. We demonstrate how only accounting for voltage scaling alone leads to a considerably larger Perror (74% on average) than in reality. Our investigation also shows that for the same Perror constraint, optimizing for combined voltage, temperature and aging effects results, on average, in 116% better energydelay product (EDP) compared to state of the art. Hussam Amrouch, Seyed Borna Ehsani, Andreas Gerstlauer, Jörg Henkel |
IEEE Trans. Computers | 1 |
| 2019 | Dynamic Guardband Selection: Thermal-Aware Optimization for Unreliable Multi-Core SystemsabstractCircuit aging has become the major reliability concern in current and upcoming technology nodes. For instance, Bias Temperature Instability (BTI) leads to an increase in the threshold voltage of a transistor. That, in turn, may prolong the critical path delay of the processor and eventually may lead to timing errors. In order to avoid aging-induced timing errors, designers employ guardbands either with respect to voltage or frequency. State-of-the-art techniques determine a guardband type at the circuit level at design time irrespective from the running workload at the system level. Our investigation revealed that generated temperatures by a running workload have the potential to play a key role in determining the appropriate guardband type with respect to system performance. Therefore, we propose a paradigm shift in designing guardbands: to select the guardband types on-the-fly with respect to the workload-induced temperatures aiming at optimizing for performance under temperature and reliability constraints. Moreover, different guardband types for different cores can be selected simultaneously when multiple applications with diverse properties suggest this to be useful. Our dynamic guardband selection allows for a higher performance compared to techniques that employ a fixed (at design time) guardband type throughout. Heba Khdr, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Computers | 2 |
| 2019 | Estimating and Mitigating Aging Effects in Routing Network of FPGAsabstractIn this paper, we present a comprehensive analysis of the impact of aging on the interconnection network of field-programmable gate arrays (FPGAs) and propose novel approaches to mitigate the aging effects on the routing network. We first show the insignificant impact of aging on data integrity of FPGAs, i.e., static noise margin and soft error rate of the configuration cells, as well as we show the negligible impact of the mentioned degradations on the FPGA performance. As such, we focus on the performance degradation of datapath transistors. In this regard, we propose a routing accompanied by a placement algorithm that prevents constant stress on transistors by evenly distributing the stress through the interconnection resources. By observing the impact of the signal probability on the aging of routing buffers, we enhance the synthesis flow as well as augment the proposed routing algorithm to converge the signal probabilities toward aging-friendly values. Experimental results over a set of industrial benchmarks and commerciallike FPGA architecture indicate the effectiveness of the proposed method with 64.3% reduction of stress duration in multiplexers and up to 45.2% improvement of the degradation of buffers. Altogether, the proposed method reduces the timing guardband by from 14.1% to 31.7%, depending on the FPGA routing architecture. Behnam Khaleghi, Behzad Omidi, Hussam Amrouch, Jörg Henkel, Hossein Asadi 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Modeling the Interdependences Between Voltage Fluctuation and BTI AgingabstractWith technology scaling, the susceptibility of circuits to different reliability degradations is steadily increasing. Aging in transistors due to bias temperature instability (BTI) and voltage fluctuation in the power delivery network of circuits due to IR-drops are the most prominent. In this paper, we are reporting for the first time that there are interdependences between voltage fluctuation and BTI aging that are nonnegligible. Modeling and investigating the joint impact of voltage fluctuation and BTI aging on the delay of circuits, while remaining compatible with the existing standard design flow, is indispensable in order to answer the vital question, “what is an efficient (i.e., small, yet sufficient) timing guardband to sustain the reliability of circuit for the projected lifetime?” This is, concisely, the key goal of this paper. Achieving that would not be possible without employing a physics-based BTI model that precisely describes the underlying generation and recovery mechanisms of defects under arbitrary stress waveforms. For this purpose, our model is validated against varied semiconductor measurements covering a wide range of voltage, temperature, frequency, and duty cycle conditions. To bring reliability awareness to existing EDA tool flows, we create standard cell libraries that contain the delay information of cells under the joint impact of aging and IR-drop. Our libraries can be directly deployed within the standard design flow because they are compatible with existing commercial tools (e.g., Synopsys and Cadence). Hence, designers can leverage the mature algorithms of these tools to accurately estimate the required timing guardbands for any circuit despite its complexity. Our investigation demonstrates that considering aging and IR-drop effects independently, as done in the state of the art, leads to employing insufficient and thus unreliable guardbands because of the nonnegligible (on average 15% and up to 25%) underestimations. Importantly, considering interdependences between aging and IR-drop does not only allow correct guardband estimations, but it also results in employing more efficient guardbands. Sami Salamin, Victor M. van Santen, Hussam Amrouch, Narendra Parihar, Souvik Mahapatra, Jörg Henkel |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Aging-constrained performance optimization for multi coresabstractCircuit aging has become a dire design concern and hence it is considered a primary design constraint. Current practice to cope with this problem is to apply (too) conservative means. Heba Khdr, Hussam Amrouch, Jörg Henkel |
DAC | 2 |
| 2018 | Estimating and optimizing BTI aging effects: from physics to CADabstractTransistor aging due to Bias Temperature Instability (BTI) is a crucial degradation that affects the reliability of circuits over time. Aging-aware circuit design flows do virtually not exist yet and even research is in its infancy. In this work, we demonstrate how the deleterious effects BTI-induced degradations can be modeled from physics, where they do occur, all the way up to the system level, where they finally take place and affect the delay and power of circuits. To achieve that, degradation-aware cell libraries, that properly capture the impact of BTI not only on the delay of standard cells but also on their static and dynamic power, are created. Unlike state of the art, which solely models the impact of BTI on the threshold voltage of transistors $(V_{th})$ , we are the first to model the other key transistor parameters degraded by BTI like carrier mobility ( $\mu$ ), sub-threshold slope ( $SS$ ), and gate-drain capacitance $(C_{gd})$ . Our cell libraries are compatible with existing commercial CAD tools. Employing the mature algorithms in such tools, enables designers – after importing our cell libraries – to accurately estimate the overall impact of aging on changing the delay and/or power of any circuit, despite its complexity. We demonstrate that $\Delta V_{th}$ alone (as done in state of the art) is insufficient to correctly model the impact of BTI either on delay or power of circuits. On the one hand, neglecting BTl-induced $\mu$ and $C_{gd}$ degradations leads to underestimating the impact that BTI has on increasing the delay of circuits. Hence, designers will employ narrower timing guardbands in which reliability of circuits during lifetime cannot be sustained. On the other hand, neglecting BTI-induced $SS$ degradation leads to overestimating the impact that BTI has on static power reduction. Hence, the potential benefit of circuits from BTI will be exaggerated. Hussam Amrouch, Victor M. van Santen, Jörg Henkel |
ICCAD | 1 |
| 2018 | Dynamic resource management for heterogeneous many-coresabstractWith the advent of many-core systems, use cases of embedded systems have become more dynamic: Plenty of applications are concurrently executed, but may dynamically be exchanged and modified even after deployment. Moreover, resources may temporally or permanently become unavailable because of thermal aspects, dynamic power management, or the occurrence of faults. This poses new challenges for reaching objectives like timeliness for real-time or performance for best-effort program execution and maximizing system utilization. In this work, we first focus on dynamic management schemes for reliability/aging optimization under thermal constraints. The reliability of on-chip systems in the current and upcoming technology nodes is continuously degrading with every new generation because transistor scaling is approaching its fundamental limits. Protecting systems against degradation effects such as circuits' aging comes with considerable losses in efficiency. We demonstrate in this work why sustaining reliability while maximizing the utilization of available resources and hence avoiding efficiency loss is quite challenging – this holds even more when thermal constraints come into play. Then, we discuss techniques for run-time management of multiple applications which sustain real-time properties. Our solution relies on hybrid application mapping denoting the combination of design-time analysis with run-time application mapping. We present a method for Real-time Mapping Reconfiguration (RMR) which enables the Run-Time Manager (RM) to execute realtime applications even in the presence of dynamic thermal- and reliability-aware resource management. This paper is paper of the ICCAD 2018 Special Session on “Managing Heterogeneous Many-cores for High-Performance and Energy-Efficiency”. The other two papers of this Special sessions are [1] and [2]. Jörg Henkel, Jürgen Teich, Stefan Wildermann, Hussam Amrouch |
ICCAD | 4 |
| 2018 | Trading Off Temperature Guardbands via Adaptive ApproximationsabstractRuntime circuit delay variations due to degradation effects like temperature are traditionally protected against using worst-case timing guardbands. Such an approach leads to a permanent performance overhead even though effects may only be transient. Recently, approximate computing has been proposed as a technique to trade off quality for various metrics. Existing approaches, however, do not target reductions in circuit delays and guardbands, or have only been applied statically. In this paper, we propose a novel design paradigm in which adaptive approximations are employed to dynamically trade off transient, degradation-induced variations in circuit delays and associated worst-case timing guardbands for permanent performance improvements with minimal quality loss. A key challenge is to design circuits that exhibit a significant delay profile across approximation levels while maintaining a high base performance. To achieve that, we introduce and implement two approaches for synthesizing arbitrary dynamic quality-versus delay-configurable circuits at fine temporal and spatial granularities while exploring associated area, speed and quality trade-offs. We apply our approach specifically to temperature variations and guardbands. Results for an IDCT image decoding example show up to 21% speedup with less than 2% area and energy impact compared to traditional guardbanding while maintaining a worst-case transient PSNR of at least 39dB. Behzad Boroujerdian, Hussam Amrouch, Jörg Henkel, Andreas Gerstlauer |
ICCD | 2 |
| 2018 | Reliability Estimations of Large Circuits in Massively-Parallel GPU-SPICEabstractSPICE simulations for reliability have special requirements. We present GPU-SPICE to serve these special requirements. First, our GPU-SPICE employs the massive parallelism found in GPUs to enable circuit simulations beyond 200K transistors. This is necessary to study reliability in microarchitecture components (e.g., multipliers, adders), as reliability estimations require full analogue SPICE simulations (instead of STA or other heuristics). Secondly, our GPU-SPICE can update transistor parameters during the circuit simulation, a feature necessary to model reliability degradation, which constantly reacts to circuit activity (e.g., Bias Temperature Instability reacting to Vgschanges by increasing/decreasing ΔVthin each transistor). Lastly, our GPU-SPICE is open-source software, this ensures that it easily can be employed, adapted and extended by other researchers. Due to the massive parallelism in a GPU and performance optimizations (convergence criteria, CUDA memory management, etc.), our GPU-SPICE is up to 218x faster than its single-threaded baseline NGSPICE. Victor M. van Santen, Hussam Amrouch, Jörg Henkel |
IOLTS | 2 |
| 2018 | Recent advances in EM and BTI induced reliability modeling, analysis and optimization (invited)
Sheldon X.-D. Tan, Hussam Amrouch, Taeyoung Kim 0001, Zeyu Sun 0001, Chase Cook, Jörg Henkel |
Integr. | 2 |
| 2018 | Aging-Aware BoostingabstractDVFS-based boosting techniques have been widely employed by commercial multi-core processors, due to their superiority in improving the performance. Boosting, however, is particularly stressing circuits and hence it significantly contributes to an accelerated aging process. Circuit aging has become a real reliability concern because it leads to an increase in transistor threshold voltage that may cause timing errors as a result of higher delays in critical paths. Thus, high performance is desirable but it shortens the circuit lifetime through aging leaving a choice to trade-off. Besides well-known long-term aging effects, recent research also reported short-term aging effects. Our claim is that DVFS-based boosting techniques should consider both long- and short-term aging effects. This can be circumvented by wider timing guardbands. But that would be more expensive. The goal of this work is therefore to analyze and optimize boosting under specific consideration of long-term and short-term aging effects. As a result of our findings, we propose the first comprehensive aging-aware, yet efficient boosting technique. The employed aging-aware cell libraries in this work are publicly available at http://ces.itec.kit.edu/dependable-hardware.php. Heba Khdr, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Computers | 2 |
| 2017 | Containing guardbandsabstractReliability concerns may overtake conventional design constraints such as cost and performance because transistors in deep nano-CMOS era are increasingly susceptible to degradation effects. This made reliability become unsustainably expensive due to need for wider and wider guardbands (i.e. safety margins). It is in fact the time to reverse this trend: instead of widening guardbands, it is inevitable to contain them. In this work, we summarize three novel means to achieve this goal. Since the causes are of physical origin, it cannot be excluded that degradation effects influence (i.e. amplify or cancel) each other. Hence, we first investigate the interdependencies of degradation effects demonstrating that they should jointly and not separately be modeled towards designing smaller, yet sufficient guardbands. Then, we show how aging-aware logic synthesis based on our so-called degradation-aware cell libraries, enables designers to employ mature optimization algorithms available in the commercial synthesis tools to obtain more resilient circuits in which guardbands are inherently contained. Finally, instantaneous transistors aging is a recent discovery that bears a large potential for reliability optimization since it is hardly explored until now. Though aging in general has been extensively studied in last decade, investigating the impact of instantaneous aging on circuits' reliability is still in its infancy. In fact, this is a paradigm shift in aging from sole long-term reliability degradation, as in the traditional view, to short-term reliability degradation. We demonstrate how employing our physics-based aging models results in considerably smaller guardbands due to the high certainty compared to empirical aging models. Hussam Amrouch, Jörg Henkel |
ASP-DAC | 1 |
| 2017 | Emerging (un-)reliability based security threats and mitigations for embedded systems: special sessionabstractThis paper addresses two reliability-based security threats and mitigations for embedded systems namely, aging and thermal side channels. Device aging can be used as a hardware attack vector by using voltage scaling or specially crafted instruction sequences to violate embedded processor guard bands. Short-term aging effects can be utilized to cause transient degradation of the embedded device without leaving any trace of the attack. (Thermal) side channels can be used as an attack vector and as a defense. Specifically, thermal side channels are an effective and secure way to remotely monitor code execution on an embedded processor and/or to possibly leak information. Although various algorithmic means to detect anomaly are available, machine learning tools are effective for anomaly detection. We will show such utilization of deep learning networks in conjunction with thermal side channels to detect code injection/modification representing anomaly. Hussam Amrouch, Prashanth Krishnamurthy, Naman Patel, Jörg Henkel, Ramesh Karri, Farshad Khorrami |
CASES | 1 |
| 2017 | Towards Aging-Induced ApproximationsabstractIn recent technology nodes, wide guardbands are needed to overcome reliability degradations due to aging. Such guardbands manifest as reduced efficiency and performance. Existing approaches to reduce guardbands trade off aging impact for increased circuit overhead. By contrast, the goal of this work is to completely remove guardbands through exploring, for the first time, application of approximate computing principles in the context of aging. As a result of naively narrowing or removing guardbands, timing errors start to appear as transistors age. We demonstrate that even in circuits that may tolerate errors, aging can be catastrophic due to unacceptable quality loss. Furthermore, quantifying such aging-induced quality loss necessitates expensive (often infeasible) gate-level simulations of the complete design. We show how nondeterministic aging-induced timing errors can be converted into deterministic and controlled approximations instead. We first translate the required guardband over time into an equivalent reduction in precision for individual RTL components. We then demonstrate how, based on pre-characterization of RTL components, we can quantify aging-induced approximation at the whole microarchitecture level without the need for further gate-level simulations. Results show that a 3 bit reduction in precision is sufficient to sustain 10 years of operation under worst-case aging in the context of an image processing circuit. This corresponds to an acceptable PSNR reduction of merely 8 dB, while at the same time increasing area and energy efficiency by 13%. Hussam Amrouch, Behnam Khaleghi, Andreas Gerstlauer, Jörg Henkel |
DAC | 1 |
| 2017 | Optimizing temperature guardbandsabstractWe introduce the first temperature guardbands optimization based on thermal-aware logic synthesis and thermal-aware timing analysis. The optimized guardbands are obtained solely due to using our so-called thermal-aware cell libraries together with existing tool flows and not due to sacrificing timing constraints (i.e. no trade-offs). We demonstrate that temperature guardbands can be optimized at design time through thermal-aware logic synthesis in which more resilient circuits against worst-case temperatures are obtained. Our static guardband optimization leads to 18% smaller guardbands on average. We also demonstrate that thermal-aware timing analysis enables designers to accurately estimate the required guardbands for a wide range of temperatures without over/under-estimations. Therefore, temperature guardbands can be optimized at operation time through employing the small, yet sufficient guardband that corresponds to the current temperature rather than employing throughout a conservative guardband that corresponds to the worst-case temperature. Our adaptive guardband optimization results, on average, in a 22% higher performance along with 9 2% less energy. Neither thermal-aware logic synthesis nor thermal-aware timing analysis would be possible without our thermal-aware cell libraries. They are compatible with use of existing commercial tools. Hence, they allow designers, for the first time, to automatically consider thermal concerns within their design tool flows even if they were not designed for that purpose. Hussam Amrouch, Behnam Khaleghi, Jörg Henkel |
DATE | 1 |
| 2017 | Ultra-low power and dependability for IoT devices (Invited paper for IoT technologies)abstractRecent advances in technologies have allowed the design of small-size low-power and low-cost devices that can be connected to the Internet, enabling the emerging paradigm of Internet-of-things (IoT). IoT covers an ever-increasing range of applications, e.g., health-care monitoring, smart homes and buildings, etc. In this invited paper, we discuss and summarize the IoT paradigm with a special focus on energy consumption and methodologies for its minimization. Furthermore, we also discuss about reliability in the context of IoT devices. In all, this paper attempts to be a starting point for readers interested in developing energy-efficient IoT devices. Jörg Henkel, Santiago Pagani, Hussam Amrouch, Lars Bauer, Farzad Samie |
DATE | 3 |
| 2016 | Power and thermal management in massive multicore chips: theoretical foundation meets architectural innovation and resource allocationabstractContinuing progress and integration levels in silicon technologies make possible complete end-user systems consisting of extremely high number of cores on a single chip targeting either embedded or high-performance computing. However, without new paradigms of energy- and thermally-efficient designs, producing information and communication systems capable of meeting the computing, storage and communication demands of the emerging applications will be unlikely. The broad topic of power and thermal management of massive multicore chips is actively being pursued by a number of researchers worldwide, from a variety of different perspectives, ranging from workload modeling to efficient on-chip network infrastructure design to resource allocation. Successful solutions will likely adopt and encompass elements from all or at least several levels of abstraction. Starting from these ideas, we consider a holistic approach in establishing the Power-Thermal-Performance (PTP) trade-offs of massive multicore processors by considering three inter-related but varying angles, viz., on-chip traffic modeling, novel Networks-on-Chip (NoC) architecture and resource allocation/mapping Paul Bogdan, Partha Pratim Pande, Hussam Amrouch, Muhammad Shafique 0001, Jörg Henkel |
CASES | 3 |
| 2016 | Reliability-aware design to suppress agingabstractDue to aging, circuit reliability has become extraordinary challenging. Reliability-aware circuit design flows do virtually not exist and even research is in its infancy. In this paper, we propose to bring aging awareness to EDA tool flows based on so-called degradation-aware cell libraries. These libraries include detailed delay information of gates/cells under the impact that aging has on both threshold voltage (Vth) and carrier mobility (μ) of transistors. This is unlike state of the art which considers Vth only. We show how ignoring μ degradation leads to underestimating guard-bands by 19% on average. Our investigation revealed that the impact of aging is strongly dependent on the operating conditions of gates (i.e. input signal slew and output load capacitance), and not solely on the duty cycle of transistors. Neglecting this fact results in employing insufficient guard-bands and thus not sustaining reliability during lifetime. Hussam Amrouch, Behnam Khaleghi, Andreas Gerstlauer, Jörg Henkel |
DAC | 1 |
| 2016 | Improving mobile gaming performance through cooperative CPU-GPU thermal managementabstractState-of-the-art thermal management techniques independently throttle the frequencies of high-performance multi-core CPU and powerful graphics processing units (GPU) on heterogeneous multiprocessor system-on-chips deployed in latest mobile devices. For graphics-intensive gaming applications, this approach is inadequate because both the CPU and the GPU contribute towards the overall application performance (frames per second or FPS) as well as the on-chip temperature. The lack of coordination between CPU and GPU induces recurrent frequency throttling to maintain on-chip temperature below the permissible limit. This leads to significantly degraded application performance and large variation in temperature over time. We propose a control-theory based dynamic thermal management technique that cooperatively scales CPU and GPU frequencies to meet the thermal constraint while achieving high performance for mobile gaming. Experimental results with six popular Android games on a commercial mobile platform show an average 19% performance improvement and over 90% reduction in temperature variance compared to the original Linux approach. Alok Prakash, Hussam Amrouch, Muhammad Shafique 0001, Tulika Mitra, Jörg Henkel |
DAC | 2 |
| 2016 | Designing guardbands for instantaneous aging effectsabstractBias Temperature Instability (BTI) is one of the key causes of reliability degradations of nano-CMOS circuits. While the long-term impact of BTI has been studied since years, the short-term implications of BTI on circuits are unexplored. In fact, in physics short-term BTI effects, i.e. instantaneous (i.e. sub μs) frequency dependent processes, have been recently reported. In order to design circuits with guardbands that are safe for long-term and instantaneous effects, new aging models are required. We are presenting the first approach that in fact considers both long-term as well as instantaneous BTI effects. It can be employed for complex circuits at the micro-architecture level. Designing guardbands based upon our physical BTI model reduces the guardbands by 41% and thus allows for the development of more cost-effective yet reliable designs. We also revisit existing state-of-the-art aging mitigation techniques to investigate how they can be properly adapted to additionally account for instantaneous aging effects. Along with our BTI model this further reduces the guardbands by up to 59%. Victor M. van Santen, Hussam Amrouch, Javier Martín-Martínez, Montserrat Nafría, Jörg Henkel |
DAC | 2 |
| 2016 | Aging-aware voltage scaling
Victor M. van Santen, Hussam Amrouch, Narendra Parihar, Souvik Mahapatra, Jörg Henkel |
DATE | 2 |
| 2016 | Stress-aware routing to mitigate aging effects in SRAM-based FPGAsabstractContinuous shrinking of transistor size to provide high computation capability along with low power consumption has been accompanied by reliability degradations due to e.g., aging phenomenon. In this regard, with huge number of configuration bits, Field-Programmable Gate Arrays (FPGAs) are more susceptible to aging since aging not only degrades the performance, it may additionally result in corrupting the configuration cells and thus causing permanent circuit malfunctioning. While several works have investigated the aging effects in Look-Up Tables (LUTs), the routing fabric of these devices is seldom studied - even though it contributes to the majority of FPGAs' resources and configuration bits. Furthermore, there is a high prospect that errors in its state to propagate to the device outputs. In this paper, we first investigate aging effects in the routing fabric of FPGAs with respect to performance and reliability degradations. Based on this investigation, we enhance the conventional routing algorithm to mitigate the impact of aging by increasing the recovery time (i.e., the mechanism used to heal aging-induced defects) of transistors used in the routing resources. We examine our proposed method as reduction in stress time and required guardband to protect against aging in the routing fabric, as well as in improving the FPGA's lifetime. Our experiments show that the proposed method reduces the average stress time and aging-induced delay of routing resources by 41% and 18.3%, respectively. This, in turn, leads to improving the device lifetime by 130% compared to baseline routing. The proposed method can be applied by simple amending of conventional routing algorithms. Thus, it incurs negligible delay overhead. Behnam Khaleghi, Behzad Omidi, Hussam Amrouch, Jörg Henkel, Hossein Asadi 0001 |
FPL | 3 |
| 2015 | Lucid infrared thermography of thermally-constrained processorsabstractThermal analysis is a prerequisite for developing reliability increasing techniques for thermally-constrained processors, i.e. processors with a high power density. For that purpose, infrared (IR) camera measurement setups have been deployed with the purpose to provide direct feedback of the impact that thermal mitigation techniques have. To obtain lucid IR images1, the IR-opaque cooling must be removed and hence, an alternative IR-transparent cooling needs to be provided to protect the chip. To this end, the majority of state-of-the-art employs an IR coolant liquid to prevent the chip from overheating. The problem is that several aspects like thermal convection may interfere with the measured IR radiations resulting in equivocal IR images. Thus, they decrease the accuracy in a way that leads to incorrectly estimating reliability. Solving this prominent problem, we introduce an IR-transparent cooling that cools the chip from its rear side allowing the camera to perspicuously capture the IR emissions as no additional layer in between impedes the radiation. It maintains the on-chip temperatures within a safe range equivalent to the original heat sink-based cooling. We demonstrate how state-of-the-art inaccurate thermal analysis results in incorrectly estimating reliability. Our setup is the most accurate, least intrusive one that has been both proposed and actually applied to state-of-the-art multi-cores (Intel 45nm dual-core and 22nm octa-core). Hussam Amrouch, Jörg Henkel |
ISLPED | 1 |
| 2014 | mDTM: Multi-objective dynamic thermal management for on-chip systemsabstractThermal hot spots and unbalanced temperatures between cores on chip can cause either degradation in performance or may have a severe impact on reliability, or both. In this paper, we propose mDTM, a proactive dynamic thermal management technique for on-chip systems. It employs multi-objective management for migrating tasks in order to both prevent the system from hitting an undesirable thermal threshold and to balance the temperatures between the cores. Our evaluation on the Intel SCC platform shows that mDTM can successfully avoid a given thermal threshold and reduce spatial thermal variation by 22%. Compared to state-of-the-art, our mDTM achieves up to 58% performance gain. Additionally, we deploy an FPGA and IR camera based setup to analyze the effectiveness of our technique. Heba Khdr, Thomas Ebi, Muhammad Shafique 0001, Hussam Amrouch, Jörg Henkel |
DATE | 4 |
| 2014 | hevcDTM: Application-driven Dynamic Thermal Management for High Efficiency Video CodingabstractThis paper presents an application-driven algorithm for Dynamic Thermal Management (DTM) for the High Efficiency Video Coding (HEVC). For efficient design of such a DTM policy, we perform an offline thermal analysis of an HEVC encoder and demonstrate the impact of different video sequences and different coding configurations on the processor temperature. Our thermal analysis is leveraged to develop an efficient application-driven DTM policy that performs temperature-aware coding along with an application-driven control of DTM knobs (e.g., frequency scaling) in order to meet the temperature constraints while still providing high video quality (i.e. PSNR loss <; 0.01dB). For accurate thermal analysis and evaluation, we deploy an infrared camera-based thermal measurement setup that, on the contrary to state-of-the-art setups, does not require adding any extra layer on top of the measured chip, thus allowing the camera to accurately capture the infrared emissions from the die. Daniel Palomino 0001, Muhammad Shafique 0001, Hussam Amrouch, Altamiro Amadeu Susin, Jörg Henkel |
DATE | 3 |
| 2014 | Towards interdependencies of aging mechanisms
Hussam Amrouch, Victor M. van Santen, Thomas Ebi, Volker Wenzel, Jörg Henkel |
ICCAD | 1 |
| 2014 | RESI: Register-Embedded Self-Immunity for Reliability EnhancementabstractTechnology scaling in the nano-CMOS era has reached a point where coping with the failures produced by soft errors has become one of the key challenges when it comes to reliability. Akin to the fact that a register file is accessed more frequently than any other architectural component, register file protection is imperative to obstruct errors from propagating throughout a computing system. Furthermore, negative bias temperature instability (NBTI) has emerged as a major concern due to its negative impact on the lifetime of pMOS devices. Indeed, many of the pMOS transistors most affected by NBTI are in the register files as they are implemented as SRAM, which are particularly vulnerable due to their small structure size. Based on our observation that some register bits are not continuously used to represent a value stored in a register, we present a technique that exploits unused bits to improve the register file immunity against soft errors and mitigate NBTI effects. We show that our technique can reduce, on average, the register file vulnerability against multiple bit upsets by 97% (up to 100%), resulting in a high system fault coverage under various scenarios, while consuming less power and still occupying a similar area footprint compared to protecting the register file against single bit upsets (SBUs) only. The achieved result is 63% better compared to the state-of-the-art in register file protection. To compare and quantify the effect of our technique, we observe its impact on the processor's temperature using an infrared thermal camera and show that, due to consuming less power per area, our technique also operates at a lower temperature compared to protecting the register file against SBUs only. Finally, we investigate how our technique additionally moderates the stress induced by NBTI in register file SRAM cells. Hussam Amrouch, Thomas Ebi, Jörg Henkel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Thermal management for dependable on-chip systemsabstractDependability has become a growing concern in the nano-CMOS era due to elevated temperatures and an increased susceptibility to temperature of the small structures. We present an overview of temperature-related effects that threaten dependability and a methodology for reducing the dependability concerns through thermal management utilizing the concept of aging budgeting. Jörg Henkel, Thomas Ebi, Hussam Amrouch, Heba Khdr |
ASP-DAC | 3 |
| 2013 | Stress balancing to mitigate NBTI effects in register filesabstractNegative Bias Temperature Instability (NBTI) is considered one of the major reliability concerns of transistors in current and upcoming technology nodes and a main cause of their diminished lifetime. We propose a new means to mitigate the effects of NBTI on SRAM-based register files, which are particularly vulnerable due to their small structure size and are under continuous voltage stress for prolonged intervals. The conducted results from our technology simulator demonstrate the severity of NBTI effects on the SRAM cells - especially when process variation is taken into account. Based on the presented analysis, we show that NBTI stress in different registers needs to be tackled using different strategies corresponding to their access patterns. To this end, we propose to selectively increase the resilience of individual registers against NBTI. Our technique balances the gate voltage stress of the two PMOS transistors of an SRAM cell such that both are under stress for approximately the same amount of time during operation - thereby minimizing the deleterious effects of NBTI. We present mitigation implementations in both hardware and in software along with the incurred overhead. Through a wide range of applications we can show that our technique reduces the NBTI-induced reliability degradation by 35% on average. This is 22% better than current State-of-the-Art. Hussam Amrouch, Thomas Ebi, Jörg Henkel |
DSN | 1 |
| 2013 | Accurate Thermal-Profile Estimation and Validation for FPGA-Mapped CircuitsabstractAccurate thermal profile estimation for FPGA, at design time, is necessary to avoid unexpected thermal hot-spots in the circuit before deploying the FPGA to the in-field operation. Both accurate dynamic and leakage power values are needed for the thermal profile estimation and they can be estimated using the FPGA vendor's tools. However these report leakage power as a single value for the whole chip, and no details are given in literature or the FPGA toolset about its distribution across the FPGA chip for the thermal simulation. To cope with this problem, we present a method for properly distributing the leakage power across the FPGA chip. The method uses a temperature-leakage loop estimation model for distributing and adapting the leakage power for more accurate thermal simulation. Furthermore, to accurately calibrate the presented method and its model and also to validate the resulting thermal profiles, we utilize an infrared thermal camera, which measures the emissions from the backside of a Virtex-5 FPGA chip. The results of testing several designs, with different sizes and frequencies, show that our approach can achieve accurate thermal-profile estimation when compared to the camera measurements, with average absolute estimation error of around 1°C across the chip. Abdulazim Amouri, Hussam Amrouch, Thomas Ebi, Jörg Henkel, Mehdi Baradaran Tahoori |
FCCM | 2 |
| 2013 | Analyzing the thermal hotspots in FPGA-based embedded systemsabstractThe rapid push towards the minimization of the feature sizes of the process technology nodes in the nano-CMOS era has significantly increased the power densities and made State-of-the-Art FPGAs vulnerable to diverse problems induced by excessive temperatures. As a result, there is a prominent need to accurately study the FPGA thermal characteristics. Our experimental setup employs a thermal camera that captures the infrared emissions from the silicon wafer of an FPGA die allowing us to evaluate the accuracy of the methods conventionally used for thermal analysis, such as thermal simulations. Based on our observation that the memory interface is the thermal hotspot of an FPGA-based embedded system, we demonstrate that the cache plays a dominant thermal role by reducing memory accesses; we carefully examine the influence of various cache parameters on the FPGA temperature and propose a model linking these quantities, with an average maximum error of only 0.64°C. Hussam Amrouch, Thomas Ebi, Josef Schneider, Sri Parameswaran, Jörg Henkel |
FPL | 1 |