VLDB 2026 Research / reviewers in the wild / expert
Florian Klemme
dblp:273/6524
· DBLP profile ↗
26ranked-venue papers
8as first author
24since 2021 · last 2026
0000-0002-0148-0523ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 8 first-author · 23 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Adaptive DLBIST for Delay Fault Testing: Minimizing PVT Variability with Zero Temperature Coefficient (ZTC) VoltageabstractAbstract Safety-critical automotive systems require on-chip testing methods to ensure high fault coverage and reliable operation. Periodic Deterministic Logic Built-In Self-Test (DLBIST) is often used to meet these demands. For automotive applications, DLBIST must operate reliably despite temperature variations, including those caused by ambient changes and self-heating in FinFET transistors. A test set effective for all temperatures typically requires a large volume, which can make DLBIST impractical. This paper proposes a robust DLBIST scheme which applies multiple voltages during power-on and power-off tests and the optimal or adapted voltage during periodic tests in system operation. If distributed sensors for on-chip temperature are available for DVFS control, they can be exploited for an adaptive DLBIST scheme. During the periodic test phase, the BIST Control Unit (BCU) dynamically selects and applies the pre-generated test set corresponding to the current operating voltage and measured temperature. This adaptive selection ensures that testing conditions precisely match the real operating points. If temperature sensors are not available, testing at the so-called Zero Temperature Coefficient (ZTC) voltage is one alternative, which is the voltage where the temperature-induced variability is minimized. This makes periodic DLBIST a feasible solution for in-field self-testing, even in cases where on-chip temperature sensors are not available. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Electron. Test. | 2 |
| 2025 | Benchmarking Cryogenic Circuits using 5 nm FinFETs for Quantum ProcessingabstractQuantum computing offers the potential to solve problems that are intractable for classical computers. A major challenge in scaling quantum computers lies in bridging the gap between cryogenic qubits, operating at millikelvin to few kelvin temperatures, and the classical CMOS-based system-on-chip (SoC) typically located at room temperature (300K). This connection introduces heat leakage, which can destabilize the qubit states. A promising solution is to relocate the control circuits and processors to the cryogenic environment, but this imposes strict constraints on power consumption due to limited cooling capacity. Additionally, the SoC must meet stringent timing requirements for qubit measurement classification. In this work, we investigate the performance of CMOS-based circuits for cryogenic operations using 5 nm FinFET technology. We begin by measuring the electrical characteristics of advanced 5 nm FinFETs at both 10K and 300K. Using these measured data, we calibrate the industry-standard compact model (BSIM-CMG) and develop two standard cell libraries for each temperature. Through the logic synthesis of six circuits from the EPFL benchmark suite, we analyze their behavior at cryogenic temperatures. Our results show that circuits at 10K achieve a 41% increase in speed compared to 300K. Further, they operate efficiently at lower supply voltages, which enables reduced power consumption while maintaining high-speed performance in cryogenic environments. Anirban Kar, Shivendra Singh Parihar, Florian Klemme, Yogesh Singh Chauhan, Hussam Amrouch |
ISCAS | 3 |
| 2025 | Small Delay Fault Testing with Multiple Voltages under Variations: Defect vs. Fault CoverageabstractAbstract It has been known and explored for many years that low voltage testing amplifies the effect of a defect, increasing the size of a Small Delay Fault (SDF) and, in the best case, turning SDFs into easily detectable stuck-at-faults. It is often overlooked that $$V_{\textrm{min}}$$ V min testing poses an additional challenge to the test pattern generation method under process variations. The standard deviation of gate delays under $$V_{\textrm{min}}$$ V min is a multiple of that under nominal voltage. The increased variation will invalidate the efficiency of test patterns generated under nominal voltage and significantly reduce fault coverage. This paper presents the first algorithm for test pattern generation specifically tuned for $$V_{\textrm{min}}$$ V min testing which obtains higher fault coverage by smaller test sets than those generated for nominal voltage. The patterns applicable to other voltage levels can be derived from the pattern set generated under extreme variations at low supply voltage. Experimental results demonstrate that the proposed method produces test patterns that outperform N-detection test sets in terms of test set volume and fault efficiency across different voltage levels. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
J. Electron. Test. | 2 |
| 2024 | Time and Space Optimized Storage-based BIST under Multiple Voltages and VariationsabstractLogic Built-In Self-Test (LBIST) with stored deterministic patterns is supported by the major CAD vendors and is gaining increasing attention, especially for safety-critical applications such as automotive. It is used for both manufacturing and periodic in-field testing. An unresolved challenge so far stems from the inevitable process variations. This paper presents the first approach for storage-based BIST addressing delay faults under process variations and multiple voltages. A unified solution for pattern generation, test set compaction and BIST hardware is presented that is compatible with commercial schemes. The solution significantly outperforms traditional N-detect for transition faults in terms of test set size, test application time and fault efficiency. Hanieh Jafarzadeh, Florian Klemme, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ETS | 2 |
| 2024 | Minimizing PVT-Variability by Exploiting the Zero Temperature Coefficient (ZTC) for Robust Delay Fault TestingabstractProcess, Voltage, Temperature (PVT) variations impede the test generation for Small Delay Faults (SDFs) significantly as test patterns effective for one circuit instance may not be valid for a different one. Temperature-induced timing variations in FinFET and Gate-All-Around (GAA) technologies are especially severe due to temperature fluctuations and self-heating. Depending on the supply voltage, they show the Temperature Effect Inversion (TEI) which describes the increase of the circuit speed with increasing temperature. The Zero Temperature Coefficient (ZTC) specifies a supply voltage where TEI approaches 0, and the optimal voltage is determined, such that the effects of temperature-induced variability are minimized. Simulation results are reported, which demonstrate that test generation at the ZTC voltage leads to higher fault coverage of SDFs while using significantly less test patterns. Hanieh Jafarzadeh, Florian Klemme, Jan Dennis Reimer, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ITC | 2 |
| 2024 | Approximation- and Quantization-Aware Training for Graph Neural NetworksabstractGraph Neural Networks (GNNs) are one of the best-performing models for processing graph data. They are known to have considerable computational complexity, despite the smaller number of parameters compared to traditional Deep Neural Networks (DNNs). Operations-to-parameters ratio for GNNs can be tens and hundreds of times higher than for DNNs, depending on the input graph size. This complexity indicates the importance of arithmetic operation optimization within GNNs through model quantization and approximation. In this work, for the first time, we combine both approaches and implementquantization-andapproximation-aware trainingfor GNNs to sustain their accuracy under the errors induced by inexact multiplications. We employ matrix multiplication CUDA kernel to speed up the simulation of approximate multiplication within GNNs. Further, we demonstrate the execution speed, accuracy, and energy efficiency of GNNs with approximate multipliers in comparison with quantized low-bit GNNs. We evaluate the performance of state-of-the-art GNN architectures (i.e., GIN, SAGE, GCN, and GAT) on various datasets and tasks (i.e., Reddit-Binary, Collab for graph classification, Cora and PubMed for node classification) with a wide range of approximate multipliers. Our framework is available online:https://github.com/TUM-AIPro/AxC-GNN. Rodion Novkin, Florian Klemme, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2023 | ML to the Rescue: Reliability Estimation from Self-Heating and Aging in Transistors All the Way up ProcessorsabstractWith increasingly confined 3D structures and newly-adopted materials of higher thermal resistance, transistor self-heating has risen to a critical reliability threat in state-of-the-art and emerging process nodes. One of the challenges of transistor self-heating is accelerated transistor aging, which leads to earlier failure of the chip if not considered appropriately. Nevertheless, adequate consideration of accelerated aging effects, induced by self-heating, throughout a large circuit design is profoundly challenging due to the large gap between where self-heating does originate (i.e., at the transistor level) and where its ultimate effect occurs (i.e., at the circuit and system levels). In this work, we demonstrate an end-to-end workflow starting from self-heating and aging effects in individual transistors all the way up to large circuits and processor designs. We demonstrate that with our accurately estimated degradations, the required timing guardband to ensure reliable operation of circuits is considerably reduced by up to 96% compared to otherwise worst-case estimations that are conventionally employed. Hussam Amrouch, Florian Klemme |
ASP-DAC | 2 |
| 2023 | Design Automation for Cryogenic CMOS CircuitsabstractCryogenic CMOS circuits operate at temperatures close to absolute zero and are essential in many applications such as controllers for quantum computing but also medical engineering, space technology, or physical instruments. However, operating circuits at cryogenic temperatures fundamentally changes the underlying semiconductor physics that governs the CMOS transistor—rendering existing design automation approaches infeasible. In this work, we propose and implement the first end-to-end approach that enables design automation for cryogenic CMOS circuits. To this end, we (1) perform the first-of-its-kind measurements of commercial 5nm FinFET transistors from 300K down to 10K, (2) use the results to validate and calibrate the first cryogenic-aware industrial-standard compact model for FinFET technology, (3) create cryogenic-aware standard cell libraries that are compatible with the existing EDA tool flows, and (4) propose an initial cryogenic-aware logic synthesis approach that re-uses established design automation expertise but optimizes it for cryogenic purposes. Evaluations, comparisons, and discussions of all these novel contributions confirm the applicability and validity of the resulting cryogenic-aware design automation flow. Victor M. van Santen, Marcel Walter, Florian Klemme, Shivendra Singh Parihar, Girish Pahwa, Yogesh Singh Chauhan, Robert Wille, Hussam Amrouch |
DAC | 3 |
| 2023 | Upheaving Self-Heating Effects from Transistor to Circuit Level using Conventional EDA Tool FlowsabstractIn this work, we are the first to demonstrate how well-established EDA tool flows can be employed to upheave Self- Heating Effects (SHE) from individual devices at the transistor level all the way up to complete large circuits at the final layout (i.e., GDS-II) level. Transistor SHE imposes an ever-growing reliability challenge due to the continuous shrinking of geometries alongside the non-ideal voltage scaling in advanced technology nodes. The challenge is largely exacerbated when more confined 3D structures are adopted to build transistors such as upcoming Nanosheet FETs and Ribbon FETs. By employing increasingly-confined structures and materials of poorer thermal conductance, heat arising within the transistor's channel is trapped inside and cannot escape. This leads to accelerated defect generation and, if not considered carefully, a profound risk to IC reliability. Due to the lack of EDA tool flows that can consider SHE, circuit designers are forced to take pessimistic worst-case assumptions (obtained at the transistor level) to ensure reliability of the complete chip for the entire projected lifetime - at the cost of sub-optimal circuit designs and considerable efficiency losses. Our work paves the way for designers to estimate less pessimistic (i.e., small yet sufficient) safety margins for their circuits leading to higher efficiency without compromising reliability. Further, it provides new perspectives and opens new doors to estimate and optimize reliability correctly in the presence of emerging SHE challenge through identifying early the weak spots and failure sources across the design. Florian Klemme, Sami Salamin, Hussam Amrouch |
DATE | 1 |
| 2023 | Robust Resistive Open Defect Identification Using Machine Learning with Efficient Feature SelectionabstractResistive open defects in FinFET circuits are reliability threats and should be ruled out before deployment. The performance variations due to these defects are similar to the effect of process variations which are mostly benign. In order not to sacrifice yield for reliability the effect of defects should be distinguished from process variations. It has been shown that machine learning (ML) schemes are able to classify defective circuits with high accuracy based on the maximum frequencies$F_{max}$obtained under multiple supply voltages$V_{dd} \in V_{op}$. The paper at hand presents a method to minimize the number of required measurements. Each supply voltage$V_{dd}$defines a feature$F_{max}(V_{dd})$. A feature selection technique is presented, which uses also the already available$F_{max}$measurements. It is shown that ML-based techniques can work efficiently and accurately with this reduced number of$F_{max}(V_{dd})$measurements. Zahra Paria Najafi-Haghi, Florian Klemme, Hanieh Jafarzadeh, Hussam Amrouch, Hans-Joachim Wunderlich |
DATE | 2 |
| 2023 | Learning-Oriented Reliability Improvement of Computing Systems From Transistor to Application LevelabstractDue to technology scaling in modern computing platforms, the safety and reliability issues have increased tremendously, which often accelerate aging, lead to permanent faults, and cause unreliable execution of applications. Failure in some computing systems like avionics may cause catastrophic consequences. Therefore, managing reliability under all circumstances of stress and environmental changes is crucial in all abstraction layers, from application to transistor levels. Machine learning techniques are recently being employed for dynamic reliability estimation and optimization. They can adapt to varying workloads and system conditions. This paper presents reliability improvement approaches from multiple perspectives-from transistor-level to application-level-and discusses their effectiveness and limitations as well as open challenges. Behnaz Ranjbar, Florian Klemme, Paul R. Genssler, Hussam Amrouch, Jinhyo Jung, Shail Dave, Hwisoo So, Kyongwoo Lee, Aviral Shrivastava, Ji-Yung Lin, Pieter Weckx, Subrat Mishra, Francky Catthoor, Dwaipayan Biswas, Akash Kumar 0001 |
DATE | 2 |
| 2023 | Robust Pattern Generation for Small Delay Faults Under Process VariationsabstractSmall Delay Faults (SDFs) introduce additional delays smaller than the capture time and require timing-aware test pattern generation. Since process variations can invalidate the effectiveness of such patterns, different circuit instances may show a different fault coverage for the same test pattern set. This paper presents a method to generate test pattern sets for SDFs which are valid for all circuit timings. The method overcomes the limitations of known timing-aware Automatic Test Pattern Generation (ATPG) which has to use fault sampling under process variations due to the computational complexity. A statistical learning scheme maximises the coverage of SDFs in circuits following the variation parameters of a calibrated industrial FinFET transistor model. The method combines efficient ATPG for Transition Faults (TFs) with fast timing-aware fault simulation on GPUs. Simulation experiments show that the size of the pattern set is significantly reduced in comparison to standard N-detection while the fault coverage even increases. Hanieh Jafarzadeh, Florian Klemme, Jan Dennis Reimer, Zahra Paria Najafi-Haghi, Hussam Amrouch, Sybille Hellebrand, Hans-Joachim Wunderlich |
ITC | 2 |
| 2023 | SyncTREE: Fast Timing Analysis for Integrated Circuit Design through a Physics-informed Tree-based Graph Neural NetworkabstractNowadays integrated circuits (ICs) are underpinning all major information technology innovations including the current trends of artificial intelligence (AI). Modern IC designs often involve analyses of complex phenomena (such as timing, noise, and power etc.) for tens of billions of electronic components, like resistance (R), capacitance (C), transistors and gates, interconnected in various complex structures. Those analyses often need to strike a balance between accuracy and speed as those analyses need to be carried out many times throughout the entire IC design cycles. With the advancement of AI, researchers also start to explore news ways in leveraging AI to improve those analyses. This paper focuses on one of the most important analyses, timing analysis for interconnects. Since IC interconnects can be represented as an RC-tree, a specialized graph as tree, we design a novel tree-based graph neural network, SyncTREE, to speed up the timing analysis by incorporating both the structural and physical properties of electronic circuits. Our major innovations include (1) a two-pass message-passing (bottom-up and top-down) for graph embedding, (2) a tree contrastive loss to guide learning, and (3) a closed formular-based approach to conduct fast timing. Our experiments show that, compared to conventional GNN models, SyncTREE achieves the best timing prediction in terms of both delays and slews, all in reference to the industry golden numerical analyses results on real IC design data. Jiajie Li 0002, Florian Klemme, Gi-Joon Nam, Tengfei Ma 0001, Hussam Amrouch, Jinjun Xiong |
NeurIPS | 3 |
| 2023 | Transistor Self-Heating-Aware Synthesis for Reliable Digital Circuit DesignsabstractWith the continuous scaling in technology nodes, the transistor self-heating effect (SHE) emerges as a growing threat to circuit reliability. Increasingly confined transistor structures and advanced materials exacerbate thermal insulation, concealing temperature hotspots in the transistor’s channel. Without the consideration of these increased temperatures, reliability effects such as aging will be underestimated, putting the circuit at risk. In this work, we propose a novel design flow that enables designers to extract accurate SHE temperatures at the circuit level and harden their design with SHE-aware synthesis. Our approach employs customized standard cell libraries to convey SHE information and guide logic synthesis toward SHE-resilient designs. Using our approach, we demonstrate effective suppression of SHE in circuits by up to 50%, trading off improved SHE resilience against timing, power, and area goals. In addition, our SHE analysis allows for accurate estimation of timing guardbands, reducing unnecessary pessimism of conventional approaches by over 50%. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Design Close to the Edge for Advanced Technology using Machine Learning and Brain-Inspired AlgorithmsabstractIn advanced technology nodes, transistor performance is increasingly impacted by different types of design-time and run-time degradation. First, variation is inherent to the manufacturing process and is constant over the lifetime. Second, aging effects degrade the transistor over its whole life and can cause failures later on. Both effects impact the underlying electrical properties of which the threshold voltage is the most important. To estimate the degradation-induced changes in the transistor performance for a whole circuit, extensive SPICE simulations have to be performed. However, for large circuits, the computational effort of such simulations can become infeasible very quickly. Furthermore, the SPICE simulations cannot be delegated to circuit designers, since the required underlying transistor models cannot be shared due to their high confidentiality for the foundry. In this paper, we tackle these challenges at multiple levels, ranging from transistor to memory to circuit level. We employ machine learning and brain-inspired algorithms to overcome computational infeasibility and confidentiality problems, paving the way towards design close to the edge. Hussam Amrouch, Florian Klemme, Paul R. Genssler |
ASP-DAC | 2 |
| 2022 | Intelligent Methods for Test and ReliabilityabstractTest methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects. Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009 |
DATE | 16 |
| 2022 | On Extracting Reliability Information from Speed BinningabstractAdaptive Voltage Frequency Scaling (AVFS) is an important means to overcome process-induced variability challenges for advanced high-performance circuits. AVFS requires and allows determining the maximum speed Fmax(Vdd) reachable under a set of certain operation voltages Vdd. In this paper, it is shown that the Fmax(Vdd) measurements contain relevant data to identify some hidden defects in a chip which are reliability threats and can cause device failures, but pass the speed binning procedure within the given specifications.Static Timing Analysis (STA) is applied to a circuit designed by using standard cell libraries in which the underlying transistors along with process variations have been carefully calibrated against industrial 14nm FinFET measurement data, and in-stances with and without injected small resistive open defects are generated. From the slope of the function Fmax(Vdd), a machine learning procedure can identify some defects with high precision and few false positives. These chips can be then discarded without any further need and cost for testing. It has to be noted that this reliability information comes for free from the data which is already generated, and does not need any additional measurements. Zahra Paria Najafi-Haghi, Florian Klemme, Hussam Amrouch, Hans-Joachim Wunderlich |
ETS | 2 |
| 2022 | Impact of NCFET Technology on Eliminating the Cooling Cost and Boosting the Efficiency of Google TPUabstractRecent breakthroughs in Neural Networks (NNs) led to significant accuracy improvements of several machine learning applications such as image classification and voice recognition. However, this accuracy improvement comes at the cost of an immense increase in computation demands. NNs became one of the most common and computationally intensive workloads in today's datacenters. To address these computational demands, Google announced in 2016 the Tensor Processing Unit (TPU), an advanced custom ASIC accelerator for NN inference. Two new TPU versions (v2 and v3) followed in 2017 and 2018 that support also training. Google TPUv3 packs an immense processing power ($\mathrm{90TFLOPS}$per chip) in a tiny and condensed area, leading to very high on-chip power densities and thus excessive temperature. In this article, superlattice thermoelectric cooling, which is one of the emerging on-chip cooling, is considered as an advanced cooling example for Google TPU and we investigate the impact of Negative Capacitance FET (NCFET), which is one of the recent emerging technologies, on the cooling and efficiency of TPU. Through full-chip design, of the computational core of the TPU, based on$14\mathrm{nm}$Intel FinFET technology and multiphysics temperature simulations, we demonstrate that NCFET can significantly minimize the required cooling-cost. More than 4000 NCFET configurations are evaluated in order to traverse the entire design space defined by the thickness of the ferroelectric layer of NCFET, the operating voltage, cooling, and the operating frequency, in addition to all possible FinFET's configurations. Moreover, our experimental evaluation shows that by eliminating the cooling cost, NCFET delivers 2.8x higher efficiency compared to the conventional FinFET baseline. Sami Salamin, Georgios Zervakis 0001, Florian Klemme, Hammam Kattan, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 3 |
| 2022 | GNN4REL: Graph Neural Networks for Predicting Circuit Reliability DegradationabstractProcess variations and device aging impose profound challenges for circuit designers. Without a precise understanding of the impact of variations on the delay of circuit paths, guardbands, which keep timing violations at bay, cannot be correctly estimated. This problem is exacerbated for advanced technology nodes, where transistor dimensions reach atomic levels and established margins are severely constrained. Hence, traditional worst-case analysis becomes impractical, resulting in intolerable performance overheads. Contrarily, process-variation/aging-aware static timing analysis (STA) equips designers with accurate statistical delay distributions. Timing guardbands that are small, yet sufficient, can then be effectively estimated. However, such analysis is costly as it requires intensive Monte-Carlo simulations. Further, it necessitates access to confidential physics-based aging models to generate the standard-cell libraries required for STA. In this work, we employ graph neural networks (GNNs) to accurately estimate the impact of process variations and device aging on the delay of any path within a circuit. Our proposed GNN4REL framework empowers designers to perform rapid and accurate reliability estimations without accessing transistor models, standard-cell libraries, or even STA; these components are all incorporated into the GNN model via training by the foundry. Specifically, GNN4REL is trained on a FinFET technology model that is calibrated against industrial 14-nm measurement data. Through our extensive experiments on EPFL and ITC-99 benchmarks, as well as RISC-V processors, we successfully estimate delay degradations of all paths—notably within seconds—with a mean absolute error down to 0.01 percentage points. Lilas Alrahis, Johann Knechtel, Florian Klemme, Hussam Amrouch, Ozgur Sinanoglu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Variability-Aware Approximate Circuit Synthesis via Genetic OptimizationabstractOne of the major barriers that CMOS devices face at nanometer scale is increasing parameter variation due to manufacturing imperfections. Process variations severely inhibit the reliable operation of circuits, as the operational frequency at the nominal process corner is insufficient to suppress timing violations across the entire variability spectrum. To avoid variability-induced timing errors, previous efforts impose pessimistic and performance-degrading timing guardbands atop the operating frequency. In this work, we employ approximate computing principles and propose a circuit-agnostic automated framework for generating variability-aware approximate circuits that eliminate process-induced timing guardbands. Variability effects are accurately portrayed with the creation of variation-aware standard cell libraries, fully compatible with standard EDA tools. The underlying transistors are fully calibrated against industrial measurements from Intel 14nm FinFET in which both electrical characteristics of transistors and variability effects are accurately captured. In this work, we explore the design space of approximate variability-aware designs to automatically generate circuits of reduced variability and increased performance without the need for timing guardbands. Experimental results show that by introducing negligible functional error of merely$\boldsymbol {5.3 \times 10^{-3}}$, our variability-aware approximate circuits can be reliably operated under process variations without sacrificing the application performance. Konstantinos Balaskas, Florian Klemme, Georgios Zervakis 0001, Kostas Siozios, Hussam Amrouch, Jörg Henkel |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | Scalable Machine Learning to Estimate the Impact of Aging on Circuits Under Workload DependencyabstractTo ensure the correct functionality of a chip throughout its entire lifetime, preliminary circuit analysis with respect to aging-induced degradation is indispensable. However, state-of-the-art techniques only allow for the consideration of uniformly applied degradations, despite the fact that different workloads will lead to different degradations due to their distinct induced activities. This imposes over-pessimism when estimating the required timing guardbands, resulting in an unnecessary loss of performance and efficiency. In this work, we propose an approach that takes real-world workload dependencies into account and generates workload-specific aging-aware standard cell libraries, allowing for accurate analysis of aging-induced degradations. We employ machine learning techniques to overcome infeasible simulation times for individual transistor aging while sustaining high prediction accuracy. We also demonstrate scalability to previously unknown workloads and discuss multiple approaches to estimate the machine learning accuracy by employing coverage metrics. In our evaluation, we achieve predictions of workload-dependent aging-aware standard cells with an average accuracy (R2 score) of 95.28%. Using predicted cell libraries in static timing analysis, timing guardbands for multiple circuits are reported with an error of less than 0.1% on average. We demonstrate that timing guardband requirements can be reduced by up to 30% when considering specific workloads over worst-case estimations as performed in state-of-the-art tool flows. Even for unknown workloads of different circuits, accurate prediction with relative errors below 1% can be achieved. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Efficient Learning Strategies for Machine Learning-Based Characterization of Aging-Aware Cell LibrariesabstractMachine learning (ML)-driven standard cell library characterization enables rapid, on-the-fly generation of cell libraries, opening the door for extensive design-space exploration and other, previously infeasible approaches. However, the benefits of ML-based cell library characterization are strongly limited by its high demand in training data and the costly SPICE simulation required to generate the training samples. Therefore, efficient learning strategies are needed to minimize the required training data for ML models while still sustaining high prediction accuracy. In this work, we explore multiple active and passive learning strategies for ML-based cell library characterization with focus on aging-induced degradation. While random sampling and greedy sampling strategies operate with low computational overhead, active learning considers the performance of ML models to find the most valuable samples for training. We also introduce a hybrid approach of active learning and greedy sampling to optimize the trade-off between reduction in training samples and computational overhead. Our experiments demonstrate an achievable training data reduction of up to 77% compared to the state of the art, depending on the targeted accuracy of the ML models. Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | Machine Learning for Circuit Aging Estimation under Workload DependencyabstractCircuit analysis with respect to aging-induced degradation is critical to ensure correct operation throughout the entire lifetime of a chip. However, state-of-the-art techniques only allow for the consideration of uniformly applied degradation, despite the fact that different workloads will lead to different degradations due to the different induced activities. This imposes over-pessimism in estimating the required timing guardbands, resulting in unnecessary losses of performance and efficiency. In this work, we propose an approach that takes real-world workload dependencies into account and generates workload-specific aging-aware standard cell libraries. This allows for accurate analysis of circuits under the actual effect of aging-induced degradation. We make use of machine learning techniques to overcome infeasible simulation times for individual transistor aging while sustaining high accuracy. In our evaluation on the PULP microprocessor, we achieve predictions of workload-dependent aging-aware standard cells with an average accuracy (R2score) of 94.7 %. Using the predicted cell libraries in Static Timing Analysis, timing guardbands are reported with an error of less than 0.1 %. We demonstrate that timing guardband requirements can be reduced by up to 21 % by considering specific workloads over worst-case analysis as performed in state-of-the-art tool flows. Florian Klemme, Hussam Amrouch |
ITC | 1 |
| 2021 | Machine Learning for On-the-Fly Reliability-Aware Cell Library CharacterizationabstractAging-induced degradation imposes a major challenge to the designer when estimating timing guardbands. This problem increases as traditional worst-case corners bring over-pessimism to designers, exacerbating competitive and close-to-the-edge designs. In this work, we present an accurate machine learning approach for aging-aware cell library characterization, enabling the designer to evaluate their circuit under the impact of precisely selected degradation. Unlike state of the art, we bring cell library characterization to the designer, empowering their capability in exploring the impact of aging while protecting confidential information from the foundry at the same time. Furthermore, the fast inference of cell libraries makes it feasible, for the first time, to examine aging-induced variability analysis in a Monte-Carlo fashion. Finally, we show that the designer is able to select a less pessimistic timing guardband by choosing adequate delta threshold voltage ( ΔVth) for their design and their needs. Our machine learning approach reaches an R2score of $>99\%$ for almost all data stored in the cell library. Only timing constraints show slightly less accuracy with an R2score around 95%. When using ML-characterized libraries in static timing analysis, we achieve errors smaller than ±0.5% and ±0.1% for path delay and dynamic power, respectively. Errors in leakage power are negligible and even smaller by orders of magnitude. Our machine learning implementation for standard cell library characterization is publicly available. Download: https://opensource.mlcad.org Florian Klemme, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | Cell Library Characterization using Machine Learning for Design Technology Co-OptimizationabstractTo explore the full potential of any circuit and ensure its functionality at run-time, cell libraries beyond the typical PVT corners are needed. This holds even more for emerging technologies like Negative Capacitance (NC)-FinFET, where research in finding the optimal set of transistor parameters is still in its infancy. Design Technology Co-Optimization (DTCO) tackles bridging the large existing gap between device physics and the figures of merit of circuits. In this paper, we propose a Machine Learning (ML) approach to rapidly generate full cell libraries on demand. This enables the designer to perform extensive design space exploration and fully automated Design Technology Co-Optimization while lowering the barrier of accessibility. We demonstrate library prediction with an R2 score of around 98% for individual values and Static Timing Analysis (STA) reports. Experimental results show that our DTCO approach overestimates the achievable improvement by around 5%, nevertheless improving upon the baseline configuration. Florian Klemme, Yogesh Singh Chauhan, Jörg Henkel, Hussam Amrouch |
ICCAD | 1 |
| 2020 | Modeling Emerging Technologies using Machine Learning: Challenges and OpportunitiesabstractCompact models of transistors act as the link between semiconductor technology and circuit design via circuit simulations. Unfortunately, compact model development and calibration is a challenging and time-intensive task, hindering rapid prototyping of a circuit (via circuit simulations) in emerging technologies. Moreover, foundries want to protect their confidential technology details to prevent reverse engineering. Hence, they limit access to compact transistor models of commercial technologies (e.g., with Non-Disclosure-Agreements). In this work, we propose Machine Learning (ML) to bridge the gap between early device measurements and later occurring compact model development. Our approach employs a Neural Network (NN) that captures the electrical response of a conventional FinFET transistor without knowledge of semiconductor physics. Additionally, our approach can be applied to emerging technologies, using Negative Capacitance FinFET (NC-FinFET) as an example for a (challenging to model) emerging technology. Inherently, the black-box nature of ML approaches keeps technology manufacturing details confidential. Furthermore, we show how using solely R2 score as our fitness function is insufficient and instead propose fitness based on key electrical characteristics or transistors like threshold voltage. Our NN-based transistor modeling can infer FinFET and NC-FinFET with an R2 score larger than 0.99 and transistor characteristics within 5% of experimental data. Florian Klemme, Jannik Prinz, Victor M. van Santen, Jörg Henkel, Hussam Amrouch |
ICCAD | 1 |