EDBT 2026 Demo / reviewers in the wild / expert
Paul R. Genssler
dblp:193/3348 · also Paul R. Genßler
· DBLP profile ↗
27ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0002-7175-7284ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 8 first-author · 25 since 2021Software engineering, systems software and programming languages · 7 · 3 first-author · 7 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Securing Hyper-Dimensional Computing: A Locking Mechanism with FPGA Implementation
Rupesh Raj Karn, Paul R. Genssler, Hussam Amrouch, Ozgur Sinanoglu |
ICISSP (2) | 2 |
| 2025 | Domain-Specific Hyperdimensional RISC-V Processor for Edge-AI TrainingabstractEdge AI has become the cornerstone of many applications. Yet, progress is limited by the large complexity of training a deep neural network (a DNN). hyperdimensional computing (HDC) is positioned as an alternative approach for Edge AI that is compact enough to enable training. The main challenge for an HDC model is to maintain its key features while balancing high inference accuracy with efficiency. A simple binary HDC model lacks accuracy, while the computational complexity of a floating-point model is too high. This work presents FixedHD, a novel 16-bit fixed-point HDC model enabling training at the Edge. FixedHD achieves an accuracy similar to floating-point model while lowering computational complexity. The model is supported by a customized RISC-V processor tailored to speedup both training and inference. The processor is extended with advanced HDC-specific instructions, a vector unit to utilize HDC’s parallel nature, and, for the first time, approximate computing to exploit its robustness. Further, memory requirements are reduced by quantizing mathematical functions and reducing the large HDC encoding matrix by up to 390 x. Compared to the baseline processor, inference and training are accelerated on average by 6.9 x and 3 x, respectively. The energy consumption is reduced by 4.6 x and 1.9 x at the cost of an increase in area by 45 %. The inference accuracy remains at the high level of floating-point models despite the heavy quantization and approximation. Sandy A. Wasif, Miran Wael, Paul R. Genssler, Eman Azab, Maggie Mashaly, Mohamed Abdelghany, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | HDCircuit: Brain-Inspired HyperDimensional Computing for Circuit RecognitionabstractCircuits possess a non-Euclidean representation, necessitating the encoding of their data structure (e.g., gate-level netlists) into fixed formats like vectors. This work is the first to propose brain-inspired hyperdimensional computing (HDC) for optimized circuit encoding. HDC does not require extensive training to encode a gate-level netlist into a hypervector and simplifies the similarity check between circuits from graph-based to the similarity between their hypervectors. We introduce a versatile HDC-based encoding method for circuit encoding. We demonstrate its effectiveness with the application of circuit recognition using ITC-99 and ISCAS-85 benchmarks. We maintain a 98.2% accuracy, even when the designs are obfuscated using logic locking. Paul R. Genssler, Lilas Alrahis, Ozgur Sinanoglu, Hussam Amrouch |
DATE | 1 |
| 2024 | DropHD: Technology/Algorithm Co-Design for Reliable Energy-Efficient NVM-Based Hyper-Dimensional Computing Under Voltage ScalingabstractBrain-inspired hyperdimensional computing (HDC) offers much more efficient computing compared to other classical deep learning and related machine learning algorithms. Unlike classical CMOS, emerging non-volatile memories (NVMs) used in the realization of HDC are susceptible to failures under voltage scaling, which is essential for energy saving. Although HDC is inherently robust against errors, this is only possible when hypervectors with a large dimension (e.g., 10,000 bits) are being used, resulting in significant energy consumption. This work demonstrates, for the first time, that different NVM technologies exhibit different error characteristics under voltage scaling. In contrast to conventional CMOS-based SRAM, we demonstrate that the error behavior is data-dependent and not captured by simple bit flips in emerging NVMs. We employ our cross-layer framework that starts from the underlying technology all the way up to the algorithm to develop the novel HDC training approach DropHD. DropHD considerably shrinks the size of hypervectors (e.g., from 10,000 bits down to merely 3000 bits), while maintaining a high inference accuracy. The use of aggressive voltage scaling reduces energy consumption by 1.6 x. DropHD further reduces it to up to 9.5 × while fully recovering the induced accuracy drop, i.e. without a tradeoff. Paul R. Genssler, Mahta Mayahinia, Simon Thomann, Mehdi Baradaran Tahoori, Hussam Amrouch |
DATE | 1 |
| 2024 | Frontiers in Edge AI with RISC-V: Hyperdimensional Computing vs. Quantized Neural NetworksabstractHyperdimensional Computing (HDC) is an emerging paradigm that stands as a compelling alternative to conventional Deep Learning algorithms. HDC holds four key promises. First, the ability to learn from little data. Second, to be robust against noise in this data. HDC also promises to be resilient against errors in the underlying hardware. This includes the memory on which the model is stored and errors in the computations of the operations, which is attributed to the encoding of information across an expansive dimensional space. Fourth, HDC can be implemented efficiently in hardware due to its lightweight and embarrassingly parallel computations. In this work, those four key promises are evaluated in a holistic way. A fixed-point and a binary HDC implementation are compared against neural network implementations. The models are executed on a RISC-V processor to ensure a fair comparison. While the results confirm the ability to learn from little data and the resiliency against errors, the higher inference accuracy of neural networks favors them in most experiments. Based on these insights, we formulate challenges and opportunities for HDC. Our implementations for QNN, binary and fixed-point HDC are available online: https://github.com/TUM-AIPro/HDC-vs-QNN Paul R. Genssler, Sandy A. Wasif, Miran Wael, Rodion Novkin, Hussam Amrouch |
DATE | 1 |
| 2024 | Algorithm to Technology Co-Optimization for CiM-Based Hyperdimensional ComputingabstractHyperdimensional computing (HDC) has been recognized as an efficient machine learning algorithm in recent years. Robustness against noise and simple computational operations, while being limited by the memory bandwidth, make it a perfect fit for the concept of computation in memory (CiM) with emerging nonvolatile memory (NVM) technologies. For an HDC accelerator based on NVM-CiM, there are different parameters from the algorithm all the way down to the technology that interact with each other and affect the overall inference accuracy as well as the energy efficiency of the accelerator. Therefore, in this paper, we propose, for the first time, a full-stack co-optimization method and use it to design an HDC accelerator based on NVM-based content addressable memory (CAM). By incorporating the device manufacturing variability and co-optimizing the algorithm and hardware design, HDC inference on our proposed NVM-based CiM accelerator can reduce the energy consumption by 3.27x, while compared to the purely software-based implementation, the inference accuracy loss is merely 0.125%. Mahta Mayahinia, Simon Thomann, Paul R. Genssler, Christopher Münch, Hussam Amrouch, Mehdi Baradaran Tahoori |
DATE | 3 |
| 2024 | In-Memory Acceleration of Hyperdimensional Genome Matching on Unreliable Emerging TechnologiesabstractNovel computer architectures like Compute-in-Memory (CiM) merge the memory and processing units, mimicking the human brain. Simultaneously, Hyperdimensional Computing (HDC) is emerging as a brain-inspired machine learning (ML) approach. Both developments hold promise for the realm of AI and computing, especially for genome-matching tasks, where large data movements overwhelm traditional von Neumann architectures. FeFET is one of the up-and-coming emerging technologies that promises to enable ultra-efficient and compact CiM architectures. However, the adoption of FeFETs is hindered by their 10 nm-thick Ferroelectric (FE) layer and process variation. Thus, calculations with FeFETs have errors (noise) that traditional ML genome-matching models cannot tolerate. To overcome these challenges, this work is the first one to i) present a reliable HDC framework (HDGIM) for highly-scaled (down to merely 3nm), multi-bit FeFET technology, ii) introduce temperature-thickness modeled noise from FeFET to the HDC system, and iii) extensively define the memorization capacity of HDC hyperparameters in order to evaluate the performance before deployment theoretically. Our novel HDC learning framework iteratively uses two models: a full-precision 32-bit HDC model, an ideal model for training, and a reduced bit-precision by a novel quantization method for validation and inference. Our results demonstrate that highly-scaled FeFET, realizing 3-bit and even 4-bit, can withstand any modeled noise given high dimensionality during inference. Considering the noise during model adjustment improves the inherent robustness by almost 9% on the 4-bit case. Hamza Errahmouni Barkam, Sanggeon Yun, Paul R. Genssler, Che-Kai Liu, Zhuowen Zou, Hussam Amrouch, Mohsen Imani |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | WaSSaBi: Wafer Selection With Self-Supervised Representations and Brain-Inspired Active LearningabstractLarge datasets are often available for machine learning tasks. However, only very few contain labels for all the samples because labeling is a very labor-intensive process. Hence, large unlabeled datasets are available but inaccessible to traditional supervised learning methods. In this work, we combine two approaches to reduce the number of required labels. First, self-supervised learning (SSL) to utilize the large unlabeled dataset. SSL creates an encoder from those unlabeled samples that transforms the input into intermediate feature representations. Second, active learning is employed for the classification where labels are required. Active learning intelligently selects the most informative samples for manual labeling. Thus, it reduces the amount the labels required to achieve a high classification accuracy. The selected samples are used to train a brain-inspired hyperdimensional computing and random forest classifier. We demonstrate the outstanding performance of our approach with the example of wafer map defect pattern classification. It is a crucial diagnostic task helping to identify systematic problems in the manufacturing and improving yield. With our proposed method, a high 97% classification accuracy is achieved with only 6% of the labeled dataset for the first time. Our approach demonstrates the potential for training a machine learning model from less labeled samples by combining SSL with active learning. Karthik Pandaram, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Beyond von Neumann Era: Brain-Inspired Hyperdimensional Computing to the RescueabstractBreakthroughs in deep learning (DL) continuously fuel innovations that profoundly improve our daily life. However, DNNs overwhelm conventional computing architectures by their massive data movements between processing and memory units. As a result, novel computer architectures are indispensable to improve or even replace the decades-old von Neumann architecture. Nevertheless, going far beyond the existing von Neumann principles comes with profound reliability challenges for the performed computations. This is due to analog computing together with emerging beyond-CMOS technologies being inherently noisy and inevitably leading to unreliable computing. Hence, novel robust algorithms become a key to go beyond the boundaries of the von Neumann era. Hyper-dimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional DL and ML algorithms. Unlike conventional DL and ML algorithms, HDC is inherently robust against errors along a much more efficient hardware implementation. In addition to these advantages at hardware level, HDC's promise to learn from little data and the underlying algebra enable new possibilities at the application level. In this work, the robustness of HDC algorithms against errors and beyond von Neumann architectures are discussed. Further, the benefits of HDC as a machine learning algorithm are demonstrated with the example of outlier detection and reinforcement learning. Hussam Amrouch, Paul R. Genssler, Mohsen Imani, Mariam Issa, Xun Jiao 0002, Wegdan Mohammad, Gloria Sepanta |
ASP-DAC | 2 |
| 2023 | Tutorial: The Synergy of Hyperdimensional and In-Memory Computing
Paul R. Genssler, Simon Thomann, Hussam Amrouch |
CODES+ISSS | 1 |
| 2023 | HDGIM: Hyperdimensional Genome Sequence Matching on Unreliable highly scaled FeFETabstractThis is the first work to present a reliable application for highly scaled (down to merely 3nm), multi-bit Ferroelectric FET (FeFET) technology. FeFET is one of the up-and-coming emerging technologies that is not only fully compatible with the existing CMOS but does hold the promise to realize ultra-efficient and compact Compute-in-Memory (CiM) architectures. Nevertheless, FeFETs struggle with the 10nm thickness of the Ferroelectric (FE) layer. This makes scaling profoundly challenging if not impossible because thinner FE significantly shrinks the memory window leading to large error probabilities that cannot be tolerated. To overcome these challenges, we propose HDGIM, a hyperdimensional computing framework catered to FeFET in the context of genome sequence matching. Genome Sequence Matching is known to have high computational costs, primarily due to huge data movement that substantially overwhelms von-Neuman architectures. On the one hand, our cross-layer FeFET reliability modeling (starting from device physics to circuits) accurately captures the impact of FE scaling on errors induced by process variation and inherent stochasticity in multi-bit FeFETs. On the other hand, our HDC learning framework iteratively adapts by using two models, a full-precision, ideal model for training and a quantized, noisy version for validation and inference. Our results demonstrate that highly scaled FeFET realizing 3-bit and even 4-bit can withstand any noise given high dimensionality during inference. If we consider the noise during model adjustment, we can improve the inherent robustness compared to adding noise during the matching process. Hamza Errahmouni Barkam, Sanggeon Yun, Paul R. Genssler, Zhuowen Zou, Che-Kai Liu, Hussam Amrouch, Mohsen Imani |
DATE | 3 |
| 2023 | Learning-Oriented Reliability Improvement of Computing Systems From Transistor to Application LevelabstractDue to technology scaling in modern computing platforms, the safety and reliability issues have increased tremendously, which often accelerate aging, lead to permanent faults, and cause unreliable execution of applications. Failure in some computing systems like avionics may cause catastrophic consequences. Therefore, managing reliability under all circumstances of stress and environmental changes is crucial in all abstraction layers, from application to transistor levels. Machine learning techniques are recently being employed for dynamic reliability estimation and optimization. They can adapt to varying workloads and system conditions. This paper presents reliability improvement approaches from multiple perspectives-from transistor-level to application-level-and discusses their effectiveness and limitations as well as open challenges. Behnaz Ranjbar, Florian Klemme, Paul R. Genssler, Hussam Amrouch, Jinhyo Jung, Shail Dave, Hwisoo So, Kyongwoo Lee, Aviral Shrivastava, Ji-Yung Lin, Pieter Weckx, Subrat Mishra, Francky Catthoor, Dwaipayan Biswas, Akash Kumar 0001 |
DATE | 3 |
| 2023 | Stress-Resiliency of AI Implementations on FPGAsabstractFPGAs have become a popular choice for machine learning acceleration for both cloud and edge devices. While traditional neural networks show impressive performance in classification tasks, Hyperdimensional Computing (HDC) is rapidly emerging as a promising novel machine learning approach for its hardware-friendly inference. In HDC, classes are embedded into high-dimensional vectors during training, and inputs can be classified by computing similarity metrics between class-vectors during inference. HDC inference is especially promoted in terms of its resiliency against errors, attributed to the large inherent redundancy. In this work, we perform a thorough experimental investigation of the fault resiliency of various FPGA-based machine learning implementations under different aspects of stress, comparing HDC with classical neural network approaches. We explore both the amount of faulty classifications as well as system crashes while subjecting the designs to timing stress using overclocking, voltage stress with excessive switching activity, and thermal stress. Jonas Krautter, Paul R. Genssler, Gloria Sepanta, Hussam Amrouch, Mehdi Baradaran Tahoori |
FPL | 2 |
| 2023 | Frontiers in AI Acceleration: From Approximate Computing to FeFET Monolithic 3D IntegrationabstractWith the rapidly expanding applications of artificial intelligence (AI), the quest for hardware acceleration to foster high-speed and energy-efficient AI computation has become ever more important. In this work, we first explore the performance and energy advantages of employing classical AI acceleration with conventional systolic multiply-accumulate (MAC) arrays. We then highlight the growing importance of monolithic 3D integration as a transformative hardware acceleration strategy, moving beyond the constraints of classical von Neumann architectures. We also discuss how brain-inspired hyperdimensional computing (HDC) offers an exciting avenue for overcoming the power-hungry requirements often associated with MAC arrays, which are inevitable in deep learning hardware. Addressing the limitations of von Neumann architectures, we present the potential of monolithic 3D integration to enable ultra-dense Processing-in-Memory (PiM) layers stacked on top of high-performance CMOS logic. This novel approach offers to enhance computational performance. Recognizing the need for compatibility with low thermal budgets, we identify ferroelectric thin-film transistors (FeTFT) as a promising candidate for back-end-ofline (BEOL) fabrication. We highlight recent advances in BEOL FeTFT technology and demonstrate how technology/algorithm co-optimization plays a crucial role in the successful realization of reliable brain-inspired HDC on potentially unreliable FeTFT-based PiM layers. Our results showcase the potential of these innovations for the development of next-generation, energy-efficient AI hardware. Paul R. Genssler, Somaya Mansour, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 2 |
| 2023 | HW/SW Co-Design for Reliable TCAM- Based In-Memory Brain-Inspired Hyperdimensional ComputingabstractBrain-inspired hyperdimensional computing (HDC) is continuously gaining remarkable attention. It is a promising alternative to traditional machine-learning approaches due to its ability to learn from little data, lightweight implementation, and resiliency against errors. However, HDC is overwhelmingly data-centric similar to traditional machine-learning algorithms. In-memory computing is rapidly emerging to overcome the von Neumann bottleneck by eliminating data movements between compute and storage units. In this work, we investigate and model the impact of imprecise in-memory computing hardware, namely TCAM cells, on the inference accuracy of HDC. Our modeling is based on 14nm FinFET technology fully calibrated with Intel measurement data. We accurately model, for the first time, the voltage-dependent error probability in SRAM-based and FeFET-based in-memory computing. Thanks to HDC's resiliency against errors, the complexity of the underlying hardware can be reduced, providing large energy savings of up to 6x. Experimental results for SRAM reveal that variability-induced errors have a probability of up to 39%. Despite such a high error probability, the inference accuracy is only marginally impacted. This opens doors to explore new tradeoffs. We also demonstrate that the resiliency against errors is application-dependent. In addition, we investigate the robustness of HDC against errors with emerging non-volatile FeFET devices instead of mature CMOS-based SRAMs. We demonstrate that inference accuracy does remain high despite the larger error probability, while large area and power savings can be obtained.All in all, HW/SW co-design is the key for efficient yet reliable in-memory HDC for both conventional CMOS technology and upcoming emerging technologies. Simon Thomann, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Computers | 2 |
| 2023 | Modeling and Predicting Transistor Aging Under Workload Dependency Using Machine LearningabstractThe pivotal issue of reliability is one of the major concerns for circuit designers. The driving force is transistor aging, dependent on operating voltage and workload. At the design time, it is difficult to estimate close-to-the-edge guardbands that keep aging effects during the lifetime at bay. This is because the foundry does not share its calibrated physics-based models, comprised of highly confidential technology and material parameters. However, the unmonitored yet necessary overestimation of degradation amounts to a performance decline, which could be preventable. Furthermore, these physics-based models are computationally complex. The costs of modeling millions of individual transistors at design time can be exorbitant. We propose the use of a machine learning model trained to replicate the physics-based model, such that no confidential parameters are disclosed. This effectual workaround is fully accessible to circuit designers for the purposes of design optimization. We demonstrate the model’s ability to generalize by training on data from one circuit and applying it successfully to a benchmark circuit. The mean relative error is as low as 1.7%, with a speedup of up to$20\times $. Circuit designers, for the first time ever, will have ease of access to a high-precision aging model, which is paramount for efficient designs. In contrast to existing work, our approach takes the full switching activity into account to model recovery effects. This work is a promising step in the direction of bridging the gap between the foundry and circuit designers. Paul R. Genssler, Hamza Errahmouni Barkam, Karthik Pandaram, Mohsen Imani, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | FDSOI-Based Analog Computing for Ultra-Efficient Hamming Distance Similarity CalculationabstractComputing the similarity between two binary strings is a frequently used operation in cryptography, machine learning, and other areas. The Hamming distance is a simple yet costly to compute similarity metric. A common way is to XOR both binary input strings and then count the number of 1s. Especially the latter popcount part is inefficient with purely digital circuits. In this paper, a novel analog circuit is proposed to compute the Hamming distance in an ultra-efficient way. Contrary to the major trend in the state of the art, no emerging technology is required. Instead, the unique feature of the mature FDSOI transistor technology is exploited for the first time to perform analog-based similarity calculation. Thanks to the additional back gate available in this technology, the transistor’s threshold voltage can be modulated by more than 1 V. Through this key feature, an ultra-efficient analog computing is realized, replacing the inefficient digital popcount traditionally built from expensive adder tree structures. The design is evaluated with an FDSOI transistor model calibrated with industrial measurements. The energy-delay product is at least 24$\times $smaller than purely digital implementations and the transistor count is reduced by over 2.6$\times $. Albi Mema, Simon Thomann, Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | Design Close to the Edge for Advanced Technology using Machine Learning and Brain-Inspired AlgorithmsabstractIn advanced technology nodes, transistor performance is increasingly impacted by different types of design-time and run-time degradation. First, variation is inherent to the manufacturing process and is constant over the lifetime. Second, aging effects degrade the transistor over its whole life and can cause failures later on. Both effects impact the underlying electrical properties of which the threshold voltage is the most important. To estimate the degradation-induced changes in the transistor performance for a whole circuit, extensive SPICE simulations have to be performed. However, for large circuits, the computational effort of such simulations can become infeasible very quickly. Furthermore, the SPICE simulations cannot be delegated to circuit designers, since the required underlying transistor models cannot be shared due to their high confidentiality for the foundry. In this paper, we tackle these challenges at multiple levels, ranging from transistor to memory to circuit level. We employ machine learning and brain-inspired algorithms to overcome computational infeasibility and confidentiality problems, paving the way towards design close to the edge. Hussam Amrouch, Florian Klemme, Paul R. Genssler |
ASP-DAC | 3 |
| 2022 | Brain-Inspired Hyperdimensional Computing for Ultra-Efficient Edge AIabstractHyperdimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional deep learning algorithms. Despite the profound success of Deep Neural Networks (DNNs) in many domains, the amount of computational power and storage that they demand during training makes deploying them in edge devices very challenging if not infeasible. This, in turn, inevitably necessitates streaming the data from the edge to the cloud which raises serious concerns when it comes to availability, scalability, security, and privacy. Further, the nature of data that edge devices often receive from sensors is inherently noisy. However, DNN algorithms are very sensitive to noise, which makes accomplishing the required learning tasks with high accuracy immensely difficult. In this paper, we aim at providing a comprehensive overview of the latest advances in HDC. HDC aims at realizing real-time performance and robustness through using strategies that more closely model the human brain. HDC is, in fact, motivated by the observation that the human brain operates on high-dimensional data representations. In HDC, objects are thereby encoded with high-dimensional vectors which have thousands of elements. In this paper, we will discuss the promising robustness of HDC algorithms against noise along with the ability to learn from little data. Further, we will present the outstanding synergy between HDC and beyond von Neumann architectures and how HDC opens doors for efficient learning at the edge due to the ultra-lightweight implementation that it needs, contrary to traditional DNNs. Hussam Amrouch, Mohsen Imani, Xun Jiao 0002, Yiannis Aloimonos, Cornelia Fermüller, Dehao Yuan, Dongning Ma, Hamza Errahmouni Barkam, Paul R. Genssler, Peter Sutor Jr. |
CODES+ISSS | 9 |
| 2022 | Intelligent Methods for Test and ReliabilityabstractTest methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects. Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009 |
DATE | 10 |
| 2022 | Wafer Map Defect Classification Based on the Fusion of Pattern and Pixel InformationabstractWith the dramatically increasing requirements on semiconductor products, improving the yield is one of the major tasks for semiconductor manufacturers. To minimize losses, automatic and efficient wafer testing tools are required to quickly notify the engineers of potential problems. One such technique is wafer map defect pattern classification, which has inspired and motivated extensive research over the last decades. Many popular studies often design novel wafer map defect identification algorithms based on manual feature extraction, statistical learning and deep neural networks, having achieved significant advancement and success. However, these methods often face challenges of training large-scale networks and few of them have noticed the full usage of the information within each wafer map. Based on the concerns above, this paper proposes a multi-task learning framework based on neural networks that fuses the information of the entire wafer map as well as the state of each individual die to enhance the defect pattern classification capability. Extensive experiments on a public real-world dataset have been conducted to justify the effectiveness of our method. Specifically, our method achieved an classification accuracy of 96.3%, which was better or comparable to other state-of-the-art approaches that required notably larger network sizes and heavy data augmentation. Yiwen Liao, Raphaël Latty, Paul R. Genssler, Hussam Amrouch, Bin Yang 0009 |
ITC | 3 |
| 2022 | Cross-layer FeFET Reliability Modeling for Robust Hyperdimensional ComputingabstractHyperdimensional computing (HDC) is an emerging learning paradigm that has gained a lot of attention due to its ability to train with fewer data, lightweight implementation, and resiliency against errors. Similar to the brain, HDC can learn patterns in one iteration from small training data by computing a similarity metric such as Hamming distance. Ferroelectric Field-Effect-Transistor (FeFET) based Ternary Content Addressable Memory (TCAM) has been demonstrated as an excellent candi-date for computing this similarity metric. However, variations in the underlying ferroelectric transistor does impact the reliable HDC operation. In this paper, we demonstrate an end-to-end cross-layer FeFET reliability modeling to obtain robust HDC across the computing stack starting from transistor physics all the way to circuits and systems. The effect of random spatial fluctuation of ferroelectric (FE) domains and other variability sources on electrical characteristics of FeFET is computed through detailed physics-based TCAD simulations. Then, the entire TCAM array is simulated in SPICE using a carefully designed and calibrated compact model to capture the effect of transistor variability on the error probability for individual Hamming distances. Finally, the error probability is employed to compute the loss of inference accuracy of HDC with a language recognition task. We observe very little loss in accuracy even with a high degree of variation. Swetaki Chatterjee, Simon Thomann, Paul R. Genssler, Yogesh Singh Chauhan, Hussam Amrouch |
VLSI-SoC | 4 |
| 2022 | Brain-Inspired Computing for Circuit Reliability CharacterizationabstractTransistor scaling steadily approaches fundamental limits. Sustaining circuit reliability becomes an overwhelming challenge for foundries. Therefore, early and rapid characterization of degradation effects impacting the circuits transistors becomes essential. Such degradation effects are caused by design-time variation due to manufacturing variability and/or run-time variation due to transistor aging. In this work, we are the first to employ Brain-Inspired Hyperdimensional computing (HDC) for circuit reliability. HDC is an emerging light-weight machine-learning solution. Nowadays, it is mainly applied to bio-signal processing. We bring the research of HDC to the next level by demonstrating how it can be applied to address the challenges in circuit reliability. This has far-reaching consequences due to the large savings achieved by 1) reducing the amount of training data, 2) removing the need to send the data to the Cloud for model training, and 3) significantly speeding up the characterization and classification tasks. We demonstrate the viability of HDC, using SRAM and other circuits as an example. HDC outperforms traditional machine learning methods, such as support vector machine, in accuracy and requires up to 20x fewer training samples. Our implementation and analysis are based on industrial 14 nm FinFET fully calibrated with Intel measurements. Paul R. Genssler, Hussam Amrouch |
IEEE Trans. Computers | 1 |
| 2022 | On the Reliability of FeFET On-Chip MemoryabstractFerroelectric Field-Effect Transistor (FeFET) is a promising future technology for non-volatile on-chip memories. It is rapidly attracting an ever-increasing attention from industry. The key advantage of FeFETs is full compatibility with the existing CMOS fabrication process beside their very low power consumption. To enable ultra-dense memories, 1-FeFET AND Arrays were proposed in which a memory cell is formed from merely a single FeFET. All access transistors, which are traditionally needed to operate memory cells, are removed. However, this imposes a new challenge ofindirect write disturbances. Neighboring memory cells are indirectly degraded whenever adirect write operationoccurs to a particular FeFET cell. Only recently the impact of such indirect disturbances on the FeFET reliability was experimentally investigated at device (i.e., transistor) level. However, to explore and properly judge the feasibility of 1-FeFET AND Arrays for on-chip memories, investigating only the reliability of individual cells is indeed insufficient. Bridging the gap between the device level and system (i.e., chip) level is inevitable. In the presence of indirect disturbances, the position of a write access within the array plays a key role, which is governed by the running workloads. In addition, whether the write operation flips the previously stored value or not also plays an important role with regards to reliability. Hence, running workloads, which determine not only the position of the memory cells to be written but also the values written to them, plays an essential role in determining 1-FeFET AND Array reliability over time. Therefore, studying the reliability of FeFETs only at the device level (as done in state of the art) is insufficient. In this work, we investigate, for the first time, the reliability of FeFET memories from device to system level. To achieve that, we develop a unified model capturing the impact of bothindirectdisturbances anddirectwrites on the reliability of FeFET cells. Our study at system level then employs the unified model in the context of application workloads. We investigate different array sizes, write voltages, write methods and a wide range of workloads using the example of CPU caches as an example of on-chip memory. We demonstrate that indirect write disturbances are the dominate effect degrading the reliability of FeFET memories. For most cells, it contributes over 90 percent to the overall induced degradation. This provides guidelines for researchers at both device and circuit level to optimize the FeFET reliability further while considering thehiddenimpact of indirect write disturbances. Paul R. Genssler, Victor M. van Santen, Jörg Henkel, Hussam Amrouch |
IEEE Trans. Computers | 1 |
| 2022 | Software-Managed Read and Write Wear-Leveling for Non-Volatile Main MemoryabstractIn-memory wear-leveling has become an important research field for emerging non-volatile main memories over the past years. Many approaches in the literature perform wear-leveling by making use of special hardware. Since most non-volatile memories only wear out from write accesses, the proposed approaches in the literature also usually try to spread write accesses widely over the entire memory space. Some non-volatile memories, however, also wear out from read accesses, because every read causes a consecutive write access. Software-based solutions only operate from the application or kernel level, where read and write accesses are realized with different instructions and semantics. Therefore different mechanisms are required to handle reads and writes on the software level. First, we design a method to approximate read and write accesses to the memory to allow aging aware coarse-grained wear-leveling in the absence of special hardware, providing the age information. Second, we provide specific solutions to resolve access hot-spots within the compiled program code (text segment) and on the application stack. In our evaluation, we estimate the cell age by counting the total amount of accesses per cell. The results show that employing all our methods improves the memory lifetime by up to a factor of 955×. Christian Hakert, Kuan-Hsun Chen, Horst Schirmeier, Lars Bauer, Paul R. Genssler, Georg von der Brüggen, Hussam Amrouch, Jörg Henkel, Jian-Jia Chen |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2021 | Brain-Inspired Computing for Wafer Map Defect Pattern ClassificationabstractBrain-Inspired hyperdimensional computing is a quickly emerging alternative machine-learning concept. Hypervectors with thousands of dimensions represent real-world data. Thanks to this redundancy, the system becomes robust against noise in the input data, but also resilient against faults, similar to the human brain. The light-weight operations with hypervectors are fully parallelizable enabling fast learning and inference at the edge. A classifier achieving high accuracies can be created through one-shot learning from few examples. Such a feature is particularly valuable in the area of semiconductor testing, where the number of training samples, especially for cutting-edge technology, is limited. In this work, we explore the applicability of brain-inspired hyperdimensional computing to the field of testing for the first time. With the example of wafer map defect pattern classification, we investigate the challenges and opportunities of this emerging concept. Paul R. Genssler, Hussam Amrouch |
ITC | 1 |
| 2020 | Impact of Self-Heating on Performance, Power and Reliability in FinFET TechnologyabstractSelf-heating is one of the biggest threats to reliability in current and advanced CMOS technologies like FinFET and Nanowire, respectively. Encapsulating the channel with the gate dielectric improved electrostatics, but also thermally insulates the channel resulting in elevated channel temperatures as the generated heat is trapped within the channel. Elevated channel temperatures lowers the performance, increases leakage power and degrades the reliability of circuits. Self-heating becomes worse in each new transistor structure (from planar transistor to FinFET to Nanowire) due to the ever-increasing thermal resistance of the transistor. This leads to elevated temperatures, which must be carefully considered while designing circuits. Otherwise, reliability cannot be ensured. This work presents a self-heating study to illustrate how self-heating matters in digital circuits. It also explores the impact of running workloads in SRAM arrays, such as register files in CPUs, and how self-heating effects in SRAM cells can be mitigated. Victor M. van Santen, Paul R. Genssler, Om Prakash 0007, Simon Thomann, Jörg Henkel, Hussam Amrouch |
ASP-DAC | 2 |