VLDB 2026 Research / reviewers in the wild / expert
Rajendra Bishnoi
dblp:142/0243
· DBLP profile ↗
76ranked-venue papers
11as first author
30since 2021 · last 2026
0000-0002-0516-7112ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 74 · 11 first-author · 28 since 2021Software engineering, systems software and programming languages · 22 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SABCIM: Self-Adaptive Biasing Scheme for Accurate and Efficient Analog Compute-in-MemoryabstractAnalog Compute-in-Memory (CIM), leveraging non-volatile memristive devices to perform in-place computations in the analog domain, holds great potential to efficiently accelerate vector-matrix multiplications (VMM) and realize AI (Artificial Intelligence) at the edge. However, the data converters in such architectures often trade-off accuracy for high energy and area overheads, practically limiting the benefits of CIM. In this work, we present SABCIM, an array-periphery co-design approach for CIM that enables accurate computation as well as digitization of analog VMM outputs with high energy efficiency and competitive area overhead. By leveraging complementary input activations and data storage, each crossbar column generates differential analog output corresponding to the vector-vector multiplication (VVM) result, while inherently addressing underlying non-idealities. This is digitized using a compact, dual-ramp voltage-to-time converter (VTC)-based analog-to-digital converter (ADC). Benchmark results indicate that our work achieves up to $19.6 \times$ higher energy efficiency compared to state-of-the-art (SOTA), while maintaining comparable accuracies. Yashvardhan Biyani, Abhairaj Singh, Rajendra Bishnoi, Said Hamdioui |
ASP-DAC | 3 |
| 2026 | Multi-Partner Project: Efficient Deep Learning Platforms for Next-Generation Embedded Edge-AI SystemsabstractThe objective of our collaborative multi-partner project is to create an open-source Deep Learning framework called AIDGE for edge and embedded Artificial Intelligence (AI), built around an established European value chain. The framework is designed to support diverse application domains that function independently while serving a broad international community. It offers an integrated, full-stack workflow from Neural Network design and optimization to AI application development and hardware-level implementation with automated code generation for specific hardware targets. The platform aims to provide researchers and developers with a flexible environment to explore novel AI concepts, rapidly prototype solutions, and ensure strong alignment between academic research and industrial requirements. This paper summarizes the progress, outcomes, and milestones achieved up to the second year of this three-year project. Rajendra Bishnoi, Mohammad Amin Yaldagard, Konstantinos Stavrakakis, Said Hamdioui, Kanishkan Vadivel, Pankaj Upadhyay, Nicolás Rodríguez 0002, Teresa van Dam, Sander Steeghs-Turchina, Agathe Archet, Prathamesh Satish Deshpande, Giovanni Grandi, Hana Krichene, William Fabre, Fabian Chersi |
DATE | 1 |
| 2026 | Multi-Partner Project: Scalable, Ferroelectric-based Accelerators for Energy Efficient Edge AI (Ferro4EdgeAI)abstractThe Computing-In-Memory (CIM) paradigm offers a promising solution to the memory-wall bottleneck that limits conventional Von Neumann architectures. By performing data processing at the same physical location where the data are stored, CIM-based architectures minimize costly data movement and drastically improve energy efficiency. When implemented with Ferroelectric Field Effect Transistors (FeFETs), additional advantages from the non-volatility, fast switching, and low operating voltage of FeFETs are added. However, the widespread adoption of FeFETs is limited by their poor endurance, which is overcome by a Back End of the Line (BEoL) integration of FeFET-2, where a ferroelectric capacitor (FeCAP) is wired to the gate of a CMOS transistor providing high endurance compatible with low-power edge applications. These properties enable dense, low-power, and high-speed matrix operations essential for AI workloads. As a result, FeFET-2-based CIM accelerators offer a promising solution for energy-efficient, high-performance AI at the edge. The Ferro4EdgeAI project aims to develop an ultra low-power, scalable edge accelerator for AI, targeting a significant gain in energy efficiency with respect to state-of-the-art AI hardware accelerators. To attain this, our project focuses on innovation all along the value chain from materials, physic concepts, device architecture, integration technologies, and accelerators in a holistic design space exploration approach. Theofilos Spyrou, Yashvardhan Biyani, Konstantinos Stavrakakis, Rajendra Bishnoi, Said Hamdioui, Joel Minguet Lopez, Louise Dumas, Jean Coignus, Denys Ly, Hugo Chazot-Ranquet, Laurent Grenouillet, Fabien Grimaud, Simon Martin 0006, Olivier Billoint, François Andrieu, Ruben Alcala, Stefan Slesazeck, Athira Sunil, Antoine Cauquil, Rosario Pronsat, Damien Deleruyelle, Cédric Marchand 0002, Alberto Bosio, Ian O'Connor, Giulio Urlini, Simon Jeannot, Mohammad Sajedi Alvar, Nima Akbari Moghaddam, Thilo Werner, Tony Schenk, Bojun Cheng, Mina Khoei, Lucía Pérez Ramírez, EunJin Koh, Somnath Kale, Nicholas Barrett |
DATE | 4 |
| 2026 | Late Breaking Results - A Systematic Vulnerability Analysis of MRAM-Based Compute-in-Memory against Side-Channel Attacks
Hossein Pourmehrani, Yashas Krishnamohan, Sumukh Prashant Bhanushali, Saurabh Dhiman, Rajendra Bishnoi, Arindam Sanyal, Farshad Firouzi, Naghmeh Karimi |
VTS | 5 |
| 2026 | Detection of Read-Disturb Effects in RRAM-Based Computation-in-Memory Architectures for Neural NetworksabstractResistive random-access memory (RRAM)-based computation-in-memory (CIM) architectures offer a promising solution to meet the stringent energy efficiency demands of executing artificial intelligence (AI) algorithms directly on edge devices. However, these architectures suffer from the read-disturb problem, which can lead to accumulated computational errors over time. To maintain the required level of computational accuracy, conventional approaches rely on a static reprogramming process after a predefined number of read cycles, necessitating large counters and resulting in inefficiencies. This paper presents experimental results using real RRAM devices to analyze the read-disturb effect and builds on these insights to propose a circuit-level detection methodology for real-time monitoring of conductance drifts. The proposed method initiates reprogramming only when the device drift exceeds a defined threshold and reprogramming is actually needed. Additionally, an analytical method is developed to determine the minimum conductance state ratio needed to meet reliable detection criteria. Based on this foundation, the proposed detection technique is further optimized for dynamic identification of read-disturb effects. Experiment-augmented SPICE simulation results, using a calibrated model implemented in TSMC 40 nm CMOS technology, validate the functionality and effectiveness of the proposed detection approach. These results demonstrate its potential to improve both the reliability and efficiency of RRAM-based CIM architectures that provide up to a 4x improvement in energy-efficiency compared to traditional periodic reprogramming methods. Mohammad Amin Yaldagard, Ankit Bende, Sumit Diware, Vikas Rana, Said Hamdioui, Rajendra Bishnoi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | PHANTOM: Power Hammering Attack and Countermeasure on Multi-Tenant ReRAM Compute-in-Memory AcceleratorsabstractThe increasing demand for efficient and low-power deep neural network (DNN) inference has advanced the adoption of ReRAM-based compute-in-memory (CiM) accelerators, which perform computations directly within memory to reduce energy consumption and enhance throughput. However, such architectures are vulnerable to security threats, especially in a multi-tenant environment where multiple users share the same physical resources. This paper introduces a new attack model for multi-tenant ReRAM-based CiM, power hammering, that exploits the temperature sensitivity of ReRAM cells, inducing local temperature increases that lead to conductance drift and ultimately result in erroneous inference outcomes. This serves as a denial-of-service (DoS) attack, where malicious co-tenants degrade inferencing accuracy and system reliability for legitimate users in a shared environment, ultimately undermining trust and causing potential losses to the service provider. Additionally, we propose a novel strategy to counter this security vulnerability. In this technique, we focus on selectively protecting important weights with error compensation hardware. These important weights are treated as faults, and their computation is offloaded to compensation hardware. Simulation results confirm the effectiveness of the proposed method in ensuring accurate classification results even under adversarial conditions, thereby enabling secure multi-tenant inference on ReRAM-based CiM accelerators. Ashish Reddy Bommana, Rajendra Bishnoi, Naghmeh Karimi, Farshad Firouzi, Krishnendu Chakrabarty |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Enhancing Parallelism and Energy-Efficiency in SOT-MRAM based CIM Architecture for On-Chip Learning
Anubha Sehgal, Alok Kumar Shukla, Sumit Diware, Sandeep Soni, Seema Dhull, Sonal Shreya, Sourajeet Roy, Rajendra Bishnoi |
DAC | 8 |
| 2025 | Multi-Partner Project: A Deep Learning Platform Targeting Embedded Hardware for Edge-AI Applications (NEUROKIT2E)abstractThe goal of the NEUROKIT2E project is to create an open-source Deep Learning framework for edge and embedded AI built around an established European value chain. This framework, called AIDGE, supports a wide range of application areas that operate independently and serve a global user community. It provides easy and fast full-stack solutions from Neural Network design and optimization to AI application development all the way down to hardware implementations while enabling code generation for application-specific targets. This platform provides flexibility for academic users in the AI domain to explore and innovate while allowing them the possibility to prototype systems, ensuring their work aligns well with industrial needs. This paper presents the results and achievements of the first part of this three-year project, along with its roadmap and expected outcomes. Rajendra Bishnoi, Mohammad Amin Yaldagard, Said Hamdioui, Kanishkan Vadivel, Manolis Sifalakis, Nicolás Rodríguez 0002, Pedro Julián, Lothar Ratschbacher, Maen Mallah, Yogesh Ramesh Patil, Fabian Chersi |
DATE | 1 |
| 2025 | C3CIM: Constant Column Current Memristor-Based Computation-in-Memory Micro-ArchitectureabstractAdvancements in Artificial Intelligence (AI) and Internet-of-Things (IoT) have increased demand for edge AI, but deployment on traditional AI accelerators, like GPUs and TPUs, using von Neumann architecture, suffer from inefficiencies due to separate memory and compute units. Computation-in-Memory (CIM), utilizing non-volatile memristor devices to leverage analog computing principles and perform in-place computations, holds great potential in improving computational efficiency by eliminating frequent data movement. However, standard implementation of CIM faces several challenges, primarily high power consumption and subsequently induced nonlinearity, debating its viability for edge devices. In this paper, we propose C3CIM, a novel memristor-based CIM micro-architecture, featuring a new bit-cell and array design, targeting efficient implementation of Neural Networks (NN). Our architecture uses a constant current source to perform Multiply-and-Accumulate (MAC) operations with a very low computation current (10 to 100 nA), thereby significantly enhancing power efficiency. We adapted C3CIM for Spiking Neural Networks (SNN) and developed a prototype using TSMC 40nm CMOS node for on-silicon validation. Furthermore, our micro-architecture was benchmarked using two SNN models based on N-MNIST and IBM-Gesture datasets, for comparison against current state-of-the-art (SOTA). Results show up to 35x reduction in power along with 6.7x saving in energy compared to SOTA, demonstrating promising potential of this work for edge AI applications. Yashvardhan Biyani, Rajendra Bishnoi, Theofilos Spyrou, Said Hamdioui |
DATE | 2 |
| 2025 | Adaptive Multi-Threshold Encoding for Energy-Efficient ECG Classification Architecture Using Spiking Neural NetworkabstractTimely identification of cardiac arrhythmia (abnormal heartbeats) is vital for early diagnosis of cardiovascular diseases. Wearable healthcare devices facilitate this process by recording heartbeats through electrocardiogram (ECG) signals and using AI-driven hardware to classify them into arrhythmia classes. Spiking neural networks (SNNs) are well-suited for such hardware as they consume low energy due to event-driven operation. However, their energy-efficiency and accuracy are constrained by encoding methods that translate real-valued ECG data into spikes. In this paper, we present an SNN-based ECG classification architecture featuring a new adaptive multi-threshold spike encoding scheme. This scheme adjusts encoding window and granularity based on the importance of ECG data samples, to capture essential information with fewer spikes. We develop a high-accuracy SNN model for such spike representation, by proposing a technique specifically tailored to our encoding. We design a hardware architecture for this model, which incorporates optimized layer post-processing for energy-efficient data-flow and employs fixed-point quantization for computational efficiency. Moreover, we integrate this architecture with our encoding scheme into a system-on-chip implementation using TSMC 40 nm technology. Our approach provides up to 5.1x energy-efficiency compared to state-of-the-art SNN-based ECG classifiers, with high accuracy. Sumit Diware, Yingzhou Dong, Mohammad Amin Yaldagard, Said Hamdioui, Rajendra Bishnoi |
DATE | 5 |
| 2025 | European Test Symposium Teams: an Anniversary SnapshotabstractThe IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing. Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand |
ETS | 66 |
| 2025 | In-Field Monitoring and Preventing Read Disturb Faults in RRAMsabstractAddressing non-idealities in Resistive Random Access Memories (RRAMs) is crucial for their successful commercialization. For example, the inherent resistance drift that occurs during consecutive read operations can induce Read Disturb Faults (RDF), leading to functional errors. This paper analyzes and characterizes the resistance drift and the RDF based on data measurements and presents a physics-based RRAM compact model that incorporates these non-idealities. Additionally, an in-field mitigation scheme is proposed, leveraging bidirectional read operations to balance the resistance. The scheme is implemented and validated through circuit simulations, both for RRAM used as memory and for RRAM-based computation-in-memory microarchitectures for deep neural networks. The results demonstrate that RRAM without any mitigation scheme can start failing after 8,000 consecutive reads, while our mitigation scheme ensures that the memory remains functional even after 106consecutive reads. Furthermore, the results indicate that using the MNIST dataset as a case study, the accuracy can drop significantly from 86% to as low as 12.5% without any mitigation scheme. In contrast, the proposed mitigation scheme improves this accuracy up to 84.2%. Hanzhi Xun, Moritz Fieback, Sicong Yuan, Erbing Hua, Hassen Aziza, Letícia Maria Veiras Bolzani, Riccardo Cantoro, Rajendra Bishnoi, Mottaqiallah Taouil, Said Hamdioui |
ETS | 9 |
| 2025 | Continuous On-Chip Learning in Neural Networks using SOT-MRAM based CIM ArchitecturesabstractComputational-In-Memory (CIM) is an energy-efficient paradigm that integrates computation directly within memory arrays, reducing the bottleneck associated with data transfer. This approach is beneficial for Artificial Intelligence (AI) applications that require on-chip learning for real-time processing. However, implementing on-chip learning in CIM architectures remains challenging due to limited throughput and energy-efficiency during both online training and inference. In conventional architectures, weight updates necessitate the inference process to halt to avoid unintended computation outcomes. To overcome this limitation, this paper presents a novel Spin-Orbit Torque (SOT)-based CIM architecture tailored for continuous on-chip learning applications, which enable weight updates without interrupting the inference. The proposed SOT bit-cell utilizes two read ports and one write port (2R1W) configuration, where one read port (1R) is dedicated to inference and one read and one write (1R1W) for on-chip learning that enables concurrent read and write operations. Our proposed architecture is evaluated at the system-level using the Generic-PDK 45 nm technology node, demonstrating 2.4× improvement in energy-efficiency and 5.4× improvement in throughput compared to state-of-the-art solutions, with minimal overhead. Anubha Sehgal, Sandeep Soni, Sumit Diware, Alok Kumar Shukla, Sourajeet Roy, Rajendra Bishnoi |
ICCAD | 6 |
| 2025 | Energy-Efficient Multi-Operand XOR Logic-Based CIM Accelerator using RRAM technologyabstractRecent advances in Resistive RAM (RRAM) based Computation-In-Memory (CIM) architectures highlight significant potential for accelerating data-intensive computing tasks. However, non-idealities in RRAM devices, such as variability, result in small sensing margins that can significantly affect the computational efficiency. This issue becomes even more pronounced when dealing with complex multi-operand logic operations. This paper introduces a circuit-level scheme for CIM-based multi-operand XOR logic operations, leveraging a Voltage-To-Time converter (VTC) to perform multi-phased XORs in a single clock cycle. In this approach, we exploit bitline capacitances for voltage-based sensing during computation, generating an output voltage that is linearly proportional to the operand values. This voltage is then converted into the desired logic output using the VTC design. Furthermore, low-power techniques are employed in the deployment of sense amplifiers, such as regulating power consumption during operation and disabling the amplifiers once the decision is made. Simulation results for a post-layout extracted 512x512 (256Kb) RRAM-based CIM array show that up to 16-operand XOR operation can be accurately and reliably performed as opposed to a maximum of three operands supported by state-of-the-art solutions, while offering up to 49× better figure-of-merit combining energy-efficiency and throughput. Abhairaj Singh, Konstantinos Stavrakakis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
ICCAD | 3 |
| 2024 | A Lightweight Architecture for Real-Time Neuronal-Spike ClassificationabstractElectrophysiological recordings of neural activity in a mouse's brain are very popular among neuroscientists for understanding brain function. One particular area of interest is acquiring recordings from the Purkinje cells in the cerebellum in order to understand brain injuries and the loss of motor functions. However, current setups for such experiments do not allow the mouse to move freely and, thus, do not capture its natural behaviour since they have a wired connection between the animal's head stage and an acquisition device. In this work, we propose a lightweight neuronalspike detection and classification architecture that leverages on the unique characteristics of the Purkinje cells to discard unneeded information from the sparse neural data in real time. This allows the (condensed) data to be easily stored on a removable storage device on the head stage, alleviating the need for wires. Synthesis results reveal a >95% overall classification accuracy while still resulting in a small-form-factor design, which allows for the free movement of mice during experiments. Moreover, the power-efficient nature of the design and the usage of STT-RAM (Spin Transfer Torque Magnetic Random Access Memory) as the removable storage allows the head stage to easily operate on a tiny battery for up to approximately 4 days. Muhammad Ali Siddiqi, David Vrijenhoek, Lennart P. L. Landsmeer, Job van der Kleij, Anteneh Gebregiorgis, Vincenzo Romano, Rajendra Bishnoi, Said Hamdioui, Christos Strydis |
CF | 7 |
| 2024 | Energy-efficient SNN Architecture using 3nm FinFET Multiport SRAM-based CIM with Online LearningabstractCurrent Artificial Intelligence (AI) computation systems face challenges, primarily from the memory-wall issue, limiting overall system-level performance, especially for Edge devices with constrained battery budgets, such as smartphones, wearables, and Internet-of-Things sensor systems. In this paper, we propose a new SRAM-based Compute-In-Memory (CIM) accelerator optimized for Spiking Neural Networks (SNNs) Inference. Our proposed architecture employs a multiport SRAM design with multiple decoupled Read ports to enhance the throughput and Transposable Read-Write ports to facilitate online learning. Furthermore, we develop an Arbiter circuit for efficient data-processing and port allocations during the computation. Results for a 128×128 array in 3nm FinFET technology demonstrate a 3.1× improvement in speed and a 2.2× enhancement in energy efficiency with our proposed multiport SRAM design compared to the traditional single-port design. At system-level, a throughput of 44 MInf/s at 607 pJ/Inf and 29mW is achieved. Lucas Huijbregts, Hsiao-Hsuan Liu, Paul Detterer, Said Hamdioui, Amirreza Yousefzadeh, Rajendra Bishnoi |
DAC | 6 |
| 2024 | Hardware-Aware Quantization for Accurate Memristor-Based Neural NetworksabstractMemristor-based Computation-In-Memory (CIM) has emerged as a compelling paradigm for designing energy-efficient neural network hardware. However, memristors suffer from conductance variation issue, which introduces computational errors in CIM hardware and leads to a degraded inference accuracy. In this paper, we present a hardware-aware quantization to mitigate the impact of conductance variation on CIM-based neural networks. We achieve this using the inherent characteristics of fixed-point arithmetic in CIM hardware. By tuning the bit-precision of weights, we align the conductance variation-induced errors with lower-order output bits. This reduces their numerical impact on the fixed-point output. We further decrease the residual errors by selectively discarding bits with low information and high error. This leads to error-free computations and a high inference accuracy. Our proposed methodology achieves 5.6× correct operations per unit energy compared to the conventional approach, while incurring very low hardware overheads. Sumit Diware, Mohammad Amin Yaldagard, Rajendra Bishnoi |
ICCAD | 3 |
| 2023 | PetaOps/W edge-AI $\mu$ Processors: Myth or reality?abstractWith the rise of deep learning (DL), our world braces for artificial intelligence (AI) in every edge device, creating an urgent need for edge-AI SoCs. This SoC hardware needs to support high throughput, reliable and secure AI processing at ultra-low power (ULP), with a very short time to market. With its strong legacy in edge solutions and open processing platforms, the EU is well-positioned to become a leader in this SoC market. However, this requires AI edge processing to become at least 100 times more energy-efficient, while offering sufficient flexibility and scalability to deal with AI as a fast-moving target. Since the design space of these complex SoCs is huge, advanced tooling is needed to make their design tractable. The CONVOLVE project (currently in Inital stage) addresses these roadblocks. It takes a holistic approach with innovations at all levels of the design hierarchy. Starting with an overview of SOTA DL processing support and our project methodology, this paper presents 8 important design choices largely impacting the energy efficiency and flexibility of DL hardware. Finding good solutions is key to making smart-edge computing a reality. Manil Dev Gomony, Floran de Putter, Anteneh Gebregiorgis, Gianna Paulin, Linyan Mei, Vikram Jain, Said Hamdioui, Victor Sanchez, Tobias Grosser, Marc Geilen, Marian Verhelst, Friedemann Zenke, Frank K. Gürkaynak, Barry de Bruin, Sander Stuijk, Simon Davidson, Sayandip De, Mounir Ghogho, Alexandra Jimborean, Sherif Eissa, Luca Benini, Dimitrios Soudris, Rajendra Bishnoi, Sam Ainsworth 0001, Federico Corradi, Ouassim Karrakchou, Tim Güneysu, Henk Corporaal |
DATE | 23 |
| 2023 | Memristor-Based Lightweight EncryptionabstractNext-generation personalized healthcare devices are undergoing extreme miniaturization in order to improve user acceptability. However, such developments make it difficult to incorporate cryptographic primitives using available target tech-nologies since these algorithms are notorious for their energy consumption. Besides, strengthening these schemes against side-channel attacks further adds to the device overheads. Therefore, viable alternatives among emerging technologies are being sought. In this work, we investigate the possibility of using memristors for implementing lightweight encryption. We propose a 40-nm RRAM-based GIFT-cipher implementation using a 1TIR configuration with promising results; it exhibits roughly half the energy consumption of a CMOS-only implementation. More importantly, its non-volatile and reconfigurable substitution boxes offer an energy-efficient protection mechanism against side-channel attacks. The complete cipher takes 0.0034 mm2of area, and encrypting a 128-bit block consumes a mere 242 pJ. Muhammad Ali Siddiqi, Jan Andrés Galvan Hernández, Anteneh Gebregiorgis, Rajendra Bishnoi, Christos Strydis, Said Hamdioui, Mottaqiallah Taouil |
DSD | 4 |
| 2023 | Dependability of Future Edge-AI Processors: Pandora's BoxabstractThis paper addresses one of the directions of the HORIZON EU CONVOLVE project being dependability of smart edge processors based on computation-in-memory and emerging memristor devices such as RRAM. It discusses how how this alternative computing paradigm will change the way we used to do manufacturing test. In addition, it describes how these emerging devices inherently suffering from many non-idealities are calling for new solutions in order to ensure accurate and reliable edge computing. Moreover, the paper also covers the security aspects for future edge processors and shows the challenges and the future directions. Manil Dev Gomony, Anteneh Gebregiorgis, Moritz Fieback, Marc Geilen, Sander Stuijk, Jan Richter-Brockmann, Rajendra Bishnoi, Sven Argo, Lara Arche Andradas, Tim Güneysu, Mottaqiallah Taouil, Henk Corporaal, Said Hamdioui |
ETS | 7 |
| 2023 | On the Reliability of RRAM-Based Neural NetworksabstractEmerging device technologies such as Resistive RAMs (RRAMs) are under investigation by many researchers and semiconductor companies; not only to realize e.g., embedded non-volatile memories, but also to enable energy-efficient computing making use of new data processing paradigms such as computation-in-memory. However, such devices suffer from various non-idealities and reliability failure mechanisms (e.g., variability, endurance, and retention); these negatively impact the memory robustness and the computation accuracy. This paper discusses the non-idealities and reliability failure mechanisms for RRAM devices, provides an overview on the most popular ones. In addition, it reports detailed anlysis of some of these based on data measurements. Finally, it presents two different mitigation schemes for RRAM based accelerators; one is based on RRAM non-ideality aware quantization and conductance control for neural network accuracy enhancement while the second is based on reliability-aware biased training technique. Hassen Aziza, Cristian Zambelli, Said Hamdioui, Sumit Diware, Rajendra Bishnoi, Anteneh Gebregiorgis |
VLSI-SoC | 5 |
| 2022 | Referencing-in-Array Scheme for RRAM-based CIM ArchitectureabstractResistive random access memory (RRAM) based computation-in-memory (CIM) architectures are attracting a lot of attention due to their potential in performing fast and energy-efficient computing. However, the RRAM variability and non-idealities limit the computing accuracy of such architectures, especially for multi-operand logic operations. This paper pro-poses a voltage-based differential referencing-in-array scheme that enables accurate two and multi-operand logic operations for RRAM-based CIM architecture. The scheme makes use of a 2T2R cell configuration to create a complementary bitcell structure that inherently acts also as a reference during the operation execution; this results in a high sensing margin. More-over, the variation-sensitive multi-operand (N)AND operation is implemented using complementary-input (N)OR operation to further improve its accuracy. Simulation results for a post-layout extracted 512x512 (256Kb) RRAM-based CIM array show that up to 56 operand (N)OR/(N)AND operation can be accurately and reliably performed as opposed to a maximum of 4 operands supported by state-of-the-art solutions, while offering up to 11.4X better energy-efficiency. Abhairaj Singh, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
DATE | 2 |
| 2022 | Accelerating RRAM Testing with a Low-cost Computation-in-Memory based DFTabstractEmerging non-volatile resistive RAM (RRAM) device technology has shown great potential to cultivate not only high-density memory storage, but also energy-efficient computing units. However, the unique challenges related to RRAM fabrication process render the traditional memory testing solutions inefficient and inadequate for high product quality. This paper presents low-cost design-for-testability (DFT) solutions that augment the testing process and improve the fault coverage. A computation-in-memory (CIM) based DFT is realized to expedite the detection and diagnosis of faults by developing logic designs involving multi-row activation. A novel addressing scheme is introduced to facilitate the diagnosis of faults. Reconfigurable logic designs are developed to detect unique RRAM faults that offer features such as programmable reference generations, period, and voltage of operation. DFT implementations are validated on a post-layout extracted platform and testing sequences are introduced by incorporating the proposed DFTs. Results show that more than 2.3× speedup and better coverage are achieved with 6× area reduction when compared with state-of-the-art solutions. Abhairaj Singh, Moritz Fieback, Rajendra Bishnoi, Filip Bradaric, Anteneh Gebregiorgis, Rajiv V. Joshi, Said Hamdioui |
ITC | 3 |
| 2022 | Dealing with Non-Idealities in Memristor Based Computation-In-Memory DesignsabstractComputation-In-Memory (CIM) using memristor devices provides an energy-efficient hardware implementation of arithmetic and logic operations for numerous applications, such as neuromorphic computing and database query. However, memristor-based CIM suffers from various non-idealities such as conductance drift, read disturb, wire parasitics, endurance and device degradation. These negatively impact the computation accuracy of CIM. It is therefore essential to deal with these non-idealities and fabrication imperfections in order to harness the full potential of CIM. This paper discusses the non-ideality challenges and provides potential solutions. Furthermore, the paper outlines the potential future directions for CIM architectures. Anteneh Gebregiorgis, Abhairaj Singh, Sumit Diware, Rajendra Bishnoi, Said Hamdioui |
VLSI-SoC | 4 |
| 2022 | Defects, Fault Modeling, and Test Development Framework for RRAMsabstractResistive RAM (RRAM) is a promising technology to replace traditional technologies such as Flash, because of its low energy consumption, CMOS compatibility, and high density. Many companies are prototyping this technology to validate its potential. Bringing this technology to the market requires high-quality tests to ensure customer satisfaction. Hence, it is of great importance to deeply understand manufacturing defects and accurately model them to develop optimal tests. This paper presents a holistic framework for defect and fault modeling that enables the development of optimal tests for RRAMs. An overview and classification of RRAM manufacturing defects are provided. Defects in contacts and interconnects are modeled as resistors. Unique RRAM defects, e.g., forming defects, require Device-Aware defect modeling which incorporates the defect’s impact on the device’s electric properties by adjusting the affected technology and electrical parameters. Additionally, a systematic approach to define the fault space is presented, followed by a methodology to validate this space. With this methodology, accurate fault modeling for contact, interconnect, and forming defects is performed and tests are developed. The tests are able to detect all faults in a time-efficient manner, thereby proving the effectiveness of the framework. Finally, an outlook on future RRAM testing is presented. Moritz Fieback, Guilherme Cardoso Medeiros, Lizhou Wu, Hassen Aziza, Rajendra Bishnoi, Mottaqiallah Taouil, Said Hamdioui |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2022 | A Survey on Memory-centric Computer ArchitecturesabstractFaster and cheaper computers have been constantly demanding technological and architectural improvements. However, current technology is suffering from three technology walls: leakage wall, reliability wall, and cost wall. Meanwhile, existing architecture performance is also saturating due to three well-known architecture walls: memory wall, power wall, and instruction-level parallelism (ILP) wall. Hence, a lot of novel technologies and architectures have been introduced and developed intensively. Our previous work has presented a comprehensive classification and broad overview of memory-centric computer architectures. In this article, we aim to discuss the most important classes of memory-centric architectures thoroughly and evaluate their advantages and disadvantages. Moreover, for each class, the article provides a comprehensive survey on memory-centric architectures available in the literature. Anteneh Gebregiorgis, Hoang Anh Du Nguyen, Rajendra Bishnoi, Mottaqiallah Taouil, Francky Catthoor, Said Hamdioui |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | A Voltage-Controlled, Oscillation-Based ADC Design for Computation-in-Memory Architectures Using Emerging ReRAMsabstractConventional von Neumann architectures cannot successfully meet the demands of emerging computation and data-intensive applications. These shortcomings can be improved by embracing new architectural paradigms using emerging technologies. In particular, Computation-In-Memory (CiM) using emerging technologies such as Resistive Random Access Memory (ReRAM) is a promising approach to meet the computational demands of data-intensive applications such as neural networks and database queries. In CiM, computation is done in an analog manner; digitization of the results is costly in several aspects, such as area, energy, and performance, which hinders the potential of CiM. In this article, we propose an efficient Voltage-Controlled-Oscillator (VCO)–based analog-to-digital converter (ADC) design to improve the performance and energy efficiency of the CiM architecture. Due to its efficiency, the proposed ADC can be assigned in a per-column manner instead of sharing one ADC among multiple columns. This will boost the parallel execution and overall efficiency of the CiM crossbar array. The proposed ADC is evaluated using a Multiplication and Accumulation (MAC) operation implemented in ReRAM-based CiM crossbar arrays. Simulations results show that our proposed ADC can distinguish up to 32 levels within 10 ns while consuming less than 5.2 pJ of energy. In addition, our proposed ADC can tolerate ≈30% variability with a negligible impact on the performance of the ADC. Mahta Mayahinia, Abhairaj Singh, Christopher Bengel, Stefan Wiefels, Muath Abu Lebdeh, Stephan Menzel, Dirk J. Wouters, Anteneh Gebregiorgis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
ACM J. Emerg. Technol. Comput. Syst. | 9 |
| 2021 | Low-Power Memristor-Based Computing for Edge-AI ApplicationsabstractWith the rise of the Internet of Things (IoT), a huge market for so-called smart edge-devices is foreseen for millions of applications, like personalized healthcare and smart robotics. These devices have to bring smart computing directly where the data is generated, while coping with the limited energy budget. Conventional von-Neumann architecture fail to meet these requirements due to e.g., memory-processor data transfer bottleneck. Memristor-based computation-in-memory (CIM) has the potential to realize smart local computing for highly parallel data-dominated AI applications by exploiting the inherent properties of the architecture and the physical characteristics of the memristors. This paper provides a broad overview of CIM architecture highlighting its potential and unique properties in enabling smart local computing. Moreover, it discusses design considerations of such architectures including both crossbar array as well as peripheral circuits; special attention is given to analog-to-digital converter (ADC), as it is the most critical unit of analog-based CIM operation e.g., vector-matrix multiplication (VMM). Finally, the paper outlines the potential future directions for CIM-based edge smart computing. Abhairaj Singh, Sumit Diware, Anteneh Gebregiorgis, Rajendra Bishnoi, Francky Catthoor, Rajiv V. Joshi, Said Hamdioui |
ISCAS | 4 |
| 2021 | A Survey of Test and Reliability Solutions for Magnetic Random Access MemoriesabstractMemories occupy most of the silicon area in nowadays' system-on-chips and contribute to a significant part of system power consumption. Though widely used, nonvolatile Flash memories still suffer from several drawbacks. Magnetic random access memories (MRAMs) have the potential to mitigate most of the Flash shortcomings. Moreover, it is predicted that they could be used for DRAM and SRAM replacement. However, they are prone to manufacturing defects and runtime failures as any other type of memory. This article provides an up-to-date and practical coverage of MRAM test and reliability solutions existing in the literature. After some background on existing MRAM technologies, defectiveness and reliability issues are discussed, as well as functional fault models used for MRAM. This article is dedicated to a summarized description of existing test and reliability improvement methods developed so far for various MRAM technologies. The last part of this article gives some perspectives on this hot topic. Patrick Girard 0001, Yuanqing Cheng, Arnaud Virazel, Wei Zhao 0010, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
Proc. IEEE | 5 |
| 2021 | SRIF: Scalable and Reliable Integrate and Fire Circuit ADC for Memristor-Based CIM ArchitecturesabstractEmerging computation-in-memory (CIM) paradigm offers processing and storage of data at the same physical location, thus alleviating critical memory-processor communication bottlenecks suffered by conventional von-Neumann architecture. Storage of data in a CIM architecture is analog in nature and therefore computation is performed in analog domain i.e. inputs and outputs are analog values. Since the outside computing environment is digital, analog-to-digital converters (ADC) are utilized to perform the output data conversion. However, ADC designs are bulky, power-hungry circuits that are prone to design variations and therefore, play an important role in determining the computing efficiency of CIM architectures. In this paper, we present a scalable and reliable integrate and fire circuit ADC (SRIF-ADC) design for CIM architectures, suitable for stringent power and area constraints. We devise a technique to stabilize the node receiving analog inputs that allows more rows to be activated at the same time, thereby increasing the operand size of input vectors. This allows better scalability in terms of higher parallelism of operations. We employ a self-timed variation-aware design approach and design measures to drastically reduce read disturb of memristor devices that address reliability issues related to the ADC design. In addition, we present a compact, built-in sample-and-hold circuit to replace the large-sized capacitance and built-in weighting technique to alleviate the need for post-processing. For multiply-and-accumulate (MAC) operation, our simulation results show that we can improve the computational parallelism by 3X as well as ADC conversion speed and energy efficiency are improved by 2X and 11.6X, respectively, compared to the state-of-the-art design. Abhairaj Singh, Muath Abu Lebdeh, Anteneh Gebregiorgis, Rajendra Bishnoi, Rajiv V. Joshi, Said Hamdioui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2020 | Tolerating Retention Failures in Neuromorphic Fabric based on Emerging Resistive MemoriesabstractIn recent years, computation is shifting from conventional high performance servers to Internet of Things (IoT) edge devices, most of which require the processing of cognitive tasks. Hence, a great effort is put in the realization of neural network (NN) edge devices and their efficiency in inferring a pretrained Neural Network. In this paper, we evaluate the retention issues of emerging resistive memories used as non-volatile weight storage for embedded NN. We exploit the asymmetric retention behavior of Spintronic based Magnetic Tunneling Junctions (MTJs), which is also present in other resistive memories like Phase-Change memory (PCM) and ReRAM, to optimize the retention of the NN accuracy over time. We propose mixed retention cell arrays and an adapted training scheme to achieve a trade-off between array size and the reliable long-term accuracy of NNs. The results of our proposed method save up to 24% of inference accuracy of an MNIST trained Multi-Layer-Perceptron on MTJ-based crossbars. Christopher Münch, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ASP-DAC | 2 |
| 2020 | Dynamic Faults based Hardware Trojan Design in STT-MRAMabstractThe emerging Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is seen as a promising candidate to replace conventional on-chip memories. It has several advantages such as high density, non-volatility, scalability, and CMOS compatibility. With this technology becoming ubiquitous, it also becomes interesting as a target for security attacks. As the fabrication process of STT-MRAM evolves, it is susceptible to various fault mechanisms which are different from those of conventional CMOS memories. These unique fault mechanisms can be exploited by an adversary to deploy hardware Trojans, which are deliberately introduced design modifications. In this work, we demonstrate how a particular stealthy circuit modification to inject a fault mechanism, namely dynamic fault, can be exploited to implement a hardware Trojan trigger which cannot be detected by standard memory testing methods. The fault mechanisms can also be used to design new payloads specific to STT-MRAM. We illustrate this by proposing a new payload by utilizing coupling faults, which leads to degraded performance and data corruption. Sarath Mohanachandran Nair, Rajendra Bishnoi, Arunkumar Vijayan, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2020 | A Universal Spintronic Technology based on Multifunctional Standardized StackabstractThe goal of the GREAT RIA project is to cointegrate multiple functions like sensors ("Sensing"), RF emitters or receivers ("Communicating") and logic/memory ("Process- ing/Storing") together within CMOS technology by adapting the Spin-Transfer Torque Magnetic Tunnel Junction (STT-MTJ), elementary constitutive cell of the MRAM memories, to a single baseline technology. Based on the STT unique set of performances (non-volatility, high speed, infinite endurance and moderate read/write power), GREAT will achieve the same goal as heterogeneous integration of devices but in a much simpler way. This will lead to a unique STT-MTJ cell technology called Multifunctional Standardized Stack (MSS). This paper presents the lessons learned in the project from the technology, compact modeling, process design kit, standard cells, as well as memory and system level design evaluation and exploration. The proposed technology and toolsets are giant leaps towards heterogeneous integrated technology and architectures for IoT. Mehdi Baradaran Tahoori, Sarath Mohanachandran Nair, Rajendra Bishnoi, Lionel Torres, Sophiane Senni, Guillaume Patrigeon, Pascal Benoit, Gregory di Pendina, Guillaume Prenat |
DATE | 3 |
| 2020 | Testing Scouting Logic-Based Computation-in-Memory ArchitecturesabstractToday's von Neumann computing systems are facing major challenges making them not suitable for evolving ultralow power (e.g., edge computing) applications. Therefore, alternative architectures that make use of post-CMOS devices are under investigation. One of these architectures is computation-in-memory (CIM) based on memristive devices; it performs (parallel) computing within the memory core, which prevents data-movement and results in low energy consumption, at the cost of some modification in memory design. Hence, a CIM die can work either in memory configuration or in computation configuration. One implementation of this architecture is based on Scouting logic; it allows the execution of logic operations within the memory. This paper discusses fault modeling and testing of CIM architectures, applied to a Scouting logic-based architecture. It demonstrates that unique faults can occur in the CIM die while in the computation configuration, and that these faults cannot be detected by just testing the CIM die in the memory configuration, thus leading to test escapes. The paper demonstrates how an efficient test can be developed that detects all faults in both configurations. Moreover, it shows that testing the die in the computation configuration reduces the overall test time while improving the outgoing product quality. Moritz Fieback, Surya Nagarajan, Rajendra Bishnoi, Mehdi Baradaran Tahoori, Mottaqiallah Taouil, Said Hamdioui |
ETS | 3 |
| 2020 | Special Session - Emerging Memristor Based Memory and CIM Architecture: Test, Repair and Yield AnalysisabstractEmerging memristor-based architectures are promising for data-intensive applications as these can enhance the computation efficiency, solve the data transfer bottleneck and at the same time deliver high energy efficiency using their normally-off/instant-on attributes. However, their storing devices are more susceptible to manufacturing defects compared to the traditional memory technologies because they are fabricated with new materials and require different manufacturing processes. Hence, in order to ensure correct functionalities for these technologies, it is necessary to have accurate fault modeling as well as proper test methodologies with high test coverage. In this paper, we propose technology specific cell-level defect modeling, accurate fault analysis and yield improvement solutions for memristor-based memory as well as Computation-In-Memory (CIM) architectures. Our overall contributions cover three abstraction levels, namely, device, architecture and system. First, we propose a device-aware test methodology in which we have introduced a key device-level characteristic to develop accurate defect model. Second, we demonstrate a yield analysis framework for memristor arrays considering reliability and permanent faults due to parametric variations and explore fault-tolerant solutions. Third, a lightweight on-line test and repair schemes is proposed for emerging CIM devices in machine learning applications. Rajendra Bishnoi, Lizhou Wu, Moritz Fieback, Christopher Münch, Sarath Mohanachandran Nair, Mehdi Baradaran Tahoori, Ying Wang 0001, Huawei Li 0001, Said Hamdioui |
VTS | 1 |
| 2020 | Mitigating Read Failures in STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is an emerging non-volatile memory technology, as a leading candidate to replace conventional on-chip memories due to its various advantages such as high density, non-volatility, scalability, high endurance and CMOS compatibility. However, read and write operations in STT-MRAM are extremely vulnerable to manufacturing variations. In particular, the read operation is becoming more susceptible to failures since the read timing and read-disturb failures have conflicting requirements of read period. To overcome this issue, we propose a technique to reduce the read period without sacrificing the target reliability requirements. The reduced read period, in turn, results in improved read performance and reduced read-disturb rates. The results show that using this technique, the read period can be reduced by 50%, and the read-disturb probability by 51%. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
VTS | 2 |
| 2020 | Crossover-aware Placement and Routing for Inkjet Printed CircuitsabstractPrinted Electronics technology is a key-enabler for smart sensors, soft robotics, and wearables. The inkjet printed electrolyte-gated field effect transistor (EGFET) technology is a promising candidate for such applications due to its low-power operation, high field-effect mobility, and on-demand fabrication. Unlike conventional silicon-based technologies, inkjet printed electronics technology is an additive manufacturing process where multiple layers are printed on top of each other to realize functional devices such as transistors and their interconnections. Due to the additive manufacturing process, the technology has limited routing layers. For routing of complex circuits, insulating crossovers are printed at the intersection of routing paths to isolate them. The crossover can alter the electrical properties of a circuit based on specific location on a routing path. In this work, we propose a crossover-aware placement and routing (COPnR) methodology for inkjet-printed circuits by integrating the crossover constraints in our design framework. Our proposed placement methodology is based on a state-of-the-art evolutionary algorithm while the routing optimization is done using a genetic algorithm. The proposed methodology is compared with the industrial standard placement and routing (PnR) tools. On average, the proposed methodology has 38% fewer crossovers and 94% fewer failing paths compared to the industrial PnR tools applied to printed circuit designs. Farhan Rasheed, Michael Hefenbrock, Rajendra Bishnoi, Michael Beigl, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2020 | Approximate Spintronic MemoriesabstractVarious applications, such as multimedia, machine learning, and signal processing, have a significant intrinsic error resilience. This makes them preferable for approximate computing as they have the ability to tolerate computations and data errors along with producing acceptable outputs. From the technology perspective, emerging technologies with inherent non-determinism and high failure rates are candidates for the realization of approximate computing. Spin Transfer Torque Magnetic Random Access Memories (STT-MRAM) is an emerging non-volatile memory technology and a potential candidate to replace SRAM due to its high density, scalability, and zero-leakage. The write operation in this technology is inherently stochastic and increases the rate of write errors. Moreover, this technology is associated with other failure mechanisms such as read-disturb and failures due to data retention. These errors are highly dependent on the STT-MRAM parameters (i.e., thermal stability, read/write current, and read/write latency), which varies with the operating temperature and the process variation effects. Fast and energy-efficient STT-MRAM designed for on-chip memories can be easily achieved by relaxing the device parameters at the cost of increased error rate, which can be addressed by approximating memory accesses. In this work, a detailed study of reliability and gains (i.e., performance and energy) tradeoff at the device and system-level of the STT-MRAM-based data cache system is presented in the scope of approximate memories. Nour Sayed, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | Secure STT-MRAM Bit-Cell Design Resilient to Differential Power Analysis AttacksabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising nonvolatile memory technology for various applications from low power to high-density memory. However, STT-MRAM is prone to power analysis attacks due to its asymmetric resistive states and switching behavior. This noninvasive class of attacks is a serious threat to system security. To reduce the correlation between the data and the power consumption of the memory, a countermeasure based on a resilient cell design with a symmetrical structure is proposed in this article. The standard cell and the proposed cell have been attacked and their resiliencies are compared. When attacked with a correlation power analysis (CPA), the proposed bit cell attacked with the hamming weight (HW) model is 100 times more resilient compared to the standard STT-MRAM cell. The proposed bit cell attacked with the hamming distance (HD) or the STT power models is more than 500 times more resilient compared to the standard STT-MRAM cell. Samir Ben Dodo, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | A Compact Low-Voltage True Random Number Generator Based on Inkjet Printing TechnologyabstractPrinted electronics (PE) is a fast-growing field with promising applications in wearables, smart sensors, and smart cards, since it provides mechanical flexibility, and low-cost, on-demand, and customizable fabrication. To secure the operation of these applications, true random number generators (TRNGs) are required to generate unpredictable bits for cryptographic functions and padding. However, since the additive fabrication process of the PE circuits results in high intrinsic variations due to the random dispersion of the printed inks on the substrate, constructing a printed TRNG is challenging. In this article, we exploit the additive customizable fabrication feature of inkjet printing to design a TRNG based on electrolyte-gated field-effect transistors (EGFETs). We also propose a printed resistor tuning flow for the TRNG circuit to mitigate the overall process variation of the TRNG so that the generated bits are mostly based on the random noise in the circuit, providing a true random behavior. The simulation results show that the overall process variation of the TRNGs is mitigated by 110 times, and the generated bitstream of the tuned TRNGs passes the National Institute of Standards and Technology - Statistical Test Suite. For the proof of concept, the proposed TRNG circuit was fabricated and tuned. The characterization results of the tuned TRNGs prove that the TRNGs generate random bitstreams at the supply voltage of down to 0.5 V. Hence, the proposed TRNG design is suitable to secure low-power applications in this domain. Ahmet Turan Erozan, Guan Ying Wang, Rajendra Bishnoi, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2020 | A Novel Printed-Lookup-Table-Based Programmable Printed Digital CircuitabstractAdvances in printed electronics (PE) enables new applications, particularly in ultra-low-cost domains. However, achieving high-throughput printing processes and manufacturing yield is one of the major challenges in the large-scale integration of PE technology. In this article, we present a programmable printed circuit based on an efficient printed lookup table (pLUT) to address these challenges by combining the advantages of the high-throughput advanced printing and maskless point-of-use final configuration printing. We propose a novel pLUT design which is more efficient in PE realization compared to existing LUT designs. The proposed pLUT design is simulated, fabricated, and programmed as different logic functions with inkjet printed conductive ink to prove that it can realize digital circuit functionality with the use of programmability features. The measurements show that the fabricated LUT design is operable at 1 V. Ahmet Turan Erozan, Dennis Weller, Farhan Rasheed, Rajendra Bishnoi, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Reliable in-memory neuromorphic computing using spintronicsabstractRecently Spin Transfer Torque Random Access Memory (STT-MRAM) technology has drawn a lot of attention for the direct implementation of neural networks, because it offers several advantages such as near-zero leakage, high endurance, good scalability, small foot print and CMOS compatibility. The storing device in this technology, the Magnetic Tunnel Junction (MTJ), is developed using magnetic layers that requires new fabrication materials and processes. Due to complexities of fabrication steps and materials, MTJ cells are subject to various failure mechanisms. As a consequence, the functionality of the neuromorphic computing architecture based on this technology is severely affected. In this paper, we have developed a framework to analyze the functional capability of the neural network inference in the presence of the several MTJ defects. Using this framework, we have demonstrated the required memory array size that is necessary to tolerate the given amount of defects and how to actively decrease this overhead by disabling parts of the network. Christopher Münch, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ASP-DAC | 2 |
| 2019 | Inkjet-Printed True Random Number Generator based on Additive Resistor TuningabstractPrinted electronics (PE) is a fast growing technology with promising applications in wearables, smart sensors and smart cards since it provides mechanical flexibility, low-cost, on-demand and customizable fabrication. To secure the operation of these applications, True Random Number Generators (TRNGs) are required to generate unpredictable bits for cryptographic functions and padding. However, since the additive fabrication process of PE circuits results in high intrinsic variation due to the random dispersion of the printed inks on the substrate, constructing a printed TRNG is challenging. In this paper, we exploit the additive customizable fabrication feature of inkjet printing to design a TRNG based on electrolyte-gated field effect transistors (EGFETs). The proposed memory-based TRNG circuit can operate at low voltages (≤ 1 V ), it is hence suitable for low-power applications. We also propose a flow which tunes the printed resistors of the TRNG circuit to mitigate the overall process variation of the TRNG so that the generated bits are mostly based on the random noise in the circuit, providing a true random behaviour. The results show that the overall process variation of the TRNGs is mitigated by 110 times, and the simulated TRNGs pass the National Institute of Standards and Technology Statistical Test Suite. Ahmet Turan Erozan, Rajendra Bishnoi, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2019 | Predictive Modeling and Design Automation of Inorganic Printed ElectronicsabstractPrinted Electronics is perceived to have a major impact in the fields of smart sensors, Internet of Things and wearables. Especially low power printed technologies such as electrolyte gated field effect transistors (EGFETs) using solution-processed inorganic materials and inkjet printing are very promising in such application domains. In this paper, we discuss a modeling approach to describe the variations of printed devices. Incorporating these models and design flows into our previously developed printed design system allows for robust circuit design. Additionally, we propose a reliability-aware routing solution for printed electronics technology based on the technology constraints in printing crossovers. The proposed methodology was validated on multiple benchmark circuits and can be easily integrated with the design automation tools-set. Farhan Rasheed, Michael Hefenbrock, Rajendra Bishnoi, Michael Beigl, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
DATE | 3 |
| 2019 | Variation-aware Fault Modeling and Test Generation for STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) offers high density, non-volatility, scalability, high endurance and CMOS compatibility, making it a promising non-volatile memory (NVM) technology. However, due to the unique magnetic fabrication processes, different bit-cell architecture and periphery circuitry, they are susceptible to different manufacturing defects and faults compared to conventional CMOS-based memories. In this paper, a detailed variation-aware defect injection is performed based on the magnetic devices and layout characteristics of STT-MRAM and unique fault models are constructed for these memories. Based on the derived fault models and behaviors, efficient test algorithms are developed to fully cover these faults. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori, Hayk T. Grigoryan, Grigor Tshagharyan |
IOLTS | 2 |
| 2019 | A Comprehensive Reliability Analysis Framework for NTC Caches: A System to Device ApproachabstractNear threshold computing (NTC) has significant role in reducing the energy consumption of modern very large scale integrated circuits designs. However, NTC designs suffer from functional failures and performance loss. Understanding the characteristics of the functional failures and variability effects is of decisive importance in order to mitigate them, and get the utmost NTC benefits. This paper presents a comprehensive cross-layer reliability analysis framework to assess the effect of soft error, aging, and process variation in the operation of near threshold voltage caches. The objective is to quantify the reliability of different SRAM designs, evaluate voltage scaling potential of caches, and to find a reliability-performance optimal cache organization for an NTC microprocessor. In this paper, the soft error rate (SER) and static noise margin (SNM) of 6T and 8T SRAM cells and their dependencies on aging and process variation are investigated by considering device, circuit, and architecture level analysis. Their experimental results show that in NTC, process variation and aging-induced SNM degradation is 2.5× higher than in the super threshold domain while SER is 8× higher. At NTC, the use of 8T instead of 6T SRAM cells can reduce the system-level SNM and SER by 14% and 22%, respectively. Besides, we observe that we can find the right balance between performance and reliability by using an appropriate cache organization at NTC which is different from the super threshold. Anteneh Gebregiorgis, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Compiler-Assisted and Profiling-Based Analysis for Fast and Efficient STT-MRAM On-Chip Cache DesignabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising candidate for large on-chip memories as a zero-leakage, high-density and non-volatile alternative to the present SRAM technology. Since memories are the dominating component of a System-on-Chip, the overall performance of the system is highly dependent on that memories. Nevertheless, the high write energy and latency of the emerging STT-MRAM are the most challenging design issues in a modern computing system. By relaxing the non-volatility of these devices, it is possible to reduce the write energy and latency costs, at the expense of reducing the retention time, which in turn may lead to loss of data. In this article, we propose a hybrid STT-MRAM design for caches with different retention capabilities. Then, based on the application requirements (i.e., execution time and memory access rate), program data layout is re-arranged at compilation time for achieving fast and energy-efficient hybrid STT-MRAM on-chip memory design with no reliability degradation. The application requirements have been defined at function granularity based on profiling and compiler-level analysis, which estimate the required retention time and memory access rate, respectively. Experimental results show that the proposed hybrid STT-MRAM cache combined with profiling-based and compiler-level analysis for the data re-arranging, on average, reduces the write energy per access by 49.7%. At system level, overall static and dynamic energy of the cache are reduced by 8.1% and 44%, respectively, whereas, the system performance has been improved up to 8.1%. Nour Sayed, Longfei Mao, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | A Spintronics Memory PUF for Resilience Against Cloning CounterfeitabstractWith the widespread use of electronic devices in embedding or processing sensitive information, new hardware security primitives have emerged to improve the shortcomings of traditional secure data storage. One of these solutions, the physically unclonable function (PUF), which is widely used for authentication and cryptography applications, extracts the unique and unclonable value from the circuit physical properties. However, a cloning attack on memory-based PUF has been demonstrated from the circuit back-side tampering, questioning its unclonable property and opening counterfeiting vulnerability in an untrusted supply chain. Spin transfer torque-magnetic random-access memory (STT-MRAM) is a promising technology due to its nonvolatility, scalability, and CMOS compatibility, therefore envisioned to be used in many embedded secure devices. In this paper, we reveal the vulnerability of existing STT-MRAM PUF solutions to back-side attacks by modeling different levels of tampering and their impact at electrical and logical levels. We propose a tamper resilient methodology for the STT-MRAM PUF design based on its switching properties. The resilience of the proposed solution to different levels of tampering and a 100% detection are confirmed with simulation results. Samir Ben Dodo, Rajendra Bishnoi, Sarath Mohanachandran Nair, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | A Comprehensive Framework for Parametric Failure Modeling and Yield Analysis of STT-MRAMabstractThe spin-transfer torque magnetic random access memory (STT-MRAM) is an emerging memory technology with several distinctive advantages such as nonvolatility, high density, scalability, and almost unlimited endurance. It is, therefore, seen as a promising candidate to replace conventional on-chip memory technologies. However, as the technology scales, yield loss due to extreme parametric variations is becoming increasingly important for STT-MRAM because of its higher sensitivity to process variation as compared to CMOS memories. In addition, the parametric variations in STT-MRAM exacerbate its stochastic switching behavior, leading to both test time fails and reliability failures in the field. Since an STT-MRAM memory array consists of both CMOS and magnetic components, the system-level failures in STT-MRAM depend on variations in both these components. In this paper, we model the system-level parametric failures of STT-MRAM considering the spatial correlation among bit cells as well as the impact of peripheral components. The proposed approach provides realistic fault distribution maps and equips the designer to investigate the efficacy of different combinations of defect tolerance techniques for an effective design-for-yield exploration. The results show that the fault distribution and yield depend on the correlation coefficient and the temperature, which, in turn, determine the correct choice of defect tolerance scheme to be adopted to mitigate them to improve the yield. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2019 | Fast and Reliable STT-MRAM Using Nonuniform and Adaptive Error Detecting and Correcting SchemeabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is an emerging nonvolatile memory technology and a potential candidate to replace CMOS-based on-chip memories. However, the bit-cell switching behavior is stochastic, which is further exacerbated due to the temperature and the process variation (PV) effects, leading to reliability failures. The conventional solution to mitigate such errors is to define the write margin based on the worst case conditions, however, this results in an excessive margin that prevents the usage of STT-MRAM in fast memories. This paper proposes to significantly reduce the write margin of STT-MRAM with no impact on the read latency, which is achieved by clustering the cache lines based on the manufactured STT-MRAM parameters and then exploiting the proposed opportunistic write approach (i.e., terminating the write process before all bit switchings are completed). This approach is supported by a novel Lazy error correcting code (Lazy-ECC), which is based on the fact that error detection is much faster than correction. Hence, the errors can be detected quickly and all erroneous data can be reverted before they arrive critical parts of the system (e.g., commit stage or memory ports). A nonuniform adaptive ECC approach to manage PV and temperature-dependent retention and read disturb failures at runtime has also been proposed. The proposed approach enables a reliable use of STT-MRAM technology for fast cache applications. Nour Sayed, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Process variation and temperature aware adaptive scrubbing for retention failures in STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is an emerging memory technology, which is seen as a promising replacement for CMOS based on-chip memories. It has several distinctive advantages such as nonvolatility, high endurance, high density, CMOS compatibility and scalability among others. However, retention failure has emerged as a major reliability concern for this technology due to the large variations in retention time because of process variations and temperature effects. The conventional solution to mitigate retention failures is to use scrubbing at regular intervals to prevent accumulation of errors, based on the worst case retention time of the memory array. But this leads to large performance and energy overheads. In this work, we propose a process variation and temperature aware scrubbing technique, where we cluster the cache lines into different groups based on their retention times and use different scrubbing intervals for each of these groups. In addition, the scrubbing interval is adjusted at run-time based on the operating temperature, to guarantee target error rate requirements. Our results show that for a 512KB cache, a group size of 4 can reduce the performance and dynamic energy overheads of scrubbing by 97%, under the same error rate constraint. Nour Sayed, Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ASP-DAC | 3 |
| 2018 | Spintronic normally-off heterogeneous system-on-chip designabstractOne of the major challenges in device down-scaling is the increase in the leakage power, which becomes a major component in the overall system power consumption. One way to deal with this problem is to introduce the concept of normally-off instant-on computing architectures, in which the system components are powered off when they are not active. An associated challenge is the back-up and restoration of system states, which in turn can introduce additional costs that erode some of the gains. A promising alternative is the use of non-volatile storage elements in the System-on-Chip (SoC) design which can instantly power-down and retain their values. In this work, we show how we can design a normally-off SoC by exploiting non-volatile latches, flip-flops and registers. The idea is to design a hybrid architecture containing conventional CMOS bistables as well as different flavors of spintronic-based non-volatile storage elements, to balance performance, area, and energy efficiency. Anteneh Gebregiorgis, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2018 | Multi-bit non-volatile spintronic flip-flopabstractAs leakage increases proportionally with the technology downscaling, it becomes extremely challenging to manage to meet the total power budget. This is because, CMOS-based logic blocks can not be completely power-gated as their flip-flops always require a retention supply to hold the system states. Alternatively, their data can be stored in a separate memory during the standby mode, however, that results in a huge area and energy overhead. Spin Transfer Torque (STT) based nonvolatile flip-flops can offer normally-off/instant-on computing features to reduce leakage by complete power shut-down without the need to transfer and restore system states separately. The non-volatile component of such flip-flops can be easily shared for the overall design optimizations. In this paper, we design a unique multi-bit non-volatile flip-flop architecture using STT devices to reduce the area and energy costs associated with nonvolatile components. This architecture is developed based on the resource sharing principle using a custom design that enables the optimization for the area and energy consumption. Moreover, we have developed a framework in which we have replaced the conventional neighbor flipflops in the layout with our proposed multi-bit non-volatile designs. Results show that using our multi-bit flip-flop architecture, we improve the system-level area and energy by 26% and 14% in average, respectively, compared to the standard single-bit non-volatile flip-flop design. Christopher Münch, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2018 | Parametric failure modeling and yield analysis for STT-MRAMabstractThe emerging Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising candidate to replace conventional on-chip memory technologies due to its advantages such as non-volatility, high density, scalability and unlimited endurance. However, as the technology scales, yield loss due to extreme parametric variations is becoming a major challenge for STT-MRAM because of its higher sensitivity to process variations as compared to CMOS memories. In addition, the parametric variations in STT-MRAM exacerbates its stochastic switching behavior, leading to both test time fails and reliability failures in the field. Since an STT-MRAM memory array consists of both CMOS and magnetic components, it is important to consider variations in both these components to obtain the failures at the system level. In this work, we model the parametric failures of STT-MRAM at the system level considering the correlation among bit-cells as well as the impact of peripheral components. The proposed approach provides realistic fault distribution maps and equips the designer to investigate the efficacy of different combinations of defect tolerance techniques for an effective design-for-yield exploration. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2018 | A cross-layer adaptive approach for performance and power optimization in STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising candidate as a universal on-chip memory technology due to non-volatility, high density and scalability. However, high write energy and latency are major challenges in this memory technology due to the asymmetry and stochastic nature of the write operation. Typically, the write current is set for the minimum energy point, which can further impact the write latency. To mitigate these issues, we propose an adaptive write current scaling technique that adjusts the write current, and hence the write latency and energy based on the performance needs at run-time. Using this technique, optimal energy and performance points for write current are obtained using detailed device and system level analysis. Furthermore, we use runtime adaptation of write current by predicting the write access rate for the next execution phase. We evaluate the efficiency of the proposed approach on SPEC2000 applications for STT-MRAM-based L1 and L2-cache levels. The results show that the effective write latency of L1 and L2 is reduced by 52.4% and 55.7% with 7.6% and 1.4% area overheads, respectively, corresponding to the overall system performance optimization of 15.5% while the total memory energy consumption is increasing by only 3.2%. Nour Sayed, Rajendra Bishnoi, Fabian Oboril, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2018 | Using multifunctional standardized stack as universal spintronic technology for IoTabstractFor monolithic heterogeneous integration, fast yet low-power processing and storage, and high integration density, the objective of the EU GREAT project is to co-integrate multiple digital and analog functions together within CMOS by adapting the Magnetic Tunneling Junctions (MTJs) into a single baseline technology enabling logic, memory, and analog functions, particularly for Internet of Things (IoT) platforms. This will lead to a unique STT-MTJ cell technology called Multifunctional Standardized Stack (MSS). This paper presents the progress in the project from the technology, compact modeling, process design kit, standard cells, as well as memory and system level design evaluation and exploration. The proposed technology and toolsets are giant leaps towards heterogeneous integrated technology and architectures for IoT. Mehdi Baradaran Tahoori, Sarath Mohanachandran Nair, Rajendra Bishnoi, Sophiane Senni, Jad Mohdad, Frédérick Mailly, Lionel Torres, Pascal Benoit, Abdoulaye Gamatié, Pascal Nouet, Frederic Ouattara, Gilles Sassatelli, Kotb Jabeur, Pierre Vanhauwaert, A. Atitoaie, I. Firastrau, Gregory di Pendina, Guillaume Prenat |
DATE | 3 |
| 2018 | Defect injection, Fault Modeling and Test Algorithm Generation Methodology for STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising alternative technology for on-chip memories due to several advantages such as high density, non-volatility, scalability, high endurance and CMOS compatibility. Due to the emerging fabrication processes of magnetic layers in the fabrication of STTMRAM devices, they are more susceptible to manufacturing defects and have different failure mechanisms compared to conventional CMOS memories. This mandates specific fault modeling and development of proper test algorithms for high coverage testing of these memories in order to ensure correct functionality in the field. In this paper, a detailed defect injection is performed based on the magnetic devices and layout characteristics of STT-MRAM and unique fault models are constructed for these memories. Based on the derived fault models and behaviors, efficient test algorithms are developed to fully cover these faults. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mehdi Baradaran Tahoori, Grigor Tshagharyan, Hayk T. Grigoryan, Gurgen Harutunyan, Yervant Zorian |
ITC | 2 |
| 2018 | Modeling and Testing of Aging Faults in FinFET Memories for Automotive ApplicationsabstractAutomotive has become one of the most prevailing sectors of the modern semiconductor industry. Due to strict requirements for safety, reliability, and security the proposed test & repair solutions for automotive applications undergo a circumstantial verification before exploitation. Traditionally production defects and soft errors occurring in the operation mode were considered to be the main source of failures for System-on-Chips (SoC). Nevertheless, aging-induced faults especially in modern technology nodes also pose certain challenges for SoC lifetime and impact the overall Failure in Time (FIT) rate of the system. Meanwhile keeping hold of low FIT rate is one of the main criteria for reliability. In this paper, a comprehensive study on aging faults is conducted for FinFET memories and an efficient test & repair methodology is proposed that meets the automotive requirements and allows decreasing the FIT rate of the system. Grigor Tshagharyan, Gurgen Harutunyan, Yervant Zorian, Anteneh Gebregiorgis, Mohammad Saber Golanbari, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ITC | 6 |
| 2018 | VAET-STT: Variation Aware STT-MRAM Analysis and Design Space Exploration ToolabstractSpin transfer torque magnetic random access memory is a promising candidate to replace CMOS based on-chip memories due to its advantages, such as nonvolatility, high density, and scalability. However, its stochastic switching and higher sensitivity to process variation compared to CMOS memories can significantly affect its performance, energy, and reliability. Although a few works exist which analyze the impact of process variation at the bit-cell level, such analysis at the system-level is missing. We have bridged this gap by developing a tool which can quantify the effect of stochasticity and process variations from the cell level to the overall memory system. The tool can perform a variation-aware design space exploration and memory configuration optimization for energy or performance while meeting reliability constraints. It also reports various failure rates and can evaluate the effectiveness of different error correcting code schemes. The results show that our framework can provide more realistic margins and the optimized variation-aware memory configuration could be significantly different from the conventional framework. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mohammad Saber Golanbari, Fabian Oboril, Fazal Hameed, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2018 | Inkjet-Printed EGFET-Based Physical Unclonable Function - Design, Evaluation, and FabricationabstractPrinted electronics (PE) is a promising technology that provides mechanical flexibility and low-cost fabrication and the key enabler for emerging applications, such as smart sensors, wearables, and Internet of Things. To use printed batteries or printed energy harvesters in the future, electrolyte-gated field-effect transistors (EGFETs) based on inorganic materials enable printed circuits requiring small supply voltage and low power. Since these applications need secure communication and/or authentication, it is imperative to embed security primitives for cryptographic key and identification purposes into the applications. Physical unclonable functions (PUFs) have been adopted widely to provide secure keys. In this paper, we present the design, simulation, fabrication, and measurements of a PUF based on EGFETs using inorganic inkjet PE. A comprehensive framework, including Monte Carlo simulations calibrated on real device measurements, is developed. Moreover, a multibit PE-PUF design is proposed to optimize area usage. Our simulation results show that the PE-PUF has ideal uniqueness (50.1%) and good reliability (89%). In addition, the proposed multibit PE-PUF reduces the area usage around 30%. The proposed PE-PUF was fabricated and the experimental results confirm that the PE-PUF can operate reliably as low as 0.5 V, and hence, it is a remarkable candidate to be utilized in low-power applications. Ahmet Turan Erozan, Gabriel Cadilha Marques, Mohammad Saber Golanbari, Rajendra Bishnoi, Simone Dehm, Jasmin Aghassi-Hagmann, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | VAET-STT: A variation aware estimator tool for STT-MRAM based memoriesabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising candidate to replace CMOS based on-chip memories due to its advantages such as non-volatility, high density and scalability. However, its stochastic switching and higher sensitivity to process variation compared to CMOS memories can significantly affect its performance, energy and reliability. Although a few works exist which analyze the impact of process variation at the bit-cell level, such analysis at the system level is missing. We have bridged this gap in our work. Specifically, we quantify the effect of stochasticity and process variations from the cell-level to the overall memory system and perform a variation-aware memory configuration optimization for energy or performance while meeting reliability constraints. Our system-level variation-aware framework has been built on top of the well-known NVSim engine. The results show that our framework can provide more realistic margins and the optimized variation-aware memory configuration could be significantly different from the conventional framework. Sarath Mohanachandran Nair, Rajendra Bishnoi, Mohammad Saber Golanbari, Fabian Oboril, Mehdi Baradaran Tahoori |
DATE | 2 |
| 2017 | Opportunistic write for fast and reliable STT-MRAMabstractDue to the stochastic switching behavior of the bit-cell in Spin Transfer Torque Magnetic Random Access Memory (STT-MRAM), an excessive write margin is required to guarantee an acceptable level of reliability and yield. This prevents the usage of STT-MRAM in fast memories such as L1 or L2 caches. The excessive write margin of STT-MRAM can be reduced to a large extent by an opportunistic write (i.e., terminating the write process before all bit switchings are completed) and by reducing thermal stability factor. The bits with unfinished writes have to be processed by robust Error Correction Codes (ECCs). However, such coding schemes have relatively large decoding latencies, which increases the overall read latency significantly. Moreover, thermally induced retention failures can limit the applicability of such schemes. In this paper, we exploit the fact that error detection is much faster than correction. Therefore, the errors can be detected quickly and all erroneous data can be reverted before they arrive critical parts of the system (e.g., commit stage or memory ports). We also provide an adaptive approach to manage temperature-dependent retention failures at runtime. Hence, our proposed approach enables the use of STT-MRAM technology for fast cache applications. Nour Sayed, Mojtaba Ebrahimi, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
DATE | 3 |
| 2017 | Exploiting STT-MRAM for approximate computingabstractSpin Transfer Torque Magnetic RAM (STT-MRAM) is an emerging non-volatile memory technology and a potential candidate to replace SRAM in processor caches. However, STT-MRAM suffers from a high write latency and high write energy consumption which have to be addressed for energy-efficient on-chip caches. The non-volatility property of STT-MRAM can be relaxed by reducing the thermal stability factor to improve both the write latency and write energy of STT-MRAM. However, this leads to increase in retention failure and read disturb rates resulting in erroneous data stored in the cache. This problem can naturally be mitigated in the scope of approximate computing in which such errors can be tolerated at the application level. In this paper, we show how STT-MRAM technology can effectively be used for approximate computing by tuning technology and application parameters to achieve an acceptable level of correctness with significant gains. Results show that using our proposed approximate computing framework, the per-access write latency and energy can be improved up to 25% and 70%, respectively. Nour Sayed, Fabian Oboril, Azadeh Shirvanian, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
ETS | 4 |
| 2017 | Leveraging Systematic Unidirectional Error-Detecting Codes for fast STT-MRAM cacheabstractSpin Transfer Torque Magnatic Random Access Memory (STT-MRAM) has the potential to become a universal memory technology due to its various attractive features such as non-volatility, high density, CMOS compatibility and zero leakage. However, STT-MRAM suffers from high write latency and poor reliability compared to SRAM. This is primarily due to its stochastic nature of switching, which makes it not suitable for fast caches such as L1. The use of robust Error Correction Coding (ECC) with multiple bit correction capability to optimize the write margin and reliability results in large decoding latencies and large number of check bits. In this paper, we propose a new solution to exclude the high costs of using ECC in terms of decoding latency and storage overhead to be able to use STT-MRAM for fast caches. We exploit the fact that STT-MRAM poses asymmetric errors due to the nature of Magnetic Tunnel Junction (MTJ) cell, which makes ECC a pessimistic solution to address such errors. Therefore, efficient Systematic Unidirectional Error-Detecting Code (SEDC) is proposed to be adopted instead of conventional ECC combined with proper cache access mechanism to fetch correct data in case of error detection. Our proposed approach provides orders of magnitude better reliability and considerable performance improvement that makes STT-MRAM viable for fast-caches compared to the existing solutions. Nour Sayed, Fabian Oboril, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
VTS | 3 |
| 2017 | Design of Defect and Fault-Tolerant Nonvolatile Spintronic Flip-FlopsabstractWith technology down scaling, static power has become one of the biggest challenges in a system on chip. Normally off computing using nonvolatile (NV) sequential elements is a promising solution to address this challenge. Recently, many NV shadow flip-flop architectures have been introduced in which magnetic tunnel junction (MTJ) cells are employed as backup storing elements. Due to the emerging fabrication processes of magnetic layers, MTJs are more susceptible to manufacturing defects than their CMOS counterparts. Moreover, unlike memory arrays that can effectively be repaired with well-established memory repair and coding schemes, flip-flops scattered in the layout are more difficult to repair. Therefore, without effective defect and fault tolerance for NV flip-flops, the manufacturing yield will be affected severely. In this paper, we propose a fault-tolerant NV latch (FTNV-L) design, in which several MTJ cells are arranged in such a way that it is resilient to various MTJ faults. The simulation results show that our proposed FTNV-L can effectively tolerate all single MTJ faults with a considerably lower overhead than traditional approaches. Rajendra Bishnoi, Fabian Oboril, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Non-Volatile Non-Shadow flip-flop using Spin Orbit Torque for efficient normally-off computingabstractWith technology scaling, conventional CMOS-based flip-flops can no longer efficiently cope with the increasing leakage power challenge. Therefore, various non-volatile flip-flop designs were recently introduced to reduce the static power consumption. However, these flipflop architectures employ non-volatile Magnetic Tunnel Junction (MTJ) storing devices only for backup, i.e. to save and restore the content before and after power gating. This limits their efficiency for aggressive power gating for effective power reduction. To overcome this limitation, we propose a novel Non-Volatile Non-Shadow flip-flop (NVNS-FF) using Spin Orbit Torque (SOT) based MTJ cells. In this design, we exploit the high speed, low energy and high-reliability features of SOT devices to employ them as active components of the flip-flop. This enables efficient normally-off computing by allowing very aggressive power gating for both short and long standby periods. Experimental results show that the NVNS-FF has similar energy and timing characteristics as conventional CMOS-based flip-flops in active mode, and at the same time it allows to reduce the static power by 5X compared to backup flip-flops. Rajendra Bishnoi, Fabian Oboril, Mehdi Baradaran Tahoori |
ASP-DAC | 1 |
| 2016 | Fault Tolerant Non-Volatile spintronic flip-flop
Rajendra Bishnoi, Fabian Oboril, Mehdi Baradaran Tahoori |
DATE | 1 |
| 2016 | A cross-layer analysis of Soft Error, aging and process variation in Near Threshold Computing
Anteneh Gebregiorgis, Saman Kiamehr, Fabian Oboril, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
DATE | 4 |
| 2016 | Low-Power Multi-Port Memory Architecture based on Spin Orbit Torque Magnetic DevicesabstractMulti-port memories are widely used as shared memory, such as register files, in a microprocessor system, and its number of ports and capacities are significantly increasing with every product generation. However, with technology advancements, multi-port memories are facing severe challenges due to their bit-cell leakage and scalability, as well as reliability issues due to increase in design complexity. In this paper, we design a novel multi-port memory architecture in which we employ emerging Spin Orbit Torque (SOT) magnetic devices as a storing component because of its several beneficial attributes such as non-volatility, scalability, zero-leakage, almost infinite endurance, low access latency, low area and immunity to soft-errors. Moreover, due to separate read and write current paths in these devices, simultaneous read and write operations can be performed on the same cell while maintaining data integrity. In our proposed architecture, we have demonstrated that with this characteristic of SOT, the read-write contention can be resolved inherently at the device-level, which can simplify the overall multi-port design. Experimental results show that our proposed multi-port design has low access latency, and high energy efficiency with negligible area overhead. Rajendra Bishnoi, Fabian Oboril, Mehdi Baradaran Tahoori |
ACM Great Lakes Symposium on VLSI | 1 |
| 2016 | Normally-OFF STT-MRAM Cache with Zero-Byte Compression for Energy Efficient Last-Level CachesabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) is a promising alternative to SRAM due to its low leakage and scalability advantages. In fact, although being more energy-efficient than SRAM, STT-MRAM caches at higher levels (e.g. L3) still incur a high energy consumption due to 1) high leakage in their read and write circuits and 2) high dynamic write energy in their bit-cells. To address this problem, we propose a novel normally-off STT-MRAM cache that exploits the fact that most applications access zero-byte patterns very frequently. In this architecture, writing of zero-bytes is avoided to reduce write energy. In addition, all read and write circuits are by default power gated (i.e. normally-off) to reduce leakage power. Then, dynamically at runtime, only those circuits required for the ongoing operation are activated. Our evaluations for an L3-cache of a multi-core microprocessor show that this approach reduces the energy consumption by 60% compared to state-of-the-art, while its impact on performance is negligible. Fabian Oboril, Fazal Hameed, Rajendra Bishnoi, Ali Ahari, Helia Naeimi, Mehdi Baradaran Tahoori |
ISLPED | 3 |
| 2016 | Layout-Based Modeling and Mitigation of Multiple Event TransientsabstractRadiation-induced multiple event transients (METs) are expected to become more frequent than single event transients (SETs) at nanoscale CMOS technology nodes. In this paper, a fast and accurate layout-based soft error rate (SER) assessment technique with consideration of both SET and MET fault models is presented. Despite existing techniques in which the adjacent MET sites are extracted from a logic-level netlist, we conduct a comprehensive layout analysis to obtain MET adjacent cells. Experimental results reveal that the layout-based technique is the only viable solution for identification of the adjacent cells as netlist-based techniques considerably underestimate the overall SER. Furthermore, by identifying the most vulnerable adjacent cells and increasing their physical distance in the layout using local adjustment rules, we are able to considerably reduce the overall SER without imposing any area and performance penalty. Mojtaba Ebrahimi, Hossein Asadi 0001, Rajendra Bishnoi, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Self-Timed Read and Write Operations in STT-MRAMabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising memory technology because of its advantageous features, such as nonvolatility, scalability, high density, zero-leakage, and CMOS compatibility. However, one of its major drawbacks is the high overall energy consumption. To make matters even worse, the write process in STT-MRAM is of stochastic nature, i.e., the completion of a write operation is nondeterministic. However, if the read/write completion could be detected on the fly (i.e., dynamically) and the respective components could be turned-OFF immediately, the energy consumption can be reduced by a large extent. Therefore, we propose a technique where the read and write completion signals are generated asynchronously on the fly using a self-timed bitwise technique. With this approach, the memory consumes power only when it is in the actual operational mode. Experimental results show that this technique significantly reduces the energy consumption of read and write operations with negligible area overhead. Rajendra Bishnoi, Fabian Oboril, Mojtaba Ebrahimi, Mehdi Baradaran Tahoori |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2015 | Evaluation of Hybrid Memory Technologies Using SOT-MRAM for On-Chip Cache HierarchyabstractMagnetic Random Access Memory (MRAM) is a very promising emerging memory technology because of its various advantages such as nonvolatility, high density and scalability. In particular, Spin Orbit Torque (SOT) MRAM is gaining interest as it comes along with all the benefits of its predecessor Spin Transfer Torque (STT) MRAM, but is supposed to eliminate some of its shortcomings. Especially the split of read and write paths in SOT-MRAM promises faster access times and lower energy consumption compared to STT-MRAM. In this paper, we provide a very detailed analysis of SOT-MRAM at both the circuit-and architecture-level. We present a detailed evaluation of performance and energy related parameters and compare the novel SOT-MRAM with several other memory technologies. Our architecture-level analysis shows that a hybrid-combination of SRAM for the L1-Data-cache, SOT-MRAM for the L1-Instruction-cache and L2-cache can reduce the energy consumption by 60% while the performance increases by 1% compared to an SRAM-only configuration. Moreover, the retention failure probability of SOT-MRAM is 27× smaller than the probability of radiation-induced Soft Errors in SRAM, for a 65 nm technology node. All of these advantages together make SOT-MRAM a viable choice for microprocessor caches. Fabian Oboril, Rajendra Bishnoi, Mojtaba Ebrahimi, Mehdi Baradaran Tahoori |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Architectural aspects in design and analysis of SOT-based memoriesabstractMagnetic Random Access Memory (MRAM) is a very promising emerging memory technology because of its various advantages such as non-volatility, high density and scalability. In particular, Spin Orbit Torque (SOT) MRAM is gaining interest as it comes along with all the benefits of its predecessor Spin Transfer Torque (STT) MRAM, but is supposed to eliminate some of its shortcomings. Especially the split of read and write paths in SOT-MRAM promises faster access times and lower energy consumption compared to STT-MRAM. In this work, we provide a very detailed analysis of SOT-MRAM at both circuit- and architecture-level. We present a detailed evaluation of performance and energy related parameters and compare the novel SOT-MRAM with several other memory technologies. Our architecture-level analysis shows that with a hybrid-combination of SRAM for the L1-cache and SOT-MRAM for the L2-cache the energy consumption can be reduced by 63 % in average while the performance can be increased by 1 %. In addition, the memory area is 43% lower compared to an SRAM-only configuration. Rajendra Bishnoi, Mojtaba Ebrahimi, Fabian Oboril, Mehdi Baradaran Tahoori |
ASP-DAC | 1 |
| 2014 | Asynchronous Asymmetrical Write Termination (AAWT) for a low power STT-MRAMabstractSpin Transfer Torque (STT) memory is an emerging and promising non-volatile storage technology. However, the high write current is still a major challenge which leads to a huge power consumption of the memory. Due to an inherent torque asymmetry of the Magnetic Tunnel Junction (MTJ) device employed in STT memories, the switching time between parallel to anti-parallel and anti-parallel to parallel magnetization is significantly different. Hence, the write latencies for writing `0' and `1' are also considerably different. In this paper, we propose a technique called Asynchronous Asymmetrical Write Termination (AAWT) which utilizes this asymmetrical behavior to terminate the write operations asynchronously and as a result significantly reduces the write power consumption. Furthermore, we present two different AAWT implementations to determine the actual write termination times. The first one makes use of a clock signal and the second one employs a self-timing approach based on an internal delay element. As shown by our experimental results, AAWT can reduce the total write energy by 30% in average with a negligible area overhead. Rajendra Bishnoi, Mojtaba Ebrahimi, Fabian Oboril, Mehdi Baradaran Tahoori |
DATE | 1 |
| 2014 | Read disturb fault detection in STT-MRAMabstractSpin Transfer Torque Magnetic Random Access Memory (STT-MRAM) has potential to become a universal memory technology because of its various advantageous features such as high density, non-volatility, scalability, high endurance and CMOS compatibility. However, read disturb is a major reliability issue in which a read operation can lead to a bitflip, because read and write current share the same path. This major reliability challenge is growing with technology scaling as read to write current ratio decreases. In this paper, we propose a circuit-level technique to detect read disturb by sensing the current during the read operation. Experimental results show that the proposed technique can effectively detect read disturb at the cost of negligible power and area overhead. Rajendra Bishnoi, Mojtaba Ebrahimi, Fabian Oboril, Mehdi Baradaran Tahoori |
ITC | 1 |