Marco Ottavi

dblp:84/5657 · DBLP profile ↗
← Back
49ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-5064-7342ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 4 first-author · 19 since 2021Software engineering, systems software and programming languages · 19 · 1 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Efficient Hash-to-Index via Rejection Sampling for Online Fault Detection with Bloom/Cuckoo Filters
abstract
Recent studies indicate that resource-efficient online fault detection in dependable computing systems may rely on probabilistic data structures such as Bloom and Cuckoo filters. To be effective and lightweight, these filters require low-latency and resource-efficient hash-to-index mappings. Existing approaches, namely, modulo, power-of-two, and multiplicative indexing, either incur high implementation cost, latency overhead, or impose rigid table size constraints that can lead to over-provisioning and suboptimal memory utilization, limiting their adoption in embedded and real-time systems. To address these limitations, this work proposes a hardware-efficient hash-to-index reduction technique based on rejection sampling. The proposed method optimizes index computation, yielding a uniform distribution for arbitrary table lengths while avoiding costly division or multiplication. Implemented as a mask-then-reject datapath on FPGA, our approach enables lightweight online checkers that maintain low area and latency footprints without sacrificing correctness. Experimental results on Bloom and Cuckoo filters demonstrate similar detection accuracy compared to canonical mappings while reducing hardware cost and latency.
Elijah Cishugi, Kuan-Hsun Chen, Marco Ottavi
ETS3
2026 Shooting Neutrons at Neurons: Radiation Testing of a Spiking Neural Network on Flash-Based FPGAs
Wim Nijsink, Bruno Endres Forlin, Amirreza Yousefzadeh, Marco Ottavi
ETS4
2026 A Single-Cycle, Area-Efficient Calibration Method for TDC-based Timing Monitors
abstract
Modern System-on-Chips (SoCs) increasingly integrate time-to-digital converter (TDC) based on-chip timing monitors to enable self-awareness for dependability-related applications. These monitors offer sub-clock cycle resolution but are highly sensitive to on-chip variations such as process, voltage, and temperature (PVT), transient voltage droop, and long-term selfaging effects. As a result, in-field calibration of these monitors becomes inevitable for achieving reliable measurements. This paper proposes a fast, single clock cycle calibration methodology that enables simultaneous calibration of multiple TDC-based timing monitors operating in the same voltage domain in an SoC. The novelty of this approach lies in the use of a shared, per-voltage-domain calibration IP instead of per-monitor selfcalibrating logic, which significantly reduces the area overhead associated with a monitoring infrastructure. Furthermore, the proposed IP is suitable for calibration against fast-changing phenomena like voltage droop as well as slower variations like PVT and aging. In addition to the conventional TDC-based positive slack monitor, design for a negative slack monitor is also discussed, which estimates negative slack during late signal transitions by sampling the transitioning signal immediately after the capturing clock edge. This enables in-field estimation of both positive and negative slack, supporting adaptive mechanisms for resilient system operation. Post layout evaluation of the proposed IP and the timing monitor is presented using the TSMC 65 nm standard cell library. Results demonstrate that the proposed calibration method reduces the average magnitude of error (RMSE) in slack measurement by 46%.
Madiha A. Sheikh, Marco Ottavi
ETS2
2026 Evaluation of FPGA Delay Injection Mechanisms for Small Delay Fault Analysis
abstract
Recent studies reveal that hyperscalers and data centres face significant silent data corruptions (SDCs) due to small delay faults (SDFs), highlighting the inadequacy of traditional fault models. SDFs originate from marginal timing degradations deeply tied to physical silicon behaviour, and evaluating their architectural impact typically requires ASIC-level modelling or time-consuming SPICE simulations. Contrary, FPGA-based emulation offers a fast and scalable platform to approximate timing-related failures by injecting controlled delay variations directly into functional paths. This paper explores various methods to introduce these delays, aiming to validate the physical primitives necessary for determining Delay Architectural Vulnerability Factors (DelayAVFs) via emulation. Multiple methods are compared, including input tap delay lines using IDELAY primitives, carry chains, component chains, fan-out, and path lengthening. The effectiveness is tested using a Ring Oscillator (RO) loop on a Xilinx 7-series FPGA at 100 MHz, determining performance based on factors like delay resolution and linearity. The findings show that the IDELAY method is the most effective, offering high precision (R2 =1) and flexibility for fault injection. Furthermore, a hybrid approach combining IDELAY and path lengthening is proposed to mitigate routing overheads, enabling precise delay injection across the full clock period. The effectiveness of the method is shown with a small injection campaign.
Tijmen T. Smit, David Postema, Pieter Paasman, Marco Ottavi
IOLTS4
2025 TrackScorer: Skyrmion Logic-in-Memory Accelerator for Tree-Based Ranking Models
abstract
Racetrack memories (RTMs) have been shown to have lower leakage power and higher density compared to traditional DRAM/SRAM technologies. However, their efficiency is often hindered by the need to shift the targeted data to access ports for read and write operations. Suitable mapping approaches are therefore essential to unleash their potential. In this work, we explore the mapping of the popular tree-based document ranking algorithm, Quickscorer, onto Skyrmion-based racetrack memories (SK-RTMs). Our approach leverages a Logic-in-Memory (LiM) accelerator, specifically designed to execute simple logic operations directly within SK-RTMs, enabling an efficient mapping of Quickscorer by exploiting its bitvector representation and inter-leaved traversal scheme of tree structures through bitwise logical operations. We present several mapping strategies, including one based on a quadratic assignment problem (QAP) optimization algorithm for optimal data placement of Quickscorer onto the racetracks. Our results demonstrate a significant reduction in read and write operations and, in certain cases, a decrease in the time spent shifting data during Quickscorer inference.
Elijah Cishugi, Sebastian Buschjäger, Martijn Noorlander, Marco Ottavi, Kuan-Hsun Chen
DATE4
2025 Bloom Filters for Soft Error Detection: Neutron and Fault Injection Validation
abstract
As memory cells continue to shrink in modern semiconductor technologies, radiation-induced Single Event Effects, such as single- and multi-bit upsets, pose growing challenges to system reliability. While effective and efficient for single and double-bit errors, traditional error detection and correction approaches, such as Error Correcting Codes (ECC), incur substantial overhead and complexity when designed to detect and correct multiple-bit errors. This study investigates the use of probabilistic data structures (PDS) as lightweight detectors for multiple-bit soft errors in memories. Leveraging the space-efficient and low-latency properties of Bloom filters, we implement a lightweight error detector (checker) within a representative memory subsystem on a flash-based FPGA. The checker's performance is validated through extensive neutron beam irradiation and fault-injection campaigns, demonstrating effective detection of multiple-bit errors with a tunable false-positive rate.
Elijah Cishugi, Tijmen T. Smit, Bruno Endres Forlin, Carlo Cazzaniga, Kuan-Hsun Chen, Marco Ottavi
IOLTS6
2025 The online reconfiguration of a distributed on-board computer: The time and network behaviour of a dependable scheduling algorithm
abstract
On-board Computers (OBCs) are at the centre of space-faring systems. With the increasing demand for cost-effective computing power in space, using high-performance commercial-off-the-shelf (COTS) components for OBCs has gained significant traction. COTS components, however, do not provide the necessary fault tolerance mechanisms. The ScOSA (Scalable On-board computing for Space Avionics) architecture uses COTS components in a distributed system to provide more computing performance and dependability. The effects of node failures are mitigated by removing the failed node from the system through reconfiguration. A reconfiguration is performed by using a set of predetermined configurations, which hinders system scalability due to exponentially increasing memory consumption depending on the number of nodes. This paper continues the work on the ScOSA online reconfiguration algorithm as a solution to this scalability problem. The online reconfiguration algorithm, which has been integrated into a scheduler, makes task scheduling decisions at run-time, eliminating the need for predetermined configurations. The six-phase scheduling mechanism uses the real-time state of the system and is a step towards higher dependability in distributed on-board computing. New test scenarios have been introduced to provide insight into the temporal and network behaviour of online reconfiguration. By evaluating in terms of time , network traffic and memory usage , it is shown that online reconfiguration is not only capable of dynamically generating configurations but also providing a solution to the scalability problem for systems with varying numbers of both nodes and tasks.
Glen te Hofsté, Andreas Lund 0001, Alexandra Coroiu, Marco Ottavi, Daniel Lüdtke
J. Syst. Archit.4
2024 Lightweight Instrumentation for Accurate Performance Monitoring in RTOSes
abstract
Evaluating performance metrics in embedded systems poses challenges, particularly due to the limited set of tools available for monitoring performance counters. In addition, performance evaluation frameworks for Real-Time Operating Systems (RTOSes) often lack the sophistication and capabilities available in general-purpose operating systems like Linux, which benefit from utilities such as perf_event. To bridge this gap, this paper presents an accurate and low-overhead instrumentation utility tailored for RTOSes. Our approach utilizes performance monitoring counters to observe individual user applications within the RTOS environment. Importantly, it enables comprehensive application monitoring by strategically placing probes at points of inherent system interference, thereby minimizing additional overhead. A pre-calibration of these probes allows for fine-grained measurements within user applications. This results in the elimination of 100 % of the overheads for most counters in our test configuration, impacting the context switch by only three additional instructions per monitored counter.
Bruno Endres Forlin, Kuan-Hsun Chen, Nikolaos Alachiotis 0001, Luca Cassano, Marco Ottavi
DATE5
2024 Neutron Beam Evaluation of Probabilistic Data Structure-based Online Checkers
abstract
High-criticality applications are vulnerable to Single Event Effects (SEEs) and require highly reliable and customizable microprocessors. Online checkers have been used to detect security and reliability issues in such systems. Popular hardware redundancy techniques such as Triple Modular Redundancy (TMR) and Dual Modular Redundancy (DMR) provide a high error coverage at the cost of substantial redundancy; therefore, there is an interest in introducing lightweight checkers that could offer the same detection ability as DMR with a much lower overhead. A possible implementation of these online checkers can be based on Probabilistic Data Structure (PDS) such as the Bloom Filter (BF). They are a form of information redundancy and an excellent complement to Single Error Correction Double Error Detection (SECDED) codes because they allow for detecting higher-order upsets. In this work, we integrate an online checker into the open-source RISC-V core NEORV32 and deploy it on a flash-based FPGA. This paper presents the evaluation of the online checker’s performance conducted under a neutron beam. The neutron beam experiments demonstrate that the real-life error rates of such structures are comparably worse than the initial simulation would indicate and that other factors can impact their performance.
Bruno Endres Forlin, Edian B. Annink, Elijah Cishugi, Carlo Cazzaniga, Paolo Rech, Gerard K. Rauwerda, Gianluca Furano, Marco Ottavi
IOLTS8
2023 Exploring Genomic Sequence Alignment for Improving Side-Channel Analysis
Heitor Uchoa, Vipul Arora 0004, Dennis Vermoen, Marco Ottavi, Nikolaos Alachiotis 0001
ESORICS (3)4
2023 DEV-PIM: Dynamic Execution Validation with Processing-in-Memory
abstract
Instruction injections or soft errors during execution on the CPU can cause serious system vulnerabilities. During the standard program flow of the processor, the injection of unauthorized instruction or the occurrence of an error in the expected instruction are the main conditions for potentially serious such vulnerabilities. With the execution of these unauthorized instructions, adversaries could exploit SoC and execute their own malicious program or get higher-level privileges on the system. On the other hand, non-intentional errors can potentially corrupt programs causing unintended executions or the cause of program crashes. Modern trusted architectures propose solutions for unauthorized execution on SoC with additional software mechanisms or extra hardware logic on the same untrusted SoC. Nevertheless, these SoCs can still be vulnerable, as long as deployed security detection mechanisms are embedded within the same SoC’s fabric. Furthermore, validation mechanisms on the SoC increase the complexity and power consumption of the SoC. This paper presents DEV-PIM, a new, high-performance, and low-cost execution validation mechanism in SoCs with external DRAM memory. The proposed approach uses processing-in-memory (PIM) method to detect instruction injections or corrupted instructions by utilising basic computing resources on a standard DRAM device. DEV-PIM transfers instructions scheduled for execution on the CPU to the DRAM and validates them by comparing content with the trusted program record on the DRAM using PIM operations. By optimising the DRAM scheduling process validation tasks are only executed when memory access is idle. The CPU retains uninterrupted memory access and can continue its normal program flow without penalty. We evaluate DEV-PIM in an end-to-end DRAM-compatible environment and run a set of software benchmarks. On average, the proposed architecture is able to detect 98.46% of instruction injections for different validation. We also measured on average only 0.346% CPU execution overhead with DEV-PIM enabled.
Alperen Bolat, Yahya Can Tugrul, Seyyid Hikmet Çelik, Sakir Sezer, Marco Ottavi, Oguz Ergin
ETS5
2023 An unprotected RISC-V Soft-core processor on an SRAM FPGA: Is it as bad as it sounds?
abstract
Fast development, low cost, and reconfigurability are becoming critical factors for aerospace applications, making SRAM FPGAs attractive. However, SRAM FPGAs are prone to errors in the user and on the configuration bits. For their correct functioning, they must be capable of withstanding failures without sacrificing much performance. When adjusting a soft core for these applications, it is essential to know where redundancies are necessary, to avoid unnecessary overhead. We characterize the reliability of an unprotected RISC-V microcontroller using an accelerated neutron beam. Our investigation shows that, for our chosen benchmark and processor, the user data in the memory banks is the leading cause of the total number of errors in the application. By reversing the benchmark operations, we could root cause the origin of the observed errors and found that most of the data corruption detected during the runs stem from previously corrupt input data or from output data that were corrupted while transmitting.
Bruno Endres Forlin, Wouter van Huffelen, Carlo Cazzaniga, Paolo Rech, Nikolaos Alachiotis 0001, Marco Ottavi
ETS6
2023 Towards Dependable RISC-V Cores for Edge Computing Devices
abstract
The migration of the computation from the cloud into edge devices, i.e., Internet-of-Things (IoTs) devices, reduces the latency and the quantity of data flowing into the network. With the emerging open-source and customizable RISC-V Instruction Set Architecture (ISA), cores based on such ISA are promising candidates for several application domains within the IoT family, such as automotive, Unnamed Aerial Vehicles (UAVs), industrial automation, healthcare, agriculture etc., where power consumption, real-time execution, security and reliability are of highest importance. In this emerging new era of connected RISC-V IoT devices, mechanisms are needed for a reliable and secure execution, still meeting area, energy consumption and computation time constraints of edge devices. We propose three mechanisms towards this goal, i.e., (i) a Root of Trust module for post-quantum secure boot, (ii) hardware checkers against hardware trojan horses and microarchitectural side-channel attacks, and (iii) a fine-grained dual core lockstep mechanism for real-time error detection and correction. The paper illustrates the proposed mechanisms with related motivations and implications, as well as a discussion on future research directions.
Pegdwende Romaric Nikiema, Alessandro Palumbo, Allan Aasma, Luca Cassano, Angeliki Kritikakou, Ari Kulmala, Jari Lukkarila, Marco Ottavi, Rafail Psiakis, Marcello Traiola
IOLTS8
2022 ERIC: An Efficient and Practical Software Obfuscation Framework
abstract
Modern cloud computing systems distribute software executables over a network to keep the software sources, which are typically compiled in a security-critical cluster, secret. However, these executables are still vulnerable to reverse engineering techniques that can extract secret information from programs (e.g., an algorithm, cryptographic keys), violating the IP rights and potentially exposing the trade secrets of the software developer. Malicious parties can (i) statically analyze the disassembly of the executable (static analysis) or (ii) dynamically analyze the software by executing it on a controlled device and observe performance counter values or exploit side-channels to reverse engineer software (dynamic analysis).We develop ERIC, a new, efficient, and general software obfuscation framework. ERIC protects software against (i) static analysis, by making only an encrypted version of software executables available to the human eye, no matter how the software is distributed, and (ii) dynamic analysis, by guaranteeing that an encrypted executable can only be correctly decrypted and executed by a single authenticated device. ERIC comprises key hardware and software components to provide efficient software obfuscation support: (i) a hardware decryption engine (HDE) enables efficient decryption of encrypted hardware in the target device, (ii) the compiler can seamlessly encrypt software executables given only a unique device identifier. Both the hardware and software components are ISA-independent, making ERIC general. The key idea of ERIC is to use physical unclonable functions (PUFs), unique device identifiers, as secret keys in encrypting software executables. Malicious parties that cannot access the PUF in the target device cannot perform static or dynamic analyses on the encrypted binary.We develop ERIC’s prototype on an FPGA to evaluate it end-to-end. Our prototype extends RISC-V Rocket Chip with the hardware decryption engine (HDE) to minimize the overheads of software decryption. We augment the custom LLVM-based compiler to enable partial/full encryption of RISC-V executables. The HDE incurs minor FPGA resource overheads, it requires 2.63% more LUTs and 3.83% more flip-flops compared to the Rocket Chip baseline. LLVM-based software encryption increases compile time by 15.22% and the executable size by 1.59%. ERIC is publicly available and can be downloaded from https://github.com/kasirgalabs/ERIC.
Alperen Bolat, Seyyid Hikmet Çelik, Ataberk Olgun, Oguz Ergin, Marco Ottavi
DSN5
2022 Yield Evaluation of Faulty Memristive Crossbar Array-based Neural Networks with Repairability
abstract
This paper evaluates the yield of a memristor-based crossbar array of artificial neural networks in the presence of stuck-at-faults (SAFs). A technique based on Markov chains is used to estimate the yield in the presence of stuck-at-faults. This method provides a high degree of accuracy. Another method that is used for analysis and comparison is the Poisson distribution, which uses the sum of all repairable fault patterns. A fault repair mechanism is also considered when evaluating the yield of the memristor crossbar array. The results demonstrate that the yield could be improved with redundancies and a higher repairable stuck-at-fault ratio.
Anu Bala, Saurabh Khandelwal, Abusaleh M. Jabir, Marco Ottavi
IOLTS4
2022 Is your FPGA bitstream Hardware Trojan-free? Machine learning can provide an answer
Alessandro Palumbo, Luca Cassano, Bruno Luzzi, José Alberto Hernández 0001, Pedro Reviriego, Giuseppe Bianchi 0001, Marco Ottavi
J. Syst. Archit.7
2022 Processor Security: Detecting Microarchitectural Attacks via Count-Min Sketches
abstract
The continuous quest for performance pushed processors to incorporate elements such as multiple cores, caches, acceleration units, or speculative execution that make systems very complex. On the other hand, these features often expose unexpected vulnerabilities that pose new challenges. For example, the timing differences introduced by caches or speculative execution can be exploited to leak information or detect activity patterns. Protecting embedded systems from existing attacks is extremely challenging, and it is made even harder by the continuous rise of new microarchitectural attacks (e.g., the Spectre and Orchestration attacks). In this article, we present a new approach based on count-min sketches for detecting microarchitectural attacks in the microprocessors featured by embedded systems. The idea is to add to the system a security checking module (without modifying the microprocessor under protection) in charge of observing the fetched instructions and identifying and signaling possible suspicious activities without interfering with the nominal activity of the system. The proposed approach can be programmed at design time (and reprogrammed after deployment) in order to always keep updated the list of the attacks that the checker is able to identify. We integrated the proposed approach in a large RISC-V core, and we proved its effectiveness in detecting several versions of the Spectre, Orchestration, Rowhammer, and Flush + Reload attacks. In its best configuration, the proposed approach has been able to detect 100% of the attacks, with no false alarms and introducing about 10% area overhead, about 4% power increase, and without working frequency reduction.
Kerem Arikan, Alessandro Palumbo, Luca Cassano, Pedro Reviriego, Salvatore Pontarelli, Giuseppe Bianchi 0001, Oguz Ergin, Marco Ottavi
IEEE Trans. Very Large Scale Integr. Syst.8
2021 A Memristive Architecture for Process Variation Aware Gas Sensing and Logic Operations
abstract
We propose novel memristive gas sensor architectures that can significantly reduce process and parametric variations in a predictable manner, while improving accuracy and overall power consumption. The proposed architecture can also be configured to realize multifunction logic operations as well as Complementary Resistive Switch with low hardware overhead in the absence of gasses. Our results show that the proposed architecture is significantly immune to process and parametric effects compared to a single sensor and almost unaffected by wire resistance, while offering much higher accuracy and much lower power consumption compared to existing techniques.
Saurabh Khandelwal, Marco Ottavi, Eugenio Martinelli, Abusaleh M. Jabir
IOLTS2
2021 Soft Error Tolerant Count Min Sketches
abstract
The estimation of the frequency of the elements on a set is needed in a wide range of computing applications. For example, to estimate the number of hits that a video gets or the number of packets in a network flow. In some cases, the number of elements in the set is very large and it is not practical to maintain a table with the exact count for each of them. Instead, simpler and more efficient data structures, commonly referred to as sketches, that provide an estimate are used. Among those structures the Count Min Sketch (CMS) is one of the most popular sketches. The CMS provides estimates that have one sided errors. In more detail, the CMS returns an estimate that is equal to or larger than the actual value. An update or check requires a small and constant number of memory accesses and the memory footprint is fixed and does not depend on the number of elements. The CMS relies on several arrays of counters that are stored in memory. Memories are prone to suffer soft errors that flip the contents of memory cells due for example to ionizing radiation. Therefore, it is of interest to study the impact that soft errors can have on the CMS estimates and to propose protection techniques that minimize their effect while requiring low overhead in terms of additional memory and circuitry. To the best of our knowledge this has not been done before. In this article, first the effect of soft errors on the CMS is evaluated by injecting errors. Then, in the second part a protection technique that does not require additional memory bits is presented and compared with the protection using a parity bit. In the last part of the article, the technique is extended to protect also against double adjacent bit errors.
Pedro Reviriego, Jorge Martínez 0001, Marco Ottavi
IEEE Trans. Computers3
2021 A Conditionally Chaotic Physically Unclonable Function Design Framework with High Reliability
abstract
Physically Unclonable Function (PUF) circuits are promising low-overhead hardware security primitives, but are often gravely susceptible to machine learning–based modeling attacks. Recently, chaotic PUF circuits have been proposed that show greater robustness to modeling attacks. However, they often suffer from unacceptable overhead, and their analog components are susceptible to low reliability. In this article, we propose the concept of a conditionally chaotic PUF that enhances the reliability of the analog components of a chaotic PUF circuit to a level at par with their digital counterparts. A conditionally chaotic PUF has two modes of operation: bistable and chaotic , and switching between these two modes is conveniently achieved by setting a mode-control bit (at a secret position) in an applied input challenge. We exemplify our PUF design framework for two different PUF variants—the CMOS Arbiter PUF and a previously proposed hybrid CMOS-memristor PUF, combined with a hardware realization of the Lorenz system as the chaotic component. Through detailed circuit simulation and modeling attack experiments, we demonstrate that the proposed PUF circuits are highly robust to modeling and cryptanalytic attacks, without degrading the reliability of the original PUF that was combined with the chaotic circuit, and incurs acceptable hardware footprint.
Saranyu Chattopadhyay, Pranesh Santikellur, Rajat Subhra Chakraborty, Jimson Mathew, Marco Ottavi
ACM Trans. Design Autom. Electr. Syst.5
2020 Yield Estimation of a Memristive Sensor Array
abstract
This paper proposes a method to calculate the yield of a memristor based sensor array considered as the probability that the chip provides acceptable sensing results when the array is affected by manufacturing defects. The modeling is based on a Markov Chain approach, in which each state represents an operating chip configuration and the state transitions take into account manufacturing defects. The proposed method is applicable to evaluate the yield with different fault models to achieve the comparative yield obtained by several redundancy allocations.
Vishal Gupta 0002, Saurabh Khandelwal, Giulio Panunzi, Eugenio Martinelli, Said Hamdioui, Abusaleh M. Jabir, Marco Ottavi
IOLTS7
2019 Fault Modeling and Simulation of Memristor based Gas Sensors
abstract
Memristors are an attractive option for use in future architectures due to their non-volatility, high density and low power operation. Gas sensing is one of the proposed application of memristive devices. In spite of these advantages, memristors are susceptible to defect densities due to the nondeterministic nature of nano-scale fabrication. In this paper, a novel spice memristor model incorporating fault models that emulates the gas sensing behaviour with/without faults is developed for simulation and integration with design automation tools. Our simulation results show that the proposed non-linear model detects the presence of the oxidising/reducing gas and analyses the defects/faults affecting the functionality of the sensor.
Saurabh Khandelwal, Anu Bala, Vishal Gupta 0002, Marco Ottavi, Eugenio Martinelli, Abusaleh M. Jabir
IOLTS4
2019 The Missing Applications Found: Robust Design Techniques and Novel Uses of Memristors
abstract
Resistive memory, also known as memristor, is an emerging potential successor to traditional CMOS charge based memories. Memristors have also recently been proposed as a promising candidate for several additional applications such as logic design, sensing, non-volatile storage, neuromorphic computing, Physically Unclonable Functions (PUFs), Content-addressable memory (CAM) and reconfigurable computing. In this paper, we explore three unique applications of memristor technology based implementations, specifically from the perspective of sensing, logic, in-memory computing and their solutions. We review solar cell health monitoring and diagnosis, describe the proposed solutions, and provide directions in memristive gas sensing and in-memory computing. For the gas sensor application, in order to determine the number of memristors to ensure a certain level of accuracy in sensitivity, a technique to optimize the sensor array based on an acceptable sensitivity variation and minimum sensitivity margin is presented. These “out-of-the-box” emerging ideas for applications of memristive devices in enhancing robustness and, at the same time, how the requirements of robust design are enabling unconventional use of the devices. To this end, the papers considers some examples of this mutual interaction.
Marco Ottavi, Vishal Gupta 0002, Saurabh Khandelwal, Shahar Kvatinsky, Jimson Mathew, Eugenio Martinelli, Abusaleh M. Jabir
IOLTS1
2017 Reliable gas sensing with memristive array
abstract
Gas sensing is one of the proposed application field of memristive devices. We used a crossbar array of memristors as gas sensor using the HP labs fabricated TiO2based memristor model in an attempt to improve sensing accuracy. We introduced the possibility of reliable multiple gases detection using multiple rows of memristors as separate sensor in a crossbar array. Our experimental results show that an array of memristors can minimise measurement errors as well as provide a good redundancy measure during gas sensing. Measurements taken from the sensors are also not affected by alternate current paths problem often experienced in crossbar architecture.
Adedotun Adeyemo, Abusaleh M. Jabir, Jimson Mathew, Eugenio Martinelli, Corrado Di Natale, Marco Ottavi
IOLTS6
2017 Guest Editorial: IEEE Transactions on Computers and IEEE Transactions on Emerging Topics in Computing Joint Special Section on Innovation in Reconfigurable Computing Fabrics from Devices to Architectures
abstract
The papers in this special section focuses on the advancement of the associated performance and reliability objectives via technology and functional heterogeneity, as well as advanced resilience properties in which reconfigurable fabrics are embarking. Emerging device characteristics such as non-volatility and new static versus dynamic energy consumption profiles, as well as novel interconnect mechanisms, and storage blocks within reconfigurable fabrics, in turn innovate architectural advances that enable new applications.
Ronald F. DeMara, Marco Platzner, Marco Ottavi
IEEE Trans. Computers3
2015 Low Delay Single Symbol Error Correction Codes Based on Reed Solomon Codes
abstract
To avoid data corruption, error correction codes (ECCs) are widely used to protect memories. ECCs introduce a delay penalty in accessing the data as encoding or decoding has to be performed. This limits the use of ECCs in high-speed memories. This has led to the use of simple codes such as single error correction double error detection (SEC-DED) codes. However, as technology scales multiple cell upsets (MCUs) become more common and limit the use of SEC-DED codes unless they are combined with interleaving. A similar issue occurs in some types of memories like DRAM that are typically grouped in modules composed of several devices. In those modules, the protection against a device failure rather than isolated bit errors is also desirable. In those cases, one option is to use more advanced ECCs that can correct multiple bit errors. The main challenge is that those codes should minimize the delay and area penalty. Among the codes that have been considered for memory protection are Reed-Solomon (RS) codes. These codes are based on non-binary symbols and therefore can correct multiple bit errors. In this paper, single symbol error correction codes based on Reed-Solomon codes that can be implemented with low delay are proposed and evaluated. The results show that they can be implemented with a substantially lower delay than traditional single error correction RS codes.
Salvatore Pontarelli, Pedro Reviriego, Marco Ottavi, Juan Antonio Maestro
IEEE Trans. Computers3
2015 A Synergetic Use of Bloom Filters for Error Detection and Correction
abstract
Bloom filters (BFs) provide a fast and efficient way to check whether a given element belongs to a set. The BFs are used in numerous applications, for example, in communications and networking. There is also ongoing research to extend and enhance BFs and to use them in new scenarios. Reliability is becoming a challenge for advanced electronic circuits as the number of errors due to manufacturing variations, radiation, and reduced noise margins increase as technology scales. In this brief, it is shown that BFs can be used to detect and correct errors in their associated data set. This allows a synergetic reuse of existing BFs to also detect and correct errors. This is illustrated through an example of a counting BF used for IP traffic classification. The results show that the proposed scheme can effectively correct single errors in the associated set. The proposed scheme can be of interest in practical designs to effectively mitigate errors with a reduced overhead in terms of circuit area and power.
Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Complementary resistive switch based stateful logic operations using material implication
abstract
Memristor based logic and memories are increasingly becoming one of the fundamental building blocks for future system design. Hence, it is important to explore various methodologies for implementing these blocks. In this paper, we present a novel Complementary Resistive Switching (CRS) based stateful logic operations using material implication. The proposed solution benefits from exponential reduction in sneak path current in crossbar implemented logic. We validated the effectiveness of our solution through SPICE simulations on a number of logic circuits. It has been shown that only 4 steps are required for implementing N input NAND gate whereas memristor based stateful logic needs N+1 steps.
Yuanfan Yang, Jimson Mathew, Dhiraj K. Pradhan, Marco Ottavi, Salvatore Pontarelli
DATE4
2013 Error detection in ternary CAMs using bloom filters
abstract
This paper presents an innovative approach to detect soft errors in Ternary Content Addressable Memories (TCAMs) based on the use of Bloom Filters. The proposed approach is described in detail and its performance results are presented. The advantages of the proposed method are that no modifications to the TCAM device are required, the checking is done on-line and the approach has low power and area overheads.
Salvatore Pontarelli, Marco Ottavi, Adrian Evans, Shi-Jie Wen
DATE2
2013 Error Detection and Correction in Content Addressable Memories by Using Bloom Filters
abstract
A content addressable memory (CAM) is an SRAM-based memory that can be accessed in parallel to search for a given search word, providing as a result the address of the matching data. Like conventional memories, a CAM can be affected by the occurrence of single event upsets (SEUs) that can alter the content of one of more memory cells causing different effects such as pseudo-HIT or pseudo-MISS events. It is well known that, because of the parallel search performed by a CAM during the query of a word, a standard error correction code could not defend it against SEU events. In this paper, we propose a method that does not require any modification to a CAM's internal structure and, therefore, can be easily applied at system level. Error detection is performed by using a probabilistic structure called "Bloom filter,” which can signal if a given data is present in the CAM. Bloom filters permit to efficiently store and query the presence of data in a set. But, while a CAM suffers from SEU induced errors, the probabilistic nature of Bloom filters has as a consequence the so called false-positive effect. This paper shows that, by combining the use of a Bloom filter with a CAM, the complementary limitations of these modules can be compensated. The combined use of a CAM and a Bloom filter is analyzed in different cases, showing that the proposed technique can be implemented with a low penalty in terms of area and power consumption.
Salvatore Pontarelli, Marco Ottavi
IEEE Trans. Computers2
2013 A Method to Construct Low Delay Single Error Correction Codes for Protecting Data Bits Only
abstract
Error correction codes (ECCs) have been used for decades to protect memories from soft errors. Single error correction (SEC) codes that can correct 1-bit error per word are a common option for memory protection. In some cases, SEC codes are extended to also provide double error detection and are known as SEC-DED codes. As technology scales, soft errors on registers also became a concern and, therefore, SEC codes are used to protect registers. The use of an ECC impacts the circuit design in terms of both delay and area. Traditional SEC or SEC-DED codes developed for memories have focused on minimizing the number of redundant bits added by the code. This is important in a memory as those bits are added to each word in the memory. However, for registers used in circuits, minimizing the delay or area introduced by the ECC can be more important. In this paper, a method to construct low delay SEC or SEC-DED codes that correct errors only on the data bits is proposed. The method is evaluated for several data block sizes, showing that the new codes offer significant delay reductions when compared with traditional SEC or SEC-DED codes. The results for the area of the encoder and decoder also show substantial savings compared to existing codes.
Pedro Reviriego, Salvatore Pontarelli, Juan Antonio Maestro, Marco Ottavi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2012 Introducing MEDIAN: A new COST Action on manufacturable and dependable multicore architectures at nanoscale
abstract
The MEDIAN (ManufacturablE and Dependable multI-core Architectures at Nanoscale) project is a EU funded COST Action aimed at creating a European network of competence and experts on all dependability aspects of future digital systems development, promoting collaboration between industry and research.
Marco Ottavi
ETS1
2011 Feedback based droop mitigation
abstract
A strong dl/dt event in a VLSI circuit can induce a temporary voltage drop and consequent malfunctioning of logic as for instance failing speed paths. This event, called power droop, usually manifests itself in at-speed scan test where a surge in switching activity (capture phase) follows a period of quiescent circuit state (shift phase). Power droop is also present during mission mode operation. However, because of the less predictable occurrence of the switching events in mission mode, usually the values of power droop measured during test are different from those measured in mission mode. To overcome the power droop problem, different mitigation techniques have been proposed. The goal of these techniques is to create a uniform current demand throughout the test. This paper proposes a feedback based droop mitigation technique which can adapt to the droop by reading the level of VDD and modifying real time the current flowing on ad-hoc droop mitigators. It is shown that the proposed solution not only can compensate for droop events occurring during test mode but also can be used as a method of mission mode droop mitigation and yield enhancement if higher power consumption is acceptable.
Salvatore Pontarelli, Marco Ottavi, Adelio Salsano, Kamran Zarrineh
DATE2
2009 Modeling and Evaluating Errors Due to Random Clock Shifts in Quantum-Dot Cellular Automata Circuits
Faizal Karim, Marco Ottavi, Hamidreza Hashempour, Vamsi Vankamamidi, Konrad Walus, André Ivanov, Fabrizio Lombardi
J. Electron. Test.2
2008 Analysis and Evaluations of Reliability of Reconfigurable FPGAs
Salvatore Pontarelli, Marco Ottavi, Vamsi Vankamamidi, Gian Carlo Cardarilli, Fabrizio Lombardi, Adelio Salsano
J. Electron. Test.2
2008 A Serial Memory by Quantum-Dot Cellular Automata (QCA)
abstract
Quantum-dot Cellular Automata (QCA) has been widely advocated as a new device architecture for nanotechnology. QCA systems require extremely low power, together with the potential for high density and regularity. These features make QCA an attractive technology for manufacturing memories in which the paradigm of memory-in-motion can be fully exploited. This paper proposes a novel serial memory architecture for QCA implementation. This architecture is based on utilizing new building blocks (referred to as tiles) in the storage and input/output circuitry of the memory. The QCA paradigm of memory-in-motion is accomplished using a novel arrangement in the storage loop and timing/clocking; a three-zone memory tile is proposed by which information is moved across a concatenation of tiles by utilizing a two-level clocking mechanism. Clocking zones are shared between memory cells and the length of the QCA line of a clocking zone is independent of the word size. QCA circuits for address decoding and input/output for simplification of the Read/Write operations are discussed in detail. An extensive comparison of the proposed architecture and previous QCA serial memories is pursued in terms of latency, timing, clocking requirements, and hardware complexity.
Vamsi Vankamamidi, Marco Ottavi, Fabrizio Lombardi
IEEE Trans. Computers2
2008 Two-Dimensional Schemes for Clocking/Timing of QCA Circuits
abstract
At nanoscale, quantum-dot cellular automata (QCA) defines a new device architecture that permits the innovative design of digital systems. Features of these systems are the allowed crossing of signal lines with different orientation in polarization on a Cartesian plane, the potential of high throughput due to efficient pipelining, fast signal switching, and propagation. However, QCA designs of even modest complexity suffer from the negative impact due to the placement of long lines of cells among clocking zones, thus resulting in increased delay, slow timing, and sensitivity to thermal fluctuations. In this paper, different schemes for clocking and timing of the QCA systems are proposed; these schemes utilize 2D techniques that permit a reduction in the longest line length in each clocking zone. The proposed clocking schemes utilize logic-propagation techniques that have been developed for systolic arrays. Placement of QCA cells is modified to ensure correct signal generation and timing. The significant reduction in the longest line length permits a fast timing and efficient pipelining to occur while guaranteeing a kink-free behavior in switching.
Vamsi Vankamamidi, Marco Ottavi, Fabrizio Lombardi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2007 QCA Circuits for Robust Coplanar Crossing
Sanjukta Bhanja, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli
J. Electron. Test.2
2006 Novel designs for thermally robust coplanar crossing in QCA
abstract
In this paper, different circuit arrangements of quantum-dot cellular automata (QCA) are proposed for the so-called coplanar crossing. These arrangements exploit the majority voting properties of QCA to allow a robust crossing of wires on the Cartesian plane. This is accomplished using enlarged lines and voting. Using a Bayesian network (BN) based simulator, new results are provided to evaluate the robustness to so-called kink of these arrangements to thermal variations. The BN simulator provides fast and reliable computation of the signal polarization versus normalized temperature. It is shown that by modifying the layout, a higher polarization level can be achieved in the routed signal by utilizing the proposed QCA arrangements
Sanjukta Bhanja, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli
DATE2
2006 Localization of Faults in Radix-n Signed Digit Adders
abstract
It is widely known that an adder can be checked by using check symbols that are residues of the numbers modulo some base. This paper extends this characteristic to a radix r signed digit (SD) representation. The confinement of the carry operation can also be exploited to localize the faulty resources in the SD adder and to reconfigure the adder in order to work with a reduced dynamic range. The fault localization procedure is presented in this paper and the reconfiguration the SD adder after the fault localization is discussed
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IOLTS2
2006 HDLQ: A HDL environment for QCA design
abstract
Emerging technologies have attracted a substantial interest in overcoming the physical limitations of CMOS as projected at the end of the Technology Roadmap; among these technologies, quantum-dot cellular automata (QCA) relies on different and novel paradigms to implement dense, low power circuits and systems for high-performance computing. As applicable to existing technologies, a hierarchical process can be utilized to facilitate the design of QCA circuits. Tools and methodologies both at system and physical levels are required to support all design phases. This article presents an HDL model to describe QCA “devices” (also referred elsewhere in the technical literature as building blocks, i.e., majority voter, inverter, wire, crossover) and facilitate the evaluation of their design. This tool, referred to as HDLQ, allows a designer to verify the logic characteristics of a QCA system, while supporting within a design environment different operational mechanisms (such as fault injection) and the unique features of QCA (such as bidirectionality and timing/clocking partitioning). The applicability of this design environment to various memory circuits for logic and timing verification is presented in detail. Various defective conditions for kinks due to thermodynamic effects and permanent faults due to manufacturing defects are considered for injection.
Marco Ottavi, Luca Schiano, Fabrizio Lombardi, Douglas Tougaw
ACM J. Emerg. Technol. Comput. Syst.1
2006 Fault Localization, Error Correction, and Graceful Degradation in Radix 2 Signed Digit-Based Adders
abstract
In this paper, a methodology for the development of fault-tolerant adders based on the radix 2 signed digit (SD) representation is presented. The use of a number representation characterized by a carry propagation confined to neighbor digits implies interesting advantages in terms of error detection, fault localization, and repair. Errors caused by faults belonging to a considered stuck-at fault set can be detected by a parity-based technique. In fact, a carry-free adder preserving the parity of the augends can be implemented allowing fault detection by using a parity checker. Regarding fault localization, the "carry-free" property of the adder ensures the confinement of the error due to a permanent fault to only few digits. The detection of the faulty digit has been obtained by using a recomputation with shifted operands method. Finally, after the fault localization, graceful degradation of the system intended as the reduction of the performances versus a correct output computation can be obtained by using two different procedures. The first one allows obtaining the correct output by recomputing the result performing two different shift operations and using the intersection of the obtained results to recover the correct output, while the second one is based on a reduced dynamic range approach, which allows us to obtain the result in only one step, but with fewer output digits.
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Computers2
2005 On the Analysis of Reed Solomon Coding for Resilience to Transient/Permanent Faults in Highly Reliable Memories
abstract
Single event upsets (SEU), as well as permanent faults, can significantly affect the correct on-line operation of digital systems, such as memories and microprocessors; a memory can be made resilient to permanent and transient faults by using modular redundancy and coding. Different memory systems are compared; these systems utilize simplex and duplex arrangements with a combination of Reed Solomon coding and scrubbing. The memory systems and their operations are analyzed by novel Markov chains to characterize the performance for dynamic reconfiguration as well as error detection and correction under the occurrence of permanent and transient faults. For a specific Reed Solomon code, the duplex arrangement is able to cope efficiently with the occurrence of permanent faults, while the use of scrubbing allows it to cope with transient faults.
Luca Schiano, Marco Ottavi, Fabrizio Lombardi, Salvatore Pontarelli, Adelio Salsano
DATE2
2005 Tile-based design of a serial memory in QCA
abstract
Quantum-dot Cellula Automata (QCA) has been widely advocated as a new device architecture fo nano technology. QCA systems require extremely low power together with the potential for high density and regularity. These features make QCA an attractive technology for manufacturing memories in which the paradigm of memory-in-motion can be fully exploited. This paper proposes a novel serial memory architecture for QCA implementation. This architecture is based on utilizing new building blocks (referred to as tiles) in the storage and input/output circuitry of the memory. The QCA paradigm of memory-in-motion is accomplished using a novel arrangement in the storage loop and timing/clocking; a three-zone memory tile is proposed by which information is moved across a concatenation of tiles by utilizing a two-level clocking mechanism. Clocking zones are shared between memory cells and the length of the QCA line in a clocking zone is independent of word size. This results in a substantial eduction in clocking zones compared with previous serial memories.
Vamsi Vankamamidi, Marco Ottavi, Fabrizio Lombardi
ACM Great Lakes Symposium on VLSI2
2005 A Comparative Evaluation of Designs for Reliable Memory Systems
Gian Carlo Cardarilli, Fabrizio Lombardi, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
J. Electron. Test.3
2005 Tile-based QCA design using majority-like logic primitives
abstract
The design of circuits and systems in Quantum-dot Cellular Automata (QCA) is still in infancy. The basic logic primitive in QCA is the majority voter (MV), that is not a universal function; so, inverters (INV) are also required. Blocks (referred to as tiles) are utilized in this article. A tile with a combined logic function of MV and INV (MV-like function) is proposed. It is shown that the MV-like tile can be effectively used in logic design as basic primitive. Tiles based on both the fully populated (FP) and non-fully populated (NFP) grids are investigated in detail. Various arrangements in inputs and outputs are also possible among the 4 sides of a grid, thus defining different tiles. Using a coherence vector simulation engine, it is shown that the 3 × 3 grid offers versatile logic operation. Different combinational functions such as majority-like and wire crossing are obtained using these tiles. Tile-based design of different circuits is compared to gate-based and SQUARES designs.
Jing Huang 0001, Mariam Momenzadeh, Luca Schiano, Marco Ottavi, Fabrizio Lombardi
ACM J. Emerg. Technol. Comput. Syst.4
2004 Simulation of reconfigurable memory core yield
abstract
We give a Markov chain model of the yield of an embedded memory core. The model allows easy inclusion of the effect of possible defects elsewhere on the chip that includes the embedded memory. We propose a reconfiguration algorithm for the case of both spare rows and columns that is simple enough that it could serve as built-in self-repair on the chip. Compared to an optimal configuration algorithm, there is no visible difference in the yield. We use parameters from an IBM embedded SRAM process to illustrate the yield calculation. We study the effect of different spare allocations. We conclude that as long as there is at least one spare of each type, the spares do not need to be balanced, once the yield impact of being part of a system-on-a-chip has been taken into account.
Marco Ottavi, Fred J. Meyer, Fabrizio Lombardi
ACM Great Lakes Symposium on VLSI1
2004 A Signed Digit Adder with Error Correction and Graceful Degradation Capabilities
Gian Carlo Cardarilli, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IOLTS2
2003 Design of a fault tolerant solid state mass memory
abstract
This paper describes a novel architecture of fault tolerant solid state mass memory (SSMM) for satellite applications. Mass memories with low-latency time, high throughput, and storage capabilities cannot be easily implemented using space qualified components, due to the inevitable technological delay of these kind of components. For this reason, the choice of commercial off the shelf (COTS) components is mandatory for this application. Therefore, the design of an electronic system for space applications, based on commercial components, must match the reliability requirements using system level methodologies. In the proposed architecture, error-correcting codes are used to strengthen the commercial dynamic random access memory (DRAM) chips, while the system controller is developed by applying fault tolerant design solutions. The main features of the SSMM are the dynamic reconfiguration capability, and the high performances which can be gracefully reduced in case of permanent faults, maintaining part of the system functionality. The paper shows the system design methodology, the architecture, and the simulation results of the SSMM. The properties of the building blocks are described in detail both in their functionality and fault tolerant capabilities. A detailed analysis of the system reliability and data integrity is reported. The graceful degradation capability of our system allows different levels of acceptable performances, in terms of active I/O link interfaces and storage capability. The results also show that the overall reliability of the SSMM is almost the same using different RS coding schemes, allowing a dynamic reconfiguration of the coding to reduce the latency (shorter codewords), or to improve the data integrity (longer codewords). The use of a scrubbing technique can be useful if a high SEU rate is expected, or if the data must be stored for a long period in the SSMM. The reported simulations show the behavior of the SSMM in presence of permanent and transient faults. In fact, we show that the SCU is able to recover from transient faults. On the other hand, using a spare microcontroller also hard faults can be tolerated. The distributed file system confines the unrecoverable fault effects only in a single I/O Interface. In this way, the SSMM maintains its capability to store and read data. The proposed system allows obtaining SSMM characterized by high reliability and high speed due the intrinsic parallelism of the switching matrix.
Gian Carlo Cardarilli, A. Leandri, P. Marinucci, Marco Ottavi, Salvatore Pontarelli, Marco Re, Adelio Salsano
IEEE Trans. Reliab.4