Martin Omaña 0001

dblp:13/5000 · also Martin Eugenio Omana · DBLP profile ↗
← Back
49ranked-venue papers
23as first author
9since 2021 · last 2026
0000-0001-8976-5365ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 49 · 23 first-author · 9 since 2021Software engineering, systems software and programming languages · 17 · 7 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Reliability of Multilevel Cell Phase-Change Memories for AI Implementations
G. Settegrani, R. Gattoni, Sara Cretí, Matteo Naldi, Martin Omaña 0001, Cecilia Metra, I. S. Troja
IOLTS5
2025 Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and Reliability
abstract
Modern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners.
Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo
DSD18
2025 On-Line Test of Fully Integrated Voltage Regulators for High Performance Microprocessors of Autonomous Systems
abstract
High performance microprocessors employed within autonomous systems usually adopt Fully Integrated Voltage Regulators (FIVRs) to enable the central power management unit to control individually the voltage of different microprocessor power domains, thus enabling a significant improvement of performance-per-Watt. However, since FIVRs are mainly implemented on the microprocessor die, they suffer from reliability problems due to the scaling of microelectronic technology. In particular, faults and aging may affect FIVRs during their in-field operation, possibly compromising the microprocessor correct operation in the field, with possible catastrophic effects if the microprocessor is executing safety critical functionalities of autonomous systems. In this paper, first we will analyze the effects of most likely faults and Bias Temperature Instability (BTI) aging mechanisms possibly affecting the FIVR during its operation in the field. We will show that almost 60% of FIVR faults may result in an incorrect output voltage, possibly compromising the microprocessor correct operation. Moreover, we will show that, due to BTI, the time it will take for the FIVR to change its output voltage in response to changes of its input reference voltage may exceed the maximum tolerable time guaranteeing the microprocessor correct operation. Based on these achieved results, we will then propose a monitor to enable the FIVR on-line test. In particular, upon the generation of an incorrect FIVR output voltage, or in case of a degraded FIVR response time to changes of the reference voltage (due to the occurrence of the considered faults or BTI), the monitor generates an output error message, that can then be adopted to activate proper recovery actions to guarantee the FIVR reliable operation, thus avoiding that faults and BTI possibly affecting FIVRs during their operation in the field can compromise the microprocessor correct operation, with possible dramatic consequences if the microprocessor is executing safety-critical functionalities.
Martin Omaña 0001, A. Menghi, A. Stefani, E. Vicini, Cecilia Metra, G. Froio, S. Petrucci
IOLTS1
2024 Silent Data Corruption and Reliability Risks due to Faults Affecting High Performance Microprocessors' Caches
abstract
Error Correcting Codes (ECCs) are frequently adopted to guarantee the correct operation in the field of caches of high performance microprocessors. They require the addition of proper encoding/decoding blocks (referred to as checkers) to the cache array. The occurrence of faults affecting such checkers has been typically neglected so far, due to their limited area compared to the cache array. This may be no longer acceptable, due to the increasing likelihood of faults possibly affecting microprocessors implemented by deeply scaled technologies, and due to the increasing requirements in terms of reliability of several applications (e.g., data centers, autonomous vehicles, unmanned robots, etc.). Based on these considerations, in this paper we analyze the effects of bridging faults possibly affecting the ECCs’ checkers, for two frequently adopted kinds of ECCs. We will show that the $68 \%$ (or the $61 \%$) of BFs possibly affecting the considered ECCs’ checkers are critical, since they may either inhibit the ECC correction ability of incorrect words read from the cache, or introduce errors in otherwise correct words read from the cache, with consequent risks for silent data corruption and microprocessor reliability. The remaining $\mathbf{3 2 \%}$ (or $39 \%$) of BFs may remain latent and accumulate with following faults or aging conditions affecting the cache, with consequent future risks for silent data corruption and microprocessor reliability. We then introduce a possible scheme to detect on line the occurrence of critical BFs that, compared to an alternate solution presented in the literature, features significantly lower impact on the ECC checker delay and area.
Martin Omaña 0001, A. Manfredi, Cecilia Metra, R. Locatelli, M. Chiavacci, S. Petrucci
IOLTS1
2024 Reliability of AI in Predicting the State of Health of Li-Ion Batteries*
Sara Cretí, Martin Omaña 0001, Cecilia Metra, Gianni Borelli
IOLTS2
2024 On the Reliability of Clock Monitoring Units for Safety Critical Applications' Microcontrollers
abstract
In multi-core microcontrollers adopted for safety critical applications, such as automotive, the frequency of clock signals is typically monitored by dedicated Clock Monitor Units (CMUs), whose correct operation is essential for the microcontroller correct operation and system’s safety. We analyse the effects of resistive bridging faults and transient faults possibly affecting a typical CMU. We will show that $39 \%$ of the considered CMU resistive bridging faults do not result in a CMU output error message, thus remaining latent. Depending on the value of their connecting resistance, up to the $49 \%$ of the latent bridging faults can make the CMU unable to indicate the presence of a monitored clock signal with an incorrect frequency, with potential catastrophic consequences for the microcontroller correct operation and system’s safety. Instead, as for transient faults, we will show that they can be reasonably considered to do not constitute a serious risk for system’s safety.
Massimiliano Zhupa, Matteo Naldi, Martin Omaña 0001, Cecilia Metra
IOLTS3
2022 Novel BTI Robust Ring-Oscillator-Based Physically Unclonable Function
abstract
Physically Unclonable Functions (PUFs) have become a promising low-cost solution for authentication and key generation in cryptosystems. However, it has been shown in the literature that the reliability of PUFs is undermined by aging mechanisms, such as Bias Temperature Instability (BTI), which may compromise their correct operation. In this paper, we present a novel ring oscillator (RO) based PUF design that is robust against BTI degradation, hereinafter referred to as Low-sensitive-BTI RO - LBTIRO. We compare our proposed LBTIRO to the standard RO-based PUF and to an alternative NBTI-robust RO-PUF recently presented in the literature, for 90 nm and 32 nm CMOS technology nodes. We show that, for considered technology nodes, our proposed LBTIRO features a higher robustness against BTI. Particularly, our LBTIRO enables a reduction of the impact of BTI on the oscillation frequency over circuit lifetime, which reaches 85.3% and 72.1% against the standard RO and the recent alternate solution, respectively, for the 90 nm technology. Moreover, we show that our proposed LBTIRO features a reduction in terms of power consumption if compared to the alternative NBTIrobust RO-PUF.
Marco Grossi, Martin Omaña 0001, Daniele Rossi 0001, Biagio Marzulli, Cecilia Metra
IOLTS2
2021 Investigation of the Impact of BTI Aging Phenomenon on Analog Amplifiers
Marco Grossi, Martin Omaña 0001
J. Electron. Test.2
2021 ST-CAC: a low-cost crosstalk avoidance coding mechanism based on three-valued numerical system
Zahra Shirmohammadi, Ata Khorami, Martin Omaña 0001
J. Supercomput.3
2019 Impact of Bias Temperature Instability (BTI) Aging Phenomenon on Clock Deskew Buffers
Marco Grossi, Martin Omaña 0001
J. Electron. Test.2
2019 Low-Cost Strategy for Bus Propagation Delay Reduction
Martin Omaña 0001, S. Govindaraj, Cecilia Metra
J. Electron. Test.1
2019 Fault-Tolerant Inverters for Reliable Photovoltaic Systems
abstract
Photovoltaic (PV) systems are increasingly adopted as a source of green energy. Due to the high economic investment that they usually involve, their reliability is becoming of concern. Recent studies have proven that faults likely to affect the inverters of PV systems in the field can dramatically reduce the energy delivered to the load. We first present a self-checking monitor to detect faults affecting the inverter in the field, as well as faults affecting the monitor itself. Then, we propose a hardware reconfiguration scheme that activated upon the monitor's generation of an alarm message after the occurrence of faults, enables to avoid their impact on the energy efficiency of the PV system. Our recovery scheme is also self-checking with respect to faults affecting itself. Therefore, our monitoring and reconfiguration schemes provide inverters with fault-tolerance ability, thus enabling to meet the increasing demand for reliable PV systems.
Martin Omaña 0001, Alessandro Fiore, Marco Mongitore, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.1
2017 New Approaches for Power Binning of High Performance Microprocessors
abstract
The significant process parameter variations occurring during fabrication of high performance sequential circuits, such as microprocessors, are posing relevant uncertainties on the power that such circuits will consume in the field, while executing workloads typical for the diverse products they are oriented to (e.g., cellular phones, notebooks, servers, etc). On the other hand, different kinds of products have different constraints on the maximal power that could be consumed during the execution of typical workloads, due to diverse needs in terms of charge autonomy, heat dissipation, etc. Consequently, the power that will be consumed by microprocessors during the execution of typical workloads in the field needs to be accurately characterized at the end of fabrication. Such a power consumption characterization (hereinafter referred to as “power binning”), will enable to classify microprocessors in “power bins”, each one containing microprocessors suitable for different kinds of products, thus enabling to introduce them all into the market for different kinds of products. Based on these considerations, in this paper we propose an approach to characterize accurately at the end of fabrication, and at low-cost (in terms of characterization time), the power that microprocessors will consume in the in-field during the execution of workloads typical for different kinds of products. Our approach exploits scan-based Logic Built-In Self-Test (LBIST) to apply to microprocessors' sequential blocks test vectors that induce on their internal nodes an activity factor (AF) similar to that experienced during the in-field execution of workloads typical for different kinds of products, thus enabling to perform power binning by simply measuring their consumed power. Our approach enables to scale the AF from 0 percent up to 97.6 percent (on average for the considered benchmark circuits) compared to conventional LBIST, with a granularity of the 2 percent, thus enabling to emulate accurately the AF induced by workloads typical of a wide range of products. We propose a hardware implementation for our approach requiring a limited area overhead (lower than 3 percent) over conventional LBIST.
Martin Omaña 0001, Marco Padovani, Kreshnik Veliu, Cecilia Metra, Juergen Alt, Rajesh Galivanche
IEEE Trans. Computers1
2017 Scalable Approach for Power Droop Reduction During Scan-Based Logic BIST
abstract
The generation of significant power droop (PD) during at-speed test performed by Logic Built-In Self Test (LBIST) is a serious concern for modern ICs. In fact, the PD originated during test may delay signal transitions of the circuit under test (CUT): an effect that may be erroneously recognized as delay faults, with consequent erroneous generation of test fails and increase in yield loss. In this paper, we propose a novel scalable approach to reduce the PD during at-speed test of sequential circuits with scan-based LBIST using the launch-on-capture scheme. This is achieved by reducing the activity factor of the CUT, by proper modification of the test vectors generated by the LBIST of sequential ICs. Our scalable solution allows us to reduce PD to a value similar to that occurring during the CUT in field operation, without increasing the number of test vectors required to achieve a target fault coverage (FC). We present a hardware implementation of our approach that requires limited area overhead. Finally, we show that, compared with recent alternative solutions providing a similar PD reduction, our approach enables a significant reduction of the number of test vectors (by more than 50%), thus the test time, to achieve a target FC.
Martin Omaña 0001, Daniele Rossi 0001, Filippo Fuzzi, Cecilia Metra, Chandra Tirumurti, Rajesh Galivanche
IEEE Trans. Very Large Scale Integr. Syst.1
2016 Inverters' self-checking monitors for reliable photovoltaic systems
Martin Omaña 0001, A. Fiore, Cecilia Metra
DATE1
2016 Low-Cost and High-Reduction Approaches for Power Droop during Launch-On-Shift Scan-Based Logic BIST
abstract
During at-speed test of high performance sequential ICs using scan-based Logic BIST, the IC activity factor (AF) induced by the applied test vectors is significantly higher than that experienced during its in field operation. Consequently, power droop (PD) may take place during both shift and capture phases, which will slow down the circuit under test (CUT) signal transitions. At capture, this phenomenon is likely to be erroneously recognized as due to delay faults. As a result, a false test fail may be generated, with consequent increase in yield loss. In this paper, we propose two approaches to reduce the PD generated at capture during at-speed test of sequential circuits with scan-based Logic BIST using the Launch-On-Shift scheme. Both approaches increase the correlation between adjacent bits of the scan chains with respect to conventional scan-based LBIST. This way, the AF of the scan chains at capture is reduced. Consequently, the AF of the CUT at capture, thus the PD at capture, is also reduced compared to conventional scan-based LBIST. The former approach, hereinafter referred to as Low-Cost Approach (LCA), enables a 50 percent reduction in the worst case magnitude of PD during conventional logic BIST. It requires a small cost in terms of area overhead (of approximately 1.5 percent on average), and it does not increase the number of test vectors over the conventional scan-based LBIST to achieve the same Fault Coverage (FC). Moreover, compared to three recent alternative solutions, LCA features a comparable AF in the scan chains at capture, while requiring lower test time and area overhead. The second approach, hereinafter referred to as High-Reduction Approach (HRA), enables scalable PD reductions at capture of up to 87 percent, with limited additional costs in terms of area overhead and number of required test vectors for a given target FC, over our LCA approach. Particularly, compared to two of the three recent alternative solutions mentioned above, HRA enables a significantly lower AF in the scan chains during the application of test vectors, while requiring either a comparable area overhead or a significantly lower test time. Compared to the remaining alternative solutions mentioned above, HRA enables a similar AF in the scan chains at capture (approximately 90 percent lower than conventional scan-based LBIST), while requiring a significantly lower test time (approximately 4.87 times on average lower number of test vectors) and comparable area overhead (of approximately 1.9 percent on average).
Martin Omaña 0001, Daniele Rossi 0001, Edda Beniamino, Cecilia Metra, Chandra Tirumurti, Rajesh Galivanche
IEEE Trans. Computers1
2015 Intermittent and Transient Fault Diagnosis on Sparse Code Signatures
abstract
Failure diagnosis of field returns typically requires high quality test stimuli and assumes that tests can be repeated. For intermittent faults with fault activation conditions depending on the physical environment, the repetition of tests cannot ensure that the behavior in the field is also observed during diagnosis, causing field returns diagnosed as no-trouble-found. In safety critical applications, self-checking circuits, which provide concurrent error detection, are frequently used. To diagnose intermittent and transient faulty behavior in such circuits, we use the stored encoded circuit outputs in case of a failure (called signatures) for later analysis in diagnosis. For the first time, a diagnosis algorithm is presented that is capable of performing the classification of intermittent or transient faults using only the very limited amount of functional stimuli and signatures observed during operation and stored on chip. The experimental results demonstrate that even with these harsh limitations it is possible to distinguish intermittent from transient faulty behavior. This is essential to determine whether a circuit in which failures have been observed should be subject to later physical failure analysis, since intermittent faulty behavior has been diagnosed. In case of transient faulty behavior, it may still be operated reliably.
Michael A. Kochte, Atefe Dalirsani, Andrea Bernabei, Martin Omaña 0001, Cecilia Metra, Hans-Joachim Wunderlich
ATS4
2015 Low-Cost On-Chip Clock Jitter Measurement Scheme
abstract
In this paper, we present a low-cost, on-chip clock jitter digital measurement scheme for high performance microprocessors. It enablesin situjitter measurement during the test or debug phase. It provides very high measurement resolution and accuracy, despite the possible presence of power supply noise (representing a major source of clock jitter), at low area and power costs. The achieved resolution is scalable with technology node and can in principle be increased as much as desired, at low additional costs in terms of area overhead and power consumption. We show that, for the case of high performance microprocessors employing ring oscillators (ROs) to measure process parameter variations (PPVs), our jitter measurement scheme can be implemented by reusing part of such ROs, thus allowing to measure clock jitter with a very limited cost increase compared with PPV measurement only, and with no impact on parameter variation measurement resolution.
Martin Omaña 0001, Daniele Rossi 0001, Daniele Giaffreda, Cecilia Metra, Asifur Rahman, Simon M. Tam
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Modeling and Detection of Hotspot in Shaded Photovoltaic Cells
abstract
In this paper, we address the problem of modeling the thermal behavior of photovoltaic (PV) cells undergoing a hotspot condition. In case of shading, PV cells may experience a dramatic temperature increase, with consequent reduction of the provided power. Our model has been validated against experimental data, and has highlighted a counterintuitive PV cell behavior, that should be considered to improve the energy efficiency of PV arrays. Then, we propose a hotspot detection scheme, enabling to identify the PV module that is under hotspot condition. Such a scheme can be used to avoid the permanent damage of the cells under hotspot, thus their drawback on the power efficiency of the entire PV system.
Daniele Rossi 0001, Martin Omaña 0001, Daniele Giaffreda, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.2
2015 Impact of Bias Temperature Instability on Soft Error Susceptibility
abstract
In this paper, we address the issue of analyzing the effects of aging mechanisms on ICs' soft error (SE) susceptibility. In particular, we consider bias temperature instability (BTI), namely negative BTI in pMOS transistors and positive BTI in nMOS transistors that are recognized as the most critical aging mechanisms reducing the reliability of ICs. We show that BTI reduces significantly the critical charge of nodes of combinational circuits during their in-field operation, thus increasing the SE susceptibility of the whole IC. We then propose a time dependent model for SE susceptibility evaluation, enabling the use of adaptive SE hardening approaches, based on the ICs lifetime.
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Alessandro Paccagnella
IEEE Trans. Very Large Scale Integr. Syst.2
2014 Clock Faults Induced Min and Max Delay Violations
Daniele Rossi 0001, Martin Omaña 0001, José Manuel Cazeaux, Cecilia Metra
J. Electron. Test.2
2013 Novel approach to reduce power droop during scan-based logic BIST
abstract
Significant peak power (PP), thus power droop (PD), during test is a serious concern for modern, complex ICs. In fact, the PD originated during the application of test vectors may produce a delay effect on the circuit under test signal transitions. This event may be erroneously recognized as presence of a delay fault, with consequent generation of an erroneous test fail, thus increasing yield loss. Several solutions have been proposed in the literature to reduce the PD during test of combinational ICs, while fewer approaches exist for sequential ICs. In this paper, we propose a novel approach to reduce peak power/power droop during test of sequential circuits with scan-based Logic GIST. In particular, our approach reduces the switching activity of the scan chains between following capture cycles. This is achieved by an original generation and arrangement of test vectors. The proposed approach presents a very low impact on fault coverage and test time, while requiring a very low cost in terms of area overhead.
Martin Omaña 0001, Daniele Rossi 0001, Filippo Fuzzi, Cecilia Metra, Chandra Tirumurti, R. Galivache
ETS1
2013 Low Cost Concurrent Error Detection Strategy for the Control Logic of High Performance Microprocessors and Its Application to the Instruction Decoder
Daniele Rossi 0001, Martin Omaña 0001, G. Garrammone, Cecilia Metra, Abhijit Jas, Rajesh Galivanche
J. Electron. Test.2
2013 Low Cost NBTI Degradation Detection and Masking Approaches
abstract
Performance degradation of integrated circuits due to aging effects, such as Negative Bias Temperature Instability (NBTI), is becoming a great concern for current and future CMOS technology. In this paper, we propose two monitoring and masking approaches that detect late transitions due to NBTI degradation in the combinational part of critical data paths and guarantee the correctness of the provided output data by adapting the clock frequency. Compared to recently proposed alternative solutions, one of our approaches (denoted as Low Area and Power (LAP) approach) requires lower area overhead and lower, or comparable, power consumption, while exhibiting the same impact on system performance, while the other proposed approach (denoted as High Performance (HP) approach) allows us to reduce the impact on system performance, at the cost of some increase in area and power consumption.
Martin Omaña 0001, Daniele Rossi 0001, Nicolò Bosio, Cecilia Metra
IEEE Trans. Computers1
2013 Faults Affecting Energy-Harvesting Circuits of Self-Powered Wireless Sensors and Their Possible Concurrent Detection
abstract
We analyze the effects of faults on an energy-harvesting circuit (EHC) providing power to a wireless biomedical multisensor node. We show that such faults may prevent the EHC from producing the power supply voltage level required by the multisensor node. Then, we propose a low-cost (in terms of power consumption and area overhead) additional circuit monitoring the voltage level produced by the EHC continuously, and concurrently with the normal operation of the device. Such a monitor gives an error indication if the generated voltage falls below the minimum value required by the sensor node to operate correctly, thus allowing the activation of proper recovery actions to guarantee system fault tolerance. The proposed monitor is self-checking with regard to the internal faults that can occur during its in-field operation, thus providing an error signal when affected by faults itself.
Martin Omaña 0001, Daniele Rossi 0001, Daniele Giaffreda, Roberto Specchia, Cecilia Metra, Marcin Marzencki, Bozena Kaminska
IEEE Trans. Very Large Scale Integr. Syst.1
2012 New Design for Testability Approach for Clock Fault Testing
abstract
We propose a new design for testability approach for testing clock faults of next generation high performance microprocessors. In fact, it has been shown that conventional manufacturing test is unable to guarantee their detection, although they could compromise the effectiveness of delay fault testing, as well as the microprocessor correct operation in the field. These conditions will of course worsen with technology scaling, due to the expected increase in fault likelihood, included clock faults. To deal with these problems we propose a design for testability approach that, by means of simple modifications to conventional clock buffers, allows clock fault detection through any conventional manufacturing test approach. This is achieved at the cost of very low increase in area and power consumption of clock buffers, and with no additional test cost or impact on the microprocessor performance and in-field operation. We then introduce a possible further modification to clock buffers that, at additional limited costs in terms of area and power consumption, allows their calibration after fabrication in order to compensate for parameter variations possibly occurring during manufacturing, thus minimizing the likelihood of either false test fails, or test misses. As an example, we show the application of our approach to the clock distribution network of the Pentium® 4 microprocessor (Other names and brands may be claimed as property of others). However, it can be applied to the clock distribution of any high performance ASIC, or microprocessor.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
IEEE Trans. Computers2
2011 Low-Cost Dynamic Compensation Scheme for Local Clocks of Next Generation High Performance Microprocessors
abstract
We propose a low cost scheme for the dynamic compensation in the field of undesired skew and duty cycle variations of local clocks of high performance microprocessors and high end ASICs. Compared to alternate approaches, our solution features lower power consumption, smaller compensation error, and a lower or comparable area overhead.
Martin Omaña 0001, Cecilia Metra, Simon M. Tam
IEEE Trans. Very Large Scale Integr. Syst.1
2010 High-Performance Robust Latches
abstract
First, a new high-performance robust latch (referred to as HiPeR latch) is presented that is insensitive to transient faults affecting its internal and output nodes by design, independently of the size of its transistors. Then, a modified version of the HiPeR latch (referred as HiPeR-CG) is proposed that is suitable to be used together with clock gating. Both proposed latches are faster than the latches most recently presented in the literature, while providing better or comparable robustness to transient faults, at comparable or lower costs in terms of area and power, respectively. Therefore, thanks to the good trade-offs in terms of performance, robustness, and cost, our proposed latches are particularly suitable to be adopted on critical paths.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers1
2009 Accurate Linear Model for SET Critical Charge Estimation
abstract
In this paper, we present an accurate linear model for estimating the minimum amount of collected charge due to an energetic particle striking a combinational circuit node that may give rise to a SET with an amplitude larger than the noise margin of the subsequent gates. This charge value will be referred to as SET critical charge (Q SET ). Our proposed model allows to calculate the Q SET of a node as a function of the size of the transistors of the gate driving the node and the fan-out gate(s), with no need for time costly electrical level simulations. This makes our approach suitable to be integrated into a design automation tool for circuit radiation hardening. The proposed model features 96% average accuracy compared to electrical level simulations performed by HSPICE. Additionally, it highlights that Q SET has a much stronger dependence on the strength of the gate driving the node, than on the node total capacitance. This property could be considered by robust design techniques in order to improve their effectiveness.
Daniele Rossi 0001, José Manuel Cazeaux, Martin Omaña 0001, Cecilia Metra, Abhijit Chatterjee
IEEE Trans. Very Large Scale Integr. Syst.3
2008 Function-Inherent Code Checking: A New Low Cost On-Line Testing Approach for High Performance Microprocessor Control Logic
abstract
We propose an on-line testing approach for the control logic of high performance microprocessors. Rather than adding information redundancy (in the form of error detecting codes), we propose to look for the information redundancy (referred to as Function-Inherent Codes) that the microprocessor control logic may inherently have, due to its required functionality. We will show that this allows to achieve on-line testing at significant savings in terms of area and power consumption, and with lower or comparable impact on system performance and design costs, compared to alternate, traditional on-line testing approaches.
Cecilia Metra, Daniele Rossi 0001, Martin Omaña 0001, Abhijit Jas, Rajesh Galivanche
ETS3
2008 Checkers' No-Harm Alarms and Design Approaches to Tolerate Them
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra
J. Electron. Test.2
2007 Novel compensation scheme for local clocks of high performance microprocessors
abstract
Clock compensation for process variations and manufacturing defects is a key strategy to achieve high performance of processors and high end ASIC. However, with the increase in process variations and defect densities, clock compensation is becoming increasingly challenging. A clock distribution system also consumes over 30% of the overall chip level power, so every little bit counts, including compensation schemes. In this paper we propose a new scheme for the compensation of undesirable skews and duty-cycle variations of local clocks of high performance microprocessors and high end ASICs. Our scheme performs compensation continuously, during the microprocessor operation, thus allowing also compensation to clock jitters due to environmental influences during operation. Compared to alternate solutions for local clock compensation, our scheme features lower power consumption, smaller compensation error, and a lower or comparable area overhead, while allowing compensation to be accomplished within the same clock cycle of skew or duty-cycle variation.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
ITC2
2007 Novel Approach to Clock Fault Testing for High Performance Microprocessors
abstract
This paper presents a novel approach for testing clock faults for high performance microprocessors. Although such faults have been shown to be likely and could compromise delay fault testing, conventional manufacturing test methodology is unable to guarantee their detection. This paper proposes a modification to the conventional clock buffers allowing standard manufacturing test to detect the faults. This is achieved at the cost of a small increase in area and power consumption of the clock buffers, but with no additional test cost or impact on the microprocessor performance and in-field operation. The approach can be applied to the clock system of any high performance chip or microprocessor.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
VTS2
2007 Latch Susceptibility to Transient Faults and New Hardening Approach
abstract
In this paper we analyze the conditions making Transient Faults (TFs) affecting the nodes of conventional latch structures generate output Soft-Errors (SEs). We investigate the susceptibility to TFs of all latch nodes and identify the most critical one(s). We show that, for standard latches using back-to-back inverters for their positive feedback, the internal nodes within their feedback path are the most critical. Such nodes will be hereafter referred to as internal feedback nodes. Based on this analysis, we first propose a low cost hardened latch that, compared to alternative hardened solutions, is able to filter out completely TFs affecting its internal feedback nodes, while presenting a lower susceptibility to TFs on the other internal nodes. This is achieved at the cost of a reduced robustness to TFs affecting the output node. To overcome this possible limitation (especially for systems for high reliability applications), we propose another version of our latch that, at the cost of a small area and power consumption increase compared to our first solution, improves also the robustness of the output node, which can be higher than that of alternative hardened solutions. Additionally, both proposed latches present a comparable or higher robustness of the input node than alternative solutions and provide a lower or comparable power-delay product and area overhead than classical implementations and alternative hardened solutions.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers1
2006 Low-cost and highly reliable detector for transient and crosstalk faults affecting FPGA interconnects
abstract
In this paper we present a novel circuit for the online detection of transient and crosstalk faults affecting the interconnects of systems implemented using Field Programmable Gate-Arrays (FPGAs). The proposed detector features self-checking ability with respect to faults possibly affecting itself, thus being suitable for systems with high reliability requirements, like those for space applications. Compared to alternate solutions, the proposed circuit requires a significantly lower area overhead, while implying a comparable, or lower, impact on system performance. We have verified our circuit operation and self-checking ability by means of post-layout simulations.
Martin Omaña 0001, José Manuel Cazeaux, Daniele Rossi 0001, Cecilia Metra
DATE1
2006 Path (Min) Delay Faults and Their Impact on Self-Checking Circuits' Operation
abstract
Min delay violations are traditionally not modeled as possible faults as a result of manufacturing defects. Usually, path delay faults are implicitly assumed to be paths' max delay violations. This, in turn, is based on the assumption that min delay violations are designed out. Most previous manufacturing defect/fault analysis works have not considered their effect on clock circuits. More recently, as burn-in becomes ineffective and process variations become more of an issue, latent defects, device degradation or wear out in the field would potentially also cripple the clock distribution network. Consequently, we should start considering also path (min) delay faults when designing on-line testable circuits, similar to what we currently do for path (max) delay faults. The challenges that this poses to the existing on-line testing strategies are discussed. Examples showing the possible incorrect behavior of a self-checking circuit as a result of this kind of faults are given. New on-line testing strategies should consequently be devised to deal with these faults
Cecilia Metra, Martin Omaña 0001, Daniele Rossi 0001, José Manuel Cazeaux
IOLTS2
2006 Checker No-Harm Alarm Robustness
abstract
In this paper we evaluate the probability that a transient fault (TF), multiple or single, affecting a checker of a self-checking circuit, gives rise to an unnecessary error indication (no-harm alarm). A new property (no-harm alarm robustness) has been defined that, in case of a fault affecting a self-checking circuit (SCC), guarantees that we can determine whether the fault is affecting the functional block, or the checker itself, and whether such a fault is a transient or a permanent fault. Finally, we propose a possible solution implementing the defined property. Its behavior has been verified by means of HSpice simulations, and we evaluate its cost in terms of area overhead and introduced delay
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Andrea Pagni
IOLTS2
2005 On Transistor Level Gate Sizing for Increased Robustness to Transient Faults
abstract
In this paper we present a detailed analysis on how the critical charge (Q/sub crit/) of a circuit node, usually employed to evaluate the probability of transient fault (TF) occurrence as a consequence of a particle hit, depends on transistors' sizing. We derive an analytical model allowing us to calculate a node's Q/sub crit/ given the size of the node's driving gate and fan-out gate(s), thus avoiding time costly electrical level simulations. We verified that such a model features an accuracy of the 97% with respect to electrical level simulations performed by HSPICE. Our proposed model shows that Q/sub crit/ depends much more on the strength (conductance) of the gate driving the node, than on the node total capacitance. We also evaluated the impact of increasing the conductance of the driving gate on TFs' propagation, hence on soft error susceptibility (SES). We found that such a conductance increase not only improves the TF robustness of the hardened node, but also that of the whole circuit.
José Manuel Cazeaux, Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Abhijit Chatterjee
IOLTS3
2005 On the Selection of Unidirectional Error Detecting Codes for Self-Checking Circuits' Area Overhead and Performance Optimization
abstract
In this paper we address the issue of optimizing the area overhead and performance of self-checking circuits using all unidirectional error detecting codes (AUEDCs), with no impact on system's reliability. In particular, we propose an error detecting code selection approach that, starting from the consideration of the functional circuit topology, allows us to identify whether or not all output bits can be simultaneously erroneous, thus actually mandating the adoption of an AUEDC. We show that, differently from common expectations, this may frequently be not the case (for approximately the 50% of the considered benchmarks) for all possible internal node stuck-ats, transistor stuck-ons, transistor stuck-opens and resistive bridgings. We then propose a tool that, starting from the (combinational or sequential) circuit high level description, allows us to identify whether or not this is the case and, in particular, which is the maximal number (t) of possibly simultaneously erroneous output bits. Based on this information, a lower redundancy error detecting code (e.g., a t-UEDC) is adopted, rather than an AUEDC, thus generally allowing reducing area overhead and impact on system's performance. Such a code is automatically implemented by our developed tool, whose effectiveness has been verified for benchmark circuits and for a FPGA implemented prototype.
Martin Omaña 0001, O. Losco, Cecilia Metra, Andrea Pagni
IOLTS1
2005 Low Cost Scheme for On-Line Clock Skew Compensation
abstract
In this paper we propose a novel buffer scheme that is able to compensate undesired skews between clocks of a synchronous system in a negligible time upon skew occurrence, thus being suitable also for on-line clock-skew correction. Clock signals are aligned one with respect to the other, starting from a reference clock, and moving forward among physically adjacent clock signals, thus creating no problem of reference clock's routing. Our solution is also able to compensate clock duty-cycle variations, which have been shown very likely in case of faults, for instance bridgings, affecting the clock distribution network. Compared to alternate solutions, our proposed scheme enables significant reductions in area overhead and power consumption, and is suitable for on-line compensation. Therefore, it allows clock skew and duty-cycle fault tolerance, thus increasing process yield and system's reliability.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
VTS1
2005 Low Cost and High Speed Embedded Two-Rail Code Checker
abstract
We propose a compact, high-speed, and highly testable parallel two-rail code checker, particularly suitable to implementing embedded checkers. In fact, it requires only two input codewords to satisfy the totally-self-checking or strongly code-disjoint property with respect to a wide set of realistic internal faults. Our checker can be employed to check the correct operation of a connected functional block using the two-rail code, to implement the output two-rail code checker of "normal" checkers for unordered codes, or to join together the error messages produced by various checkers (possibly using different codes) present within the same self-checking system. The behavior of our checker has been verified by means of electrical level simulations (performed using HSPICE), considering both nominal values and statistical variations of electrical parameters. We also propose a possible modification to our checker internal structure that makes it able to provide an output error indication remaining latched until the application of a proper reset signal. Depending on the considered application and recovery technique to be employed upon the generation of an error indication at the checker output, one proposed solution or the other may be preferable.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers1
2004 Are Our Design for Testability Features Fault Secure?
abstract
We analyze the risks associated with faults affecting some common design for testability (DFT) features employed within digital products. We will show that some DFT structures may become useless, with consequent dramatic impact on test effectiveness and product quality. We borrow the fault secure property and we will show that it guarantees that no escapes or false acceptance of faulty products may occur because of faults within the DFT structures.
Cecilia Metra, Martin Omaña 0001
DATE3
2004 Low-Area On-Chip Circuit for Jitter Measurement in a Phase-Locked Loop
José Manuel Cazeaux, Martin Omaña 0001, Cecilia Metra
IOLTS2
2004 Hardware Reconfiguration Scheme for High Availability Systems
Cecilia Metra, Martin Omaña 0001, Andrea Pagni
IOLTS3
2004 Risks Associated with Faults within Test Pattern Compactors and Their Implications on Testing
abstract
We analyze the risks associated with faults affecting a key component block of today's DFT structures, that is the compactor. We show that, because of compactors' internal faults, DFT structures may become useless, with consequent dramatic impact on test effectiveness, product quality and defect level. We borrow the well-known fault secure property for DFT compactors and we show that it guarantees that no escapes or false acceptance of faulty products may occur because of faults within compactors. We discuss the fault secureness of some recently proposed compactors and we provide general design rules to be followed to guarantee fault secureness.
Cecilia Metra, Martin Omaña 0001
ITC3
2004 Model for Transient Fault Susceptibility of Combinational Circuits
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
J. Electron. Test.1
2003 High Speed and Highly Testable Parallel Two-Rail Code Checker
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
DATE1
2003 A Model for Transient Fault Propagation in Combinatorial Logic
abstract
Transient faults (TFs) are increasingly affecting micro-electronic devices as their size decreases. During the design phase, the robustness of circuits for high reliability applications with respect to this kind of faults is generally validated through simulations. However, traditional HSPICE like simulators are too slow for the task of simulating the effects of TFs on large circuits. In this paper, we present a novel mathematical model to accurately estimate the possible propagation of transient fault-due glitches through a CMOS combinational circuit, which is suitable to be used into a new simulation tool able to provide good accuracy, while significantly speeding up simulations, with respect to HPSICE. In particular, our model allows approximately 90% accuracy with respect to HSPICE simulations.
Martin Omaña 0001, Giacinto Papasso, Daniele Rossi 0001, Cecilia Metra
IOLTS1
2003 Novel Transient Fault Hardened Static Latch
abstract
University of Bologna
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
ITC1