Cecilia Metra

dblp:87/6317 · DBLP profile ↗
← Back
96ranked-venue papers
31as first author
7since 2021 · last 2026
0000-0002-1408-5725ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 94 · 30 first-author · 7 since 2021Software engineering, systems software and programming languages · 34 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2026 Reliability of Multilevel Cell Phase-Change Memories for AI Implementations
G. Settegrani, R. Gattoni, Sara Cretí, Matteo Naldi, Martin Omaña 0001, Cecilia Metra, I. S. Troja
IOLTS6
2025 Non-Functional Properties in HPC Systems: Design Exploration of Energy, Power, and Reliability
abstract
Modern HPC systems must be designed considering different parameters, which include cost, performance, and throughput, as well as non-functional properties, such as power/energy consumption and reliability. This paper describes the work performed and the results achieved by the partners of the Italian National Research Center for HPC, Big Data and Quantum Computing in the frame of the sub-project dealing with Future HPC architectures and solutions. The work in this subproject focused on advanced design and monitoring techniques for devising energy- and power-efficient, reliable parallel architectures based on open standards (e.g., RISC-V) and design space exploration techniques and tools. This paper provides a summary of the achieved results and developed products stemming from the activities of the different partners.
Giovanni Agosta, Enrico Bini, Davide Baroffio, Carlo Brandolese, Michele Castrovilli, Daniele Cattaneo 0002, Daniele Cesarini, William Fornaciari, Andrea Galimberti, Alberto Garfagnini, Arsenii Gavrikov, Francesco Iannone, Marco Lapegna, Tomas Antonio López, Gabriele Magnani, Gabriele Mencagli, Cecilia Metra, Martin Omaña 0001, Filippo Palombi, Federico Reghenzani, Josie E. Rodriguez Condia, A. Serafini, Matteo Sonza Reorda, Davide Zoni, Giuseppe Zummo
DSD17
2025 On-Line Test of Fully Integrated Voltage Regulators for High Performance Microprocessors of Autonomous Systems
abstract
High performance microprocessors employed within autonomous systems usually adopt Fully Integrated Voltage Regulators (FIVRs) to enable the central power management unit to control individually the voltage of different microprocessor power domains, thus enabling a significant improvement of performance-per-Watt. However, since FIVRs are mainly implemented on the microprocessor die, they suffer from reliability problems due to the scaling of microelectronic technology. In particular, faults and aging may affect FIVRs during their in-field operation, possibly compromising the microprocessor correct operation in the field, with possible catastrophic effects if the microprocessor is executing safety critical functionalities of autonomous systems. In this paper, first we will analyze the effects of most likely faults and Bias Temperature Instability (BTI) aging mechanisms possibly affecting the FIVR during its operation in the field. We will show that almost 60% of FIVR faults may result in an incorrect output voltage, possibly compromising the microprocessor correct operation. Moreover, we will show that, due to BTI, the time it will take for the FIVR to change its output voltage in response to changes of its input reference voltage may exceed the maximum tolerable time guaranteeing the microprocessor correct operation. Based on these achieved results, we will then propose a monitor to enable the FIVR on-line test. In particular, upon the generation of an incorrect FIVR output voltage, or in case of a degraded FIVR response time to changes of the reference voltage (due to the occurrence of the considered faults or BTI), the monitor generates an output error message, that can then be adopted to activate proper recovery actions to guarantee the FIVR reliable operation, thus avoiding that faults and BTI possibly affecting FIVRs during their operation in the field can compromise the microprocessor correct operation, with possible dramatic consequences if the microprocessor is executing safety-critical functionalities.
Martin Omaña 0001, A. Menghi, A. Stefani, E. Vicini, Cecilia Metra, G. Froio, S. Petrucci
IOLTS5
2024 Silent Data Corruption and Reliability Risks due to Faults Affecting High Performance Microprocessors' Caches
abstract
Error Correcting Codes (ECCs) are frequently adopted to guarantee the correct operation in the field of caches of high performance microprocessors. They require the addition of proper encoding/decoding blocks (referred to as checkers) to the cache array. The occurrence of faults affecting such checkers has been typically neglected so far, due to their limited area compared to the cache array. This may be no longer acceptable, due to the increasing likelihood of faults possibly affecting microprocessors implemented by deeply scaled technologies, and due to the increasing requirements in terms of reliability of several applications (e.g., data centers, autonomous vehicles, unmanned robots, etc.). Based on these considerations, in this paper we analyze the effects of bridging faults possibly affecting the ECCs’ checkers, for two frequently adopted kinds of ECCs. We will show that the $68 \%$ (or the $61 \%$) of BFs possibly affecting the considered ECCs’ checkers are critical, since they may either inhibit the ECC correction ability of incorrect words read from the cache, or introduce errors in otherwise correct words read from the cache, with consequent risks for silent data corruption and microprocessor reliability. The remaining $\mathbf{3 2 \%}$ (or $39 \%$) of BFs may remain latent and accumulate with following faults or aging conditions affecting the cache, with consequent future risks for silent data corruption and microprocessor reliability. We then introduce a possible scheme to detect on line the occurrence of critical BFs that, compared to an alternate solution presented in the literature, features significantly lower impact on the ECC checker delay and area.
Martin Omaña 0001, A. Manfredi, Cecilia Metra, R. Locatelli, M. Chiavacci, S. Petrucci
IOLTS3
2024 Reliability of AI in Predicting the State of Health of Li-Ion Batteries*
Sara Cretí, Martin Omaña 0001, Cecilia Metra, Gianni Borelli
IOLTS3
2024 On the Reliability of Clock Monitoring Units for Safety Critical Applications' Microcontrollers
abstract
In multi-core microcontrollers adopted for safety critical applications, such as automotive, the frequency of clock signals is typically monitored by dedicated Clock Monitor Units (CMUs), whose correct operation is essential for the microcontroller correct operation and system’s safety. We analyse the effects of resistive bridging faults and transient faults possibly affecting a typical CMU. We will show that $39 \%$ of the considered CMU resistive bridging faults do not result in a CMU output error message, thus remaining latent. Depending on the value of their connecting resistance, up to the $49 \%$ of the latent bridging faults can make the CMU unable to indicate the presence of a monitored clock signal with an incorrect frequency, with potential catastrophic consequences for the microcontroller correct operation and system’s safety. Instead, as for transient faults, we will show that they can be reasonably considered to do not constitute a serious risk for system’s safety.
Massimiliano Zhupa, Matteo Naldi, Martin Omaña 0001, Cecilia Metra
IOLTS4
2022 Novel BTI Robust Ring-Oscillator-Based Physically Unclonable Function
abstract
Physically Unclonable Functions (PUFs) have become a promising low-cost solution for authentication and key generation in cryptosystems. However, it has been shown in the literature that the reliability of PUFs is undermined by aging mechanisms, such as Bias Temperature Instability (BTI), which may compromise their correct operation. In this paper, we present a novel ring oscillator (RO) based PUF design that is robust against BTI degradation, hereinafter referred to as Low-sensitive-BTI RO - LBTIRO. We compare our proposed LBTIRO to the standard RO-based PUF and to an alternative NBTI-robust RO-PUF recently presented in the literature, for 90 nm and 32 nm CMOS technology nodes. We show that, for considered technology nodes, our proposed LBTIRO features a higher robustness against BTI. Particularly, our LBTIRO enables a reduction of the impact of BTI on the oscillation frequency over circuit lifetime, which reaches 85.3% and 72.1% against the standard RO and the recent alternate solution, respectively, for the 90 nm technology. Moreover, we show that our proposed LBTIRO features a reduction in terms of power consumption if compared to the alternative NBTIrobust RO-PUF.
Marco Grossi, Martin Omaña 0001, Daniele Rossi 0001, Biagio Marzulli, Cecilia Metra
IOLTS5
2019 Low-Cost Strategy for Bus Propagation Delay Reduction
Martin Omaña 0001, S. Govindaraj, Cecilia Metra
J. Electron. Test.3
2019 Fault-Tolerant Inverters for Reliable Photovoltaic Systems
abstract
Photovoltaic (PV) systems are increasingly adopted as a source of green energy. Due to the high economic investment that they usually involve, their reliability is becoming of concern. Recent studies have proven that faults likely to affect the inverters of PV systems in the field can dramatically reduce the energy delivered to the load. We first present a self-checking monitor to detect faults affecting the inverter in the field, as well as faults affecting the monitor itself. Then, we propose a hardware reconfiguration scheme that activated upon the monitor's generation of an alarm message after the occurrence of faults, enables to avoid their impact on the energy efficiency of the PV system. Our recovery scheme is also self-checking with respect to faults affecting itself. Therefore, our monitoring and reconfiguration schemes provide inverters with fault-tolerance ability, thus enabling to meet the increasing demand for reliable PV systems.
Martin Omaña 0001, Alessandro Fiore, Marco Mongitore, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.4
2017 New Approaches for Power Binning of High Performance Microprocessors
abstract
The significant process parameter variations occurring during fabrication of high performance sequential circuits, such as microprocessors, are posing relevant uncertainties on the power that such circuits will consume in the field, while executing workloads typical for the diverse products they are oriented to (e.g., cellular phones, notebooks, servers, etc). On the other hand, different kinds of products have different constraints on the maximal power that could be consumed during the execution of typical workloads, due to diverse needs in terms of charge autonomy, heat dissipation, etc. Consequently, the power that will be consumed by microprocessors during the execution of typical workloads in the field needs to be accurately characterized at the end of fabrication. Such a power consumption characterization (hereinafter referred to as “power binning”), will enable to classify microprocessors in “power bins”, each one containing microprocessors suitable for different kinds of products, thus enabling to introduce them all into the market for different kinds of products. Based on these considerations, in this paper we propose an approach to characterize accurately at the end of fabrication, and at low-cost (in terms of characterization time), the power that microprocessors will consume in the in-field during the execution of workloads typical for different kinds of products. Our approach exploits scan-based Logic Built-In Self-Test (LBIST) to apply to microprocessors' sequential blocks test vectors that induce on their internal nodes an activity factor (AF) similar to that experienced during the in-field execution of workloads typical for different kinds of products, thus enabling to perform power binning by simply measuring their consumed power. Our approach enables to scale the AF from 0 percent up to 97.6 percent (on average for the considered benchmark circuits) compared to conventional LBIST, with a granularity of the 2 percent, thus enabling to emulate accurately the AF induced by workloads typical of a wide range of products. We propose a hardware implementation for our approach requiring a limited area overhead (lower than 3 percent) over conventional LBIST.
Martin Omaña 0001, Marco Padovani, Kreshnik Veliu, Cecilia Metra, Juergen Alt, Rajesh Galivanche
IEEE Trans. Computers4
2017 Scalable Approach for Power Droop Reduction During Scan-Based Logic BIST
abstract
The generation of significant power droop (PD) during at-speed test performed by Logic Built-In Self Test (LBIST) is a serious concern for modern ICs. In fact, the PD originated during test may delay signal transitions of the circuit under test (CUT): an effect that may be erroneously recognized as delay faults, with consequent erroneous generation of test fails and increase in yield loss. In this paper, we propose a novel scalable approach to reduce the PD during at-speed test of sequential circuits with scan-based LBIST using the launch-on-capture scheme. This is achieved by reducing the activity factor of the CUT, by proper modification of the test vectors generated by the LBIST of sequential ICs. Our scalable solution allows us to reduce PD to a value similar to that occurring during the CUT in field operation, without increasing the number of test vectors required to achieve a target fault coverage (FC). We present a hardware implementation of our approach that requires limited area overhead. Finally, we show that, compared with recent alternative solutions providing a similar PD reduction, our approach enables a significant reduction of the number of test vectors (by more than 50%), thus the test time, to achieve a target FC.
Martin Omaña 0001, Daniele Rossi 0001, Filippo Fuzzi, Cecilia Metra, Chandra Tirumurti, Rajesh Galivanche
IEEE Trans. Very Large Scale Integr. Syst.4
2016 Inverters' self-checking monitors for reliable photovoltaic systems
Martin Omaña 0001, A. Fiore, Cecilia Metra
DATE3
2016 Low-Cost and High-Reduction Approaches for Power Droop during Launch-On-Shift Scan-Based Logic BIST
abstract
During at-speed test of high performance sequential ICs using scan-based Logic BIST, the IC activity factor (AF) induced by the applied test vectors is significantly higher than that experienced during its in field operation. Consequently, power droop (PD) may take place during both shift and capture phases, which will slow down the circuit under test (CUT) signal transitions. At capture, this phenomenon is likely to be erroneously recognized as due to delay faults. As a result, a false test fail may be generated, with consequent increase in yield loss. In this paper, we propose two approaches to reduce the PD generated at capture during at-speed test of sequential circuits with scan-based Logic BIST using the Launch-On-Shift scheme. Both approaches increase the correlation between adjacent bits of the scan chains with respect to conventional scan-based LBIST. This way, the AF of the scan chains at capture is reduced. Consequently, the AF of the CUT at capture, thus the PD at capture, is also reduced compared to conventional scan-based LBIST. The former approach, hereinafter referred to as Low-Cost Approach (LCA), enables a 50 percent reduction in the worst case magnitude of PD during conventional logic BIST. It requires a small cost in terms of area overhead (of approximately 1.5 percent on average), and it does not increase the number of test vectors over the conventional scan-based LBIST to achieve the same Fault Coverage (FC). Moreover, compared to three recent alternative solutions, LCA features a comparable AF in the scan chains at capture, while requiring lower test time and area overhead. The second approach, hereinafter referred to as High-Reduction Approach (HRA), enables scalable PD reductions at capture of up to 87 percent, with limited additional costs in terms of area overhead and number of required test vectors for a given target FC, over our LCA approach. Particularly, compared to two of the three recent alternative solutions mentioned above, HRA enables a significantly lower AF in the scan chains during the application of test vectors, while requiring either a comparable area overhead or a significantly lower test time. Compared to the remaining alternative solutions mentioned above, HRA enables a similar AF in the scan chains at capture (approximately 90 percent lower than conventional scan-based LBIST), while requiring a significantly lower test time (approximately 4.87 times on average lower number of test vectors) and comparable area overhead (of approximately 1.9 percent on average).
Martin Omaña 0001, Daniele Rossi 0001, Edda Beniamino, Cecilia Metra, Chandra Tirumurti, Rajesh Galivanche
IEEE Trans. Computers4
2015 Intermittent and Transient Fault Diagnosis on Sparse Code Signatures
abstract
Failure diagnosis of field returns typically requires high quality test stimuli and assumes that tests can be repeated. For intermittent faults with fault activation conditions depending on the physical environment, the repetition of tests cannot ensure that the behavior in the field is also observed during diagnosis, causing field returns diagnosed as no-trouble-found. In safety critical applications, self-checking circuits, which provide concurrent error detection, are frequently used. To diagnose intermittent and transient faulty behavior in such circuits, we use the stored encoded circuit outputs in case of a failure (called signatures) for later analysis in diagnosis. For the first time, a diagnosis algorithm is presented that is capable of performing the classification of intermittent or transient faults using only the very limited amount of functional stimuli and signatures observed during operation and stored on chip. The experimental results demonstrate that even with these harsh limitations it is possible to distinguish intermittent from transient faulty behavior. This is essential to determine whether a circuit in which failures have been observed should be subject to later physical failure analysis, since intermittent faulty behavior has been diagnosed. In case of transient faulty behavior, it may still be operated reliably.
Michael A. Kochte, Atefe Dalirsani, Andrea Bernabei, Martin Omaña 0001, Cecilia Metra, Hans-Joachim Wunderlich
ATS5
2015 Low-Cost On-Chip Clock Jitter Measurement Scheme
abstract
In this paper, we present a low-cost, on-chip clock jitter digital measurement scheme for high performance microprocessors. It enablesin situjitter measurement during the test or debug phase. It provides very high measurement resolution and accuracy, despite the possible presence of power supply noise (representing a major source of clock jitter), at low area and power costs. The achieved resolution is scalable with technology node and can in principle be increased as much as desired, at low additional costs in terms of area overhead and power consumption. We show that, for the case of high performance microprocessors employing ring oscillators (ROs) to measure process parameter variations (PPVs), our jitter measurement scheme can be implemented by reusing part of such ROs, thus allowing to measure clock jitter with a very limited cost increase compared with PPV measurement only, and with no impact on parameter variation measurement resolution.
Martin Omaña 0001, Daniele Rossi 0001, Daniele Giaffreda, Cecilia Metra, Asifur Rahman, Simon M. Tam
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Modeling and Detection of Hotspot in Shaded Photovoltaic Cells
abstract
In this paper, we address the problem of modeling the thermal behavior of photovoltaic (PV) cells undergoing a hotspot condition. In case of shading, PV cells may experience a dramatic temperature increase, with consequent reduction of the provided power. Our model has been validated against experimental data, and has highlighted a counterintuitive PV cell behavior, that should be considered to improve the energy efficiency of PV arrays. Then, we propose a hotspot detection scheme, enabling to identify the PV module that is under hotspot condition. Such a scheme can be used to avoid the permanent damage of the cells under hotspot, thus their drawback on the power efficiency of the entire PV system.
Daniele Rossi 0001, Martin Omaña 0001, Daniele Giaffreda, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Impact of Bias Temperature Instability on Soft Error Susceptibility
abstract
In this paper, we address the issue of analyzing the effects of aging mechanisms on ICs' soft error (SE) susceptibility. In particular, we consider bias temperature instability (BTI), namely negative BTI in pMOS transistors and positive BTI in nMOS transistors that are recognized as the most critical aging mechanisms reducing the reliability of ICs. We show that BTI reduces significantly the critical charge of nodes of combinational circuits during their in-field operation, thus increasing the SE susceptibility of the whole IC. We then propose a time dependent model for SE susceptibility evaluation, enabling the use of adaptive SE hardening approaches, based on the ICs lifetime.
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Alessandro Paccagnella
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Clock Faults Induced Min and Max Delay Violations
Daniele Rossi 0001, Martin Omaña 0001, José Manuel Cazeaux, Cecilia Metra
J. Electron. Test.4
2013 Novel approach to reduce power droop during scan-based logic BIST
abstract
Significant peak power (PP), thus power droop (PD), during test is a serious concern for modern, complex ICs. In fact, the PD originated during the application of test vectors may produce a delay effect on the circuit under test signal transitions. This event may be erroneously recognized as presence of a delay fault, with consequent generation of an erroneous test fail, thus increasing yield loss. Several solutions have been proposed in the literature to reduce the PD during test of combinational ICs, while fewer approaches exist for sequential ICs. In this paper, we propose a novel approach to reduce peak power/power droop during test of sequential circuits with scan-based Logic GIST. In particular, our approach reduces the switching activity of the scan chains between following capture cycles. This is achieved by an original generation and arrangement of test vectors. The proposed approach presents a very low impact on fault coverage and test time, while requiring a very low cost in terms of area overhead.
Martin Omaña 0001, Daniele Rossi 0001, Filippo Fuzzi, Cecilia Metra, Chandra Tirumurti, R. Galivache
ETS4
2013 Low Cost Concurrent Error Detection Strategy for the Control Logic of High Performance Microprocessors and Its Application to the Instruction Decoder
Daniele Rossi 0001, Martin Omaña 0001, G. Garrammone, Cecilia Metra, Abhijit Jas, Rajesh Galivanche
J. Electron. Test.4
2013 Low Cost NBTI Degradation Detection and Masking Approaches
abstract
Performance degradation of integrated circuits due to aging effects, such as Negative Bias Temperature Instability (NBTI), is becoming a great concern for current and future CMOS technology. In this paper, we propose two monitoring and masking approaches that detect late transitions due to NBTI degradation in the combinational part of critical data paths and guarantee the correctness of the provided output data by adapting the clock frequency. Compared to recently proposed alternative solutions, one of our approaches (denoted as Low Area and Power (LAP) approach) requires lower area overhead and lower, or comparable, power consumption, while exhibiting the same impact on system performance, while the other proposed approach (denoted as High Performance (HP) approach) allows us to reduce the impact on system performance, at the cost of some increase in area and power consumption.
Martin Omaña 0001, Daniele Rossi 0001, Nicolò Bosio, Cecilia Metra
IEEE Trans. Computers4
2013 Faults Affecting Energy-Harvesting Circuits of Self-Powered Wireless Sensors and Their Possible Concurrent Detection
abstract
We analyze the effects of faults on an energy-harvesting circuit (EHC) providing power to a wireless biomedical multisensor node. We show that such faults may prevent the EHC from producing the power supply voltage level required by the multisensor node. Then, we propose a low-cost (in terms of power consumption and area overhead) additional circuit monitoring the voltage level produced by the EHC continuously, and concurrently with the normal operation of the device. Such a monitor gives an error indication if the generated voltage falls below the minimum value required by the sensor node to operate correctly, thus allowing the activation of proper recovery actions to guarantee system fault tolerance. The proposed monitor is self-checking with regard to the internal faults that can occur during its in-field operation, thus providing an error signal when affected by faults itself.
Martin Omaña 0001, Daniele Rossi 0001, Daniele Giaffreda, Roberto Specchia, Cecilia Metra, Marcin Marzencki, Bozena Kaminska
IEEE Trans. Very Large Scale Integr. Syst.5
2012 New Design for Testability Approach for Clock Fault Testing
abstract
We propose a new design for testability approach for testing clock faults of next generation high performance microprocessors. In fact, it has been shown that conventional manufacturing test is unable to guarantee their detection, although they could compromise the effectiveness of delay fault testing, as well as the microprocessor correct operation in the field. These conditions will of course worsen with technology scaling, due to the expected increase in fault likelihood, included clock faults. To deal with these problems we propose a design for testability approach that, by means of simple modifications to conventional clock buffers, allows clock fault detection through any conventional manufacturing test approach. This is achieved at the cost of very low increase in area and power consumption of clock buffers, and with no additional test cost or impact on the microprocessor performance and in-field operation. We then introduce a possible further modification to clock buffers that, at additional limited costs in terms of area and power consumption, allows their calibration after fabrication in order to compensate for parameter variations possibly occurring during manufacturing, thus minimizing the likelihood of either false test fails, or test misses. As an example, we show the application of our approach to the clock distribution network of the Pentium® 4 microprocessor (Other names and brands may be claimed as property of others). However, it can be applied to the clock distribution of any high performance ASIC, or microprocessor.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
IEEE Trans. Computers1
2011 Error correcting code analysis for cache memory high reliability and performance
abstract
In this paper we address the issue of improving ECC correction ability beyond that provided by the standard SEC/DED Hsiao code. We analyze the impact of the standard SEC/DED Hsiao ECC and for several double error correcting (DEC) codes on area overhead and cache memory access time for different codeword sizes and code-segment sizes, as well as their correction ability as a function of codeword/code-segment sizes. We show the different trade-offs that can be achieved in terms of impact on area overhead, performance and correction ability, thus giving insight to designers for the selection of the optimal ECC and codeword organization/code-segment size for a given application.
Daniele Rossi 0001, N. Timoncini, M. Spica, Cecilia Metra
DATE4
2011 Guest Editors' Introduction: Special Section on Concurrent On-Line Testing and Error/Fault Resilience of Digital Systems
abstract
THE continuous scaling of microelectronic technology, while allowing to integrate increasingly complex and high performance systems on a die, poses new challenges to their reliable operation in the field, due to the increased likelihood of faults and aging phenomena possibly occurring in the field and compromising the system’s correct operation. Several on-line testing and error/fault resilience techniques have been employed in the past to implement highly reliable, fault tolerant systems for mission critical applications, in areas like space, military, automotive, medical, banking, etc. However, new faults and aging phenomena occurring in the field are posing unique on-line testing and error/fault resilience challenges even for mainstream applications, where cost is a crucial factor. This mandates the development and adoption of innovative solutions optimized for cost, power and area.
Cecilia Metra, Rajesh Galivanche
IEEE Trans. Computers1
2011 Low-Cost Dynamic Compensation Scheme for Local Clocks of Next Generation High Performance Microprocessors
abstract
We propose a low cost scheme for the dynamic compensation in the field of undesired skew and duty cycle variations of local clocks of high performance microprocessors and high end ASICs. Compared to alternate approaches, our solution features lower power consumption, smaller compensation error, and a lower or comparable area overhead.
Martin Omaña 0001, Cecilia Metra, Simon M. Tam
IEEE Trans. Very Large Scale Integr. Syst.2
2010 High-Performance Robust Latches
abstract
First, a new high-performance robust latch (referred to as HiPeR latch) is presented that is insensitive to transient faults affecting its internal and output nodes by design, independently of the size of its transistors. Then, a modified version of the HiPeR latch (referred as HiPeR-CG) is proposed that is suitable to be used together with clock gating. Both proposed latches are faster than the latches most recently presented in the literature, while providing better or comparable robustness to transient faults, at comparable or lower costs in terms of area and power, respectively. Therefore, thanks to the good trade-offs in terms of performance, robustness, and cost, our proposed latches are particularly suitable to be adopted on critical paths.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers3
2009 Detecting Multiple Faults in One-Dimensional Arrays of Reversible QCA Gates
Xiaojun Ma 0002, Jing Huang 0001, Cecilia Metra, Fabrizio Lombardi
J. Electron. Test.3
2009 Testing Resistive Opens and Bridging Faults Through Pulse Propagation
abstract
This paper addresses the problems related to resistive opens and bridging faults that lie out of the most critical paths. These faults cannot be detected by traditional delay fault testing because the induced delay defects are not large enough to result in timing violations when the test rate is equal to the nominal operating frequency. In spite of this problem, resistive opens and bridgings should be detected because they may give rise to reliability problems. To detect them, we propose a testing method that is based on the propagation of pulses within the faulty circuit and that exploits the degraded capability of faulty paths to propagate pulses. The effectiveness of our method is analyzed at the transistor level and compared with the use of reduced clock periods to detect the same class of faults. Results show similar performance in the case of resistive opens and better performance in the case of bridgings. Moreover, the proposed approach is not affected by possible problems in the clock distribution.
Michele Favalli, Cecilia Metra
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2009 Accurate Linear Model for SET Critical Charge Estimation
abstract
In this paper, we present an accurate linear model for estimating the minimum amount of collected charge due to an energetic particle striking a combinational circuit node that may give rise to a SET with an amplitude larger than the noise margin of the subsequent gates. This charge value will be referred to as SET critical charge (Q SET ). Our proposed model allows to calculate the Q SET of a node as a function of the size of the transistors of the gate driving the node and the fan-out gate(s), with no need for time costly electrical level simulations. This makes our approach suitable to be integrated into a design automation tool for circuit radiation hardening. The proposed model features 96% average accuracy compared to electrical level simulations performed by HSPICE. Additionally, it highlights that Q SET has a much stronger dependence on the strength of the gate driving the node, than on the node total capacitance. This property could be considered by robust design techniques in order to improve their effectiveness.
Daniele Rossi 0001, José Manuel Cazeaux, Martin Omaña 0001, Cecilia Metra, Abhijit Chatterjee
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Function-Inherent Code Checking: A New Low Cost On-Line Testing Approach for High Performance Microprocessor Control Logic
abstract
We propose an on-line testing approach for the control logic of high performance microprocessors. Rather than adding information redundancy (in the form of error detecting codes), we propose to look for the information redundancy (referred to as Function-Inherent Codes) that the microprocessor control logic may inherently have, due to its required functionality. We will show that this allows to achieve on-line testing at significant savings in terms of area and power consumption, and with lower or comparable impact on system performance and design costs, compared to alternate, traditional on-line testing approaches.
Cecilia Metra, Daniele Rossi 0001, Martin Omaña 0001, Abhijit Jas, Rajesh Galivanche
ETS1
2008 Risks for Signal Integrity in System in Package and Possible Remedies
abstract
We analyze the electrical phenomena that can affect the integrity of the communication among different chips within a System in Package (SiP). We address these issues for a real case, for which electrical parameters are extracted from layout and used to build a netlist employed for electrical characterization. We show that crosstalk, and in particular inductive crosstalk, is the electrical phenomenon mainly affecting signal transmission within the SiP. Then,we evaluate the kinds of errors that can be originated. We show that errors caused by inductive coupling among SiP interconnects can be unidirectional only, thus allowing designers to implement error control coding techniques based on All Unidirectional Error Detecting codes. This allows significant cost reduction over the alternate use of non-unidirectional error detecting codes.
Daniele Rossi 0001, Paolo Angelini, Cecilia Metra, Giovanni Campardo, Gian Pietro Vanalli
ETS3
2008 Reversible Gates and Testability of One Dimensional Arrays of Molecular QCA
Xiaojun Ma 0002, Jing Huang 0001, Cecilia Metra, Fabrizio Lombardi
J. Electron. Test.3
2008 Checkers' No-Harm Alarms and Design Approaches to Tolerate Them
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra
J. Electron. Test.3
2008 Power Consumption of Fault Tolerant Busses
abstract
On-chip interconnects in very deep submicrometer technology are becoming more sensitive and prone to errors caused by power supply noise, crosstalk, delay variations and transient faults. Error-correcting codes (ECCs) can be employed in order to provide signal transmission with the necessary data integrity. In this paper, the impact of ECCs to encode the information on a very deep submicrometer bus on bus power consumption is analyzed. To fulfill this purpose, both the bus wires (with mutual capacitances, drivers, repeaters and receivers) and the encoding–decoding circuitry are accounted for. After a detailed analysis of power dissipation in deep submicrometer fault-tolerant busses using Hamming single ECCs, it is shown that no power saving is possible by choosing among different Hamming codes. A novel scheme, called Dual Rail, is then proposed. It is shown that Dual Rail, combined with a proper bus layout, can provide a reduction of energy consumption. In particular, it is shown how the passive elements of the bus (bottom and mutual wire capacitances), active elements of the bus (buffers) and error-correcting circuits contribute to power consumption, and how different tradeoffs can be achieved. The analysis presented in this paper has been performed considering a realistic bus structure, implemented in a standard 0.13- $\mu{\hbox{m}}$ CMOS technology.
Daniele Rossi 0001, André K. Nieuwland, Steven V. E. S. van Dijk, Richard P. Kleihorst, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.5
2007 Interactive presentation: Pulse propagation for the detection of small delay defects
Michele Favalli, Cecilia Metra
DATE2
2007 Configurable Error Control Scheme for NoC Signal Integrity
abstract
In this paper we propose a novel error control scheme to cope with errors affecting the communication links of a NoC. Our scheme can be configured in Correction Mode, Detection Mode, and Mixed Mode, depending on the particular application, thus allowing to meet different Quality of Service (QoS) levels in terms of error control. For each configuration mode, we propose different error control policies and we consider SEC Hamming codes, SEC/DED Hsiao codes, and Symbol Error Correcting codes. We evaluate advantages and drawbacks of each approach, in terms of signal integrity, area overhead and impact on performance.
Daniele Rossi 0001, Paolo Angelini, Cecilia Metra
IOLTS3
2007 Novel compensation scheme for local clocks of high performance microprocessors
abstract
Clock compensation for process variations and manufacturing defects is a key strategy to achieve high performance of processors and high end ASIC. However, with the increase in process variations and defect densities, clock compensation is becoming increasingly challenging. A clock distribution system also consumes over 30% of the overall chip level power, so every little bit counts, including compensation schemes. In this paper we propose a new scheme for the compensation of undesirable skews and duty-cycle variations of local clocks of high performance microprocessors and high end ASICs. Our scheme performs compensation continuously, during the microprocessor operation, thus allowing also compensation to clock jitters due to environmental influences during operation. Compared to alternate solutions for local clock compensation, our scheme features lower power consumption, smaller compensation error, and a lower or comparable area overhead, while allowing compensation to be accomplished within the same clock cycle of skew or duty-cycle variation.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
ITC1
2007 Novel Approach to Clock Fault Testing for High Performance Microprocessors
abstract
This paper presents a novel approach for testing clock faults for high performance microprocessors. Although such faults have been shown to be likely and could compromise delay fault testing, conventional manufacturing test methodology is unable to guarantee their detection. This paper proposes a modification to the conventional clock buffers allowing standard manufacturing test to detect the faults. This is achieved at the cost of a small increase in area and power consumption of the clock buffers, but with no additional test cost or impact on the microprocessor performance and in-field operation. The approach can be applied to the clock system of any high performance chip or microprocessor.
Cecilia Metra, Martin Omaña 0001, Simon M. Tam
VTS1
2007 Won't On-Chip Clock Calibration Guarantee Performance Boost and Product Quality?
abstract
In today's high performance (multi-GHz) microprocessors' design, on-chip clock calibration features are needed to compensate for electrical parameter variations as a result of manufacturing process variations. The calibration features allow performance boost after manufacturing test and maintain such performance levels during normal operation, thus preserving product quality. This strategy has been proven successful commercially. In this paper, we discuss the impact on performance and product quality of both permanent and transient faults possibly affecting these calibration circuits during manufacturing and normal operation, respectively. In particular, we consider the case of an on-chip clock calibration feature of a commercial high performance microprocessor. We will show that some possible permanent faults may render the on-chip clock calibration schemes useless (in process variations' compensation), while it is impossible for common manufacturing testing to detect this incorrect behavior. This means that a faulty operating microprocessor may pass the testing phase and be put onto the market, with a consequent impact on product quality and increase in Defect Level. Similarly, we will show that some possible transient faults occurring during the microprocessor in-field operation could defeat the purpose of on-chip clock calibration, again resulting in faulty operation of the microprocessor. This has long range implications to microprocessors' design as well, considering that process variations on die, as well as across the process, would worsen with continued scaling. Proper strategies to test these clock calibration features and to guarantee their correct operation in the field cannot be ignored. Possible design approaches to solve this problem will be discussed.
Cecilia Metra, Daniele Rossi 0001
IEEE Trans. Computers1
2007 Latch Susceptibility to Transient Faults and New Hardening Approach
abstract
In this paper we analyze the conditions making Transient Faults (TFs) affecting the nodes of conventional latch structures generate output Soft-Errors (SEs). We investigate the susceptibility to TFs of all latch nodes and identify the most critical one(s). We show that, for standard latches using back-to-back inverters for their positive feedback, the internal nodes within their feedback path are the most critical. Such nodes will be hereafter referred to as internal feedback nodes. Based on this analysis, we first propose a low cost hardened latch that, compared to alternative hardened solutions, is able to filter out completely TFs affecting its internal feedback nodes, while presenting a lower susceptibility to TFs on the other internal nodes. This is achieved at the cost of a reduced robustness to TFs affecting the output node. To overcome this possible limitation (especially for systems for high reliability applications), we propose another version of our latch that, at the cost of a small area and power consumption increase compared to our first solution, improves also the robustness of the output node, which can be higher than that of alternative hardened solutions. Additionally, both proposed latches present a comparable or higher robustness of the input node than alternative solutions and provide a lower or comparable power-delay product and area overhead than classical implementations and alternative hardened solutions.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers3
2006 Low-cost and highly reliable detector for transient and crosstalk faults affecting FPGA interconnects
abstract
In this paper we present a novel circuit for the online detection of transient and crosstalk faults affecting the interconnects of systems implemented using Field Programmable Gate-Arrays (FPGAs). The proposed detector features self-checking ability with respect to faults possibly affecting itself, thus being suitable for systems with high reliability requirements, like those for space applications. Compared to alternate solutions, the proposed circuit requires a significantly lower area overhead, while implying a comparable, or lower, impact on system performance. We have verified our circuit operation and self-checking ability by means of post-layout simulations.
Martin Omaña 0001, José Manuel Cazeaux, Daniele Rossi 0001, Cecilia Metra
DATE4
2006 Analysis of the impact of bus implemented EDCs on on-chip SSN
abstract
In this paper, we analyze the impact of error detecting codes, implemented on an on-chip bus, on the on-chip simultaneous switching noise (SSN). First, we analyze in detail how SSN is impacted by different bus transitions, pointing out its dependency on the number and placement of switching wires. Afterwards, we present an analytical model that we have developed in order to estimate the SSN, and that we prove to be very accurate in SSN prediction. Finally, by employing the developed model, we estimate the SSN due to different EDCs implemented on an on-chip bus. In particular, we highlight how their differences in the number of switching wires, bus parallelism and codewords influence the on-chip SSN.
Daniele Rossi 0001, Carlo Steiner, Cecilia Metra
DATE3
2006 Path (Min) Delay Faults and Their Impact on Self-Checking Circuits' Operation
abstract
Min delay violations are traditionally not modeled as possible faults as a result of manufacturing defects. Usually, path delay faults are implicitly assumed to be paths' max delay violations. This, in turn, is based on the assumption that min delay violations are designed out. Most previous manufacturing defect/fault analysis works have not considered their effect on clock circuits. More recently, as burn-in becomes ineffective and process variations become more of an issue, latent defects, device degradation or wear out in the field would potentially also cripple the clock distribution network. Consequently, we should start considering also path (min) delay faults when designing on-line testable circuits, similar to what we currently do for path (max) delay faults. The challenges that this poses to the existing on-line testing strategies are discussed. Examples showing the possible incorrect behavior of a self-checking circuit as a result of this kind of faults are given. New on-line testing strategies should consequently be devised to deal with these faults
Cecilia Metra, Martin Omaña 0001, Daniele Rossi 0001, José Manuel Cazeaux
IOLTS1
2006 Checker No-Harm Alarm Robustness
abstract
In this paper we evaluate the probability that a transient fault (TF), multiple or single, affecting a checker of a self-checking circuit, gives rise to an unnecessary error indication (no-harm alarm). A new property (no-harm alarm robustness) has been defined that, in case of a fault affecting a self-checking circuit (SCC), guarantees that we can determine whether the fault is affecting the functional block, or the checker itself, and whether such a fault is a transient or a permanent fault. Finally, we propose a possible solution implementing the defined property. Its behavior has been verified by means of HSpice simulations, and we evaluate its cost in terms of area overhead and introduced delay
Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Andrea Pagni
IOLTS3
2006 Guest Editors' Introduction: Special Section on Design and Test of Systems-on-Chip (SoC)
abstract
IT is with great pleasure that we introduce the special section on Design and Test of Systems-on-Chips (SoC) to the readership of the IEEE Transactions on Computers. This special section consists of eight papers that have been selected to cover a wide spectrum of techniques and applications which are encountered in the design, manufacturing, assembly, and test of today’s SoC. These papers are authored by outstanding researchers and cover experimental and speculative topics. As with all special sections, these topics are only representative of the publically available literature currently provided by the technical community. Systems-on-a-Chip (SoC) represent a rapidly growing and promising field in the electronic and computer industry. Such tremendous growth is the result of significant advances in microelectronic technology that make it possible to build on the same silicon substrate complex systems, including electronic (analog, digital, and mixed mode), mechanical, optical, RF, and microwave cores, as well as sensors, actuators, and software-based systems. As a result, SoCs are very complex hardware/software systems, offering high-performance features. Examples of possible applications include wireless systems, real-time control systems, space exploration systems, and others. SoCs offer the inherent advantages in which computers and their digital domain can be merged to a variety of technologies and applications which, in the past, were attained at boardlevel. The design and test of such complex systems, however, still constitutes a major challenge. From a design point of view, the ability to have a correctly functioning SOC depends on the ability to design and analyze a mixed-technology and to properly account for the interactions among the various cores. Proper management of interfaces between diverse cores and synchronization are examples of the problems to be faced. The unavailability of proper design and simulation tools for such complex systems and the limited resources for different technologies (such as mixed-signal systems) make such an effort rather difficult. The relation between the different modules of an SoC must be properly established using advanced techniques whose technological basis is just emerging. It is expected that these techniques will be highly inter and intradisciplinary in nature, thus involving designers with different backgrounds. Configurability and programmability of the cores in an SoC suggest that wide applicability of these systems is indeed possible with great flexibility in integration. A further issue that designers are confronting is the evaluation of different configurations associated with the high density integration of cores. The merging of different technologies (such as digital and analog) on a single chip is also of high speculative interest because the manufacturing and organization of these systems is in its infancy. These new features must be evaluated at the early design stages because they have a considerable effect on the performance of SoCs as well as their viability for cost-effective implementation. From a testing point of view, the development of proper test access mechanisms is one of the major challenges to be faced in the near future, as indicated also by the 1999 International Technology Roadmap for Semiconductors. The test access mechanisms must be compatible with the IEEE P1500 standard, which was developed for embedded core testing and which leaves the problem of the Test Access Mechanism (TAM) design to the system integrator. Several test access mechanisms have been proposed, including dedicated test bus, multiplexed access, etc., but the goal is to find a solution allowing the best trade-off between the test quality and cost (including the testing time). To evaluate test quality, however, complex failure mechanisms which might occur in such complex systems should also be evaluated. As an example, the possible occurrence of undesired coupling, noise, and skews between clock signals of diverse cores are some of the simplest failures which may affect the operation of the interfaces among the cores. Proper models, fault simulation tools, and test quality figures are needed. Testing time should also be reduced. In fact, while the design and test of SoCs requires a long time (similar to any computer system), the time-to-market of an SoC should be kept as short as possible to meet today’s consumers’ changing requirements. The high performance possibly offered by the integration of complex systems on the same chip makes SoCs very promising for real-time applications, like control systems for automotive, avionic, space, chemical plants, etc. As an example, the first prototypes of SOCs, including electronic and mechanical cores, have already been employed by NASA for space exploration missions. Such a promising application potential, however, poses the problem of reliable design and verification as well as correct design and test. In the past, fault tolerance has been generally adopted to electronic systems for many critical mission applications. The possible adoption of fault tolerance techniques for SoCs (to include not only electronics, digital IEEE TRANSACTIONS ON COMPUTERS, VOL. 55, NO. 2, FEBRUARY 2006 97
Jien-Chung Lo, Cecilia Metra, Fabrizio Lombardi
IEEE Trans. Computers2
2005 On Transistor Level Gate Sizing for Increased Robustness to Transient Faults
abstract
In this paper we present a detailed analysis on how the critical charge (Q/sub crit/) of a circuit node, usually employed to evaluate the probability of transient fault (TF) occurrence as a consequence of a particle hit, depends on transistors' sizing. We derive an analytical model allowing us to calculate a node's Q/sub crit/ given the size of the node's driving gate and fan-out gate(s), thus avoiding time costly electrical level simulations. We verified that such a model features an accuracy of the 97% with respect to electrical level simulations performed by HSPICE. Our proposed model shows that Q/sub crit/ depends much more on the strength (conductance) of the gate driving the node, than on the node total capacitance. We also evaluated the impact of increasing the conductance of the driving gate on TFs' propagation, hence on soft error susceptibility (SES). We found that such a conductance increase not only improves the TF robustness of the hardened node, but also that of the whole circuit.
José Manuel Cazeaux, Daniele Rossi 0001, Martin Omaña 0001, Cecilia Metra, Abhijit Chatterjee
IOLTS4
2005 Load and Logic Co-Optimization for Design of Soft-Error Resistant Nanometer CMOS Circuits
abstract
Technology scaling has led to reduced noise margins and increased susceptibility of logic circuits to transient errors. In this paper, a novel methodology to increase the robustness of combinational circuits to transient errors is proposed. The number of errors propagated to the primary outputs (POs) is minimized by adding optimal amounts of capacitive loading to the POs of the logic circuit. Using a novel delay-assignment-variation (DAV) based optimization methodology, the sizes, supply voltages and threshold voltages of internal gates (not primary outputs) are chosen to minimize the energy and delay overhead due to the added loads. Experiments on ISCAS'85 benchmarks show that 79.3% soft-error reduction can be obtained on the average with modest increase in circuit delay and energy. Comparison with other techniques shows that our technique has a much better energy-delay-reliability trade-off compared to others.
Yuvraj Singh Dhillon, Abdulkadir Utku Diril, Abhijit Chatterjee, Cecilia Metra
IOLTS4
2005 Coding Techniques for Low Switching Noise in Fault Tolerant Busses
abstract
As device geometries shrink, power supply voltage decreases, and chip complexity increases, the noise induced by the increased amount of simultaneously switching devices (especially the strong bus drivers (SSN)), is becoming crucial in determining the signal integrity of a system. In this paper we propose ways of merging transition reducing coding techniques with coding techniques for fault tolerant busses (implementing either error detecting codes and error recovery, or correcting codes). In particular, we focus on merging bus-invert code along with the employed error detection or correction coding technique, and show that the maximum number of simultaneous switching drivers can be drastically reduced, thus reducing the SSN and increasing signal integrity. Furthermore, we show how, by properly merging the bus invert encoder and the check bit generator, the latency introduced by the proposed coding techniques can be minimized and the number of additional wires can be kept minimal.
André K. Nieuwland, Atul Katoch, Daniele Rossi 0001, Cecilia Metra
IOLTS4
2005 On the Selection of Unidirectional Error Detecting Codes for Self-Checking Circuits' Area Overhead and Performance Optimization
abstract
In this paper we address the issue of optimizing the area overhead and performance of self-checking circuits using all unidirectional error detecting codes (AUEDCs), with no impact on system's reliability. In particular, we propose an error detecting code selection approach that, starting from the consideration of the functional circuit topology, allows us to identify whether or not all output bits can be simultaneously erroneous, thus actually mandating the adoption of an AUEDC. We show that, differently from common expectations, this may frequently be not the case (for approximately the 50% of the considered benchmarks) for all possible internal node stuck-ats, transistor stuck-ons, transistor stuck-opens and resistive bridgings. We then propose a tool that, starting from the (combinational or sequential) circuit high level description, allows us to identify whether or not this is the case and, in particular, which is the maximal number (t) of possibly simultaneously erroneous output bits. Based on this information, a lower redundancy error detecting code (e.g., a t-UEDC) is adopted, rather than an AUEDC, thus generally allowing reducing area overhead and impact on system's performance. Such a code is automatically implemented by our developed tool, whose effectiveness has been verified for benchmark circuits and for a FPGA implemented prototype.
Martin Omaña 0001, O. Losco, Cecilia Metra, Andrea Pagni
IOLTS3
2005 Low Cost Scheme for On-Line Clock Skew Compensation
abstract
In this paper we propose a novel buffer scheme that is able to compensate undesired skews between clocks of a synchronous system in a negligible time upon skew occurrence, thus being suitable also for on-line clock-skew correction. Clock signals are aligned one with respect to the other, starting from a reference clock, and moving forward among physically adjacent clock signals, thus creating no problem of reference clock's routing. Our solution is also able to compensate clock duty-cycle variations, which have been shown very likely in case of faults, for instance bridgings, affecting the clock distribution network. Compared to alternate solutions, our proposed scheme enables significant reductions in area overhead and power consumption, and is suitable for on-line compensation. Therefore, it allows clock skew and duty-cycle fault tolerance, thus increasing process yield and system's reliability.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
VTS3
2005 Self-Checking Voter for High Speed TMR Systems
José Manuel Cazeaux, Daniele Rossi 0001, Cecilia Metra
J. Electron. Test.3
2005 Low Cost and High Speed Embedded Two-Rail Code Checker
abstract
We propose a compact, high-speed, and highly testable parallel two-rail code checker, particularly suitable to implementing embedded checkers. In fact, it requires only two input codewords to satisfy the totally-self-checking or strongly code-disjoint property with respect to a wide set of realistic internal faults. Our checker can be employed to check the correct operation of a connected functional block using the two-rail code, to implement the output two-rail code checker of "normal" checkers for unordered codes, or to join together the error messages produced by various checkers (possibly using different codes) present within the same self-checking system. The behavior of our checker has been verified by means of electrical level simulations (performed using HSPICE), considering both nominal values and statistical variations of electrical parameters. We also propose a possible modification to our checker internal structure that makes it able to provide an output error indication remaining latched until the application of a proper reset signal. Depending on the considered application and recovery technique to be employed upon the generation of an error indication at the checker output, one proposed solution or the other may be preferable.
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
IEEE Trans. Computers3
2004 Are Our Design for Testability Features Fault Secure?
abstract
We analyze the risks associated with faults affecting some common design for testability (DFT) features employed within digital products. We will show that some DFT structures may become useless, with consequent dramatic impact on test effectiveness and product quality. We borrow the fault secure property and we will show that it guarantees that no escapes or false acceptance of faulty products may occur because of faults within the DFT structures.
Cecilia Metra, Martin Omaña 0001
DATE1
2004 Low-Area On-Chip Circuit for Jitter Measurement in a Phase-Locked Loop
José Manuel Cazeaux, Martin Omaña 0001, Cecilia Metra
IOLTS3
2004 New High Speed CMOS Self-Checking Voter
José Manuel Cazeaux, Daniele Rossi 0001, Cecilia Metra
IOLTS3
2004 Hardware Reconfiguration Scheme for High Availability Systems
Cecilia Metra, Martin Omaña 0001, Andrea Pagni
IOLTS1
2004 Impact of ECCs on Simultaneously Switching Output Noise for On-Chip Busses of High Reliability Systems
Daniele Rossi 0001, A. Muccio, André K. Nieuwland, Atul Katoch, Cecilia Metra
IOLTS5
2004 Risks Associated with Faults within Test Pattern Compactors and Their Implications on Testing
abstract
We analyze the risks associated with faults affecting a key component block of today's DFT structures, that is the compactor. We show that, because of compactors' internal faults, DFT structures may become useless, with consequent dramatic impact on test effectiveness, product quality and defect level. We borrow the well-known fault secure property for DFT compactors and we show that it guarantees that no escapes or false acceptance of faulty products may occur because of faults within compactors. We discuss the fault secureness of some recently proposed compactors and we provide general design rules to be followed to guarantee fault secureness.
Cecilia Metra, Martin Omaña 0001
ITC1
2004 Guest Editorial
Cecilia Metra, Matteo Sonza Reorda
J. Electron. Test.1
2004 Model for Transient Fault Susceptibility of Combinational Circuits
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
J. Electron. Test.3
2004 Implications of Clock Distribution Faults and Issues with Screening Them during Manufacturing Testing
abstract
Based on real process data of a reference microprocessor, fault models are derived for the manufacturing defects most likely to affect signals of the clock distribution network. Their probability is estimated with Inductive Fault Analysis performed on the actual layout of the reference microprocessor. The effects of the most likely faults have been evaluated by electrical level simulations. We have found that, contrary to common assumptions, only a small percentage of such faults result in catastrophic failures easily detected during manufacturing testing. On the contrary, the majority of such faults lead to local failures not likely to be detected during manufacturing testing, despite their possibly compromising the microprocessor operation and reliability. In particular, we have found that the clock faults can be detected during manufacturing testing in only 12 percent of cases. Even more surprisingly, we have also found that, in 10 percent of cases, the undetected clock faults also invalidate the testing procedure itself.
Cecilia Metra, Stefano Di Francescantonio
IEEE Trans. Computers1
2004 TMR voting in the presence of crosstalk faults at the voter inputs
abstract
In high reliability systems, the effectiveness of fault tolerant techniques, such as Triple-Modular-Redundancy (TMR), must be validated with respect to the faults that are likely in the current technology. In todays' Integrated Circuits (IC), this is the case of crosstalks, whose importance is growing because of device & interconnect scaling. This paper analyzes the problem of crosstalk faults at the inputs of voters in TMR systems. In particular, possible problems are illustrated, and it is shown that such crosstalk may invalidate the reliability of both voting, and diagnosing operations. The problem is analyzed from a probabilistic point of view. Its occurrence is estimated by using a set of TMR systems obtained with combinational benchmarks as functional modules. The possible problems of such operations are discussed in the presence of crosstalk faults. It is shown that crosstalk may invalidate the reliability of both voting, and diagnosis operations. A probabilistic model of the voting & diagnosis operations in the presence of crosstalk has been developed. Finally, such a model has been used to estimate the probability of voting & diagnosis failures in a set of TMR systems obtained by using combinational benchmarks as functional modules. We have shown that the presence of crosstalk faults at voter inputs may impair both the voting, and the diagnosis mechanisms. This problem has been quantified by applying a probabilistic model of crosstalk fault effects on voting and diagnosis to a set of benchmark circuits. Results show that crosstalk may create a reliability problem for TMR systems. Such a problem can be solved by using on-line testing or design for testability providing additional controllability & observability to the replicated functional units.
Michele Favalli, Cecilia Metra
IEEE Trans. Reliab.2
2003 High Speed and Highly Testable Parallel Two-Rail Code Checker
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
DATE3
2003 A Model for Transient Fault Propagation in Combinatorial Logic
abstract
Transient faults (TFs) are increasingly affecting micro-electronic devices as their size decreases. During the design phase, the robustness of circuits for high reliability applications with respect to this kind of faults is generally validated through simulations. However, traditional HSPICE like simulators are too slow for the task of simulating the effects of TFs on large circuits. In this paper, we present a novel mathematical model to accurately estimate the possible propagation of transient fault-due glitches through a CMOS combinational circuit, which is suitable to be used into a new simulation tool able to provide good accuracy, while significantly speeding up simulations, with respect to HPSICE. In particular, our model allows approximately 90% accuracy with respect to HSPICE simulations.
Martin Omaña 0001, Giacinto Papasso, Daniele Rossi 0001, Cecilia Metra
IOLTS4
2003 Power Consumption of Fault Tolerant Codes: the Active Elements
abstract
On-chip global interconnections in very deep submicron technology (VDSM) ICs are becoming more sensitive and prone to errors caused by power supply noise, crosstalk noise, delay variations and transient faults. Error correcting codes can be employed in order to provide signal transmission with the necessary data integrity. We compared Dual Rail encoding versus Hamming with respect to power consumption of the bus wires themselves (passive capacity model) [Rossi et al., 2002]. In this paper we analyze the contribution of the active elements of both coding schemes. We first present a detailed analysis of the power consumption of an encoded bus, taking into account the bus wires (with mutual capacitances, drivers, repeaters and receivers), as well as the encoding/decoding circuitry. Then we compare the two considered coding technique with respect to the power consumption, and we show how different tradeoffs can be achieved. Our analysis is based on a realistic bus structure, implemented in a 0.13/spl mu/m CMOS technology.
Daniele Rossi 0001, Steven V. E. S. van Dijk, Richard P. Kleihorst, André K. Nieuwland, Cecilia Metra
IOLTS5
2003 Crosstalk Effect Minimization for Encoded Busses
abstract
In this paper we present a technique which allows to reduce the crosstalk-induced delay within busses implementing an error detecting/correcting code. This technique is based on the observation that the maximum delay on an encoded bus is usually due to the check bits that are added to provide the desired error detection/ tolerance ability. These bits, in fact, are computed from the bus information bits by an ad hoc encoder, which adds an extra delay to the crosstalk-induced bus delay. We will show that, by proper placement of the lines carrying the information with respect to those carrying the check bits, it is possible to reduce the effective coupling capacitance due to the Miller effect among adjacent lines. This allows a reduction of propagation delay which, depending on the implemented code, can overcome the 20% with respect to the conventional placement of encoded busses.
L. Di Silvio, Daniele Rossi 0001, Cecilia Metra
IOLTS3
2003 Novel Transient Fault Hardened Static Latch
abstract
University of Bologna
Martin Omaña 0001, Daniele Rossi 0001, Cecilia Metra
ITC3
2003 Guest Editorial
Cecilia Metra, Matteo Sonza Reorda
J. Electron. Test.1
2003 Error Correcting Strategy for High Speed and High Density Reliable Flash Memories
Daniele Rossi 0001, Cecilia Metra
J. Electron. Test.2
2003 Concurrent detection of power supply noise
abstract
We propose a methodology for the concurrent detection of power supply noise affecting a general synchronous system and exceeding a tolerance bound to be chosen according to the system's constraints. Our solution is based on a suitable self-checking scheme which concurrently monitors a signal of the system clock distribution network and which is, by design, able to provide an output error message upon the occurrence of power supply noise. The produced error indication can then be exploited to recover from the detected noise (thus guaranteeing system's correct operation), or to accomplish diagnosis. Our scheme negligibly impacts system's performance, features self-checking ability with respect to a wide set of possible internal faults and keeps on revealing concurrently the occurrence of power supply noise, despite the possible presence of noise affecting also ground.
Cecilia Metra, Luca Schiano, Michele Favalli
IEEE Trans. Reliab.1
2002 Problems Due to Open Faults in the Interconnections of Self-Checking Data-Paths
abstract
In this work, the problem of open faults affecting the interconnections of SC circuits composed by data-path and control is analyzed. In particular it is shown that, in case opens affect control signals, some problems may arise even if both control and data-path signals are concurrently checked. In particular, wrong codewords may be generated at the outputs of multiplexers and registers. To address this problem, new registers and multiplexers are proposed which allow the design data-paths which are TSC with respect to opens (and resistive opens). These components are also TSC with respect to stuck-at, transistor and gross delay faults. They present a good testability with respect to resistive bridgings.
Michele Favalli, Cecilia Metra
DATE2
2002 Self-Checking Scheme for the On-Line Testing of Power Supply Noise
abstract
We propose a self-checking scheme for the on-line testing of power supply noise, exceeding a tolerance bound, to be chosen according to system constraints. Upon the occurrence of such a noise, our scheme provides an output error message, which can be exploited for diagnostic purposes or to recover from the detected noise (thus guaranteeing correct system operation). As far as we are aware, no on-line testing scheme for power supply noise has been proposed up to now. Our scheme has negligible impact upon system performance, features a self-checking ability (with respect to a wide set of possible internal faults) and reveals, on-line, the occurrence of power supply noise, despite the possible presence of noise affecting ground.
Cecilia Metra, Luca Schiano, Bruno Riccò, Michele Favalli
DATE1
2002 Clock Faults? Impact on Manufacturing Testing and Their Possible Detection Through On-Line Testing
abstract
This paper investigates the impact of faults affecting the clock distribution network of synchronous systems on manufacturing testing. Previous researches based on real process data and inductive fault analysis of a reference microprocessor showed that, contrary to common expectations, the majority of clock faults leads to local failures not likely to be detected by manufacturing testing, despite their ability to compromise the microprocessor operation and reliability. This paper shows that clock faults can be detected by means of conventional stuck-at, delay and transition testing in only 12% of cases and that in 10% of cases the undetected clock faults invalidate the testing procedures themselves. In addition, in the 29% of cases, clock faults are likely to cause race conditions that, although generally not considered by delay fault testing, might as well compromise system's correct operation. The possible adoption of on-line testing techniques to avoid such dangerous conditions is finally discussed.
Cecilia Metra, Stefano Di Francescantonio
ITC1
2002 Single Output Distributed Two-Rail Checker with Diagnosing Capabilities for Bus Based Self-Checking Architectures
Michele Favalli, Cecilia Metra
J. Electron. Test.2
2002 On-Chip Clock Faults' Detector
Cecilia Metra, Michele Favalli, Stefano Di Francescantonio, Bruno Riccò
J. Electron. Test.1
2002 Guest Editorial
Dimitris Nikolos, John P. Hayes, Michael Nicolaidis, Cecilia Metra
J. Electron. Test.4
2001 Optimization of error detecting codes for the detection of crosstalk originated errors
abstract
This work applies weight based codes to the detection of crosstalk originated errors. This type of fault, whose importance grows with device scaling may originate errors that are undetectable by the commonly used error detecting codes in VLSI ICs. Conversely, such errors can be easily detected by weight based codes that, however, have smaller encoding capabilities. In order to reduce the cost of these codes, a graph theoretic optimization is used. Moreover new applications of these codes are explored regarding the synthesis of self-checking FSMs, and the detection of errors related to the clock distribution network.
Michele Favalli, Cecilia Metra
DATE2
2001 On-line testing of transient and crosstalk faults affecting interconnections of FPGA-implemented systems
abstract
In this paper we propose a self-checking scheme for the on-line testing of transient and crosstalk faults affecting the interconnections of synchronous systems implemented using Field-Programmable Gate-Arrays (FPGAs). An FPGA prototype has been implemented, whose correct operation has been verified by means of post-layout simulations and experimental measurements.
Cecilia Metra, Andrea Pagano, Bruno Riccò
ITC1
2000 On-Line Testing and Diagnosis of Bus Lines with respect to Intermediate Voltage Values
abstract
Summary form only given. This paper presents a self-checking, on-line testing and diagnosis scheme for bus lines affected by intermediate voltage values possibly due to bridging faults, or to different kinds of faults affecting the bus connected units.
Cecilia Metra, Michele Favalli, Bruno Riccò
DATE1
2000 Bridging Faults in Pipelined Circuits
Michele Favalli, Cecilia Metra
J. Electron. Test.2
2000 Intermediacy Prediction for High Speed Berger Code Checkers
Cecilia Metra, Jien-Chung Lo
J. Electron. Test.1
2000 Self-Checking Detection and Diagnosis of Transient, Delay, and Crosstalk Faults Affecting Bus Lines
abstract
We present a self-checking detection and diagnosis scheme for transient, delay, and crosstalk faults affecting bus lines of synchronous systems. Faults that are likely to result in the connected logic sampling incorrect bus data are on-line detected. The position of the affected line(s) within the considered bus is identified and properly encoded. The proposed scheme is self-checking with respect to a realistic set of possible internal faults, including node stuck-ats, transistor stuck-ons, transistor stuck-opens, resistive bridgings, transient faults, delays and crosstalks.
Cecilia Metra, Michele Favalli, Bruno Riccò
IEEE Trans. Computers1
1999 On the Design of Self-Checking Functional Units Based on Shannon Circuits
abstract
This paper investigates the application of Shannon (BDD) circuits, that feature interesting low-power capabilities, to the design of self-checking functional units. A technique is proposed that, by using a time redundancy approach, makes this kind of circuits totally self-checking with respect to stuck-at-faults. For a set of possibly used pass-transistor-based CMOS implementations, we show that the totally self-checking or the strongly fault secure properties hold for a wider set of realistic faults, including transistors stuck-open/on and bridgings.
Michele Favalli, Cecilia Metra
DATE2
1999 Self-checking scheme for very fast clocks' skew correction
abstract
This paper presents a digital scheme to correct undesired skews between couples of clocks of synchronous systems. Correction is automatically and very fastly performed during system run-time. The proposed scheme is self-checking with respect to a wide set of possible internal (permanent as well as temporary) faults, and is easily scalable to account for different skew tolerance/sensitivity requirements. It is suitable to be implemented in VLSI, very deep submicron technology, as well as using field programmable gate arrays.
Cecilia Metra, Flavio Giovanelli, Mani Soma, Bruno Riccò
ITC1
1999 Bus crosstalk fault-detection capabilities of error-detecting codes for on-line testing
abstract
This paper analyses some of the most common error-detecting codes used in self-checking circuits with respect to the errors induced by crosstalk faults (CFs). The electrical-level behavior of circuits in the presence of CFs has been analyzed by considering these faults as parametric. A logic-level model providing the probability of errors has been abstracted and applied to the case of functional unit outputs (buses). Finally, the probability of detectable and undetectable errors has been evaluated for the parity, two-rail, m-out-of-n, and Berger codes, thus providing some design hint.
Michele Favalli, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.2
1998 Highly Testable and Compact 1-out-of-n Code Checker with Single Output
abstract
This paper presents a novel 1-out-of-n checker that, compared to the other implementations up to now presented, features the advantages of: (i) satisfying the TSC or SCD property with respect to all possible internal faults representative of realistic failures; (ii) presenting a single output line; (iii) requiring significantly lower area overhead.
Cecilia Metra, Michele Favalli, Bruno Riccò
DATE1
1998 Novel Technique for Testing FPGAs
abstract
This paper presents a novel technique for testing Field Programmable Gate Arrays (FPGAs), suitable for use in the case of frequent FPGA reuse and rapid dynamic modifiability of the implemented function.
Cecilia Metra, Michel Renovell, Giovanni A. Mojoli, Jean-Michel Portal, Sandro Pastore, Joan Figueras, Yervant Zorian, Davide Salvi, Giacomo R. Sechi
DATE1
1998 On-line detection of logic errors due to crosstalk, delay, and transient faults
abstract
This paper analyses the problem of systems' on-line testing with respect to logic errors due to crosstalk, delay and transient faults. In particular we show that logic errors due to crosstalk noise between internal, adjacent lines may not be on-line detectable by conventional concurrent error detection techniques using error detecting codes. Hence, a detector is proposed that allows the on-line detection of such logic errors, and that is self-checking with respect to a wide set of possible internal faults representative of realistic failures, including crosstalk, delay, and transient faults.
Cecilia Metra, Michele Favalli, Bruno Riccò
ITC1
1997 On-Line Testing Scheme for Clock's Faults
abstract
This paper proposes an on-line testing scheme for permanent and temporary faults which affect signals of the clock distribution network of synchronous systems, and which make them stuck-at, or change with incorrect frequency or duty-cycle. By means of straightforward modifications, the proposed scheme can be also used to detect on-line undesired skews between couples of clock signals.
Cecilia Metra, Michele Favalli, Bruno Riccò
ITC1
1997 Highly testable and compact single output comparator
abstract
In this paper a single output self-checking n-input comparator is presented. The proposed circuit, which can be used as n-variable two-rail checker or as equality checker features a compact structure, is Totally-Self-Checking or Strongly Code-Disjoint with respect to a wide set of realistic faults, and requires a limited set of input code words for fault detection (thus it can be used to implement also embedded comparators).
Cecilia Metra, Michele Favalli, Bruno Riccò
VTS1
1997 On-line detection of bridging and delay faults in functional blocks of CMOS self-checking circuits
abstract
This paper investigates the detection of parametric bridging and delay faults affecting the functional block of CMOS self-checking circuits (SCCs). As far as these faults are concerned, classical definitions are shown to become ambiguous because they are entirely based on logic considerations. Thus, new definitions are proposed here to consider the analog and dynamic effects of such faults, and to ensure that they do not produce any problem at the system level. Moreover, electrical level design rules aimed at satisfying these conditions are proposed for self-checking circuits with combinational functional blocks. The problem of their practicability and effectiveness is analyzed in detail, and is shown by means of significant examples.
Cecilia Metra, Michele Favalli, Piero Olivo, Bruno Riccò
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1996 Embedded two-rail checkers with on-line testing ability
abstract
This paper addresses the problem of the design of embedded two-rail checkers. In particular a simple additional circuit is proposed which can be used to make a two-rail checker receive all the codewords of the two-rail code, independently of which and how many codewords are produced by its driving functional block or checkers. The proposed circuit features a high online self-testing ability with respect to possible internal faults and a compact structure.
Cecilia Metra, Michele Favalli, Bruno Riccò
VTS1
1996 Sensing circuit for on-line detection of delay faults
abstract
A sensing circuit for on-line testing of delay faults is presented. It can be used to monitor the outputs of circuits that are either general, or designed to be self-checking with respect to steady-state errors. Detailed analyses of the proposed circuit have shown that it is preferable to alternate solutions from the point of view of both the accuracy and the self-testing capability that make it suitable for self-checking applications. Checking architectures for delay faults, making use of the proposed sensing circuit and of standard checkers, are presented.
Michele Favalli, Cecilia Metra
IEEE Trans. Very Large Scale Integr. Syst.2
1995 Design of CMOS checkers with improved testability of bridging and transistor stuck-on faults
Cecilia Metra, Michele Favalli, Piero Olivo, Bruno Riccò
J. Electron. Test.1
1992 CMOS Checkers with Testable Bridging and Transistor Stuck-on Faults
Cecilia Metra, Michele Favalli, Piero Olivo, Bruno Riccò
ITC1