Adrian Evans

dblp:15/4817 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 7 first-author · 9 since 2021Software engineering, systems software and programming languages · 19 · 4 first-author · 5 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 ALIFE-BCI: An Adaptive Low-power Integrated Feature Extractor for Brain-Computer Interfaces
abstract
Brain-Computer Interfaces (BCIs) have the potential to restore motion for patients suffering from spinal cord injuries. Making such systems embedded, or even implantable, imposes strict low power constraints. Feature extraction, which transforms brain signals into intermediate representations before decoding motor intent, is typically the most compute intensive step. In this work, we introduce ALIFE-BCI, an Adaptive Quality Feature Extractor (AQFE), based on a Continuous Wavelet Transform (CWT) that captures the signal dynamics in both the time and frequency domains. The system is optimized with a top-down approach: (i) At the algorithmic level, it implements a piecewise linear approximation of the CWT that allows real-time energy-accuracy trade-offs. (ii) At the architectural level, memory reuse and parallelism are used to balance area and compute performance. (iii) At the circuit level, low-power techniques are used in a 22 nm FDSOI technology physical implementation flow. Three variants, with different levels of parallelism, are explored to extract 960 features at a rate of 10 Hz for a BCI motor application. The optimal variant, with an area of only 0.061 mm2, achieves 0.27 μW/feature at maximum quality, and 0.13 μW/feature at minimum quality, resulting in 8× lower power than existing digital solutions. Combined, these characteristics make the system well-suited for ultra-low-power implantable BCI decoders.
Joe Saad, Ivan Miro Panades, Adrian Evans, Lorena Anghel
DATE3
2025 Enabling a Portable Brain Computer Interface for Rehabilitation of Spinal Cord Injuries
abstract
In clinical trials, brain signal decoders combined with spinal stimulation have shown to be a promising means to restore mobility to paraplegic and tetraplegic patients. To make this technology available for home use, the complex brain signal decoding must be performed using a low-power, portable battery operated system. This case study shows how the decoding algorithm for a Brain-Computer Interface (BCI) system was ported to an embedded platform, resulting in an over 25 x power reduction, compared to the previous implementation, while respecting real-time and accuracy constraints.
Adrian Evans, Victor Roux-Sibillon, Joe Saad, Ivan Miro Panades, Tetiana Aksenova, Lorena Anghel
DATE1
2025 Dose Effects and Mitigation in 28nm FD-SOI Advanced Interface Bus Chiplet Interconnect
abstract
Chiplet technology for 2.5-D/3-D integration is being rapidly adopted, including for System-on-Chips used in mission critical applications. We present a Total Ionizing Dose Effect study of a prototype Advanced Interface Bus (AIB) die-to-die interface in 28nm FD-SOI technology using pulsed x-rays and in-situ monitoring of the dose. Degradation of the maximum working frequency of the interface was observed due to the deposited dose. Using either the core voltage or the body bias voltage, we demonstrate that it is possible to partially recover this frequency loss. In space applications, this compensation, combined with in-situ dose monitoring, can be used to extend the life of 2.5-D/3-D circuits using high-speed die-to-die interfaces.
Antoine Rouget, Fady Abouzeid, Adrian Evans, Victor Malherbe, Aleksandra Chumakova, Philippe Roche, Fabien Clermidy
IOLTS3
2024 SpDCache: Region-Based Reduction Cache for Outer-Product Sparse Matrix Kernels
abstract
Improvements in computer performance depend increasingly on specialized accelerators and recently, numerous architectures optimized for sparse matrix kernels have been proposed, however, they do not exploit the structural properties of the matrices. SpDCache is a cache for outer-product Sparse Matrix-Vector Multiplication (SpMV) which has storage strategies optimized for both dense and sparse regions and which performs reductions locally in this cache. Real world matrices typically have a dense band which benefits from being blocked in the dense region of our cache, while the sparse regions benefit from fine-grained storage and a shift of the computation close to the main memory. We present the architectural principals of SpDCache and show that it reduces main memory traffic by -8x and increases the cache utilization by - 2x for banded matrices.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ASAP2
2024 A Scalable Low-Latency FPGA Architecture for Spin Qubit Control Through Direct Digital Synthesis
abstract
Scaling qubit control is a key issue for Large Scale Quantum (LSQ) computing and hardware control systems are increasingly costly in logic and memory resources. We present a newly developed compact Direct Digital Synthesis (DDS) architecture for signal generation for spin qubits that is scalable in terms of waveform accuracy and the number of synchronized channels. Fine control of gate voltages is achieved by on-the-fly generation of very precise ramps. Embedded memory requirements are reduced by orders of magnitude compared to current Arbitrary Waveform Generator (AWG) architectures, removing a major scalability barrier for quantum computing.
Mathieu Toubeix, Eric Guthmuller, Adrian Evans, Tristan Meunier
DATE3
2024 New Standard-under-Development for Chiplet Interconnect Test and Repair: IEEE Std P3405
abstract
IEEE Std P3405 is a new standardization activity under the umbrella of TTTC’s Test Technology Standardization Committee (TTSC). In 2023, a Study Group formulated a Project Authorization Request (PAR), which was approved and since December 1, 2023, the P3405 Working Group is active under elected chair Sreejit Chakravarty. This standardization activity focuses exclusively on the test and repair of chiplets’ inter-die interconnects. In the PAR, the scope of the activity is described as follows. "Chiplet-based designs contain dies using proprietary interconnect technology. These dies might come from multiple design groups. Inter-chiplet interconnects are dense, large in number, and prone to manufacturing defects. For cost-effective chiplet packaging, an effective and efficient mechanism to test and repair chiplet interconnects is required. The chiplet interconnect test and repair infrastructure is spread across chiplets and designed by multiple design groups, necessitating the need for a standard for chiplet interconnect test and repair. The purpose of IEEE Std P3405 is to enable interoperability of interconnect test and repair infrastructure of chiplets from multiple design groups. Chiplet-based designs involve multiple parties: Chiplet Maker(s), Packagers, and End User(s). Features supporting the test and repair of chiplet interconnects are part of individual chiplets, which are implemented by individual Chiplet Makers. These features are needed to serve the Chiplet Makers’ (prepackaging), Packagers’, and End Users’ test and repair objectives." In this special session, a handful prominent members of the Working Group express their personal views on the outcome of the standardization work. The views expressed are from the authors alone and do not necessarily align with the view of the IEEE Std P3405 Working Group.
Erik Jan Marinissen, Adrian Evans, Po-Yao Chuang, Martin Keim, Anshuman Chandra
ETS2
2024 OpenSource Heterogeneous Chiplet-based Computing Architectures
abstract
Leading edge processors, such as AMD's MI300 and Intel's Ponte Vecchio, rely on 3D integration of heterogeneous architectures including CPUs and GPUs or vector processors to provide the highest performance. There are many models for memory coherency within such chips and there is a need for research on the best way to map various kernels to these architectures including studying how best to share data. Unfortunately, the hardware in commercial processors is closed, which limits research opportunities, particularly for hardware/software co-design. Recent projects such as OpenPiton, and numerous follow-on projects, have made it possible for the research community to develop coherent, multi-core systems based on RISC-V. The next step is to enhance the support in open source, multi-core platforms for heterogeneous computing (GPUs, FPGA accelerators) and to explore 3D partitioning of such systems. In this paper, we present the state-of-the art of such open source platforms and sketch a roadmap which we hope will enable the research community to continue to contribute to the development of today's advanced heterogeneous architectures.
Adrian Evans, César Fuguet Tortolero, Davy Million
ICCAD1
2024 Dedicated Hardware Accelerators for Processing of Sparse Matrices and Vectors: A Survey
abstract
Performance in scientific and engineering applications such as computational physics, algebraic graph problems or Convolutional Neural Networks (CNN), is dominated by the manipulation of large sparse matrices—matrices with a large number of zero elements. Specialized software using data formats for sparse matrices has been optimized for the main kernels of interest: SpMV and SpMSpM matrix multiplications, but due to the indirect memory accesses, the performance is still limited by the memory hierarchy of conventional computers. Recent work shows that specific hardware accelerators can reduce memory traffic and improve the execution time of sparse matrix multiplication, compared to the best software implementations. The performance of these sparse hardware accelerators depends on the choice of the sparse format, COO , CSR , etc, the algorithm, inner-product , outer-product , Gustavson , and many hardware design choices. In this article, we propose a systematic survey which identifies the design choices of state-of-the-art accelerators for sparse matrix multiplication kernels. We introduce the necessary concepts and then present, compare, and classify the main sparse accelerators in the literature, using consistent notations. Finally, we propose a taxonomy for these accelerators to help future designers make the best choices depending on their objectives.
Valentin Isaac-Chassande, Adrian Evans, Yves Durand, Frédéric Rousseau 0001
ACM Trans. Archit. Code Optim.2
2021 MOZART: Masking Outputs with Zeros for Architectural Robustness and Testing of DNN Accelerators
abstract
Deep Neural Networks (DNNs) are increasingly used in safety critical autonomous systems. In this paper, we present MOZART, a DNN accelerator architecture which provides fault detection and fault tolerance. MOZART is a systolic architecture based on the Output Stationary (OS) variant, as it is the one that inherently limits fault propagation. In addition, MOZART achieves fault detection with on-line functional testing of the Processing Elements (PEs). Faulty PEs are swiftly taken off-line with minimal classification impact. The implementation of our approach on Squeezenet results in a loss of accuracy of less than 3% in the presence of a single faulty PE, compared to 15-33% without mitigation. The area overhead for the test logic does not exceed 8%. Dropout during training further improves fault tolerance, without a priori knowledge of the faults.
Stéphane Burel, Adrian Evans, Lorena Anghel
IOLTS2
2017 EDA support for functional safety - How static and dynamic failure analysis can improve productivity in the assessment of functional safety
abstract
Integrated circuits used in high-reliability applications must demonstrate low failure rates and high-levels of fault detection coverage. Safety Integrity Level (SIL) metrics indicated by the general IEC 61508 standard and the derived Automotive Safety Integrity Level (ASIL) specified by the ISO 26262 standard specify specific failure (FIT) rates and fault coverage metrics (e.g. SPFM and LFM) that must met. To demonstrate that an integrated circuit meets these expectations requires a combination of expert design analysis combined with fault injection (FI) simulations. During FI simulations, specific hardware faults (e.g. transients, stuck-at) are injected in specific nodes of the circuits (e.g. flip flops or logic gates). Designing an effective fault-injection platform is challenging, especially designing a platform that can be re used effectively across designs. We propose an architecture for a complete FI platform, easily integrated into a general-purpose design verification environment (DVE) that is implemented using UVM. The proposed fault simulation methodology is augmented using static analysis techniques based on fault propagation probability assessment and clustering approaches accelerating the fault simulation campaigns. The overall framework aims to: identify safety-threatening device features, provide objective failure metric and support design improvement efforts. We present a worked example where a 32-bit RISC V CPU has been subjected to an extensive static and dynamic failure analysis process, as a part of a standard-mandated functional safety assessment.
Dan Alexandrescu, Adrian Evans, Maximilien Glorieux, Issam Nofal
IOLTS2
2017 BPPT - Bulk potential protection technique for hardened sequentials
abstract
In this paper, we present a method for hardening memory and sequential cells against soft errors. The effect of the ionizing particle on the bulk potential is exploited to prevent the induced SET from propagating in the circuit. The proposed method requires a minimum number of extra transistors. The solution is applied to D Flip-Flop design, and alpha and heavy-ions test results are presented.
Issam Nofal, Adrian Evans, Anlin He, Li Chen 0001, Rui Liu 0011, Mo Chen 0008, Sang H. Baeg, Shi-Jie Wen, Richard Wong
IOLTS2
2016 RIIF-2: Toward the next generation reliability information interchange format
abstract
This paper describes the joint effort of the two FP7 EU projects CLERECO and MoRV toward the definition of an extended reliability information exchange format able to manage reliability information for the full system stack, from technology up to the software level. The paper starts from the RIIF language initiative, proposing a set of new features to improve the expression power of the language and to extend it to the software layer of a system. The proposed extended reliability information exchange format named RIIF-2 has the potential to support the development of next generation reliability analysis tools that will help to fully include reliability evaluation into an automated design flow, pushing cross-layer reliability considerations at the same level of importance as area, timing and power consumption when performing design exploration for new products.
Alessandro Savino 0001, Stefano Di Carlo, Alessandro Vallero, Gianfranco Politano, Dimitris Gizopoulos, Adrian Evans
IOLTS6
2015 A call for cross-layer and cross-domain reliability analysis and management
abstract
For many applications, reliability, availability and trustability are key factors, requiring careful design to meet the end users expectations. The complex ASICs, which are now ubiquitous, often embed tens of millions of flip-flops, hundreds of megabits of embedded SRAM, and hundreds of millions of combinatorial cells. These designs integrate IP from multiple providers and are implemented in advanced process technologies, making it challenging to evaluate their reliability. Initiatives such as RIIF (Reliability Information Interchange Format) allow the formalization, specification and modeling of extra-functional, reliability properties for technology, circuits and systems. Continuing these efforts, we propose RAFT (Reliability Architect Framework and Toolset) - a reliability-centric framework including reliability data and models, methodologies and tools allowing system reliability exploration and optimization using mathematical models and high-level tools. The proposed approach can be combined with performance management methodologies aiming at reducing the engineering effort devoted to reliability analysis and improvement.
Dan Alexandrescu, Adrian Evans, Enrico Costenaro, Maximilien Glorieux
IOLTS2
2015 Flip-flop SEU reduction through minimization of the temporal vulnerability factor (TVF)
abstract
The effects of soft-errors in flip-flops remains a concern in large designs. There exist many radiation hardened flip-flops, however, these are custom cells and not available to all designers. In this paper, we explore a technique for the mitigation of flip-flop soft-errors through an optimization of the temporal vulnerability factor (TVF). By selectively inserting delay on the input or output of flip-flops, the probability of propagation of single event upsets (SEUs) can be minimized. The selection of where to insert the added delay is formulated as a linear programming problem. In this way, the flip-flop soft-error rate (SER) can be minimized subject to overhead constraints.
Adrian Evans, Enrico Costenaro, Arkady Bramnik
IOLTS1
2015 Comprehensive Analysis of Sequential and Combinational Soft Errors in an Embedded Processor
abstract
Radiation-induced soft errors have become a key challenge in advanced commercial electronic components and systems. We present the results of a soft error rate (SER) analysis of an embedded processor. Our SER analysis platform accurately models generation, propagation, and masking effects starting from a technology response model derived using TCAD simulations at the device level all the way to application masking. The platform employs a combination of accurate models at the device level, analytical error propagation at gate level, and fault emulation at the architecture/application level to provide the detailed contribution of each component (flip-flops, combinational gates, and SRAMs) to the overall SER. At each stage in the modeling hierarchy, an appropriate level of abstraction is used to propagate the effect of errors to the next higher level. Unlike previous studies which are based on very simple test chips, analyzing the entire processor gives more insight into the relative contributions of combinational and sequential SER. The results of this analysis can assist circuit designers to adopt effective hardening techniques to reduce the overall SER while meeting the required power and performance constraints.
Mojtaba Ebrahimi, Adrian Evans, Mehdi Baradaran Tahoori, Enrico Costenaro, Dan Alexandrescu, Vikas Chandra, Razi Seyyedi
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 A Class of SEC-DED-DAEC Codes Derived From Orthogonal Latin Square Codes
abstract
Radiation-induced soft errors are a major reliability concern for memories. To ensure that memory contents are not corrupted, single error correction double error detection (SEC-DED) codes are commonly used, however, in advanced technology nodes, soft errors frequently affect more than one memory bit. Since SEC-DED codes cannot correct multiple errors, they are often combined with interleaving. Interleaving, however, impacts memory design and performance and cannot always be used in small memories. This limitation has spurred interest in codes that can correct adjacent bit errors. In particular, several SEC-DED double adjacent error correction (SEC-DED-DAEC) codes have recently been proposed. Implementing DAEC has a cost as it impacts the decoder complexity and delay. Another issue is that most of the new SEC-DED-DAEC codes miscorrect some double nonadjacent bit errors. In this brief, a new class of SEC-DED-DAEC codes is derived from orthogonal latin squares codes. The new codes significantly reduce the decoding complexity and delay. In addition, the codes do not miscorrect any double nonadjacent bit errors. The main disadvantage of the new codes is that they require a larger number of parity check bits. Therefore, they can be useful when decoding delay or complexity is critical or when miscorrection of double nonadjacent bit errors is not acceptable. The proposed codes have been implemented in Hardware Description Language and compared with some of the existing SEC-DED-DAEC codes. The results confirm the reduction in decoder delay.
Pedro Reviriego, Salvatore Pontarelli, Adrian Evans, Juan Antonio Maestro
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Comprehensive analysis of alpha and neutron particle-induced soft errors in an embedded processor at nanoscales
abstract
Radiation-induced soft errors have become a key challenge in advanced commercial electronic components and systems. We present results of Soft Error Rate (SER) analysis of an embedded processor. Our SER analysis platform accurately models all generation, propagation and masking effects starting from a technology response model derived using TCAD simulations at the device level all the way to application masking. The platform employs a combination of empirical models at the device level, analytical error propagation at logic level and fault emulation at the architecture/application level to provide the detailed contribution of each component (flip-flops, combinational gates, and SRAMs) to the overall SER. At each stage in the modeling hierarchy, an appropriate level of abstraction is used to propagate the effect of errors to the next higher level. Unlike previous studies which are based on very simple test chips, analyzing the entire processor gives more insight into the contributions of different components to the overall SER. The results of this analysis can assist circuit designers to adopt effective hardening techniques to reduce the overall SER while meeting required power and performance constraints.
Mojtaba Ebrahimi, Adrian Evans, Mehdi Baradaran Tahoori, Razi Seyyedi, Enrico Costenaro, Dan Alexandrescu
DATE2
2014 Connecting different worlds - Technology abstraction for reliability-aware design and Test
abstract
The rapid shrinking of device geometries in the nanometer regime requires new technology-aware design methodologies. These must be able to evaluate the resilience of the circuit throughout all System on Chip (SoC) abstraction levels. To successfully guide design decisions at the system level, reliability models, which abstract technology information, are required to identify those parts of the system where additional protection in the form of hardware or software coun-termeasures is most effective. Interfaces such as the presented Resilience Articulation Point (RAP) or the Reliability Interchange Information Format (RIIF) are required to enable EDA-assisted analysis and propagation of reliability information. The models are discussed from different perspectives, such as design and test.
Ulf Schlichtmann, Veit Kleeberger, Jacob A. Abraham, Adrian Evans, Christina Gimmler-Dumont, Michael Glaß, Andreas Herkersdorf, Sani R. Nassif, Norbert Wehn
DATE4
2014 Managing SER costs of complex systems through Linear Programming
abstract
Single Event Effects negatively impact the reliability of complex electronic devices and systems. System architects, reliability engineers and digital designers have to invest considerable resources to successfully meet the reliability goals set by the final user or application. The cost of SER mitigation techniques (e.g. additional power and reduced performance) may render the product less competitive. This paper proposes an approach that allows a system architect to select the best SEE management techniques subject to given cost and performance constraints. In this methodology, the costs of SER protection (area, power, engineering effort, IP costs) are expressed as a cost function depending on the selected protection schemes. A separate function expresses the reliability and/or availability as a function of the protection schemes. Then, Linear Programming techniques are used to select a set of protection techniques that minimizes the costs, subject to the reliability constraints being met. This systematic approach enables system-architects to find a minimal-cost SER protection strategy and thus reducing over-design and unnecessary overheads.
Dan Alexandrescu, Nematollah Bidokhti, Andy Yu, Adrian Evans, Enrico Costenaro
IOLTS4
2014 New approaches for synthesis of redundant combinatorial logic for selective fault tolerance
abstract
With shrinking process technologies, the likelihood of mid-life faults in combinatorial logic is increasing. Approximate logic functions are a promising approach to mitigate such faults as the technique can be applied to any digital circuit, it protects against multiple fault models and offers a trade-off between increased area and fault coverage. In this paper we present a new algorithm for generating approximate logic functions. The algorithm considers the failure probabilities of the gates and it uses a sum of product (SOP) representation. The results on some circuits show that FIT rate can be reduced by 75% with an area penalty of 46% and inserting only two additional layers of logic.
Li Chen 0001, Rui Liu 0011, Adrian Evans, Dan Alexandrescu, Shi-Jie Wen, Richard Wong
IOLTS4
2013 Error detection in ternary CAMs using bloom filters
abstract
This paper presents an innovative approach to detect soft errors in Ternary Content Addressable Memories (TCAMs) based on the use of Bloom Filters. The proposed approach is described in detail and its performance results are presented. The advantages of the proposed method are that no modifications to the TCAM device are required, the checking is done on-line and the approach has low power and area overheads.
Salvatore Pontarelli, Marco Ottavi, Adrian Evans, Shi-Jie Wen
DATE3
2013 State-aware single event analysis for sequential logic
abstract
Single Event Effects in sequential logic cells represent the current target for analysis and improvement efforts in both industry and academia. We propose a state-aware analysis methodology that improves the accuracy of Soft Error Rate data for individual sequential instances based on the circuit and application. Furthermore, we exploit the intrinsic imbalance between the SEU susceptibility of different flip-flop states to implement a low-cost SER improvement strategy. Careful, per-state SEE analysis of sequential cells also highlights SET phenomena in flip-flops. We apply de-rating techniques to accurately evaluate their contribution to the overall flip-flop SEE sensitivity.
Dan Alexandrescu, Enrico Costenaro, Adrian Evans
IOLTS3
2013 Hierarchical RTL-based combinatorial SER estimation
abstract
With increased device integration and a gradual trend toward higher operating frequencies, the effect of radiation induced transients in combinatorial logic (SETs) can no longer be ignored. Electrical, logical and temporal masking prevent the majority of SETs from becoming functional failures. Current work on SET analysis starts from a gate-level circuit representation, however, in an industrial design cycle, by the time a gate-level netlist is available, it is too late to make design changes. We propose a hierarchical SET analysis methodology that can be applied at the RTL level. The SET sensitivity of the cell library and the masking characteristics of standard combinatorial design blocks are pre-characterized and stored in compact models. The SET sensitivity of a complex circuit is then calculated by decomposing it into blocks and combining the compact SET models. Experimental results are presented for an ALU implemented in the NanGate library.
Adrian Evans, Dan Alexandrescu, Enrico Costenaro
IOLTS1
2013 Synthesis of Redundant Combinatorial Logic for Selective Fault Tolerance
abstract
With shrinking process technologies, the likelihood of mid-life logic faults is increasing. In this paper, we present an approach for mitigating the effects of faults in combinatorial logic through the selective addition of redundant logic. This approach can be applied to a generic digital circuit, protects against multiple fault models and offers a trade-off between area and fault coverage. The results show that fault coverage can be improved by 4x with an area penalty of 50% and only two additional layers of logic.
Li Chen 0001, Adrian Evans, Shi-Jie Wen, Richard Wong
PRDC3
2013 Hot topic session 4A: Reliability analysis of complex digital systems
abstract
Today, there are several trends that are making the reliability analysis of complex integrated circuits an important challenge in industry. As transistor geometries shrink, the number of physical failure mechanisms is increasing while at the same time the number of transistors per chip is still growing. The rollout of new services is pushing compute demands both in handheld devices and in the data center which is driving up complexity and the level of integration. People are becoming critically dependent on mobile services and expect high availability. Looking forward to the deployment of the Internet of Things (IoT) where processors and routers will be embedded in billions of end-points, we are only going to see an increased demand for reliable computing. In this session, we bring together three different industrial perspectives on reliability. The first looks at the end-points, the second looks at the servers and the last looks at the economic drivers for reliability and the demand for new EDA tools for reliability analysis. In the first talk, Rob Aitken from ARM will discuss the reliability challenges in mobile applications. As mobile systems continue to increase in size and complexity, and user requirements are also becoming more stringent, it is important for designers of mobile systems to be aware of reliability issues, and to adapt their methodologies accordingly. This talk discusses the issues involved, from latent defects, through soft errors, aging and wearout, and shows how to consider these as part of the design process, how to quantify their effects, and how to mitigate them through design changes. In the second presentation, Burcin Aktan from Intel is going to discuss the evolution of the reliability features that are found in server applications. With so many processing units packed in data centers the reliability requirements on an individual device is growing, especially with integrated memory controllers and very high bandwidth data pathways. What was an “add-on” to a device function, 10–15 years ago, now needs to be considered carefully with stringent budgets distributed to each functional block that contribute to overall error rates. This talk will focus on the evolution of reliability features in a number of server products leading into the current state and look at how today's designers are dealing with the challenges of gathering requirements, translating these to design implementation and delivering quality features to customers. Finally we will close with remarks on future directions and possible research areas. In the final presentation, Olivier Lauzeral from iROC Technologies will discuss the importance of methodologies for the reliability analysis of complex SoCs. There is an inherent cost to adding reliability features in a complex IC and designers need to be able to make informed decisions about how much hardware to allocate for mitigation (redundancy, error correction, repair). A prerequisite to make such choices is clearly defined targets and this requires an economic framework where the cost of failures is understood. Once the reliability targets for a system and individual devices are established, there is a need for EDA tools which allow designers to compute the failure rate and failure modes within the device. This analysis must include all failure mechanisms (radiation effects, lifetime effects, manufacturing detects) and take into account the relevant de-ratings between faults and observed errors. This new EDA infra-structure is key for designers to make effective trade-offs in order to arrive at a cost effective design.
Adrian Evans, Michael Nicolaidis, Robert C. Aitken, Burcin Aktan, Olivier Lauzeral
VTS1
2012 RIIF - Reliability information interchange format
abstract
In this paper, a new standard language called RIIF (Reliability Information Interchange Format) is defined which enables designers to specify the failure characteristics and reliability requirements for simple and complex components. This language enables EDA tools to analyze reliability models and to compute the failure rates for complex systems. A formal language makes it possible for suppliers and consumers to exchange reliability information in a consistent fashion and to use this information to build accurate reliability models. The RIIF language is a general purpose reliability modeling language and is not tied to a specific application domain or implementation technology.
Adrian Evans, Michael Nicolaidis, Shi-Jie Wen, Dan Alexandrescu, Enrico Costenaro
IOLTS1
1998 Functional Verification of Large ASICs
abstract
This paper describes the functional verification effort during a specific hardware development program that included three of the largest ASICs designed at Nortel. These devices marked a transition point in methodology as verification took front and centre on the critical path of the ASIC schedule. Both the simulation and emulation strategies are presented. The simulation methodology introduced new techniques such as ASIC sub-system level behavioural modeling, large multi-chip simulations, and random pattern simulations. The emulation strategy was based on a plan that consisted of integrating parts of the real software on the emulated system. This paper describes how these technologies were deployed, analyzes the bugs that were found and highlights the bottlenecks in functional verification as systems become more complex.
Adrian Evans, Allan Silburt, Gary Vrckovnik, Thane Brown, Mario Dufresne, Geoffrey Hall, Tung Ho
DAC1