Mihalis Psarakis

dblp:80/1382 · DBLP profile ↗
← Back
52ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-5359-619XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 9 first-author · 7 since 2021Software engineering, systems software and programming languages · 13 · 3 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 HA-CAAP: Hardware-Assisted Continuous Authentication and Attestation Protocol for IoT Based on Blockchain
abstract
The increasing integration of Internet of Things (IoT) devices in various sectors has created complex and dynamically changing interconnected systems. In several multiauthority and multidomain applications, IoT devices may continuously change their connectivity status, leading to dynamic topologies; an IoT device may be connected to different gateways at different times, to support the provisioning of a distributed service. However, these complex environments increase exposure to security threats, such as device spoofing or cloning attacks. Even worse, without continuous device inventorying, an adversary may easily duplicate a cloned device and concurrently connect compromised Sybil nodes at different gateways, without getting noticed. This article proposes Hardware-assisted, continuous authentication and attestation protocol (HA-CAAP), a hardware-assisted continuous authentication and attestation protocol for IoT devices. By using a physically unclonable function-based periodic authentication mechanism, gateways can continuously authenticate their connected devices and detect modifications in their connectivity status, in nearly real-time. Through a private blockchain, gateways are able to continuously exchange information about their connectivity state and securely share a dynamic device inventory, to detect possible Sybil attacks. In addition, by integrating a continuous gateway attestation mechanism in the blockchain, the protocol prevents nontrusted gateways from joining in and assures their integrity. To evaluate HA-CAAP, security is formally analyzed, while a proof-of-concept implementation is used to analyze the protocol’s performance, for realistic application scenarios.
Vangelis Malamas, Panayiotis Kotzanikolaou, Konstantinos Nomikos, Christos Zonios, Vasileios Tenentes, Mihalis Psarakis
IEEE Internet Things J.6
2024 Automated Hardware Security Countermeasure Integration Inside High Level Synthesis
abstract
High-level Synthesis (HLS) methodology has revolutionized the development of complex hardware designs. It enables the rapid conversion of algorithmic descriptions of functionalities to highly optimized hardware equivalents. While modern HLS tools excel in addressing classic design constraints, such as area, latency and power requirements, they fall short regarding security considerations. Security's role is significantly emphasized in today's digital environment, given the existence of powerful hardware attacks, such as Fault Injection (FI) and Side-Channel Analysis (SCA) attacks. HLS methodology can theoretically facilitate the integration of security measures from the high level, yet its core mechanisms do not actively address the preservation or the improvement of security levels of any countermeasure described. Instead, it may sacrifice security enhancement entirely in circuits of high optimization goals. In this work, first, we propose the automatic countermeasure insertion in a way so that both HLS optimization efforts and the secure addition of the countermeasure are implemented effectively. Secondly, we modify the internal mechanisms of the HLS scheduling algorithm and operation chaining to reduce vulnerable points of the design. We demonstrate our methodology by performing fault injection experiments and comparing the results with a straightforward countermeasure integration technique in terms of hardware security and traditional design metrics.
Amalia-Artemis Koufopoulou, Athanasios Papadimitriou, Mihalis Psarakis, David Hély
DATE3
2024 Considerations on LAMDA Controller With BB-BC Tuning Toward Hardware Implementation
abstract
In real-world industrial processes, achieving self-tuning of control unit parameters using an intelligent optimization algorithm offers significant advantages in adapting to the diverse characteristics of dynamic systems. Conventional controllers often fall short in real-time application scenarios and reliability requirements. In response, adaptable learning algorithms have gained acceptance across various industrial control applications. Handling complex computational tasks, extensive training data, and multivariable demands requires a robust platform to achieve optimal results. In this article, we present a multi-input, multioutput (MIMO) control system that utilizes the Big Bang–Big Crunch optimization algorithm alongside the learning algorithm for multivariate data analysis (LAMDA). Our modifications to the LAMDA controller involve using the ordered weighted average instead ofT-norm orS-norm as aggregate functions. Additionally, we adopted field-programmable gate array (FPGA) technology to improve execution speed and energy efficiency compared to conventional platforms (e.g., microcontrollers). We modeled the MIMO system for heating, ventilation, and air conditioning system using state space equations. The hardware system was designed with (very high-speed integrated circuit) hardware description language (VHDL). To demonstrate our approach, we implemented the entire system on an AMD/Xilinx Virtex-7 FPGA, and experimental results confirmed the feasibility of our proposed optimization strategy.
Almabrok Abdoalnasir, Mihalis Psarakis, Anastasios I. Dounis
IEEE Trans. Ind. Informatics2
2024 Single Event Effects Assessment of UltraScale+ MPSoC Systems Under Atmospheric Radiation
abstract
The AMD UltraScale+ XCZU9EG, a multiprocessor system-on-chip (MPSoC) with integrated programmable logic (PL), is vulnerable to the effects of atmospheric radiation due to its large SRAM count. This article explores the effectiveness of the MPSoC's embedded soft-error mitigation mechanisms through accelerated atmospheric-like neutron radiation testing and dependability analysis. We test the device on a broad range of workloads, such as multithreaded software for pose estimation and weather prediction and a software/hardware codesign image classification application running on the AMD deep-learning processing unit (DPU). We found that for a one-node MPSoC system in New York City at 40 k feet (e.g., avionics), software applications demonstrate a mean time to failure (MTTF) of over 121 months, evidencing effective upset recovery. However, specific workloads, such as the DPU, displayed an MTTF of 4 months, which is attributed to the high failure rate of its PL accelerator. Yet, we show the DPU's MTTF can be extended to 87 months with no extra overhead by ignoring the failure rate of tolerable errors since these do not affect the DPU results.
Dimitris Agiakatsikas, Nikos Foutris, Aitzan Sari, Vasileios Vlagkoulis, Ioanna Souvatzoglou, Mihalis Psarakis, Ruiqi Ye, John Goodacre, Mikel Luján, Maria Kastriotou, Carlo Cazzaniga, Christopher Frost 0002
IEEE Trans. Reliab.6
2023 Detecting Hardware Faults in Approximate Adders via Minimum Redundancy
abstract
Approximate Computing (AC) is an emerging design paradigm that exploits the error-resiliency of specific applications to trade-off between accuracy, performance, area, and power. Nonetheless, fault tolerance remains an open issue in AC since hardware (HW) faults that are caused, for example, by radiation-induced effects, environmental disturbances, or aging/wear-out phenomena, can lead to an arithmetic error out of application specification boundaries. In this work, we guard approximate adders against HW faults by selectively inserting Hardware Fault Detection (HFD) redundancy, i.e., parity code/parity prediction or Double Modular Redundancy (DMR) into the Approximate Arithmetic Circuits (AACs). Specifically, we insert HFD only to the 1-bit adder cells of AACs that can cause an arithmetic error out of their specifications when corrupted by HW faults. Therefore, our proposed approach introduces less area and delay overheads than blindly duplicating the whole AAC. We employ our methodology to state-of-the-art approximate adder models (either low-latency approximate adder or approximate full adder models) to prove that our proposed technique inserts HFD into the AACs without negating their original approximation gains.
Ioannis Tsounis, Dimitris Agiakatsikas, Mihalis Psarakis
IOLTS3
2023 Impact of Voltage Scaling on Soft Errors Susceptibility of Multicore Server CPUs
abstract
Microprocessor power consumption and dependability are both crucial challenges that designers have to cope with due to shrinking feature sizes and increasing transistor counts in a single chip. These two challenges are mutually destructive: microprocessor reliability deteriorates at lower supply voltages that save power. An important dependability metric for microprocessors is their radiation-induced soft error rate (SER). This work goes beyond state-of-the-art by assessing the trade-offs between voltage scaling and soft error rate (SER) on a microprocessor system executing workloads on real hardware and a full software stack setup. We analyze data from accelerated neutron radiation testing for nominal and reduced microprocessor operating voltages. We perform our experiments on a 64-bit Armv8 multicore microprocessor built on 28 nm process technology. We show that the SER of SRAM arrays can increase up to 40.4% when the device operates at reduced supply voltage levels. To put our findings into context, we also estimate the radiation-induced Failures in Time (FIT) rate of various workloads for all the studied voltage levels. Our results show that the total and the Silent Data Corruptions (SDC) FIT of the microprocessor operating at voltage-scaled conditions can be 6.6 × and 16 × larger than at the nominal voltage, respectively. Moreover, changes in the microprocessor’s clock frequency do not have a noticeable impact on its soft error susceptibility. The findings of this work can aid computer architects in striking a balance between power and dependability, thus, designing more robust and efficient microprocessors.
Dimitris Agiakatsikas, George Papadimitriou 0001, Vasileios Karakostas, Dimitris Gizopoulos, Mihalis Psarakis, Camille Bélanger-Champagne, Ewart Blackmore
MICRO5
2023 A Methodology for Fault-tolerant Pareto-optimal Approximate Designs of FPGA-based Accelerators
abstract
Approximate Computing Techniques (ACTs) take advantage of resilience computing applications to trade off among output precision, area, power, and performance. ACTs can lead to significant gains at affordable costs when efficiently implemented on Field Programmable Gate Array– (FPGA) based accelerators. Although several novel ACTs works have been proposed for FPGA accelerators, their applicability to high-assurance systems has not been explored as much. ACTs are becoming necessary in many critical Edge computing systems, such as self-driving cars and Earth observation satellites, to increase computational efficiency. However, an important question comes to mind when targeting critical systems: Does ACT optimization negatively affect the reliability of the system and how can one find optimal design architectures that blend classic mitigation techniques like Triple Modular Redundancy with approximation- and precise-based arithmetic hardware units to achieve the best possible computational efficiency without compromising dependability? This work aims to solve this research problem by introducing a Design Space Exploration (DSE) methodology that employs ACTs in arithmetic units of the design and identifies Pareto-optimal microarchitectures that balance all relevant gains of ACTs, such as area, speed, power, failure rate, and precision, by inserting the correct amount of approximation in the design. In a nutshell, our DSE methodology has formulated the DSE with a Multi-Objective Optimization Problem (MOP). Each Pareto-optimal solution of our tool finds which arithmetic units of the design to implement with precise and approximate circuits and which units to selectively triplicate to remove single points of failure that compromise system reliability below acceptable thresholds. We also suggest another formulation of the DSE into a Single-Objective constraint Optimization Problem (ScOP) producing a single optimal point, and that the user may demand, as a less time-consuming alternative to the MOP if a complete Pareto-front is not needed. Our methodology generates fault-tolerant versions of the Pareto-optimal approximate designs (or simple optimized approximate designs if the ScOP choice is picked) by selectively applying mitigation techniques in a way that the overheads of redundant resources for fault-tolerance do not negate the gains of approximation in comparison to the fault-tolerant versions of the precise design. We evaluate our method on two FPGA-based accelerators: a JPEG encoder and an H.264/Advanced Video Coding decoder. Our experimental results show significant gains in area, frequency, and power consumption without compromising output quality and system reliability compared to classic solutions that replicate all or a part of the resources of the precise design to increase dependability metrics.
Ioannis Tsounis, Dimitris Agiakatsikas, Mihalis Psarakis
ACM Trans. Embed. Comput. Syst.3
2022 The Impact of Hardware Folding on Dependability in Spaceborne FPGA-based Neural Networks
abstract
Commercial SRAM-based field-programmable gate arrays (FPGAs) are becoming popular computing platforms for building efficient Neural Network (NN) accelerators for space missions. FPGAs can implement custom NN architectures that are tailored to the requirements of the mission to improve the performance-to-watt ratio of the design. However, SRAM FPGAs are vulnerable to radiation-induced Single Event Upsets (SEUs), imposing significant design-for-reliability challenges. In this work, we study the impact of hardware folding on the dependability of Binarised NN (BNN) FPGA accelerators. Hard-ware folding configures the level of resource sharing in the design. We implemented three design versions of a BNN that performs image classification. The BNNs were generated with FINN, an open-source framework for developing quantised NNs on AMD-Xilinx FPGAs. The BNNs were implemented on a Zynq-7020 system-on-chip FPGA and tested with configuration memory fault injection experiments to estimate their SEU vulnerability. The three BNN design versions have a maximum (Max), medium (Med), and minimum (Min) folding factor, respectively. Assuming a Low Earth Orbit (LEO), our results show that the Med BNN has the highest Mean Time Between Failure (MTBF) and the Min has the lowest MTBF. However, Min has the highest Mean Executions Between Failure (MEBF) due to its high computational performance.
Ioanna Souvatzoglou, Dimitris Agiakatsikas, George Antonopoulos, Vasileios Vlagkoulis, Aitzan Sari, Athanasios Papadimitriou, Mihalis Psarakis
FPT7
2022 Security and Reliability Evaluation of Countermeasures implemented using High-Level Synthesis
abstract
As the complexity of digital circuits increases, High-Level Synthesis (HLS) is becoming a valuable tool to increase productivity and design reuse by utilizing relevant Electronic Design Automation (EDA) flows, either for Application-Specific Integrated Circuits (ASIC) or for Field Programmable Gate Arrays (FPGA). Side Channel Analysis (SCA) and Fault Injection (FI) attacks are powerful hardware attacks, capable of greatly weakening the theoretical security levels of secure implementations. Furthermore, critical applications demand high levels of reliability including fault tolerance. The lack of security and reliability driven optimizations in HLS tools makes it necessary for the HLS-based designs to validate that the properties of the algorithm and the countermeasures have not been compromised due to the HLS flow. In this work, we provide results on the resilience evaluation of HLS-based FPGA implementations for the aforementioned threats. As a test case, we use multiple versions of an on-the-fly SBOX algorithm integrating different countermeasures (hiding and masking), written in C and implemented using Vivado HLS. We perform extensive evaluations for all the designs and their optimization scenarios. The results provide evidence of issues arising from HLS optimizations on the security and reliability of cryptographic implementations. Furthermore, the results put HLS algorithms to the test of designing secure accelerators and can lead to improving them towards the goal of increasing productivity in the domain of secure and reliable cryptographic implementations.
Amalia-Artemis Koufopoulou, Kalliopi Xevgeni, Athanasios Papadimitriou, Mihalis Psarakis, David Hély
IOLTS4
2021 Analyzing the Impact of Approximate Adders on the Reliability of FPGA Accelerators
abstract
In this paper, we evaluate the impact of approximate adders on the reliability of FPGA-based accelerators for applications that present inherent error resilience. We perform an exhaustive fault injection campaign to examine the effects of single bit upsets (SEUs) in the adders in the DCT block of a JPEG encoder IP core. We analyse how much the reliability of the JPEG encoder deteriorates with the use of approximate instead of accurate adders.
Ioannis Tsounis, Athanasios Papadimitriou, Mihalis Psarakis
ETS3
2020 On a Security-oriented Design Framework for Medical IoT Devices: The Hardware Security Perspective
abstract
As medical devices more and more use Internet of Things based technologies, serious concerns are raised about their security and the privacy of patient's personal health data. To address these concerns, while maintaining reasonable overheads, designers of medical devices need to take security into account from the beginning until the completion of their designs. In this work we identify the relevant security domains and focus to the Hardware Security perspective. Additionally, we present a secure design and evaluation framework which can assist designers towards more secure medical devices. The framework integrates a complete insulin pump architecture containing all the basic components used in such applications. To illustrate the advantages of the proposed framework we perform a Side Channel Analysis attack against the embedded encryption algorithm of the device to obtain the secret encryption key. Then, we make use of the framework to identify all the components of the system which are either directly or indirectly affected by the attack. This analysis leads us to determine more complex combined attacks which may complement the SCA attack into compromising the overall security of the system.
Konstantinos Nomikos, Athanasios Papadimitriou, George Stergiopoulos, Dimitris Koutras, Mihalis Psarakis, Panayiotis Kotzanikolaou
DSD5
2019 Guest Editorial Special Issue on Secure Embedded IoT Devices for Resilient Critical Infrastructures
abstract
The Internet of Things (IoT) creates new technological opportunities for a wide range of systems, such as industrial control systems, smart power grids, vehicular networks (VNs) and intelligent transportation systems, body area networks and healthcare monitoring and control systems and smart homes. At the same time, IoT also increases the threat surface for potential adversaries targeting critical interconnected systems that may depend on embedded IoT systems, such as supervisory control and data acquisition (SCADA) systems. At this point, attackers could take advantage from the incorporation of the paradigm to exploit new security gaps, probably caused by unforeseen interoperability and adaptability problems. Indeed, the deployment of Internet-enabled embedded devices that are distributed over major critical domains may create indirect and nonobvious interconnections with the underlying critical infrastructures (CIs). There is a need to further explore the security issues related to IoT technologies and to assure the resilience of CIs against advanced IoT-enabled attacks. The focus of this special issue is therefore to provide readers with the latest advances in securing the interaction between embedded IoT devices and CIs in order to increase their resilience to advanced IoT-enabled threats.
Cristina Alcaraz, Mike Burmester, Jorge Cuéllar, Xinyi Huang 0001, Panayiotis Kotzanikolaou, Mihalis Psarakis
IEEE Internet Things J.6
2016 A fault injection platform for the analysis of soft error effects in FPGA soft processors
abstract
Soft processors in SRAM-based FPGAs are gaining acceptance as enabling technology for building embedded systems in several market domains, even for critical applications such as space, transportation and medical devices. However, due to the high vulnerability of SRAM-based FPGAs to single-event upsets (SEUs), which is expected to be aggravated in the future, as FPGA devices are moving aggressively to the nanometer regime, the hardening of soft processors against soft errors will become a major design issue especially for critical applications. Most SEU mitigation approaches proposed in the past are based on the triplication or duplication techniques, thus imposing significant area and performance overheads. A more detailed analysis of the soft error sensitivity of FPGA soft processor and their faulty behavior will enable the development of efficient, low-cost hardening techniques. To this end, we present a fault injection platform based on an open-source CAD framework (RapidSmith) for the analysis of soft error effects in Xilinx FPGA soft processors. Our platform supports the estimation of soft error sensitivity per configuration bit/frame, processor component and benchmark. An on-chip microcontroller is used to inject and correct soft errors in the configuration memory and monitor target processor behavior. It includes a custom peripheral to monitor and record specific processor signals (e.g. exception signals, performance counters) which may manifest the effects of soft errors. The proposed platform is demonstrated through an extensive fault injection campaign in the Leon3 soft processor. The novelty of the framework is it's availability as open-source fault-injection tool designed to target soft processors and the introduction of fault identification by using performance counter.
Aitzan Sari, Mihalis Psarakis
DDECS2
2014 A soft error vulnerability analysis framework for Xilinx FPGAs
abstract
Today's SRAM-based FPGAs provide a reach set of computing resources which makes them attractive in demanding and critical application domains, such as avionics and space. Unfortunately, their high reliance on SRAM configuration memory arise reliability issues due to the single-event upsets (SEUs). Considering the criticality of these applications, the vulnerability analysis of FPGA designs to SEUs becomes essential part of the design flow. In this context, we present an open-source framework for the soft error vulnerability analysis of Xilinx FPGA devices. The proposed framework will allow researchers to evaluate their reliability-aware CAD algorithms and estimate the soft error susceptibility of the designs at early stages of the implementation flow for the latest Xilinx architectures.
Aitzan Sari, Dimitris Agiakatsikas, Mihalis Psarakis
FPGA3
2014 Accelerated online error detection in many-core microprocessor architectures
abstract
Forthcoming many-core processors are expected to be highly unreliable due to their high design complexity and aggressive manufacturing technology scaling. Online functional testing is an attractive low-cost error detection solution. A functional error detection scheme for many-core architectures can easily employ existing techniques from single-core microprocessors and exploit the available massive parallelism to reduce the total test execution time. However, the straightforward execution of test programs on such parallel architectures does not achieve the maximum theoretical speedup due to severe congestion on common hardware resources, especially the shared memory and the interconnection network. In this paper, we first identify the memory hierarchy parameters of many-core architectures that slow down the execution of parallel test programs. Then, we study typical test programs to identify which of their parts can be parallelized to improve performance. Finally, we propose a test program parallelization methodology for many-core architectures to accelerate online detection of permanent faults. We evaluate the proposed methodology on a popular many-core architecture, Intel's Single-chip Cloud Computer (SCC) showing an up to 47.6X speedup compared to a serial test program execution approach.
Manolis Kaliorakis, Mihalis Psarakis, Nikos Foutris, Dimitris Gizopoulos
VTS2
2013 Online error detection in multiprocessor chips: A test scheduling study
abstract
Multicore architectures are employed in the majority of computing domains (general-purpose microprocessors as well as specialized high-performance architectures such as network processors). Online error detection in such chips can employ effective techniques from single core microprocessors, however, effective test scheduling should be employed to minimize the overall chip test execution time which can significantly increase due to congestion on common hardware resources used by the cores. In this paper, we analyze the most important aspects of online error detection and scheduling in multiprocessor chips and evaluate test execution time in several different configurations of Intel's SCC architecture.
Manolis Kaliorakis, Nikos Foutris, Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis
IOLTS4
2013 Combining checkpointing and scrubbing in FPGA-based real-time systems
abstract
SRAM-based FPGAs provide an attractive solution for building high-performance embedded computing systems. Fault tolerant mechanisms are usually implemented in FPGA-based critical systems to improve their vulnerability to transient faults. Most fault tolerant approaches proposed so far in the literature for FPGA systems utilize checkpointing and scrubbing techniques for the fault recovery and repair operations, respectively, and rely on redundancy-based fault detection solutions. In this paper, we study the feasibility of building a low-cost fault-tolerant approach for FPGA-based realtime systems that combines checkpointing and scrubbing, the latter for both fault detection and repair. We calculate the checkpoint frequencies that guarantee the execution of the tasks within their deadlines in the presence of transient faults, taking into consideration the scrubbing time of the FPGA processor. Furthermore, we propose a selective scrubbing approach to reduce the scrubbing time and make feasible the fault tolerant execution of tasks with tight deadlines. We demonstrate the proposed approach in a Leon-3-based SoC in a Virtex-5 FPGA.
Aitzan Sari, Mihalis Psarakis, Dimitris Gizopoulos
VTS2
2013 A Fault Tolerant Approach for FPGA Embedded Processors Based on Runtime Partial Reconfiguration
Alexandros Vavousis, Andreas Apostolakis, Mihalis Psarakis
J. Electron. Test.3
2012 Fault tolerant FPGA processor based on runtime reconfigurable modules
abstract
The increasing use of field programmable devices for the implementation of embedded processors and systems-on-chip even in mission-critical applications demands for fault tolerant techniques to improve reliability and extend system lifetime. Furthermore, the runtime partial reconfiguration potentials of the latest FPGA devices along with the availability of unused programmable resources in most FPGA designs provide interesting opportunities to build fault tolerant mechanisms. In this paper, we exploit the latest dynamic reconfiguration advances and propose a fault-tolerant FPGA processor architecture based on runtime reconfigurable modules. We partition the processor core into reconfigurable modules and duplicate these modules to implement a concurrent error detection mechanism. For every duplicated module we generate precompiled configurations which include spare resources and are used to runtime repair the defective module. The processor freezes upon the detection of an error and an on-chip controller coordinates the processor recovery and repair in a reconfiguration process transparent to the processor. We demonstrate the proposed approach in OpenRISC core, a widely-used open-source soft processor.
Mihalis Psarakis, Andreas Apostolakis
ETS1
2012 FPGA-based Acceleration for Tracking Audio Effects in Movies
abstract
In this paper we propose an FPGA-based hardware platform to accelerate an audio tracking method. Our tracking approach is inspired by the problem of molecular sequence alignment and adopts a well-known dynamic programming algorithm (Smith-Waterman algorithm) from the area of bioinformatics. However, the high computational complexity of such algorithms imposes a significant barrier to their adoption by audio tracking systems. To alleviate the time-consuming problem and achieve realistic response times, we propose the acceleration of computationally intensive parts of our tracking method using an FPGA-based platform. Our FPGA accelerator is actually based on the systolization of the Smith-Waterman algorithm proposed in previous approaches for the acceleration of bio-sequence scanning but the special requirements of the audio tracking method impose significant design challenges in the accelerator architecture. The accelerator has been implemented in a Xilinx Virtex-5 device and the experimental results show that it achieves significant speedup compared with the software implementation of the tracking method. The proposed approach has been tested in the context of detecting animal sounds in audio streams from movies, where a basic requirement is to reduce the noisiness of the detection results by means of exploiting the statistical nature of the scores that are generated by the dynamic programming algorithm.
Mihalis Psarakis, Aggelos Pikrakis, Giannis Dendrinos
FCCM1
2011 Architectures for online error detection and recovery in multicore processors
abstract
The huge investment in the design and production of multicore processors may be put at risk because the emerging highly miniaturized but unreliable fabrication technologies will impose significant barriers to the life-long reliable operation of future chips. Extremely complex, massively parallel, multi-core processor chips fabricated in these technologies will become more vulnerable to: (a) environmental disturbances that produce transient (or soft) errors, (b) latent manufacturing defects as well as aging/wearout phenomena that produce permanent (or hard) errors, and (c) verification inefficiencies that allow important design bugs to escape in the system. In an effort to cope with these reliability threats, several research teams have recently proposed multicore processor architectures that provide low-cost dependability guarantees against hardware errors and design bugs. This paper focuses on dependable multicore processor architectures that integrate solutions for online error detection, diagnosis, recovery, and repair during field operation. It discusses taxonomy of representative approaches and presents a qualitative comparison based on: hardware cost, performance overhead, types of faults detected, and detection latency. It also describes in more detail three recently proposed effective architectural approaches: a software-anomaly detection technique (SWAT), a dynamic verification technique (Argus), and a core salvaging methodology.
Dimitris Gizopoulos, Mihalis Psarakis, Sarita V. Adve, Pradeep Ramachandran, Siva Kumar Sastry Hari, Daniel J. Sorin, Albert Meixner, Arijit Biswas, Xavier Vera
DATE2
2011 Scrubbing-based SEU mitigation approach for Systems-on-Programmable-Chips
abstract
The extensive use of Systems-on-Programmable-Chips (SoPCs) in many application domains emphasizes the importance of analysing the vulnerability of the designs to single-event upsets (SEUs) and proposing efficient low-cost mitigation approaches. Most SEU mitigation approaches proposed so far in the literature for SoPCs are based on the use of popular hardware redundant techniques combined with memory scrubbing to avoid fault accumulation. However, all these approaches do not exploit the fact that a large portion of the configuration bits for a particular mapped design are non-sensitive, i.e. do not affect the circuit behaviour in the presence of upsets. In this paper, we first analyse the sensitivity of the configuration bits of a SoPC mapped in a Xilinx Virtex-5 device relying on the vendor implementation tools. The sensitivity analysis showed that the configuration memory scrubbing “wastes” time scanning a large number of frames containing a disproportionately small number of configuration bits due to their high dispersion. To resolve this, we propose a constraint-driven re-placement method to reduce the number of sensitive configuration frames and consequently the scrubbing time. Finally, we present a low-cost SEU mitigation approach for SoPCs which uses configuration memory scan and scrubbing as fault detection and fault repair mechanisms combined with checkpointing and rollback for fault recovery. We demonstrate the efficiency of the proposed mitigation approach in a Leon3-based SoPC implemented in a Xilinx Virtex-5 device.
Aitzan Sari, Mihalis Psarakis
FPT2
2011 Accelerating microprocessor silicon validation by exposing ISA diversity
abstract
Microprocessor design validation is a time consuming and costly task that tends to be a bottleneck in the release of new architectures. The validation step that detects the vast majority of design bugs is the one that stresses the silicon prototypes by applying huge numbers of random tests. Despite its bug detection capability, this step is constrained by extreme computing needs for random tests simulation to extract the bug-free memory image for comparison with the actual silicon image.
Nikos Foutris, Dimitris Gizopoulos, Mihalis Psarakis, Xavier Vera, Antonio González 0001
MICRO3
2011 Chip Self-Organization and Fault Tolerance in Massively Defective Multicore Arrays
abstract
We study chip self-organization and fault tolerance at the architectural level to improve dependable continuous operation of multicore arrays in massively defective nanotechnologies. Architectural self-organization results from the conjunction of self-diagnosis and self-disconnection mechanisms (to identify and isolate most permanently faulty or inaccessible cores and routers), plus self-discovery of routes to maintain the communication in the array. In the methodology presented in this work, chip self-diagnosis is performed in three steps, following an ascending order of complexity: interconnects are tested first, then routers through mutual test, and cores in the last step. The mutual testing of routers is especially important as faulty routers are disconnected by good ones with no assumption on the behavior of defective elements. Moreover, the disconnection of faulty routers is not physical (“hard”) but logical (“soft”) in that a good router simply stops communicating with any adjacent router diagnosed as defective. There is no physical reconfiguration in the chip and no need for spare elements. Ultimately, the multicore array may be viewed as a black box, which incorporates protection mechanisms and self-organizes, while the external control reduces to a simple chip validation test which, in the simplest cases, reduces to counting the number of valid and accessible cores.
Jacques Henri Collet, Piotr Zajac, Mihalis Psarakis, Dimitris Gizopoulos
IEEE Trans. Dependable Secur. Comput.3
2010 MT-SBST: Self-test optimization in multithreaded multicore architectures
abstract
Instruction-based or software-based self-testing (SBST) is a scalable functional testing paradigm that has gained increasing acceptance in testing of single-threaded uniprocessors. Recent computer architecture trends towards chip multiprocessing and multithreading have raised new challenges in the test process. In this paper, we present a novel self-test optimization strategy for multithreaded, multicore microprocessor architectures and apply it to both manufacturing testing (execution from on-chip cache memory) and post-silicon validation (execution from main memory) setups. The proposed self-test program execution optimization aims to: (a) take maximum advantage of the available execution parallelism provided by multiple threads and multiple cores, (b) preserve the high fault coverage that single-thread execution provides for the processor components, and (c) enhance the fault coverage of the thread-specific control logic of the multithreaded multiprocessor. The proposed multithreaded (MT) SBST methodology generates an efficient multithreaded version of the test program and schedules the resulting test threads into the hardware threads of the processor to reduce the overall test execution time and on the same time to increase the overall fault coverage. We demonstrate our methodology in the OpenSPARC T1 processor model which integrates eight CPU cores, each one supporting four hardware threads. MT-SBST methodology and scheduling algorithm significantly speeds up self-test time at both the core level (3.6 times) and the processor level (6.0 times) against single-threaded execution, while at the same time it improves the overall fault coverage. Compared with straightforward multithreaded execution, it reduces the self-test time at both the core level and the processor level by 33% and 20%, respectively. Overall, MT-SBST reaches more than 91% stuck-at fault coverage for the functional units and 88% for the entire chip multiprocessor, a total of more than 1.5M logic gates.
Nikos Foutris, Mihalis Psarakis, Dimitris Gizopoulos, Andreas Apostolakis, Xavier Vera, Antonio González 0001
ITC2
2009 Exploiting Thread-Level Parallelism in Functional Self-Testing of CMT Processors
abstract
Major microprocessor vendors have integrated functional software-based self-testing in their manufacturing test flows during the last decade. Functional self-testing is performed by test programs that the processor executes at-speed from on-chip memory. Multiprocessors and multithreaded architectures are constantly becoming the typical general-purpose computing paradigm, and thus the various existing uniprocessor functional self-testing schemes must be adopted and adjusted to meet the testing requirements of complex multiprocessors. A major challenge in porting a functional self-testing approach from the uniprocessor to the multiprocessor case is to take advantage of the inherent execution parallelism offered by the multiple cores and the multiple threads in order to reduce test execution time. In this paper, we study the application of functional self-testing to chip multithreaded (CMT) processors. We propose a method that exploits thread-level parallelism (TLP) to speed up the execution of self-test routines in every physical core of a multiprocessor chip. The proposed method effectively splits the self-test routines into shorter ones, assigns the new routines to the hardware threads of the core and schedules their execution in order to minimize the core idle intervals due to cache misses or long latency operations and maximize the utilization of core computing resources. We demonstrate our method in the open-source CMT multiprocessor model, Sunpsilas OpenSPARC T1, which contains eight CPU cores, each one supporting four hardware threads. Our experimental results show a self-test execution speedup of more than three times compared to the single thread execution.
Andreas Apostolakis, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis, Ishwar Parulkar
ETS2
2009 Enhanced self-configurability and yield in multicore grids
abstract
As we move deeper in the nanotechnology era, computer architecture is solicited to manipulate tremendous numbers of devices per chip with high defect densities. These trends provide new computing opportunities but efficiently exploiting them will require a shift towards novel, highly parallel architectures. Fault tolerant mechanisms will have to be integrated to the design to deal with the low yield of future nanofabrication processes. In this paper we consider multi processor grid (MPG) architectures that assure scalability beyond hundreds of cores per chip. We study self-diagnosis and self-configuration methods at the architectural level and propose an enhanced self-configuration methodology that enables usage of a maximum percentage of available fault-free cores in MPGs with high defect densities. We show that our approach achieves usability of all fault-free cores for the case of fault-free routers whereas previous work was efficient for defect densities of up to 20-25% of defective cores. We also address the case of faulty routers, achieving usability of almost all fault-free nodes (fault-free cores having a fault-free router) for very high defect densities both in the cores and in the routers.
Eleftherios Kolonis, Michael Nicolaidis, Dimitris Gizopoulos, Mihalis Psarakis, Jacques Henri Collet, Piotr Zajac
IOLTS4
2009 Software-Based Self-Testing of Symmetric Shared-Memory Multiprocessors
abstract
Software-based or instruction-based self-testing has recently emerged as an effective alternative for the manufacturing and online testing of microprocessors, and is progressively adopted by major microprocessor manufacturers mainly as a supplement to other mature and well-established testing approaches to reach higher test quality. Thus far, software-based self-test approaches presented in the literature have focused almost exclusively on uniprocessors. With the continuing prevalence of multiprocessors, the focus of such research approaches moves from the uniprocessor to the multiprocessor case. In this paper, we study the application of software-based self-testing on symmetric shared-memory multiprocessors (SMP) considering the most common interconnection architectures, shared bus and crossbar switch. We focus on the impact of the shared-memory system architecture, the cache coherence mechanisms, and the interconnection architecture on the execution time of self-test programs running on each separate core and exploit the SMP's parallelism during testing to reduce the test execution time. We propose a generic methodology that allocates the test programs and test responses into the shared on-chip memory and schedules the test routines among the cores aiming at the reduction of the total test application time, and thus, test cost, for the SMP, by increasing the execution parallelism and reducing both bus contentions and data cache invalidations. We demonstrate the proposed solutions with detailed experiments on several two-core, four-core, and eight-core SMP benchmarks based on a popular RISC benchmark processor using both the shared bus and the crossbar switch interconnection architectures.
Andreas Apostolakis, Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis
IEEE Trans. Computers3
2009 Instruction-Based Online Periodic Self-Testing of Microprocessors with Floating-Point Units
abstract
Online periodic testing of microprocessors is a valuable means to increase the reliability of a low-cost system, when neither hardware nor time redundant protection schemes can be applied. This is particularly valid for floating-point (FP) units, which are becoming more common in embedded systems and are usually protected from operational faults through costly hardware redundant approaches. In this paper, we present scalable instruction-based self-test program development for both single and double precision FP units considering different instruction sets (MIPS, PowerPC, and Alpha), different microprocessor architectures (32/64-bit architectures) and different memory configurations. Moreover, we introduce bit-level manipulation instruction sequences that are essential for the development of FP unit's self-test programs. We developed self-test programs for single and double precision FP units on 32-bit and 64-bit microprocessor architectures and evaluated them with respect to the requirements of low-cost online periodic self-testing: fault coverage, memory footprint, execution time, and power consumption, assuming different memory hierarchy configurations. Our comprehensive experimental evaluations reveal that the instruction set architecture plays a significant role in the development of self-test programs. Additionally, we suggest the most suitable self-test program development approach when memory footprint or low power consumption is of paramount importance.
George Xenoulis, Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis
IEEE Trans. Dependable Secur. Comput.3
2008 Functional Self-Testing for Bus-Based Symmetric Multiprocessors
abstract
Functional, instruction-based self-testing of microprocessors has emerged as an effective alternative or supplement to other testing approaches, and is progressively adopted by major microprocessor manufacturers. In this paper, we study, for first time, the applicability of functional self-testing on bus-based symmetric multiprocessors (SMP) and the exploitation of SMPs parallelism during testing. We focus on the impact of the memory system architecture and the cache coherency mechanisms on the execution of self-test programs on the processor cores. We propose a generic self-test routines scheduling algorithm aiming at the reduction of the total test application time for the SMP by reducing both bus contention and data cache coherency invalidation. We demonstrate the proposed solutions with detailed experiments in two-core and four-core SMP benchmarks based on a RISC processor core.
Andreas Apostolakis, Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis
DATE3
2008 Systematic Software-Based Self-Test for Pipelined Processors
abstract
Software-based self-test (SBST) has recently emerged as an effective methodology for the manufacturing test of processors and other components in systems-on-chip (SoCs). By moving test related functions from external resources to the SoC's interior, in the form of test programs that the on-chip processor executes, SBST significantly reduces the need for high-cost, big-iron testers, and enables high-quality at-speed testing and performance binning. Thus far, SBST approaches have focused almost exclusively on the functional (programmer visible) components of the processor. In this paper, we analyze the challenges involved in testing an important component of modern processors, namely, the pipelining logic, and propose a systematic SBST methodology to address them. We first demonstrate that SBST programs that only target the functional components of the processor are not sufficient to test the pipeline logic, resulting in a significant loss of overall processor fault coverage. We further identify the testability hotspots in the pipeline logic using two fully pipelined reduced instruction set computer (RISC) processor benchmarks. Finally, we develop a systematic SBST methodology that enhances existing SBST programs so that they comprehensively test the pipeline logic. The proposed methodology is complementary to previous SBST techniques that target functional components (their results can form the input to our methodology, and thus we can reuse the test development effort behind preexisting SBST programs). We automate our methodology and incorporate it in an integrated software environment (developed using Java, XML, and archC) for the automatic generation of SBST routines for microprocessors. We apply the methodology to the two complex benchmark RISC processors with respect to two fault models: stuck-at fault model and transition delay fault model. Simulation results show that our methodology provides significant improvements for the two fault models, both for the entire processor (12% fault coverage improvement on average) and for the pipeline logic itself (19% fault coverage improvement on average), compared to a conventional SBST approach.
Dimitris Gizopoulos, Mihalis Psarakis, Miltiadis Hatzimihail, Michail Maniatakos, Antonis M. Paschalis, Anand Raghunathan, Srivaths Ravi 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2007 A Functional Self-Test Approach for Peripheral Cores in Processor-Based SoCs
abstract
Functional Software-Based Self-Testing (SBST) of microprocessors and processor-based testing of Systems-on- Chip (SoCs) have recently attracted the attention of test technology research community because they provide an effective alternative to other traditional testing and self- testing approaches. SBST allows at-speed functional testing of SoC cores with virtually no circuit overhead and limited dependence on external testers. Despite the importance of peripheral control cores in SoCs (in terms of functionality and size), the applicability and limits of SBST on them have not been comprehensively studied. In this paper, we study the effectiveness of SBST on a broad class of communication peripherals. A systematic application of SBST on two popular peripheral cores (UART and Ethernet) demonstrates the effectiveness of SBST on SoCs with such cores.
Andreas Apostolakis, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
IOLTS2
2007 A methodology for detecting performance faults in microprocessors via performance monitoring hardware
abstract
Speculative execution of instructions boosts performance in modern microprocessors. Control and data flow dependencies are overcome through speculation mechanisms, such as branch prediction or data value prediction. Because of their inherent self-correcting nature, the presence of defects in speculative execution units does not affect their functionality (and escapes traditional functional testing approaches) but impose severe performance degradation. In this paper, we investigate the effects of performance faults in speculative execution units and propose a generic, software-based test methodology, which utilizes available processor resources: hardware performance monitors and processor exceptions, to detect these faults in a systematic way. We demonstrate the methodology on a publicly available fully pipelined RISC processor that has been enhanced with the most common speculative execution unit, the branch prediction unit. Two popular schemes of predictors built around a Branch Target Buffer have been studied and experimental results show significant improvements on both cases fault coverage of the branch prediction units increased from 80% to 97%. Detailed experiments for the application of a functional self-testing methodology on a complete RISC processor incorporating both a full pipeline structure and a branch prediction unit have not been previously given in the literature.
Miltiadis Hatzimihail, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
ITC2
2007 Functional Processor-Based Testing of Communication Peripherals in Systems-on-Chip
abstract
Software-based self-testing (SBST) of microprocessors and processor-based testing of systems-on-chip (SoCs) recently captured intense test technology research efforts because they provide an effective alternative or supplement to other classic testing and self-testing approaches. Despite the importance of peripheral cores in SoCs (in terms of functionality and size), the applicability and limits of SBST on them have not been comprehensively studied. In this paper, we study the effectiveness of SBST on a broad class of communication peripheral cores in SoCs. A systematic application of SBST on two popular peripherals (universal asynchronous receiver transmitter and Ethernet) demonstrates the effectiveness of SBST on SoCs with such cores.
Andreas Apostolakis, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Systematic software-based self-test for pipelined processors
abstract
Software-based self-test (SBST) has recently emerged as an effective methodology for the manufacturing test of processors and other components in Systems-on-Chip (SoCs). By moving test related functions from external resources to the SoC’s interior, in the form of test programs that the on-chip processor executes, SBST eliminates the need for high-cost testers, and enables high-quality at-speed testing. Thus far, SBST approaches have focused almost exclusively on the functional (directly programmer visible) components of the processor. In this paper, we analyze the challenges involved in testing an important component of modern processors, namely, the pipelining logic, and propose a systematic SBST methodology to address them. We first demonstrate that SBST programs that only target the functional components of the processor are insufficient to test the pipeline logic, resulting in a significant loss of fault coverage. We further identify the testability hotspots in the pipeline logic. Finally, we develop a systematic SBST methodology that enhances existing SBST programs to comprehensively test the pipeline logic. The proposed methodology is complementary to previous SBST techniques that target functional components (their results can form the input to our methodology), and can reuse the test development effort behind existing SBST programs. We applied the methodology to two complex, fully pipelined processors. Results show that our methodology provides fault coverage improvements of up to 15% (12 % on average) for the entire processor, and fault coverage improvements of 22 % for the pipeline logic, compared to a conventional SBST approach.
Mihalis Psarakis, Dimitris Gizopoulos, Miltiadis Hatzimihail, Antonis M. Paschalis, Anand Raghunathan, Srivaths Ravi 0001
DAC1
2006 A Low-Cost SEU Fault Emulation Platform for SRAM-Based FPGAs
abstract
In this paper, we introduce a fully automated low cost hardware/software platform for efficiently performing fault emulation experiments targeting SEUs in the configuration bits of FPGA devices, without the need for expensive radiation experiments. We propose a method for significantly reducing the fault list by removing the faults on unused LUT bit positions. We also target the design flip-flops found in the configurable logic blocks (CLBs) inside the FPGA. Run-time reconfigurability of Virtex devices using JBits is exploited to provide the means not only for fault injection but fault detection as well. First, we consider five possible application scenarios for evaluating different self-test schemes. Then, we apply the least favorite and most time consuming of these scenarios on two 32times32 multiplier designs, demonstrating that transferring the simulation processing workload to FPGA hardware can allow for acceleration of simulation time of more than two orders of magnitude
P. Kenterlis, Nektarios Kranitis, Antonis M. Paschalis, Dimitris Gizopoulos, Mihalis Psarakis
IOLTS5
2006 Testability Analysis and Scalable Test Generation for High-Speed Floating-Point Units
abstract
High-speed datapaths in microprocessors and embedded processors contain complex floating-point (FP) arithmetic units which have a critical role in the processor's performance. Although the FP units' complex structure consists of classic integer arithmetic components, the embedded components encounter serious testability problems due to their limited accessibility from the FP unit ports and testability loss due to FP unit inherent operations, such as rounding and normalization. In this paper, we analyze the testability problems and present scalable test generation for FP units using as a demonstration vehicle the popular, high-speed, two-path architecture of the most complex unit, the FP adder. The key feature of the presented methodology is the identification of testability conditions that guarantee effective test pattern application and fault propagation for each of the components of the FP adder. The identified test conditions can be utilized with respect to any fault model and are independent of the internal structure and the size of the components. Thus, they can be applied to FP adders of various exponent and significant sizes (single, double, and custom precision), as well as to other types of FP units, which also consist of classic integer arithmetic components similarly interconnected
George Xenoulis, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
IEEE Trans. Computers2
2005 Test Generation Methodology for High-Speed Floating Point Adders
abstract
High performance real number operations in embedded processors' and microprocessors' datapaths are realized by floating point (FP) arithmetic units. FP units have a complex structure which although consisting of classic integer arithmetic components faces serious testability problems due to the limited accessibility of the components from the FP unit ports. In this paper we present a test generation methodology for FP adders based on the high-speed, two-path architecture. The key feature of the presented methodology is the identification of testability conditions that guarantee effective test pattern application and fault propagation for each of the components of the FP adder. According to our test methodology, the testability conditions guide test generation process. The identified test conditions are independent of the internal structure and the size of the components. Thus, they can be applied to floating point adders of various exponent and significand sizes built with components of different architectures.
George Xenoulis, Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
IOLTS2
2005 Built-in sequential fault self-testing of array multipliers
abstract
Microprocessor datapath architectures operate on signed numbers usually represented in two's-complement or sign-magnitude formats. The multiplication operation is performed by optimized array multipliers of various architectures which are often produced by automatic module generators. Array multipliers have either a standard, nonrecoded signed (or unsigned) architecture or a recoded (modified Booth's algorithm) architecture. High-quality testing of array multipliers based on a comprehensive sequential fault model and not affecting their well-optimized structure has not been proposed in the past. In this paper, we present a built-in self-testing (BIST) architecture for signed and unsigned array multipliers with respect to a comprehensive sequential fault model. The BIST architecture does not alter the well-optimized multiplier structure. The proposed test sets can be applied externally but their regular nature makes them very suitable for embedded, self-test application by simple specialized hardware which imposes small overheads. Two different implementations of the BIST architecture are proposed. The first implementation focuses on the test invalidation problem and targets robust sequential fault testing, while the second one focuses on test cost reduction (test time and hardware overhead).
Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2003 Easily Testable Cellular Carry Lookahead Adders
Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis, Yervant Zorian
J. Electron. Test.2
2001 Deterministic software-based self-testing of embedded processor cores
abstract
A deterministic software-based self-testing methodology for processor cores is introduced that efficiently tests the processor datapath modules without any modification of the processor structure. It provides a guaranteed high fault coverage without repetitive fault simulation experiments which is necessary in pseudorandom software-based processor self-testing approaches. Test generation and output analysis are performed by utilizing the processor functional modules like accumulators (arithmetic part of ALU) and shifters (if they exist) through processor instructions. No extra hardware is required and there is no performance degradation.
Antonis M. Paschalis, Dimitris Gizopoulos, Nektarios Kranitis, Mihalis Psarakis, Yervant Zorian
DATE4
2001 Robust and Low-Cost BIST Architectures for Sequential Fault Testing in Datapath Multipliers
abstract
The modified Booth array multiplier is the most ubiquitous multiplier architecture in the datapaths of either general purpose microprocessors or specialized Digital Signal Processors. Sequential fault testing for Booth array multipliers has never been proposed in the past. In this paper, we present two BIST architectures for modified Booth array multipliers with respect to the Realistic Sequential Cell Fault model (RS-CFM). The first BIST architecture aims to resolve the test invalidation problem to the largest possible extent, while the second one aims to test cost reduction. Both BIST architectures achieve very high sequential fault coverage and impose moderate hardware and delay overhead. Simplified variations of the two BIST architectures are also presented for the non-recoded signed array multipliers. Thus, the proposed BIST architectures offer a universal BIST solution that covers the totality of signed array multipliers: non-recoded and recoded.
Mihalis Psarakis, Antonis M. Paschalis, Nektarios Kranitis, Dimitris Gizopoulos, Yervant Zorian
VTS1
2001 An Effective Deterministic BIST Scheme for Shifter/Accumulator Pairs in Datapaths
Nektarios Kranitis, Antonis M. Paschalis, Dimitris Gizopoulos, Mihalis Psarakis, Yervant Zorian
J. Electron. Test.4
2000 Effective Low Power BIST for Datapaths
abstract
Power in processing cores (microprocessors, DSPs) is primarily consumed in the datapath part. Among the datapath functional modules, multipliers consume the largest amount of power due to their size and complexity. We propose a low power BIST scheme for datapaths built around multiplier-accumulator pairs. The target is low average power dissipation between successive test vectors. This is achieved by taking advantage of the regularity of multiplier modules and achieving very high fault coverage by a linear-sized test set with as small as possible input switching activity. The proposed BIST scheme is more efficient than pseudorandom BIST for the same high fault coverage target. Up to 77.2% power saving is achieved in the set of experimental results provided in the paper.
Dimitris Gizopoulos, Nektarios Kranitis, Mihalis Psarakis, Antonis M. Paschalis, Yervant Zorian
DATE3
2000 Low Power/Energy BIST Scheme for Datapaths
abstract
Power in processing cores (microprocessors, DSPs) is primarily consumed in the functional modules of the datapath. Among these modules, multipliers consume the largest amount of power due to their size and complexity. We propose low power BIST schemes for datapath architectures built around multiplier-accumulator pairs, based on deterministic test patterns. Two alternatives are proposed depending on whether the target is low energy dissipation during a BIST session or low power dissipation (i.e. average energy dissipation between successive test vectors). The proposed BIST schemes are more efficient than pseudorandom BIST for the same high fault coverage target. Up to 78.33% energy saving is achieved by the proposed low energy BIST scheme and up to 82.22% power saving is achieved by the proposed low power BIST scheme, compared with pseudorandom BIST.
Dimitris Gizopoulos, Nektarios Kranitis, Mihalis Psarakis, Antonis M. Paschalis, Yervant Zorian
VTS3
2000 Sequential Fault Modeling and Test Pattern Generation for CMOS Iterative Logic Arrays
abstract
Iterative Logic Arrays (ILAs) are widely used in the datapath parts of digital circuits, like general purpose microprocessors, embedded processors, and digital signal processors. Testing strategies based on more comprehensive fault models than the traditional combinational fault models have become an imperative need in CMOS technology. In this paper, first, we introduce a comprehensive, cell-level, sequential fault model suitable for ILAs, termed Realistic Sequential Cell Fault Model (RS-CFM). RS-CFM drastically reduces test complexity compared to exhaustive two-pattern testing proposed so far in the literature for sequential ILA testing, without sacrificing test quality. In addition, it favors robustness of sequential test sets both at the cell and the array levels. Second, a new Automatic Test Pattern Generator (ILA-ATPG) based on RS-CFM for the case of one-dimensional ILAs is presented. ILA-ATPG can handle all classes of one-dimensional ILAs: unilateral or bilateral ILAs, with or without vertical inputs/outputs. Based on a graph model, ILA-ATPG explores the C-testability and linear-testability of the ILA under test and resolves the test invalidation problem constructing robust test sequences. The efficiency of ILA-ATPG is demonstrated through a comprehensive set of experimental results over all classes of one-dimensional ILAs, including all practical one-dimensional ILAs, as well as a number of more complex benchmarks.
Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis, Yervant Zorian
IEEE Trans. Computers1
1999 An Effective BIST Architecture for Fast Multiplier Cores
abstract
Wallace free summation in conjunction with Booth encoding are well known techniques to design fast multiplier cores widely used as embedded cores in the design of complex systems on chip. Testing of such multiplier cores deeply embedded in complex ICs requires the utilization of a BIST architecture that can be easily synthesized along with the multiplier by the module generator. In this paper we introduce an effective BIST architecture for fast multipliers that completely complies with this requirement. The algorithmic BIST patterns that this architecture generates guarantee a fault coverage higher than 99%. The required test pattern generator consists of a simple fixed-size binary counter, independent of the multiplier size. Accumulator-based compaction is adopted since multipliers and adders co-exist in most datapath architectures.
Antonis M. Paschalis, Nektarios Kranitis, Mihalis Psarakis, Dimitris Gizopoulos, Yervant Zorian
DATE3
1999 An Effective BIST Architecture for Sequential Fault Testing in Array Multipliers
abstract
Sequential fault testing approaches for array multipliers proposed in the past target only external testing and impose significant hardware overhead due to excessive DFT modifications. In this paper we present, for the first time, a BIST architecture which does not require any DFT modifications in the multiplier structure and provides a fault coverage larger than 99% for a comprehensive sequential fault model (RS-CFM) for any multiplier size. Both robust and non-robust testing are considered. The applicability of the BIST architecture is further justified considering the case of the transistor stuck-open fault model, where a fault coverage larger than 99% is also achieved in any case.
Mihalis Psarakis, Antonis M. Paschalis, Dimitris Gizopoulos, Yervant Zorian
VTS1
1998 Robustly Testable Array Multipliers under Realistic Sequential Cell Fault Model
abstract
Traditional combinational fault models are not sufficient to detect to most common failure mechanisms in CMOS Iterative Logic Arrays (ILAs). The Realistic Sequential Cell Fault Model (RS-CFM) provides a comprehensive, robust test methodology for ILAs. It also satisfies the requirements for low test complexity and cell implementation independence. Adopting RS-CFM, we first provide sufficient conditions for two-dimensional (2D) ILAs to be robustly testable for first time in the literature. Then, we propose sufficient modifications to the carry-save and carry-propagate array multipliers so that they can be treated robustly with respect to RS-CPM with a test set of linear size.
Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis, Yervant Zorian
VTS1
1998 Test Generation and Fault Simulation for Cell Fault Model using Stuck-at Fault Model based Test Tools
Mihalis Psarakis, Dimitris Gizopoulos, Antonis M. Paschalis
J. Electron. Test.1
1997 An Effective BIST Scheme for Arithmetic Logic Units
abstract
Multifunction arithmetic logic units (ALUs) that realize complex arithmetic and logic operations (like the operations of the 74/spl times/181 family) are widely used in today's complex integrated circuits, such as commercial microprocessors and digital signal processors. These ALUs are built around either ripple-carry (RC) adders, carry-lookahead (CLA) adders or mixed CLA/RC adders depending on area and performance requirements. In this paper, first, we introduce novel C-testable multifunction ALUs built around RC adders and linear-testable multifunction ALUs built around CLA adders and mined CLA/RC adders with respect to CFM. Then, we introduce an effective ALU BIST scheme for all three types of ALUs (RC, CLA, mixed CLA/RC) that hits the target of a unified datapath BIST architecture, since it is compatible to an effective BIST scheme for datapaths. Complete CFM testability is achieved with a reasonable number of deterministic test patterns in all cases. The scheme imposes reasonable area overhead and negligible delay overhead and owing to its inherent high regularity can be easily adopted for automatic BIST synthesis of datapaths.
Dimitris Gizopoulos, Antonis M. Paschalis, Yervant Zorian, Mihalis Psarakis
ITC4
1997 Robust Sequential Fault Testing of Iterative Logic Arrays
abstract
Technology advances provide today the capability of integrating large Iterative Logic Arrays (ILAs) in the same chip. Traditional combinational fault models are not sufficient to detect all failures in CMOS ILAs. Robust test generation for sequential faults in ILAs has not been considered in the literature. Two-pattern tests for sequential fault detection in ILAs can be invalidated either at the cell level due to arbitrary delays inside the cells or at the array level due to the appearance of glitches at the cell inputs. A realistic sequential fault model for any type of ILA is introduced. The fault model along with algorithms for the elimination of glitches in one-dimensional ILAs provide a comprehensive methodology for robust sequential fault testing. C-testability and linear-testability are seeked to provide efficient test sets. Results of the implementation of the method on a comprehensive set of benchmark one-dimensional ILAs are provided. Test complexity and thus test cost is greatly reduced compared to exhaustive two-pattern testing proposed in the past for sequential fault testing in ILAs.
Dimitris Gizopoulos, Mihalis Psarakis, Antonis M. Paschalis
VTS2