Angeliki Kritikakou

dblp:77/8933 · DBLP profile ↗
← Back
60ranked-venue papers
11as first author
32since 2021 · last 2026
0000-0002-9293-469XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 9 first-author · 26 since 2021Software engineering, systems software and programming languages · 15 · 3 first-author · 11 since 2021Computer networks · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Project Highlights - Reliability Evaluation for ARCHYTAS AI hardware accelerators
Angeliki Kritikakou, Fernando Santos 0001, Marcello Traiola, Rafael Billig Tonetto, Olivier Sentieys, Paolo Rech, Haralampos-G. D. Stratigopoulos, Georgios Keramidas
IOLTS1
2026 ENFOR-SA: End-to-end Cross-layer Transient Fault Injector for Efficient and Accurate DNN Reliability Assessment on Systolic Arrays
abstract
Recent advances in deep learning have produced highly accurate but increasingly large and complex DNNs, making traditional fault-injection techniques impractical. Accurate fault analysis requires RTL-accurate hardware models. However, this significantly slows evaluation compared with software-only approaches, particularly when combined with expensive HDL instrumentation. In this work, we show that such high-overhead methods are unnecessary for systolic array (SA) architectures and propose ENFOR-SA, an end-to-end framework for DNN transient fault analysis on SAs. Our two-step approach employs cross-layer simulation and uses RTL SA components only during fault injection, with the rest executed at the software level. Experiments on CNNs and Vision Transformers demonstrate that ENFOR-SA achieves RTL-accurate fault injection with only 6% average slowdown compared to software-based injection, while delivering at least two orders of magnitude speedup (average $569\times$) over full-SoC RTL simulation and a $2.03\times$ improvement over a state-of-the-art cross-layer RTL injection tool. ENFOR-SA code is publicly available at https://github.com/rafaabt/ENFOR-SA.
Rafael Billig Tonetto, Marcello Traiola, Fernando Santos 0001, Angeliki Kritikakou
VTS4
2026 Energy-Efficient and Reliable Task Mapping and Offloading for Multicore Edge Devices With DVFS
abstract
Multicore platforms based on NoC are promising architectures for safety-critical applications. Application execution performance is determined by task mapping, with reliable execution, real-time response, and energy efficiency as requirements. We can perform task duplication, DVFS, and multipath routing to meet these requirements during task mapping. Furthermore, the computation platforms have limited computation capacity and energy supply in several application domains. Some complex tasks can be offloaded from the edge device to the cloud for execution. However, such task offloading influences task mapping on the edge device. Existing approaches seldom consider the correlation of task offloading to the cloud and task mapping on the edge device. To address this limitation, we jointly consider task mapping inside the NoC-based multicore edge device and task offloading to the cloud to optimize energy consumption while satisfying reliability and real-time constraints. This problem is formulated as a mixed-integer nonlinear programming and linearized to find the optimal solution. We propose a novel three-step heuristic with a feedback mechanism to enhance task schedulability and reduce computation time. We evaluate the behavior of our approaches through exhaustive simulations. The results show that our approaches outperform existing methods in terms of energy efficiency, task reliability, and schedulability.
Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Olivier Sentieys, Shibo He
IEEE Internet Things J.4
2026 QoS-Aware Approximate Task Mapping on Heterogeneous Multicore Platforms with DVFS and Task Migration
abstract
Heterogeneous Multicore Platforms (HMPs) have been widely adopted to execute tasks across a range of applications. Under limited system resources and diverse application requirements, allocating and executing dependent Approximate Computing (AC) tasks on these platforms to achieve high Quality-of-Service (QoS) is challenging. Dynamic Voltage and Frequency Scaling (DVFS) and task migration have proven effective for improving QoS while balancing time and energy consumption. However, existing approaches often overlook the migration overhead and the resulting dynamic changes in task dependencies, which can adversely affect mapping outcomes. To address these issues, this article presents a novel AC task mapping method that maximizes system QoS under multiple constraints on HMPs, accounting for task migration overhead, DVFS, and changes in Directed Acyclic Graph (DAG) topology. We first formulate this joint design problem as a complex nonlinear programming problem. Next, we linearize the nonlinear terms without performance loss by introducing auxiliary variables and additional constraints. Building on this formulation, we propose an optimal (OPT) and a low-complexity Heuristic Algorithm (HEU), derived from problem decomposition and a greedy strategy, which divides the Mixed-Integer Non-Linear Programming (MINLP) problem into two smaller subproblems with fewer variables and constraints, solving them sequentially. The simulation results show that the proposed OPT method achieves higher QoS performance, measured at about 2.389 times on average and up to 4.115 times, while its feasibility is increased to about 3.263 times on average and up to 9.667 times, compared to other state-of-the-art methods. In addition, the average QoS of the proposed HEU method is about 0.577 times that of the proposed method, but its computation time is over a thousand times shorter.
Hengyan Song, Lei Mo, Tamim M. Al-Hasan, Angeliki Kritikakou, Xiaojun Zhai, Shibo He, Olivier Sentieys
ACM Trans. Embed. Comput. Syst.4
2025 European Test Symposium Teams: an Anniversary Snapshot
abstract
The IEEE European Test Symposium (ETS) has been facilitating progress in electronic systems testing since its launch in 1996. On the occasion of its 30th anniversary, this collaborative paper gathers sections by 21 ETS teams to outline their influential ideas and milestones. Each team’s section highlights historical perspective, current research, frameworks and projects as well as forward-looking research agendas in the area of electronic-based circuits and systems testing, reliability, safety, security and validation. This anniversary summary documents how research of various ETS teams, exemplifying the test community, has been evolving and transitioning from concepts to practical standards and Electronic Design Automation (EDA) tools and flows. This legacy is a strong base to drive the next generation of advances in electronic systems testing.
Maksim Jenihhin, Jaan Raik, Artur Jutman, Natalia Cherezova, Raimund Ubar, Liviu Miclea, Szilárd Enyedi, Iulia Stefan, Ovidiu Stan, Cosmina Corches, Zebo Peng, Petru Eles, Rolf Drechsler, S. Eggersglüß, Görschwin Fey, Andreas Glowatz, Daniel Tille, Georges Gielen, Anthony Coyette, Wim Dobbelaere, Ronny Vanhooren, Po-Yao Chuang, Erik Jan Marinissen, Giorgio Di Natale, M. Barragan, Paolo Maistri, S. Mir, Vatajelu I. Vatajelu, Paolo Bernardi 0002, Stefano Di Carlo, Paolo Prinetto, Matteo Sonza Reorda, Massimo Violante, Haralampos-G. D. Stratigopoulos, M. K. Michael, Stelios Neophytou, Stavros Hadjitheophanous, Kyriakos Christou, M. Skitsas, Alberto Bosio, Bastien Deveautour, Patrick Girard 0001, Marcello Traiola, Arnaud Virazel, Fernando Santos 0001, Angeliki Kritikakou, Gioele Casagranda, Marzio Vallero, Flavio Vella, Paolo Rech, Letícia Maria Veiras Bolzani, Milos Krstic, Marko S. Andjelkovic, Fabian Vargas 0001, Grigor Tshagharyan, Gurgen Harutunyan, Valery A. Vardanian, Samvel K. Shoukourian, Yervant Zorian, Jennifer Dworak, Kundan Nepal, Theodore W. Manikas, Mottaqiallah Taouil, Moritz Fieback, Anteneh Gebregiorgis, Rajendra Bishnoi, Said Hamdioui, Abhijit Chatterjee, Anurup Saha, Suhasini Komarraju, K. Ma, Chandramouli N. Amarnath, Mehdi Baradaran Tahoori, Mahta Mayahinia, Maryam Rajabalipanah, Katayoon Basharkhah, N. Nosrati, Zahra Jahanpeima, Zainalabedin Navabi, Hans-Joachim Wunderlich, Sybille Hellebrand
ETS46
2025 SERA-Float: A Soft Error Resilient Approximate Floating-Point Computing Format
abstract
Approximate computing (AxC) reduces power consumption with minimal accuracy loss, benefiting error-tolerant, compute-intensive tasks such as machine learning, deep learning, and image processing. However, existing AxC methods often ignore the vulnerability to soft errors. Such errors can interact with approximation-induced errors, causing system failures or unexpected exceptions. To our knowledge, no work has addressed both soft error resilience and exception avoidance in approximate floating-point computing. This gap is particularly critical in deep neural network (DNN) inference, where soft error-induced errors or exceptions can significantly affect the stability and accuracy of computations.In this paper, we introduce SERA-Float, an approximate floating-point format resilient to soft errors. Specifically, it is designed to protect floating-point computations from soft error-induced errors and exception-triggering bit-flips. Unlike prior floating-point formats, SERA-Float protects the sign and exponent bits using error-correcting codes and relies on storing 8 valid bits of mantissa rather than performing coarse truncation. Additionally, by tracking critical bits in the floating-point representation, SERA-Float prevents overflow, underflow, and NaN exceptions. Our evaluation demonstrates that SERA-Float improves the reliability of floating-point operations during DNN inference by significantly reducing exceptions and ensuring the stability of computations. Moreover, it enables energy-efficient arithmetic by leveraging narrower arithmetic units, yielding up to 80.3% energy savings per multiplication with a 0.9% reduction in DNN inference accuracy.
Vishesh Mishra, Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys, Urbi Chatterjee
ICCAD3
2025 Fault Tolerance in Quantized and Pruned Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNN), particularly those used in critical applications, such as autonomous driving, medical systems, and aerospace, require high reliability. While these algorithms exhibit inherent resilience, they remain sus-ceptible to Single-Event Effects (SEE) occurring at the hard-ware and impacting the model execution. These effects, usually induced by interactions with radiation particles, can lead to errors in electronic components, potentially causing incorrect inferences and increasing the risk of mispredictions. Meanwhile, quantization and pruning are widely employed to reduce the hardware footprint of CNN models, facilitating their deployment on embedded systems. Even when the models are reduced, CNN remain too large for an exhaustive fault injection campaign to assess their resilience. To address these challenges, we propose SFI4NN, a Statistical Fault Injection (SFI) framework specifically designed to evaluate the fault sensitivity of fixed-point quantized and pruned CNN architectures. Furthermore, we analyze the model resilience as a function of the pruning rate, showing that CNN sensitivity increases as pruning becomes more aggressive. The obtained results enable the development of hardware hardening strategies with reduced costs that are tailored to the reliability requirements of targeted applications. Experimental results demonstrate a 96 % improvement in resilience, with minimal hardware overhead compared to conventional hardening techniques such as triplication.
Wilfread Guillemé, Angeliki Kritikakou, Youri Helen, Cédric Killian, Daniel Chillet
IOLTS2
2025 Contention and Reliability-Aware Energy Efficiency Task Mapping on NoC-Based MPSoCs
abstract
Recently, network-on-chip (NoC)-based multiprocessor system-on-chips (MPSoCs) have become popular computing platforms for real-time applications due to high communication performance and energy efficiency over traditional bus-based MPSoCs. Due to the nature of network structures, network congestion along with transient faults, can significantly affect communication efficiency and system reliability. Most existing works have rarely focused on the concurrent optimization of network contention, reliability, and energy consumption. Here, we study the problem of contention and reliability-aware task mapping under real-time constraints for dynamic voltage and frequency scaling-enabled NoC. The problem entails optimizing voltage/frequency on cores and links to reduce energy consumption and ensure system reliability, while task mapping and slack time are adopted to alleviate network contention and reduce latency. We aim to minimize computation and communication energy and balance workload. This problem is formulated as a mixed-integer nonlinear programming, and we present an effective linearization scheme that equivalently transforms it into a mixed-integer linear programming to find the optimal solution. To reduce computation time, we propose a three-step heuristic, including task allocation, frequency scaling and edge scheduling, and communication contention management. Finally, we perform extensive simulations to evaluate the proposed method. The results show we can achieve 31.6% and 21.7% energy savings, with 95.5% and 98.6% less contention than the existing methods.
Lei Mo, Xinmei Li, Angeliki Kritikakou, Xiaojun Zhai
IEEE Trans. Reliab.3
2024 Efficient Neural Networks: from SW optimization to specialized HW accelerators
abstract
Artificial Neural Networks (ANNs) appear to be one of the technological revolutions of recent human history. The capability of such systems does not come at a low cost, which led researchers to develop more and more efficient techniques to implement them. Optimization approaches have been developed, such as pruning and quantization, leading to reduced memory and computation requirements. Furthermore, such approaches are adapted to the specific hardware platform features to further increase efficiency. To improve it further, the HW programmability can be traded off in favor of more specialized custom HW ANN accelerators. In this education abstract, we illustrate how optimizing operations execution at different levels, from SW to HW, can improve the efficiency of ANN execution.
Marcello Traiola, Angeliki Kritikakou, Silviu-Ioan Filip, Olivier Sentieys
CASES2
2024 HTAG-eNN: Hardening Technique with AND Gates for Embedded Neural Networks
abstract
Embedded Neural Networks (NNs) face significant challenges due to Single-Event Upsets (SEUs), compromising their reliability. To address this challenge, previous works study SEU layers sensitivity of AI models. Contrary to these techniques, remaining at high level, we propose a more accurate analysis, highlighting that, except for the last layer, faults transitioning from 0 to 1 significantly impact classification outcomes. Based on this specific behavior, we propose a simple hardware block able to detect and mitigate the SEU impact. Obtained results show that HTAG protection efficiency is near 96.85% for the LeNet-5 CNN inference model, suitable for an embedded system. This result can be improved with other protection methods for the classification layer. Additionally, it significantly reduces area overhead and critical path compared to existing approaches.
Wilfread Guillemé, Angeliki Kritikakou, Youri Helen, Cédric Killian, Daniel Chillet
DAC2
2024 Cross-Layer Reliability Evaluation and Efficient Hardening of Large Vision Transformers Models
abstract
Vision Transformers (ViTs) are highly accurate Machine Learning (ML) models. However, their large size and complexity increase the expected error rate due to hardware faults. Measuring the error rate of large ViT models is challenging, as conventional microarchitectural fault simulations can take years to produce statistically significant data. This paper proposes a two-level evaluation based on data collected through more than 70 hours of neutron beam experiments and more than 600 hours of software fault simulation. We consider 12 ViT models executed in 2 NVIDIA GPU architectures. We first characterize the fault model in ViT's kernels to identify the faults more likely to propagate to the output. We then design dedicated procedures efficiently integrated into the ViT to locate and correct these faults. We propose Maximum corrupted Malicious values (MaxiMals), an experimentally tuned low-cost mitigation solution to reduce the impact of transient faults on ViTs. We demonstrate that MaxiMals can correct 90.7% of critical failures, with execution time overheads as low as 5.61%.
Lucas Roquet, Fernando Santos 0001, Paolo Rech, Marcello Traiola, Olivier Sentieys, Angeliki Kritikakou
DAC6
2024 Reliability and Security of AI Hardware
abstract
In recent years, Artificial Intelligence (AI) systems have achieved revolutionary capabilities, providing intelligent solutions that surpass human skills in many cases. However, such capabilities come with power-hungry computation workloads. Therefore, the implementation of hardware acceleration becomes as fundamental as the software design to improve energy efficiency, silicon area, and latency of AI systems. Thus, innovative hardware platforms, architectures, and compiler-level approaches have been used to accelerate AI workloads. Crucially, innovative AI acceleration platforms are being adopted in application domains for which dependability must be paramount, such as autonomous driving, healthcare, banking, space exploration, and industry 4.0. Unfortunately, the complexity of both AI software and hardware makes the dependability evaluation and improvement extremely challenging. Studies have been conducted on both the security and reliability of AI systems, such as vulnerability assessments and countermeasures to random faults and analysis for side-channel attacks. This paper describes and discusses various reliability and security threats in AI systems, and presents representative case studies along with corresponding efficient countermeasures.
Dennis Gnad, Martin Gotthard, Jonas Krautter, Angeliki Kritikakou, Vincent Meyers, Paolo Rech, Josie E. Rodriguez Condia, Annachiara Ruospo, Ernesto Sánchez 0001, Fernando Santos 0001, Olivier Sentieys, Mehdi Baradaran Tahoori, Russell Tessier, Marcello Traiola
ETS4
2024 VANDOR: Mitigating SEUs into Quantized Neural Networks
abstract
Embedded neural networks are increasingly deployed in critical applications, such as avionics and autonomous vehicle control. However, their reliability is challenged by various sources of soft errors, including radiation-induced faults from cosmic ray strikes, leading to Single Event Upsets (SEUs). To ensure the reliability of such systems, we present a novel hardware-based fault protection strategy tailored for embedded neural networks. The idea is based on mitigating faults by adapting at run-time any erroneous values (parameters, intermediate data) due to SEU towards zero upon fault detection. As neural networks exhibit heterogeneous sensitivity to fault direction, our hardware-based approach triplicates the sign bit (TMR) and uses a Voter block based on logical AND/OR gates to handle fault directionality. Through a comprehensive and exhaustive fault injection study, conducted on a Convolutional Neural Network (CNN) model, implemented on FPGA using fixed-point quantization, we show that our method is applicable to various hardware architectures while optimizing hardware cost, a crucial aspect in the context of embedded systems. Obtained results show that VANDOR protection efficiency is near ${9 0 . 9 7 \%}$ for the LeNet-5 CNN inference model, suitable for an embedded system. Additionally, it significantly reduces area overhead compared to existing approaches.
Wilfread Guillemé, Angeliki Kritikakou, Youri Helen, Cédric Killian, Daniel Chillet
IOLTS2
2024 Combining Fault Simulation and Beam Data for CNN Error Rate Estimation on RISC-V Commercial Platforms
abstract
Thanks to the RISC-V open-source Instruction Set Architecture, researchers and developers can efficiently propose new solutions at a low cost and low power consumption. RISCV-based architectures can then be customized to run Machine Learning (ML) algorithms efficiently and inserted in safety and mission-critical domains, where the execution must be reliable. However, a fault in the hardware resources can compromise the system’s ability to operate correctly. Thus, it is necessary to characterize the ML applications’ vulnerabilities on RISCV processors and how errors in those operations impact the Convolutional Neural Network (CNN) misclassification rate. In this research paper, we assess the error rate induced by neutrons on the basic operations of a CNN running on a RISC-V-based processor (GAP8) and how each operation contributes to the entire CNN error rate. Our findings indicate that memory errors are the primary contributors to the system’s error rate. Furthermore, we present a case study demonstrating how the CNN microbenchmarks can be used to estimate the error rate of an entire CNN. By combining data from fault simulation and beam experiments, our error rate estimation led to a result that closely matches those obtained solely from beam experiments.
Fernando Santos 0001, Marcello Traiola, Angeliki Kritikakou
IOLTS3
2024 Reliability Assessment of Large DNN Models: Trading Off Performance and Accuracy
abstract
The adoption of Deep Neural Networks (DNNs) in several domains allows for increased effectiveness in applications that deal with massive data-intensive and complex data inputs. When employed in safety-critical scenarios, such as automotive, aerospace, healthcare, and autonomous robotics, assessing the DNNs' reliability and functional safety is crucial to ensure their correct in-field operation, even in the presence of hardware faults. However, the system complexity and the massive amounts of data to be processed by DNNs prevent the effective adoption of traditional strategies for reliability characterization and for identifying the most fault-sensitive structures. Accurate fault assessment strategies usually require unacceptable computational power and large evaluation times. On the other hand, faster strategies commonly lack accuracy in correctly representing system faults. Consequently, it is necessary to develop effective strategies that trade-off between performance and accuracy. This work analyses three reliability assessment strategies for deep neural networks and their underlying hardware, highlighting the main solutions and challenges in terms of evaluation performance and fault characterization accuracy. We overview different solutions to evaluate the hardware accelerators implementing DNNs at three abstraction levels:$i$) by physically injecting faults on a GPU running DNNs, ii) by performing microarchitectural characterization of GPUs to develop application-accurate error models, and iii) by using structure-aware cross-layer error modeling on DNN hardware accelerators. Our experimental results indicate that accurate error representation requires structural features from the targeted hardware.
Junchao Chen 0001, Giuseppe Esposito, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Angeliki Kritikakou, Milos Krstic, Robert Limas Sierra, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Marcello Traiola, Alessandro Veronesi
VLSI-SoC5
2024 Special Session: Reliability Assessment Recipes for DNN Accelerators
abstract
Reliability assessment is mandatory to guarantee the correct behavior of Deep Neural Network (DNN) hardware accelerators in safety-critical applications. While fault injection stands out as a well-established, practical and robust method for reliability assessment, it is still a very time-consuming process. This paper contributes with three recipes for optimizing the efficiency of the reliability assessment: a) hybrid analytical and hierarchical FI-based reliability assessment for systolic-array-based DNN accelerators; b) mixing techniques for the reliability assessment of in-chip AI accelerators in GPUs; c) reliability assessment of DNN hardware accelerators through physical fault injection. The experimental results demonstrate the efficiency of the proposed methods applied to their target DNN HW accelerator platforms.
Mohammad Hasan Ahmadilivani, Alberto Bosio, Bastien Deveautour, Fernando Santos 0001, Juan-David Guerrero-Balaguera, Maksim Jenihhin, Angeliki Kritikakou, Robert Limas Sierra, Salvatore Pappalardo, Jaan Raik, Josie E. Rodriguez Condia, Matteo Sonza Reorda, Mahdi Taheri, Marcello Traiola
VTS7
2023 A machine-learning-guided framework for fault-tolerant DNNs
abstract
Deep Neural Networks (DNNs) show promising per-formance in several application domains. Nevertheless, DNN results may be incorrect, not only because of the network intrinsic inaccuracy, but also due to faults affecting the hardware. Ensuring the fault tolerance of DNN is crucial, but common fault tolerance approaches are not cost-effective, due to the prohibitive overheads for large DNNs. This work proposes a comprehensive framework to assess the fault tolerance of DNN parameters and cost-effectively protect them. As a first step, the proposed framework performs a statistical fault injection. The results are used in the second step with classification-based machine learning methods to obtain a bit-accurate prediction of the criticality of all network parameters. Last, Error Correction Codes (ECCs) are selectively inserted to protect only the critical parameters, hence entailing low cost. Thanks to the proposed framework, we explored and protected two Convolutional Neural Networks (CNNs), each with four different data encoding. The results show that it is possible to protect the critical network parameters with selective ECCs while saving up to 79% memory w.r.t. conventional ECC approaches.
Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys
DATE2
2023 Impact of Transient Faults on Timing Behavior and Mitigation with Near-Zero WCET Overhead
abstract
As time-critical systems require timing guarantees, Worst-Case Execution Times (WCET) have to be employed. However, WCET estimation methods usually assume fault-free hardware. If proper actions are not taken, such fault-free WCET approaches become unsafe, when faults impact the hardware during execution. The majority of approaches, dealing with hardware faults, address the impact of faults on the functional behavior of an application, i.e., denial of service and binary correctness. Few approaches address the impact of faults on the application timing behavior, i.e., time to finish the application, and target faults occurring in memories. However, as the transistor size in modern technologies is significantly reduced, faults in cores cannot be considered negligible anymore. This work shows that faults not only affect the functional behavior, but they can have a significant impact on the timing behavior of applications. To expose the overall impact of faults, we enhance vulnerability analysis to include not only functional, but also timing correctness, and show that faults impact WCET estimations. As common techniques to deal with faults, such as watchdog timers and re-execution, have large timing overhead for error detection and correction, we propose a mechanism with near-zero and bounded timing overhead. A RISC-V core is used as a case study. The obtained results show that faults can lead up to almost 700% increase in the maximum observed execution time between fault-free and faulty execution without protection, affecting the WCET estimations. On the contrary, the proposed mechanism is able to restore fault-free WCET estimations with a bounded overhead of 2 execution cycles.
Pegdwende Romaric Nikiema, Angeliki Kritikakou, Marcello Traiola, Olivier Sentieys
ECRTS2
2023 harDNNing: a machine-learning-based framework for fault tolerance assessment and protection of DNNs
abstract
Deep Neural Networks (DNNs) show promising performance in several application domains, such as robotics, aerospace, smart healthcare, and autonomous driving. Nevertheless, DNN results may be incorrect, not only because of the network intrinsic inaccuracy, but also due to faults affecting the hardware. Indeed, hardware faults may impact the DNN inference process and lead to prediction failures. Therefore, ensuring the fault tolerance of DNN is crucial. However, common fault tolerance approaches are not cost-effective for DNNs protection, because of the prohibitive overheads due to the large size of DNNs and of the required memory for parameter storage. In this work, we propose a comprehensive framework to assess the fault tolerance of DNNs and cost-effectively protect them. As a first step, the proposed framework performs data-type-and-layer-based fault injection, driven by the DNN characteristics. As a second step, it uses classification-based machine learning methods in order to predict the criticality, not only of network parameters, but also of their bits. Last, dedicated Error Correction Codes (ECCs) are selectively inserted to protect the critical parameters and bits, hence protecting the DNNs with low cost. Thanks to the proposed framework, we explored and protected two Convolutional Neural Networks (CNNs), each with four different data encoding. The results show that it is possible to protect the critical network parameters with selective ECCs while saving up to 83% memory w.r.t. conventional ECC approaches.
Marcello Traiola, Angeliki Kritikakou, Olivier Sentieys
ETS2
2023 Towards Dependable RISC-V Cores for Edge Computing Devices
abstract
The migration of the computation from the cloud into edge devices, i.e., Internet-of-Things (IoTs) devices, reduces the latency and the quantity of data flowing into the network. With the emerging open-source and customizable RISC-V Instruction Set Architecture (ISA), cores based on such ISA are promising candidates for several application domains within the IoT family, such as automotive, Unnamed Aerial Vehicles (UAVs), industrial automation, healthcare, agriculture etc., where power consumption, real-time execution, security and reliability are of highest importance. In this emerging new era of connected RISC-V IoT devices, mechanisms are needed for a reliable and secure execution, still meeting area, energy consumption and computation time constraints of edge devices. We propose three mechanisms towards this goal, i.e., (i) a Root of Trust module for post-quantum secure boot, (ii) hardware checkers against hardware trojan horses and microarchitectural side-channel attacks, and (iii) a fine-grained dual core lockstep mechanism for real-time error detection and correction. The paper illustrates the proposed mechanisms with related motivations and implications, as well as a discussion on future research directions.
Pegdwende Romaric Nikiema, Alessandro Palumbo, Allan Aasma, Luca Cassano, Angeliki Kritikakou, Ari Kulmala, Jari Lukkarila, Marco Ottavi, Rafail Psiakis, Marcello Traiola
IOLTS5
2023 Near-optimal energy-efficient partial-duplication task mapping of real-time parallel applications
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
J. Syst. Archit.2
2023 Approximation-Aware Task Deployment on Heterogeneous Multicore Platforms With DVFS
abstract
Heterogeneous (HE) multicore platforms, such as ARM big.LITTLE, are widely used to execute embedded applications under multiple and contradictory constraints, such as energy consumption and real-time (RT) execution. To fulfill these constraints and optimize system performance, application tasks should be efficiently mapped on multicore platforms. Embedded applications are usually tolerant to approximated results but acceptable quality of service (QoS). Modeling-embedded applications by using the elastic task model, namely, imprecise computation (IC) task model, can balance system QoS, energy consumption, and RT performance during task deployment. However, state-of-the-art approaches seldom consider the problem of IC task deployment on HE multicore platforms. They typically neglect task migration, which can improve the solutions due to its flexibility during the task deployment process. This article proposes a novel QoS-aware task deployment method to maximize system QoS under energy and RT constraints, where the frequency assignment (FA), task allocation (TA), scheduling, and migration are optimized simultaneously. The task deployment problem is formulated as mixed-integer nonlinear programming. Then, it is linearized to mixed-integer linear programming to find the optimal (OPT) solution. Furthermore, based on the problem structure and problem decomposition, we propose a novel heuristic (HEU) with low computational complexity. The subproblems regarding FA, TA, scheduling, and adjustment are considered and solved in sequence. Finally, the simulation results show that the proposed task deployment method improves the system QoS by 31.2% on average (up to 112.8%) compared to the state-of-the-art methods and the designed HEU achieves about 53.9% (on average) performance of the OPT solution with a negligible computing time.
Xinmei Li, Lei Mo, Angeliki Kritikakou, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Mitigating Mode-switch through Run-time Computation of Response Time
abstract
Mixed-critical systems consist of applications with different criticality. In these systems, different confidence levels of Worst-Case Execution Time (WCET) estimations are used. Dual criticality systems use a less pessimistic, but with lower level of assurance, WCET estimation, and a safe, but pessimistic, WCET estimation. Initially, both high and low criticality tasks are executed. When a high criticality task exceeds its less pessimistic WCET, the system switches mode and low criticality tasks are usually dropped, reducing the overall system Quality of Service (QoS). To postpone mode switch, and thus, improve QoS, existing approaches explore the slack, created dynamically, when the actual execution of a task is faster than its WCET. However, existing approaches observe this slack only after the task has finished execution. To enhance dynamic slack exploitation, we propose a fine-grained approach that is able to expose the slack during the progress of a task, and safely uses it to postpone mode switch. The evaluation results show that the proposed approach has lower cost and achieves significant improvements in avoiding mode-switch, compared to existing approaches.
Angeliki Kritikakou, Stefanos Skalistis
ACM Trans. Design Autom. Electr. Syst.1
2023 Energy Optimized Task Mapping for Reliable and Real-Time Networked Systems
abstract
Energy efficiency, real-time response, and data transmission reliability are important objectives during networked systems design. This paper aims to develop an efficient task mapping scheme to balance these important but conflicting objectives. To achieve this goal, tasks are triplicated to enhance reliability and mapped on the wireless nodes of the networked systems with Dynamic Voltage and Frequency Scaling (DVFS) capabilities to reduce energy consumption while still meeting real-time constraints. Our contributions include the mathematical formulation of this task mapping problem as mixed-integer programming that balances node energy consumption, enhancing data reliability, under real-time and energy constraints. Compared with the State-of-the-Art (SoA) , a joint-design problem is considered in this paper, where DVFS, task triplication, task allocation, and task scheduling are optimized concurrently. To find the optimal solution, the original problem is linearized, and a decomposition-based method is proposed. The optimality of the proposed method is proved rigorously. Furthermore, a heuristic based on the greedy algorithm is designed to reduce the computation time. The proposed methods are evaluated and compared through a series of simulations. The results show that the proposed triplication-based task mapping method on average achieves 24.84% runtime reduction and 28.62% energy saving compared to the SoA methods.
Lei Mo, Angeliki Kritikakou, Xianghui Cao
ACM Trans. Sens. Networks3
2022 Flodam: Cross-Layer Reliability Analysis Flow for Complex Hardware Designs
abstract
Modern technologies make hardware designs more and more sensitive to radiation particles and related faults. As a result, analysing the behavior of a system under radiation-induced faults has become an essential part of the system design process. Existing approaches either focus on analysing the radiation impact at the lower hardware design layers, without further propagating any radiation-induced fault to the system execution, or analyse system reliability at higher hardware or application layers, based on fault models that are agnostic of the fabrication technology and the radiation environment. Flodam combines the benefits of existing approaches by providing a novel cross-layer reliability analysis from the semiconductor layer up to the application layer, able to quantify the risks of faults under a given context, taking into account the environmental conditions, the physical hardware design and the application under study.
Angeliki Kritikakou, Olivier Sentieys, Guillaume Hubert, Youri Helen, Jean-Francois Coulon, Patrice Deroux-Dauphin
DATE1
2022 Energy Efficient, Real-time and Reliable Task Deployment on NoC-based Multicores with DVFS
abstract
Task deployment plays an important role in the overall system performance, especially for complex architectures, including several cores with Dynamic Voltage and Frequency Scaling (DVFS) and Network-on-Chips (NoC). Task deployment affects not only the energy consumption but also the real-time response and reliability of the system. In this work, a task deployment approach is proposed to optimize the overall system energy consumption, including computation of the cores and communication of the NoC, under task reliability and real-time constraints. More precisely, the task deployment approach combines task allocation and scheduling, frequency assignment, task duplication, and multi-path data routing. The task deployment problem is formulated using mixed-integer non-linear programming. To find the optimal solution, the original problem is equivalently transformed to mixed-integer linear programming, and solved by state-of-the-art solvers. Furthermore, a decomposition-based heuristic, with low computational complexity, is proposed to deal with scalability. Finally, extended simulations evaluate the proposed methods.
Lei Mo, Angeliki Kritikakou, Ji Liu 0003
DATE3
2022 Functional and Timing Implications of Transient Faults in Critical Systems
abstract
Embedded systems in critical domains, such as auto-motive, aviation, space domains, are often required to guarantee both functional and temporal correctness. Considering transient faults, fault analysis and mitigation approaches are implemented at various levels of the system design, in order to maintain the functional correctness. However, transient faults and their mitigation methods have a timing impact, which can affect the temporal correctness of the system. In this work, we expose the functional and the timing implications of transient faults for critical systems. More precisely, we initially highlight the timing effect of transient faults occurring in the combinational and sequential logic of a processor. Furthermore, we propose a full stack vulnerability analysis that drives the design of selective hardware-based mitigation for real-time applications. Last, we study the timing impact of software-based reliability mitigation methods applied in a COTS GPU, using a fault tolerant middleware.
Angeliki Kritikakou, Panagiota Nikolaou, Ivan Rodriguez-Ferrandez, Joseph Paturel, Leonidas Kosmidis, Maria K. Michael, Olivier Sentieys, David Steenari
IOLTS1
2022 Experimental evaluation of neutron-induced errors on a multicore RISC-V platform
abstract
RISC-V architectures have gained importance in the last years due to their flexibility and open-source Instruction Set Architecture (ISA), allowing developers to efficiently adopt RISC-V processors in several domains with a reduced cost. For application domains, such as safety-critical and mission-critical, the execution must be reliable as a fault can compromise the system’s ability to operate correctly. However, the application’s error rate on RISC-V processors is not significantly evaluated, as it has been done for standard x86 processors. In this work, we investigate the error rate of a commercial RISC-V ASIC platform, the GAP8, exposed to a neutron beam. We show that for computing-intensive applications, such as classification Convolutional Neural Networks (CNN), the error rate can be $3.2 \times$ higher than the average error rate. Additionally, we find that the majority (96.12%) of the errors on the CNN do not generate misclassifications. Finally, we also evaluate the events that cause application interruption on GAP8 and show that the major source of incorrect interruptions is application hangs (i.g., due to an infinite loop or a racing condition)
Fernando Santos 0001, Angeliki Kritikakou, Olivier Sentieys
IOLTS2
2022 BiSuT: A NoC-Based Bit-Shuffling Technique for Multiple Permanent Faults Mitigation
abstract
Since several decades, fault tolerance has become a major research field due to transistor shrinking and core number increasing in system-on-chip (SoC). Especially, faults occurring to the network-on-chips (NoCs) of those systems have a significant impact, due to the high amount of data, crossing the NoC, for the communication among intellectual properties (IPs). Furthermore, existing fault-tolerant approaches cannot efficiently deal with several permanent faults, which occur in NoC routers. To address these limitations, we propose the bit shuffling method (BiSuT) for fault-tolerant NoCs that reduces the impact of faults on data communications. To achieve that, the proposed approach exploits, at runtime, the position of permanent faults and changes the order of bits inside a flit. Our method reduces, as much as possible, the impact of faults by transferring the faults on least significant bits (LSBs), instead of keeping them on most significant bits (MSBs). The results obtained by extensive evaluations show that BiSuT can reduce the impact of multiple permanent faults, with low hardware costs, compared to the existing approaches, like the Hamming code.
Romain Mercier, Cédric Killian, Angeliki Kritikakou, Youri Helen, Daniel Chillet
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Fault-Tolerant Mapping of Real-Time Parallel Applications under multiple DVFS schemes
abstract
On multicore platforms, reliable task execution, as well as low energy consumption, are essential. Dynamic Voltage/Frequency Scaling (DVFS) is typically used for energy saving, but with a negative impact on reliability, especially when the frequency is low. Using high frequencies to meet reliability constraints is not always feasible, while multiple replicas increase energy consumption. To minimize energy consumption, enhancing reliability, without violating real-time constraints, we propose an approach that combines distinct reliability enhancement techniques, under task-level, processor-level and system-level DVFS. Our task mapping problem jointly optimizes task allocation, task frequency assignment, and task duplication, under multiple constraints. This is achieved by formulating the task mapping problem as a mixed integer non-linear programming problem and equivalently transforming it into a mixed integer linear programming, that is optimally solved. From the obtained results, the proposed approach achieves better energy consumption and finds solutions, when other approaches fail.
Minyu Cui, Angeliki Kritikakou, Lei Mo, Emmanuel Casseau
RTAS2
2021 Real-Time Imprecise Computation Tasks Mapping for DVFS-Enabled Networked Systems
abstract
Networked systems are useful for a wide range of applications, many of which require distributed and collaborative data processing to satisfy real-time requirements. On one hand, networked systems are usually resource constrained, mainly regarding the energy supply of the nodes and their computation and communication abilities. On the other hand, many real-time applications can be executed in an imprecise way, where an approximate result is acceptable as long as the baseline Quality of Service (QoS) is satisfied. Such applications can be modeled through imprecise computation (IC) tasks. To achieve a better tradeoff between QoS and limited system resources, while meeting application requirements, the IC-tasks must be efficiently mapped to the system nodes. To tackle this problem, we first construct an IC-task mapping problem that aims to maximize system QoS subject to real-time and energy constraints. Dynamic voltage and frequency scaling (DVFS) and multipath routing are explored to further enhance real-time performance and reduce energy consumption. Second, based on the problem structure, we propose an optimal approach to perform IC-task mapping and prove its optimality. Furthermore, to enhance the scalability of the proposed approach, we present a heuristic IC-task mapping method with low computation time. Finally, the simulation results demonstrate the effectiveness of the proposed methods in terms of the solution quality and the computation time.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys, Xianghui Cao
IEEE Internet Things J.2
2021 Software and hardware co-design for sustainable cyber-physical systems
abstract
This special issue aims to provide a platform for the researchers, academia, and industry to present their novel solutions, applications, tools, software, hardware, and algorithms designed for addressing various sustainability challenges in CPS.The response from the CPS community was enthusiastic: the special issue received 34 manuscripts submitted by the authors from China, United States, India, Korea, Lebanon, and so on.According to the Journal of Software: Practice and Experience (SPE) review standards, this special issue accepted 14 high-quality research articles that cover a wide range of topics.These articles provide the software and hardware co-design solutions to improve dependability, energy efficiency, quality of service (QoS) of CPS, and also to introduce the methodologies for specific CPS applications.
Junlong Zhou, Angeliki Kritikakou, Dakai Zhu 0001, José L. Martínez Lastra, Shiyan Hu 0001
Softw. Pract. Exp.2
2020 Dynamic Interference-Sensitive Run-time Adaptation of Time-Triggered Schedules
abstract
Over-approximated Worst-Case Execution Time (WCET) estimations for multi-cores lead to safe, but over-provisioned, systems and underutilized cores. To reduce WCET pessimism, interference-sensitive WCET (isWCET) estimations are used. Although they provide tighter WCET bounds, they are valid only for a specific schedule solution. Existing approaches have to maintain this isWCET schedule solution at run-time, via time-triggered execution, in order to be safe. Hence, any earlier execution of tasks, enabled by adapting the isWCET schedule solution, is not possible. In this paper, we present a dynamic approach that safely adapts isWCET schedules during execution, by relaxing or completely removing isWCET schedule dependencies, depending on the progress of each core. In this way, an earlier task execution is enabled, creating time slack that can be used by safety-critical and mixed-criticality systems to provide higher Quality-of-Services or execute other best-effort applications. The Response-Time Analysis (RTA) of the proposed approach is presented, showing that although the approach is dynamic, it is fully predictable with bounded WCET. To support our contribution, we evaluate the behavior and the scalability of the proposed approach for different application types and execution configurations on the 8-core Texas Instruments TMS320C6678 platform, obtaining significant performance improvements compared to static approaches.
Stefanos Skalistis, Angeliki Kritikakou
ECRTS2
2020 Progress-aware Dynamic Slack Exploitation in Mixed-critical Systems: Work-in-Progress
abstract
Mixed-critical systems consist of high criticality and low criticality applications. When a high criticality task exceeds its less pessimistic Worst Case Execution Time (WCET), the system switches mode and low criticality tasks are usually dropped. To postpone mode switch, existing approaches explore the slack, created dynamically due to an actual execution of a task that is faster than its WCET. However, the slack is visible only after the task has finished. To further enable dynamic slack exploitation, we propose a fine-grained approach that exposes the slack created due to the progress of tasks, during execution, and safely uses it to postpone mode switch.
Angeliki Kritikakou, Stefanos Skalistis
EMSOFT1
2020 Multiple Permanent Faults Mitigation Through Bit-Shuffling for Network-an-Chip Architecture
abstract
Since several decades, fault tolerance has become a major research field, due to transistor shrinking and core number increasing in System-on-Chip (SoC). Especially, faults occurring at the Network-on-Chips (NoCs) of those systems have a significant impact, since NoCs are the key component of on-chip communication. Several fault tolerant approaches have been proposed, which are, however, limited against multiple permanent faults. To reduce the impact of these faults on the data communications, we propose a bit-shuffling method for fault tolerant NoCs. The proposed approach exploits, at runtime, the position of the permanent faults and changes the order of bits inside a flit. Our bit-shuffling method reduces as much as possible the fault impact, by transferring the faults from Most Significant Bits (MSBs) towards Least Significant Bits (LSBs). With this technique, we show that, in presence of multiple permanent faults, the Mean Square Error (MSE) on the payload transmission is reduce from 1017to 105under three permanent fault for 32-bit unsigned integers. This technique also ensures the correct transmission of headers under multiple permanent faults.
Romain Mercier, Cédric Killian, Angeliki Kritikakou, Youri Helen, Daniel Chillet
ICCD3
2019 Approximation-aware Task Deployment on Asymmetric Multicore Processors
abstract
Asymmetric Multicore Processors (AMP) are a very promising architecture to deal efficiently with the wide diversity of applications. In real-time application domains, in-time approximated results are preferred than accurate - but too late - results. In this work, we propose a deployment approach that exploits the heterogeneity provided by AMP architectures and the approximation tolerance provided by the applications, so as to increase as much as possible the quality of the results under given energy and timing constraints. Initially, an optimal approach is proposed based on problem linearization and decomposition. Then, a heuristic approach is developed based on iteration relaxation of the optimal version. The obtained results show 16.3% reduction in the computation time for the optimal approach compared to the conventional optimal approaches. The proposed heuristic approach is about 100 times faster at the cost of a 29.8% QoS degradation in comparison with the optimal solution.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
DATE2
2019 Fine-Grained Hardware Mitigation for Multiple Long-Duration Transients on VLIW Function Units
abstract
Technology scaling makes hardware more susceptible to radiation, which can cause multiple transient faults with long duration. In these cases, the affected function unit is usually considered as faulty and is not further used. To reduce this performance degradation, the proposed hardware mechanism detects the faults that are still active during execution and reschedules the instructions to use the fault-free components of the affected function units. The results show multiple long-duration fault mitigation with low performance, area, and power overhead.
Rafail Psiakis, Angeliki Kritikakou, Olivier Sentieys
DATE2
2019 Timely Fine-Grained Interference-Sensitive Run-Time Adaptation of Time-Triggered Schedules
abstract
In time-critical systems, run-time adaptation is required to improve the performance of time-triggered execution, derived based on Worst-Case Execution Time (WCET) of tasks. By improving performance, the systems can provide higher Quality-of-Service, in safety-critical systems, or execute other best-effort applications, in mixed-critical systems. To achieve this goal, we propose a parallel interference-sensitive run-time adaptation mechanism that enables a fine-grained synchronisation among cores. Since the run-time adaptation of offline solutions can potentially violate the timing guarantees, we present the Response-Time Analysis (RTA) of the proposed mechanism showing that the system execution is free of timing-anomalies. The RTA takes into account the timing behavior of the proposed mechanism and its associated WCET. To support our contribution, we evaluate the behavior and the scalability of the proposed approach for different application types and execution configurations on the 8-core Texas Instruments TMS320C6678 platform. The obtained results show significant performance improvement compared to state-of-the-art centralized approaches.
Stefanos Skalistis, Angeliki Kritikakou
RTSS2
2019 Energy-Aware Multiple Mobile Chargers Coordination for Wireless Rechargeable Sensor Networks
abstract
Wireless charging provides dynamic power supply for wireless sensor networks (WSNs). Such systems, are typically considered under the scenario of wireless rechargeable sensor networks (WRSNs). With the use of mobile chargers (MCs), the flexibility of WRSNs is further enhanced. However, the use of MCs poses several challenges during the system design. The coordination process has to simultaneously optimize the scheduling, the moving time, and the charging time of multiple MCs under limited system resources (time and energy). Efficient methods that jointly solve these challenges are generally lacking in the literature. In this paper, we address the multiple MCs coordination problem under multiple system requirements. First, we aim at minimizing the energy consumption of MCs, guaranteeing that every sensor will not run out of energy. We formulate the multiple MCs coordination problem as a mixed-integer linear programming and derive a set of desired network properties. Second, we propose a novel decomposition method to optimally solve the problem, as well as to reduce the computation time. Our approach divides the problem into a subproblem for the MC scheduling and a subproblem for the MC moving time and charging time, and solves them iteratively by utilizing the solution of one into the other. The convergence of proposed method is analyzed theoretically. Simulation results demonstrate the effectiveness and scalability of the proposed method in terms of solution quality and computation time.
Lei Mo, Angeliki Kritikakou, Shibo He
IEEE Internet Things J.2
2019 Mapping imprecise computation tasks on cyber-physical systems
Lei Mo, Angeliki Kritikakou
Peer-to-Peer Netw. Appl.2
2019 Event-Driven Joint Mobile Actuators Scheduling and Control in Cyber-Physical Systems
abstract
In cyber-physical systems, mobile actuators can enhance system's flexibility and scalability, but at the same time incurs complex couplings in the scheduling and controlling of the actuators. In this paper, we propose a novel event-driven method aiming at satisfying a required level of control accuracy and saving energy consumption of the actuators, while guaranteeing a bounded action delay. We formulate a joint-design problem of both actuator scheduling and output control. To solve this problem, we propose a two-step optimization method. In the first step, the problem of actuator scheduling and action time allocation is decomposed into two subproblems. They are solved iteratively by utilizing the solution of one in the other. The convergence of this iterative algorithm is proved. In the second step, an online method is proposed to estimate the error and adjust the outputs of the actuators accordingly. Through simulations and experiments, we demonstrate the effectiveness of the proposed method.
Lei Mo, Pengcheng You, Xianghui Cao, Yeqiong Song, Angeliki Kritikakou
IEEE Trans. Ind. Informatics5
2018 Distributed Node Coordination for Real-Time Energy-Constrained Control in Wireless Sensor and Actuator Networks
abstract
Wireless sensor and actuator networks (WSANs) are emerging as a new generation of wireless sensor networks. Due to the coupling between the sensing areas of the sensors and the action areas of the actuators, the efficient coordination among the nodes is a great challenge. In this paper, we address the problem of distributed node coordination in WSANs aiming at meeting the user's requirements on the states of the points of interest (POIs) in a real-time and energy-efficient manner. The node coordination problem is formulated as a nonlinear program. To solve it efficiently, the problem is divided into two correlated subproblems: 1) the sensor-actuator (S-A) coordination and 2) the actuator-actuator (A-A) coordination. In the S-A coordination, a distributed federated Kalman filter-based estimation approach is applied for the actuators to collaborate with their ambient sensors to estimate the states of the POIs. In the A-A coordination, a distributed Lagrange-based control method is designed for the actuators to optimally adjust their outputs, based on the estimated results from the S-A coordination. The convergence of the proposed method is proved rigorously. As the proposed node coordination scheme is distributed, we find the optimal solution while avoiding high computational complexity. The simulation results also show that the proposed distributed approach is an efficient and practically applicable method with reasonable complexity.
Lei Mo, Xianghui Cao, Yeqiong Song, Angeliki Kritikakou
IEEE Internet Things J.4
2018 Energy-Quality-Time Optimized Task Mapping on DVFS-Enabled Multicores
abstract
Multicore architectures have great potential for energy-constrained embedded systems, such as energy-harvesting wireless sensor networks. Some embedded applications, especially the real-time ones, can be modeled as imprecise computation tasks. A task is divided into a mandatory subtask that provides a baseline quality-of-service (QoS) and an optional subtask that refines the result to increase the QoS. Combining dynamic voltage and frequency scaling, task allocation, and task adjustment, we can maximize the system QoS under real-time and energy supply constraints. However, the nonlinear and combinatorial nature of this problem makes it difficult to solve. This paper first formulates a mixed-integer nonlinear programming problem to concurrently carry out task-to-processor allocation, frequency-to-task assignment and optional task adjustment. We provide a mixed-integer linear programming form of this formulation without performance degradation and we propose a novel decomposition algorithm to provide an optimal solution with reduced computation time compared to state-of-the-art optimal approaches (22.6% in average). We also propose a heuristic version that has negligible computation time.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 DYNASCORE: DYNAmic Software COntroller to Increase REsource Utilization in Mixed-Critical Systems
abstract
In real-time mixed-critical systems, Worst-Case Execution Time (WCET) analysis is required to guarantee that timing constraints are respected—at least for high-criticality tasks. However, the WCET is pessimistic compared to the real execution time, especially for multicore platforms. As WCET computation considers the worst-case scenario, it means that whenever a high-criticality task accesses a shared resource in multicore platforms, it is considered that all cores use the same resource concurrently. This pessimism in WCET computation leads to a dramatic underutilization of the platform resources, or even failing to meet the timing constraints. In order to increase resource utilization while guaranteeing real-time guarantees for high-criticality tasks, previous works proposed a runtime control system to monitor and decide when the interferences from low-criticality tasks cannot be further tolerated. However, in the initial approaches, the points where the controller is executed were statically predefined. In this work, we propose a dynamic runtime control which adapts its observations to online temporal properties, further increasing the dynamism of the approach, and mitigating the unnecessary overhead implied by existing static approaches. Our dynamic adaptive approach allows one to control the ongoing execution of tasks based on runtime information, and further increases the gains in terms of resource utilization compared with static approaches.
Angeliki Kritikakou, Thibaut Marty, Matthieu Roy
ACM Trans. Design Autom. Electr. Syst.1
2017 WCET-aware parallelization of model-based applications for multi-cores: The ARGO approach
abstract
Parallel architectures are nowadays not only confined to the domain of high performance computing, they are also increasingly used in embedded time-critical systems. The ARGO H2020 project1provides a programming paradigm and associated tool flow to exploit the full potential of architectures in terms of development productivity, time-to-market, exploitation of the platform computing power and guaranteed real-time performance. In this paper we give an overview of the objectives of ARGO and explore the challenges introduced by our approach.
Steven Derrien, Isabelle Puaut, Panayiotis Alefragis, Marcus Bednara, Harald Bucher, Clément David, Yann Debray, Umut Durak, Imen Fassi, Christian Ferdinand, Damien Hardy, Angeliki Kritikakou, Gerard K. Rauwerda, Simon Reder, Martin Sicks, Timo Stripf, Kim Sunesen, Timon D. ter Braak, Nikos S. Voros, Jürgen Becker 0001
DATE12
2017 Decomposed Task Mapping to Maximize QoS in Energy-Constrained Real-Time Multicores
abstract
Multicore architectures are now widely used in energy-constrained real-time systems, such as energy-harvesting wireless sensor networks. To take advantage of these multicores, there is a strong need to balance system energy, performance and Quality-of-Service (QoS). The Imprecise Computation (IC) model splits a task into mandatory and optional parts allowing to tradeoff QoS. The problem of mapping, i.e. allocating and scheduling, IC-tasks to a set of processors to maximize system QoS under real-time and energy constraints can be formulated as a Mixed Integer Linear Programming (MILP) problem. However, state-of-the-art solving techniques either demand high complexity or can only achieve feasible (suboptimal) solutions. In this paper, we develop an effective decomposition-based approach to achieve an optimal solution while reducing computational complexity. It decomposes the original problem into two smaller easier-to-solve problems: a master problem for IC-tasks allocation and a slave problem for IC-tasks scheduling. We also provide comprehensive optimality analysis for the proposed method. Through the simulations, we validate and demonstrate the performance of the proposed method, resulting in an average 55% QoS improvement with regards to published techniques.
Lei Mo, Angeliki Kritikakou, Olivier Sentieys
ICCD2
2016 A high-performance matrix-matrix multiplication methodology for CPU and GPU architectures
Vasilios I. Kelefouras, Angeliki Kritikakou, Iosif Mporas, Vasileios Kolonias
J. Supercomput.2
2016 Array Size Computation under Uniform Overlapping and Irregular Accesses
abstract
The size required to store an array is crucial for an embedded system, as it affects the memory size, the energy per memory access, and the overall system cost. Existing techniques for finding the minimum number of resources required to store an array are less efficient for codes with large loops and not regularly occurring memory accesses. They have to approximate the accessed parts of the array leading to overestimation of the required resources. Otherwise, their exploration time is increased with an increase over the number of the different accessed parts of the array. We propose a methodology to compute the minimum resources required for storing an array which keeps the exploration time low and provides a near-optimal result for regularly and non-regularly occurring memory accesses and overlapping writes and reads.
Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis
ACM Trans. Design Autom. Electr. Syst.1
2015 A methodology for speeding up loop kernels by exploiting the software information and the memory architecture
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis
Comput. Lang. Syst. Struct.2
2015 A methodology for speeding up matrix vector multiplication for single/multi-core architectures
Vasilios I. Kelefouras, Angeliki Kritikakou, Elissavet Papadima, Constantinos E. Goutis
J. Supercomput.2
2014 Run-Time Control to Increase Task Parallelism In Mixed-Critical Systems
abstract
Although multi/many-core platforms enable the parallel execution of tasks, the sharing of resources may lead to long WCETs that fail to meet the real-time constraints of the system. Then, a safe solution is the execution of the most critical tasks in isolation followed by the execution of the remaining tasks. To improve the system performance, we propose an approach where a critical task can run in parallel with less critical tasks, as long as the real-time constraints are met. When no further interferences can be tolerated, the proposed run-time control suspends the low critical tasks until the termination of the critical task. In this paper, we describe the design and prove the correctness of our approach. To do so, a graph grammar is defined to formally model the critical task as a set of control flow graphs on which a safe partial WCET analysis is applied and used at run-time to control the safe execution of the critical task.
Angeliki Kritikakou, Claire Pagetti, Olivier Baldellon, Matthieu Roy, Christine Rochange
ECRTS1
2014 A scalable and near-optimal representation of access schemes for memory management
abstract
Memory management searches for the resources required to store the concurrently alive elements. The solution quality is affected by the representation of the element accesses: a sub-optimal representation leads to overestimation and a non-scalable representation increases the exploration time. We propose a methodology to near-optimal and scalable represent regular and irregular accesses. The representation consists of a set of pattern entries to compactly describe the behavior of the memory accesses and of pattern operations to consistently combine the pattern entries. The result is a final sequence of pattern entries which represents the global access scheme without unnecessary overestimation.
Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis
ACM Trans. Archit. Code Optim.1
2014 A methodology for speeding up edge and line detection algorithms focusing on memory architecture utilization
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis
J. Supercomput.2
2014 A Matrix-Matrix Multiplication methodology for single/multi-core architectures using SIMD
Vasilios I. Kelefouras, Angeliki Kritikakou, Constantinos E. Goutis
J. Supercomput.2
2013 Near-Optimal Microprocessor and Accelerators Codesign with Latency and Throughput Constraints
abstract
A systematic methodology for near-optimal software/hardware codesign mapping onto an FPGA platform with microprocessor and HW accelerators is proposed. The mapping steps deal with the inter-organization, the foreground memory management, and the datapath mapping. A step is described by parameters and equations combined in a scalable template. Mapping decisions are propagated as design constraints to prune suboptimal options in next steps. Several performance-area Pareto points are produced by instantiating the parameters. To evaluate our methodology we map a real-time bio-imaging application and loop-dominated benchmarks.
Angeliki Kritikakou, Francky Catthoor, George Athanasiou, Vasilios I. Kelefouras, Constantinos E. Goutis
ACM Trans. Archit. Code Optim.1
2013 Near-optimal and scalable intrasignal in-place optimization for non-overlapping and irregular access schemes
abstract
Storage-size management techniques aim to reduce the resources required to store elements and to concurrently provide efficient addressing during element accessing. Existing techniques are less appropriate for large iteration spaces with increased numbers of irregularly spread holes. They either have to approximate the accessed regions, leading to overestimation of the final resources, or they require prohibited exploration time to find the storage size. In this work, we present a near-optimal and scalable methodology for storage-size, intrasignal, in-place optimization, that is, to compute the minimum amount of resources required to store the elements of a group (array), for irregular complex access schemes in the target domain of non-overlapping store and load accesses.
Angeliki Kritikakou, Francky Catthoor, Vasilios I. Kelefouras, Constantinos E. Goutis
ACM Trans. Design Autom. Electr. Syst.1
2012 Priority Handling Aggregation Technique (PHAT) for Wireless Sensor Networks
abstract
Wireless Sensor Networks (WSNs) have limited power capabilities, whereas they serve applications which usually require specific packets, i.e. High Priority Packets (HPP), to be delivered before a deadline. Hence, it is essential to reduce the energy consumption and to have real-time behavior. To achieve this goal we propose a hybrid technique which explores the benefits of data aggregation without data size reduction in combination with prioritized queues. The energy consumption is reduced by appending data from incoming packets with already buffered Low Priority Packets (LPP). The real-time behavior is achieved by directly forwarding the HPP to the next node. Our study explores the impact of the proposed hybrid technique in several all-to-one data flow scenarios with various traffic loads, wait time intervals and percentage of HPP. Our results show gain up to 23,3% in packet loss and 36,6% in energy consumption compared with the direct forwarding of packets.
Dimitris Tsitsipis, Sofia-Maria Dima, Angeliki Kritikakou, Christos Panagiotou, John V. Gialelis, Harris E. Michail, Stavros A. Koubias
ETFA3
2012 A data locality methodology for matrix-matrix multiplication algorithm
Nikolaos Alachiotis 0002, Vasilios I. Kelefouras, George Athanasiou, Harris E. Michail, Angeliki Kritikakou, Constantinos E. Goutis
J. Supercomput.5
2011 Data merge: A data aggregation technique for wireless sensor networks
abstract
The Wireless Sensor Networks (WSNs) have limited power and communication capabilities, combined with the requirement for long network lifetime. To increase it, methods to reduce energy consumption are highly required. To achieve this goal, we study a data aggregation technique without size reduction, i.e. data merge. It is a generic technique, since it is also usable in applications with heterogeneous data and requirements for high accuracy. This study presents the impact of the data merge technique on WSNs applications executed under various realistic data flow scenarios, traffic loads and wait time intervals. Our results show significant reductions in both packet loss and radio energy consumption.
Dimitris Tsitsipis, Sofia-Maria Dima, Angeliki Kritikakou, Christos Panagiotou, Stavros A. Koubias
ETFA3
2010 Ultra High Speed SHA-256 Hashing Cryptographic Module for IPSec Hardware/Software Codesign
Harris E. Michail, George Athanasiou, Angeliki Kritikakou, Constantinos E. Goutis, Andreas Gregoriades, Vicky Papadopoulou Lesta
SECRYPT3