Leonidas Kosmidis

dblp:118/0818 · DBLP profile ↗
← Back
54ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0001-9751-1058ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 47 · 10 first-author · 17 since 2021Software engineering, systems software and programming languages · 21 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Introduction to Post Quantum Cryptography for Resource Constrained Systems
Leonidas Kosmidis, Sergi Alcaide, Matina Maria Trompouki, Antonio Escobar-Molero, Adam Anglart, Daniel Van Geest, Martin Zuber, Jean-Paul Truong, Andrej Grebenc, Detlef Houdeau
ETS1
2025 Hardware support for contention tracking in CPU and GPU last-level cache
abstract
Modern MPSoCs increasingly rely on resource utilization to improve application performance with different computation needs. The last-level cache (LLC) is one of the main shared resources, contributing to the improvement of aggregated performance. However, LLC sharing also increases individual application performance variability, which is undesirable in scenarios where performance guarantees are required. While deploying cache partitioning mechanisms allows regaining predictability, they negatively affect aggregated performance. This confronts system designers with the dire conundrum of choosing between aggregated performance and predictability. We contend that adding hardware support to track contention among tasks (kernels) in the LLC enables it to be shared, removing shortcomings brought by partitioning while providing a clear view of how tasks (kernels) affect each other in the LLC of the CPUs and GPUs. This approach enables achieving the desired balance between performance and predictability. Thus, we propose a low-overhead hardware mechanism, called demotion counters (DC), that tightly estimates the contention tasks (kernels) generate on each other in the shared LLC, outperforming other solutions that build on existing hardware contention-tracking proposals which suffer an average workload breakdown deviation (wbd) over 0.13. Our results also show that DC introduces 0.66% area overhead. 1
Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
J. Syst. Archit.2
2024 Formal Methods for High Integrity GPU Software Development and Verification
abstract
Modern safety critical systems require high levels of performance for advanced functionalities, which are not possible with the simple conventional architectures currently used in them. Embedded General Purpose Graphics Processing Units (GPG-PUs) are among the hardware technologies which can provide the high performance required in these domains. However, their massively parallel nature complicates the verification of their software and increases its cost because it usually involves code coverage through extensive human-driven testing. The Ada SPARK language has traditionally been used in highly-critical environments for its formal verification capabilities and powerful type system. The use of such tools which are backed up by theorem provers, has significantly lowered the amount of effort needed to validate functionality of safety-critical systems. In this European Space Agency (ESA) funded project, we utilize AdaCore's CUDA backend for Ada in conjunction with the SPARK language subset to assess the state of static verification for GPU kernels. We assess the error detection capabilities of the available tools and we formulate a methodology to maximise their effectiveness. Moreover, our project results using ESA's open source GPU4S Benchmarking suite show that common programming mistakes in GPU software development can be prevented.
Dimitris Aspetakis, Leonidas Kosmidis, Matina Maria Trompouki, Jose Ruiz, Gábor Marosy
DATE2
2024 The METASAT Model-Based Engineering Workflow and Digital Twin Concept
abstract
Considering the complexity in the design of future satellites and the need for compliance with ECSS standards, the METASAT project proposes a novel design methodology based on model-based engineering and supported by open architecture hardware. The initiative highlights the potential of software vir-tualisation layers, like hypervisors, to meet standards compliance on high-performance computing platforms. Our focus is on the development of a toolchain tailored for these advanced hard-ware/software layers. In our view, without such innovation, the satellite industry may face high and unsustainable development costs and timelines, which could affect its competitiveness and reliability. This paper provides an overview of the model-based engineering toolchain, the workflow, and the digital twin concept proposed for the METASAT project. We present the purpose of each component and how they work together in the broader METASAT vision, with the objective of showing how this approach can make the development process more efficient and enhance dependability in the sector.
Alejandro J. Calderón, Irune Yarza, Stefano Sinisi, Lorenzo Lazzara, Valerio Di Valerio, Giulia Stazi, Leonidas Kosmidis, Matina Maria Trompouki, Alessandro Ulisse, Aitor Amonarriz, Peio Onaindia
DATE7
2024 The METASAT Modelling and Code Generation Toolchain for XtratuM and Hardware Accelerators
abstract
Given the complex design requirements of future satellites and the need to comply with ECSS standards, the METASAT project introduces an innovative design approach that uses model-based engineering and is supported by open architecture hardware. The project highlights the importance of software virtualisation layers, such as hypervisors, in achieving compliance with standards on high-performance computing plat-forms. The project is dedicated to creating a specialised toolchain for these advanced hardware and software layers. We believe that without such innovations, the satellite industry could face increased and unsustainable development costs and timelines, potentially compromising its competitiveness and dependability. This paper provides a general overview of the METASAT project and introduces the approaches being used to add support for the XtratuM hypervisor and code generation for hardware accelerators in a model-based engineering workflow.
Alejandro J. Calderón, Aitor Amonarriz, Mar Hernández, Leonidas Kosmidis, Jannis Wolf, Marc Solé, Matina Maria Trompouki, Mikel Segura, Peio Onaindia
DSD4
2024 Proton Evaluation of Single Event Effects in the NVIDIA GPU Orin SoM: Understanding Radiation Vulnerabilities Beyond the SoC
abstract
In this paper, we investigate the single event effects under proton irradiation for the state-of-the-art embedded GPU NVIDIA Jetson Orin NX System-on-Module (SoM). Designed for deployment across safety critical domains and in particularly automotive, this system represents a cutting-edge advancement in high performance embedded computing with functional safety features, which makes it an ideal candidate for use in space systems. Our study evaluates the Single-Event Effects (SEE) manifested within the SoM’s central processing unit (CPU), graphics processing unit (GPU), and associated peripherals. Through our analysis, we aim to delineate the origins of Single Event Functional Interrupts (SEFI) occurring at the SoM level. Furthermore, we provide a detailed exposition on the errors observed within the GPU complex, elucidating the requisite conditions for their manifestation. Unlike previous works which treat embedded GPUs under irradiation as black box, we are able to identify the source of SEEs through ARM’s RAS subsystem, and observe for the first time in literature GPU SEEs. Our investigation culminates in a comprehensive assessment of the SoM’s susceptibility, identifying particularly sensitive components.
Ivan Rodriguez-Ferrandez, Leonidas Kosmidis, Maris Tali, David Steenari, Alex Hands, Camille Bélanger-Champagne
IOLTS2
2023 Unraveling the Mystery of NVIDIA's Unified Memory for Safety-Critical GPU Systems
abstract
In the domain of safety-critical systems there is an increasing need for more compute-capable and higher performance devices. This comes from the dramatic increase on the software complexity caused by the newest intelligent and autonomous systems. Graphics Processing Units (GPUs), as multi-processing accelerators, are an ideal choice in this aspect due to their ability to handle big amount of data and computations. In order to ease the challenging task of programming such devices, vendors are continuously adding features, such as Unified Memory (UM), which allow the programmers to reduce their developing time on GPU applications. However, the use of GPU poses several challenges on safety-critical systems due to its close source nature and proprietary implementation. Therefore, this paper shows a deeper insight on how this feature works and present a way of exploiting this knowledge to reduce the execution time of applications using NVIDIA's UM feature. We demonstrate that these optimizations can make the data migrations predictable and reduce the required time.
Xabier Arauzo, Irune Yarza, Leonidas Kosmidis, Alejandro J. Calderón, Marcos Rodriguez
DSD3
2023 Space Shuttle: A Test Vehicle for the Reliability of the SkyWater 130nm PDK for Future Space Processors
abstract
Recently the ASIC industry experiences a massive change with more and more small and medium businesses entering the custom ASIC development. This trend is fueled by the recent open hardware movement and relevant government and privately funded initiatives. These new developments can open new opportunities in the space sector – which is traditionally characterised by very low volumes and very high non-recurrent (NRE) costs – if we can show that the produced chips have favourable radiation properties. In this paper, we describe the design and tape-out of Space Shuttle, the first test chip for the evaluation of the suitability of the SkyWater 130nm PDK and the OpenLane EDA toolchain using the Google/E-fabless shuttle run for future space processors.
Ivan Rodriguez-Ferrandez, Leonidas Kosmidis, Maris Tali, David Steenari
IOLTS2
2023 Vector Extensions in COTS Processors to Increase Guaranteed Performance in Real-Time Systems
abstract
The need for increased application performance in high-integrity systems such as those in avionics is on the rise as software continues to implement more complex functionalities. The prevalent computing solution for future high-integrity embedded products is multi-processor systems-on-chip (MPSoC) processors. MPSoCs include central processing unit (CPU) multicores that enable improving performance via thread-level parallelism. MPSoCs also include generic accelerators (graphics processing units [GPUs]) and application-specific accelerators. However, the data processing approach (DPA) required to exploit each of these underlying parallel hardware blocks carries several open challenges to enable the safe deployment in high-integrity domains. The main challenges include the qualification of its associated runtime system and the difficulties in analyzing programs deploying the DPA with out-of-the-box timing analysis and code coverage tools. In this work, we perform a thorough analysis of vector extensions (VExts) in current commercial off-the-shelf (COTS) processors for high-integrity systems. We show that VExts prevent many of the challenges arising with parallel programming models and GPUs. Unlike other DPAs, VExts require no runtime support, prevent design race conditions that might arise with parallel programming models, and have minimum impact on the software ecosystem, enabling the use of existing code coverage and timing analysis tools. We develop vectorized versions of neural network kernels and show that the NVIDIA Xavier VExts provide a reasonable increase in guaranteed application performance of up to 2.7x. Our analysis contends that VExts are the DPA approach with arguably the fastest path for adoption in high-integrity systems.
Roger Pujol, Josep Jorba 0002, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
ACM Trans. Embed. Comput. Syst.4
2022 SPARROW: A Low-Cost Hardware/Software Co-designed SIMD Microarchitecture for AI Operations in Space Processors
abstract
Recently there is an increasing interest in the use of artificial intelligence for on-board processing as indicated by the latest space missions, which cannot be satisfied by existing low-performance space-qualified processors. Although COTS AI accelerators can provide the required performance, they are not designed to meet space requirements. In this work, we co-design a low-cost SIMD micro-architecture integrated in a space qualified processor, which can significantly increase its performance. Our solution has no impact on the processor's 100 MHz frequency and consumes minimal area thanks to its innovative design compared to conventional vector micro-architectures. For the minimum configuration of our baseline space processor, our results indicate a performance boost of up to 9.3× for commonly used AI-related and image processing algorithms and 5.5× faster for a complex, space-relevant inference application with just 30% area increase.
Marc Solé, Leonidas Kosmidis
DATE2
2022 Contention Tracking in GPU Last-Level Cache
abstract
The Last-level cache (LLC) is one of the main GPU’s shared resources that contributes to improve performance but also increases individual kernel’s performance variability. This is detrimental in scenarios in which some level of performance predictability is required. While predictability can be regained by deploying cache partitioning (isolation) mechanisms, isolation negatively affects performance efficiency. This work shows that not partitioning the LLC and providing the ability to track the contention that kernels generate on each other allows them to share LLC space, hence increasing efficiency, while the system designer obtains a clear view of how each kernel affects each other in the LLC so as to balance performance and predictability goals. In this line, we propose GPU demotion counters (GDC), a low-overhead hardware mechanism to track contention that kernels generate on each other in the shared LLC.
Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla
ICCD2
2022 Functional and Timing Implications of Transient Faults in Critical Systems
abstract
Embedded systems in critical domains, such as auto-motive, aviation, space domains, are often required to guarantee both functional and temporal correctness. Considering transient faults, fault analysis and mitigation approaches are implemented at various levels of the system design, in order to maintain the functional correctness. However, transient faults and their mitigation methods have a timing impact, which can affect the temporal correctness of the system. In this work, we expose the functional and the timing implications of transient faults for critical systems. More precisely, we initially highlight the timing effect of transient faults occurring in the combinational and sequential logic of a processor. Furthermore, we propose a full stack vulnerability analysis that drives the design of selective hardware-based mitigation for real-time applications. Last, we study the timing impact of software-based reliability mitigation methods applied in a COTS GPU, using a fault tolerant middleware.
Angeliki Kritikakou, Panagiota Nikolaou, Ivan Rodriguez-Ferrandez, Joseph Paturel, Leonidas Kosmidis, Maria K. Michael, Olivier Sentieys, David Steenari
IOLTS5
2022 Sources of Single Event Effects in the NVIDIA Xavier SoC Family under Proton Irradiation
abstract
In this paper we characterise two embedded GPU devices from the NVIDIA Xavier family System-on-Chip (SoC) using a proton beam. We compare the NVIDIA Xavier NX and Industrial devices, that respectively target commercial and automotive applications. We evaluate the Single-Event Effect (SEE) rate of both modules and their sub-components, both the CPU and GPU, using different power modes, and we try for the first time to identify their exact sources using the on-line testing facilities included in their ARM based system. Our conclusion is that the most sensitive part of the CPU complex of the SoC is the tag array of the various cache structures, while no errors were observed in the GPU, probably because of its fast execution compared to the CPU part of the application during the radiation campaign.
Ivan Rodriguez-Ferrandez, Maris Tali, Leonidas Kosmidis, Marta Rovituso, David Steenari
IOLTS3
2021 Comparison of GPU Computing Methodologies for Safety-Critical Systems: An Avionics Case Study
abstract
Introducing advanced functionalities in safety-critical systems requires using more powerful architectures such as GPUs. However software in safety-critical industries is subject to functional certification, which cannot be achieved using standard GPU programming languages such as CUDA and OpenCL. Fortunately, GPUs are already used in certified critical systems for display tasks, using safety-certified solutions such as OpenGL SC 2.0. In this paper, we compare two state-of-the art graphics-based methodologies, OpenGL SC 2.0 and Brook Auto/BRASIL for the implementation of a prototype avionics case study. We evaluate both methods on a realistic industrial setup, composed by an avionics-grade GPU and a safety-certified GPU driver in terms of development metrics and performance, showing their feasibility.
Marc Benito, Matina Maria Trompouki, Leonidas Kosmidis, Juan David Garcia, Sergio Carretero, Ken Wenger
DATE3
2021 The UP2DATE Baseline Research Platforms
abstract
The UP2DATE H2020 project focuses on highperformance heterogeneous embedded platforms for critical systems. We will develop observability and controllability solutions to support online updates while ensuring safety and security for mixed-criticality tasks. In this paper, we describe the rationale behind the selection of the baseline research platforms which will be used to develop and demonstrate the project concepts, including a performance comparison to identify the most efficient one.
Alvaro Jover-Alvarez, Alejandro J. Calderón, Iván Rodriguez, Leonidas Kosmidis, Kazi Asifuzzaman, Patrick Uven, Kim Grüttner, Tomaso Poggi, Irune Agirre
DATE4
2021 GPU4S: Major Project Outcomes, Lessons Learnt and Way Forward
abstract
Embedded GPUs have been identified from both private and government space agencies as promising hardware technologies to satisfy the increased needs of payload processing. The GPU4S (GPU for Space) project funded from the European Space Agency (ESA) has explored in detail the feasibility and the benefit of using them for space workloads. Currently at the closing phases of the project, in this paper we describe the main project outcomes and explain the lessons we learnt. In addition, we provide some guidelines for the next steps towards their adoption in space.
Leonidas Kosmidis, Iván Rodriguez, Alvaro Jover-Alvarez, Sergi Alcaide, Jérôme Lachaize, Olivier Notebaert, Antoine Certain, David Steenari
DATE1
2021 Security, Reliability and Test Aspects of the RISC-V Ecosystem
abstract
RISC-V has emerged as a viable solution on academia and industry. However, to use open source hardware for safety-critical applications, we need a deep understanding of the way in which well established mechanisms for testing and reliability could be integrated and deployed on the RISC-V ecosystem, and we need a clear knowledge on how such an ecosystem can be leveraged to improve security. This paper includes four contributions presenting the potential of RISC-V in security research, the way in which RISC-V can be hardened against power analysis attacks, how to implement, using RISC-V, software and hardware/software solutions for dual core lock step, and how to perform system-level testing in the RISC-V ecosystem.
Jaume Abella 0001, Sergi Alcaide, Jens Anders, Francisco Bas, Steffen Becker 0001, Elke De Mulder, Nourhan Elhamawy, Frank K. Gürkaynak, Helena Handschuh, Carles Hernández 0001, Michael Hutter, Leonidas Kosmidis, Ilia Polian, Matthias Sauer 0002, Stefan Wagner 0001, Francesco Regazzoni 0001
ETS12
2020 An On-board Algorithm Implementation on an Embedded GPU: A Space Case Study
abstract
On-board processing requirements of future space missions are constantly increasing, calling for new hardware than the traditional ones used in space. Embedded GPUs are an attractive candidate offering both high performance capabilities and low power consumption, but there are no complex industrial case studies from the space domain demonstrating these advantages. In this paper we present the GPU parallelisation of an on-board algorithm, as well as its performance on a promising embedded GPU COTS platform targeting critical systems.
Iván Rodriguez, Leonidas Kosmidis, Olivier Notebaert, Francisco J. Cazorla, David Steenari
DATE2
2020 UP2DATE: Safe and secure over-the-air software updates on high-performance mixed-criticality systems
abstract
Following the same trend of consumer electronics, safety-critical industries are starting to adopt Over-The-Air Software Updates (OTASU) on their embedded systems. The motivation behind this trend is twofold. On the one hand, OTASU offer several benefits to the product makers and users by improving or adding new functionality and services to the product without a complete redesign. On the other hand, the increasing connectivity trend makes OTASU a crucial cyber-security demand to download latest security patches. However, the application of OTASU in the safety-critical domain is not free of challenges, specially when considering the dramatic increase of software complexity and the resulting high computing performance demands. This is the mission of UP2DATE, a recently launched project funded within the European H2020 programme focused on new software update architectures for heterogeneous high-performance mixed-criticality systems. This paper gives an overview of UP2DATE and its foundations, which seeks to improve existing OTASU solutions by considering safety, security and availability from the ground up in an architecture that builds around composability and modularity.
Irune Agirre, Peio Onaindia, Tomaso Poggi, Irune Yarza, Francisco J. Cazorla, Leonidas Kosmidis, Kim Grüttner, Mohammed Abuteir, Jan Loewe, Juan M. Orbegozo, Stefania Botta
DSD6
2020 Timing of Autonomous Driving Software: Problem Analysis and Prospects for Future Solutions
abstract
The software used to implement advanced functionalities in critical domains (e.g. autonomous operation) impairs software timing. This is not only due to the complexity of the underlying high-performance hardware deployed to provide the required levels of computing performance, but also due to the complexity, non-deterministic nature, and huge input space of the artificial intelligence (AI) algorithms used. In this paper, we focus on Apollo, an industrial-quality Autonomous Driving (AD) software framework: we statistically characterize its observed execution time variability and reason on the sources behind it. We discuss the main challenges and limitations in finding a satisfactory software timing analysis solution for Apollo and also show the main traits for the acceptability of statistical timing analysis techniques as a feasible path. While providing a consolidated solution for the software timing analysis of Apollo is a huge effort far beyond the scope of a single research paper, our work aims to set the basis for future and more elaborated techniques for the timing analysis of AD software.
Miguel Alcon, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
RTAS3
2020 GMAI: Understanding and Exploiting the Internals of GPU Resource Allocation in Critical Systems
abstract
Critical real-time systems require strict resource provisioning in terms of memory and timing. The constant need for higher performance in these systems has led industry to recently include GPUs. However, GPU software ecosystems are by their nature closed source, forcing system engineers to consider them as black boxes, complicating resource provisioning. In this work, we reverse engineer the internal operations of the GPU system software to increase the understanding of their observed behaviour and how resources are internally managed. We present our methodology that is incorporated in GMAI (GPU Memory Allocation Inspector), a tool that allows system engineers to accurately determine the exact amount of resources required by their critical systems, avoiding underprovisioning. We first apply our methodology on a wide range of GPU hardware from different vendors showing its generality in obtaining the properties of the GPU memory allocators. Next, we demonstrate the benefits of such knowledge in resource provisioning of two case studies from the automotive domain, where the actual memory consumption is up to 5.6× more than the memory requested by the application.
Alejandro J. Calderón, Leonidas Kosmidis, Carlos F. Nicolás, Francisco J. Cazorla, Peio Onaindia
ACM Trans. Embed. Comput. Syst.2
2019 Assessing the Adherence of an Industrial Autonomous Driving Framework to ISO 26262 Software Guidelines
abstract
The complexity and size of Autonomous Driving (AD) software are comparably higher than that of software implementing other (standard) functionalities in the car. To make things worse, a big fraction of AD software is not specifically designed for the automotive (or any other critical) domain, but the mainstream market. This brings uncertainty on to which extent AD software adheres to guidelines in safety standards. In this paper, we present our experience in applying ISO 26262 -- the applicable functional safety standard for road vehicles -- software safety guidelines to industrial AD software, in particular, Apollo, a heterogeneous Autonomous Driving framework used extensively in industry. We provide quantitative and qualitative metrics of compliance for many ISO 26262 recommendations on software design, implementation, and testing.
Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla, Guillem Bernat
DAC2
2019 High-Integrity GPU Designs for Critical Real-Time Automotive Systems
abstract
Autonomous Driving (AD) imposes the use of high-performance hardware, such as GPUs, to perform object recognition and tracking in real-time. However, differently to the consumer electronics market, critical real-time AD functionalities require a high degree of resilience against faults, in line with the automotive ISO26262 functional safety standard requirements. ISO26262 imposes the use of some source of independent redundancy for the most critical functionalities so that a single fault cannot lead to a failure, being dual core lockstep (DCLS) with diversity the preferred choice for computing devices. Unfortunately, GPUs do not support diverse DCLS by construction, thus failing to meet ISO26262 requirements efficiently.In this paper we propose lightweight modifications to GPUs to enable diverse DCLS for critical real-time applications without diminishing their performance for non-critical applications. In particular, we show how enabling specific mechanisms for software-controlled kernel scheduling in the GPU, allows guaranteeing that redundant kernels can be executed in different resources so that a single fault cannot lead to a failure, as imposed by ISO26262. Our results on a GPU simulator and an NVIDIA GPU prove the viability of the approach and its effectiveness on high-performance GPU designs needed for AD systems.
Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001
DATE2
2019 GPU4S: Embedded GPUs in Space
abstract
Following the same trend of automotive and avionics, the space domain is witnessing an increase in the on-board computing performance demands. This raise in performance needs comes from both control and payload parts of the spacecraft and calls for advanced electronics able to provide high computational power under the constraints of the harsh space environment. On the non-technical side, for strategic reasons it is mandatory to get European independence on the used computing technology. In this project, which is still in its early phases, we study the applicability of embedded GPUs in space, which have shown a dramatic improvement of their performance per-watt ratio coming from their proliferation in consumer markets based on competitive European technology. To that end, we perform an analysis of the existing space application domains to identify which software domains can benefit from their use. Moreover, we survey the embedded GPU domain in order to assess whether embedded GPUs can provide the required computational power and identify the challenges which need to be addressed for their adoption in space. In this paper, we describe the steps to be followed in the project, as well as the results of our preliminary analyses in the first months of the project.
Leonidas Kosmidis, Jérôme Lachaize, Jaume Abella 0001, Olivier Notebaert, Francisco J. Cazorla, David Steenari
DSD1
2019 Generating and Exploiting Deep Learning Variants to Increase Heterogeneous Resource Utilization in the NVIDIA Xavier
abstract
Deep learning-based solutions and, in particular, deep neural networks (DNNs) are at the heart of several functionalities in critical-real time embedded systems (CRTES) from vision-based perception (object detection and tracking) systems to trajectory planning. As a result, several DNN instances simultaneously run at any time on the same computing platform. However, while modern GPUs offer a variety of computing elements (e.g. CPUs, GPUs, and specific accelerators) in which those DNN tasks can be executed depending on their computational requirements and temporal constraints, current DNNs are mainly programmed to exploit one of them, namely, regular cores in the GPU. This creates resource imbalance and under-utilization of GPU resources when executing several DNN instances, causing an increase in DNN tasks' execution time requirements. In this paper, (a) we develop different variants (implementations) of well-known DNN libraries used in the Apollo Autonomous Driving (AD) software for each of the computing elements of the latest NVIDIA Xavier SoC. Each variant can be configured to balance resource requirements and performance: the regular CPU core implementation that can run on 2, 4, and 6 cores; the GPU regular and Tensor core variants that can run in 4 or 8 GPU’s Streaming Multiprocessors (SM); and 1 or 2 NVIDIA’s Deep Learning Accelerators (NVDLA); (b) we show that each particular variant/configuration offers a different resource utilization/performance point; finally, (c) we show how those heterogeneous computing elements can be exploited by a static scheduler to sustain the execution of multiple and diverse DNN variants on the same platform.
Roger Pujol, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
ECRTS3
2019 Understanding and Exploiting the Internals of GPU Resource Allocation for Critical Systems
abstract
Critical real-time systems require strict resource provisioning in terms of memory and timing. The constant need for higher performance in these systems has led industry to recently include GPUs. However, GPU software ecosystems are by their nature closed source, forcing system engineers to consider them as black boxes, complicating resource provisioning. In this work we reverse engineer the internal operations of the GPU system software to increase the understanding of their observed behaviour and how resources are internally managed. This way, we allow system engineers to accurately determine the exact amount of resources required by their critical systems, avoiding underprovisioning. We first apply our methodology on a wide range of GPU hardware showing its generality in obtaining the properties of the GPU memory allocators. Next, we demonstrate the benefits of such knowledge in resource provisioning of two case studies from the automotive domain, where the actual memory consumption is up to 5.6 × more than the memory requested by the application.
Alejandro J. Calderón, Leonidas Kosmidis, Carlos F. Nicolás, Francisco J. Cazorla, Peio Onaindia
ICCAD2
2019 BRASIL: A High-Integrity GPGPU Toolchain for Automotive Systems
abstract
Embedded General Purpose Graphics Processing Units (GPGPUs) are increasingly used in automotive to enable Advanced Driving Assistance (ADAS) and Autonomous driving. However, their functional safety certification has been addressed so far only at the application or hardware level. In this work, we demonstrate how the entire GPGPU software stack can be developed and certified up to ASIL-D, including its toolchain. We introduce BRASIL, a GPGPU toolchain targeting high-integrity GPGPU computing and we explain our methodology to achieve ISO 26262 compliance, as well as our tool qualification strategy.
Matina Maria Trompouki, Leonidas Kosmidis
ICCD2
2019 Software-only Diverse Redundancy on GPUs for Autonomous Driving Platforms
abstract
Autonomous driving (AD) builds upon high-performance computing platforms including (1) general purpose CPUs as well as (2) specific accelerators, being GPUs one of the main representatives. Microcontrollers have reached ASIL-D compliance by implementing diverse redundancy with lockstep execution. However, ASIL-D compliant GPUs rely on either fully redundant lockstep GPUs (i.e. 2 GPUs), which doubles hardware costs, or fully redundant systems with a GPU and another accelerator, which virtually doubles design and validation/verification (V&V) costs. In this paper we analyze the degree of diversity achieved when implementing redundancy on a single GPU, showing that diverse redundancy is not achieved in many cases, and propose software strategies that guarantee achieving diverse redundancy for any kernel on systems using commercial off-the-shelf (COTS) GPUs, thus showing how to achieve ASIL-D compliance on a single COTS GPU in controlled scenarios.
Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001
IOLTS2
2019 Performance Analysis and Optimization of Automotive GPUs
abstract
Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) have drastically increased the performance demands of automotive systems. Suitable high-performance platforms building upon Graphic Processing Units (GPUs) have been developed to respond to this demand, being NVIDIA Jetson TX2 a relevant representative. However, whether high-performance GPU configurations are appropriate for automotive setups remains as an open question. This paper aims at providing light on this question by modelling an automotive GPU (Jetson TX2), analyzing its microarchitectural parameters against relevant benchmarks, and identifying specific configurations able to meaningfully increase performance within similar cost envelopes, or to decrease costs preserving original performance levels. Overall, our analysis opens the door to the optimization of automotive GPUs for further system efficiency.
Fabio Mazzocchetti, Pedro Benedicte, Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla
SBAC-PAD4
2018 Modelling multicore contention on the AURIXTM TC27x
abstract
Multicores are becoming ubiquitous in automotive. Yet, the expected benefits on integration are challenged by multicore contention concerns on timing V&V. Worst-case execution time (WCET) estimates are required as early as possible in the software development, to enable prompt detection of timing misbehavior. Factoring in multicore contention necessarily builds on conservative assumptions on interference, independent of co-runners load on shared hardware. We propose a contention model for automotive multicores that balances time-composability with tightness by exploiting available information on contenders. We tailor the model to the AURIX TC27x and provide tight WCET estimates using information from performance monitors and software configurations.
Enrique Díaz, Enrico Mezzetti, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla
DAC3
2018 Brook auto: high-level certification-friendly programming for GPU-powered automotive systems
abstract
Modern automotive systems require increased performance to implement Advanced Driving Assistance Systems (ADAS). GPU-powered platforms are promising candidates for such computational tasks, however current low-level programming models challenge the accelerator software certification process, while they limit the hardware selection to a fraction of the available platforms. In this paper we present Brook Auto, a high-level programming language for automotive GPU systems which removes these limitations. We describe the challenges and solutions we faced in its implementation, as well as a complete evaluation in terms of performance and productivity, which shows the effectiveness of our method.
Matina Maria Trompouki, Leonidas Kosmidis
DAC2
2018 Industrial experiences with resource management under software randomization in ARINC653 avionics environments
abstract
Injecting randomization in different layers of the computing platform has been shown beneficial for security, resilience to software bugs and timing analysis. In this paper, with focus on the latter, we show our experience regarding memory and timing resource management when software randomization techniques are applied to one of the most stringent industrial environments, ARINC653-based avionics. We describe the challenges in this task, we propose a set of solutions and present the results obtained for two commercial avionics applications, executed on COTS hardware and RTOS.
Leonidas Kosmidis, Cristian Maxim, Victor Jégu, Francis Vatrinet, Francisco J. Cazorla
ICCAD1
2017 Dynamic software randomisation: Lessons learnec from an aerospace case study
abstract
Timing Validation and Verification (V&V) is an important step in real-time system design, in which a system's timing behaviour is assessed via Worst Case Execution Time (WCET) estimation and scheduling analysis. For WCET estimation, measurement-based timing analysis (MBTA) techniques are widely-used and well-established in industrial environments. However, the advent of complex processors makes it more difficult for the user to provide evidence that the software is tested under stress conditions representative of those at system operation. Measurement-Based Probabilistic Timing Analysis (MBPTA) is a variant of MBTA followed by the PROXIMA European Project that facilitates formulating this representativeness argument. MBPTA requires certain properties to be applicable, which can be obtained by selectively injecting randomisation in platform's timing behaviour via hardware or software means. In this paper, we assess the effectiveness of the PROXIMA's dynamic software randomisation (DSR) with a space industrial case study executed on a real unmodified hardware platform and an industrial operating system. We present the challenges faced in its development, in order to achieve MBPTA compliance and the lessons learned from this process. Our results, obtained using a commercial timing analysis tool, indicate that DSR does not impact the average performance of the application, while it enables the use of MBPTA. This results in tighter pWCET estimates compared to current industrial practice.
Fabrice Cros, Leonidas Kosmidis, Franck Wartel, David Morales, Jaume Abella 0001, Ian Broster, Francisco J. Cazorla
DATE2
2017 Probabilistic timing analysis on time-randomized platforms for the space domain
abstract
Timing Verification is a fundamental step in real-time embedded systems, with measurement-based timing analysis (MBTA) being the most common approach used to that end. We present a Space case study on a real platform that has been modified to support a probabilistic variant of MBTA called MBPTA. Our platform provides the properties required by MBPTA with the predicted WCET estimates with MBPTA being competitive to those with current MBTA practice while providing more solid evidence on their correctness for certification.
Mikel Fernández, David Morales, Leonidas Kosmidis, Alen Bardizbanyan, Ian Broster, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla, Paulo Machado, Luca Fossati
DATE3
2017 Optimisation opportunities and evaluation for GPGPU applications on low-end mobile GPUs
abstract
Previous works in the literature have shown the feasibility of general purpose computations for non-visual applications on low-end mobile graphics processors using graphics APIs. These works focused only on the functional aspects of the software, ignoring the implementation details and therefore their performance implications due to their particular micro-architecture. Since various steps in such applications can be implemented in multiple ways, we identify optimisation opportunities, explore the different options and evaluate them. We show that the implementation details can significantly affect the obtained performance with discrepancies up to 3 orders of magnitude and we demonstrate the effectiveness of our proposal on two embedded platforms, obtaining more than 16× speedup over benchmarks designed following OpenGL ES 2 best practices.
Matina Maria Trompouki, Leonidas Kosmidis
DATE2
2017 An open benchmark implementation for multi-CPU multi-GPU pedestrian detection in automotive systems
abstract
Modern and future automotive systems incorporate several Advanced Driving Assistance Systems (ADAS). Those systems require significant performance that cannot be provided with traditional automotive processors and programming models. Multicore CPUs and Nvidia GPUs using CUDA are currently considered by both automotive industry and research community to provide the necessary computational power. However, despite several recent published works in this domain, there is an absolute lack of open implementations of GPU-based ADAS software, that can be used for benchmarking candidate platforms. In this work, we present a multi-CPU and GPU implementation of an open implementation of a pedestrian detection benchmark based on the Viola-Jones image recognition algorithm. We present our optimization strategies and evaluate our implementation on a multiprocessor system featuring multiple GPUs, showing an overall 88.5× speedup over the sequential version.
Matina Maria Trompouki, Leonidas Kosmidis, Nacho Navarro
ICCAD2
2016 Towards general purpose computations on low-end mobile GPUs
Matina Maria Trompouki, Leonidas Kosmidis
DATE2
2016 PROXIMA: Improving Measurement-Based Timing Analysis through Randomisation and Probabilistic Analysis
abstract
The use of increasingly complex hardware and software platforms in response to the ever rising performance demands of modern real-time systems complicates the verification and validation of their timing behaviour, which form a time-and-effort-intensive step of system qualification or certification. In this paper we relate the current state of practice in measurement-based timing analysis, the predominant choice for industrial developers, to the proceedings of the PROXIMA (Probabilistic real-time control of mixed-criticality multicore systems) project in that very field. We recall the difficulties that the shift towards more complex computing platforms causes in that regard. Then we discuss the probabilistic approach proposed by PROXIMA to overcome some of those limitations. We present the main principles behind the PROXIMA approach as well as the changes it requires at hardware or software level underneath the application. We also present the current status of the project against its overall goals, and highlight some of the principal confidence-building results achieved so far.
Francisco J. Cazorla, Jaume Abella 0001, Jan Andersson, Tullio Vardanega, Francis Vatrinet, Iain Bate, Ian Broster, Mikel Azkarate-askatsua, Franck Wartel, Liliana Cucu-Grosjean, Fabrice Cros, Glenn Farrall, Adriana Gogonel, Andrea Gianarro, Benoit Triquet, Carles Hernández 0001, Code Lo, Cristian Maxim, David Morales, Eduardo Quiñones, Enrico Mezzetti, Leonidas Kosmidis, Irune Agirre, Mikel Fernández, Mladen Slijepcevic, Philippa Conmy, Walid Talaboulma
DSD22
2016 TASA: toolchain-agnostic static software randomisation for critical real-time systems
abstract
Measurement-Based Probabilistic Timing Analysis (MBPTA) derives WCET estimates for tasks running on processors comprising high-performance features such as caches. MBPTA's correct application requires the system to exhibit certain timing properties, which can be achieved by injecting randomisation in the timing behaviour of the task under analysis. However, existing software-randomisation techniques require costly modifications in the industrial production toolchain (compiler, linker, runtime or hardware) in terms of development and certification. In this paper we present TASA, a new software randomisation tool that relies on source-code transformations of the application (i) requiring no changes in existing toolchains, which heavily reduces tool qualification and implementation costs; and (ii) achieving competitive WCET estimates that we assess on a gcc- and a llvm-based compilation toolchain on a real board.
Leonidas Kosmidis, Roberto Vargas, David Morales, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla
ICCAD1
2016 A confidence assessment of WCET estimates for software time randomized caches
abstract
Obtaining Worst-Case Execution Time (WCET) estimates is a required step in real-time embedded systems during software verification. Measurement-Based Probabilistic Timing Analysis (MBPTA) aims at obtaining WCET estimates for industrial-size software running upon hardware platforms comprising high-performance features. MBPTA relies on the randomization of timing behavior (functional behavior is left unchanged) of hard-to-predict events like the location of objects in memory - and hence their associated cache behavior - that significantly impact software's WCET estimates. Software time-randomized caches (sTRc) have been recently proposed to enable MBPTA on top of Commercial off-the-shelf (COTS) caches (e.g. modulo placement). However, some random events may challenge MBPTA reliability on top of sTRc. In this paper, for sTRc and programs with homogeneously accessed addresses, we determine whether the number of observations taken at analysis, as part of the normal MBPTA application process, captures the cache events significantly impacting execution time and WCET. If this is not the case, our techniques provide the user with the number of extra runs to perform to guarantee that cache events are captured for a reliable application of MBPTA. Our techniques are evaluated with synthetic benchmarks and an avionics application.
Pedro Benedicte, Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla
INDIN2
2015 Timing analysis of an avionics case study on complex hardware/software platforms
Franck Wartel, Leonidas Kosmidis, Adriana Gogonel, Andrea Baldovin, Zoë Stephenson, Benoit Triquet, Eduardo Quiñones, Code Lo, Enrico Mezzetti, Ian Broster, Jaume Abella 0001, Liliana Cucu-Grosjean, Tullio Vardanega, Francisco J. Cazorla
DATE2
2014 Containing Timing-Related Certification Cost in Automotive Systems Deploying Complex Hardware
abstract
Measurement-Based Probabilistic Timing Analysis (MBPTA) techniques simplify deriving tight and trustworthy WCET estimates for industrial-size programs running on complex processors. MBPTA poses some requirements on the timing behaviour of the hardware/software platform: execution times of end-to-end runs have to be independent and identically distributed (i.i.d.). Hardware and software solutions have been deployed to accomplish MBPTA requirements. The latter has achieved the i.i.d. properties running on some commercial off-the-shelf (COTS) processor designs. Unfortunately, software randomisation challenges functional verification needed for certification since it introduces indirections through pointers in the code. In this paper we propose a new approach to software randomisation able to contain its functional verification costs. Our approach performs software randomisation statically, as opposed to current dynamic approaches. We carefully review the requirements of the new approach and prove its feasibility.
Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Glenn Farrall, Franck Wartel, Francisco J. Cazorla
DAC1
2014 Time-Analysable Non-Partitioned Shared Caches for Real-Time Multicore Systems
abstract
Shared caches in multicores challenge Worst-Case Execution Time (WCET) estimation due to inter-task interferences. Hardware and software cache partitioning address this issue although they complicate data sharing among tasks and the Operating System (OS) task scheduling and migration. In the context of Probabilistic Timing Analysis (PTA) time-randomised caches are used. We propose a new hardware mechanism to control inter-task interferences in shared time-randomised caches without the need of any hardware or software partitioning. Our proposed mechanism effectively bounds inter-task interferences by limiting the cache eviction frequency of each task, while providing tighter WCET estimates than cache partitioning algorithms. In a 4-core multicore processor setup our proposal improves cache partitioning by 56% in terms of guaranteed performance and 16% in terms of average performance.
Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
DAC2
2014 Bus designs for time-probabilistic multicore processors
abstract
Probabilistic Timing Analysis (PTA) reduces the amount of information needed to provide tight WCET estimates in real-time systems with respect to classic timing analysis. PTA imposes new requirements on hardware design that have been shown implementable for single-core architectures. However, no support has been proposed for multicores so far. In this paper, we propose several probabilistically-analysable bus designs for multicore processors ranging from 4 cores connected with a single bus, to 16 cores deploying a hierarchical bus design. We derive analytical models of the probabilistic timing behaviour for the different bus designs, show their suitability for PTA and evaluate their hardware cost. Our results show that the proposed bus designs (i) fulfil PTA requirements, (ii) allow deriving WCET estimates with the same cost and complexity as in single-core processors, and (iii) provide higher guaranteed performance than single-core processors, 3.4x and 6.6x on average for an 8-core and a 16-core setup respectively.
Javier Jalle, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
DATE2
2014 Measurement-Based Probabilistic Timing Analysis and Its Impact on Processor Architecture
abstract
Critical Real-Time Embedded Systems (CRTES) industry needs increasingly complex hardware to attain the performance/cost ratio required to keep competitive edge in the market. Worst-case execution time (WCET) analysis is central to CRTES development. Whereas current timing analysis techniques are sound, their viability is hampered by the soaring cost of acquiring detailed knowledge of the internal operation and state of the system, at both software and hardware level. This is a major hurdle to using them for increasingly complex hardware platforms. Measurement-Based Probabilistic Timing Analysis (PTA) reduces the cost of acquiring the knowledge needed for computing trustworthy WCET bounds. This paper presents the changes required to hardware design to facilitate the use of the PTA techniques.
Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Ian Broster, Francisco J. Cazorla
DSD1
2014 PUB: Path Upper-Bounding for Measurement-Based Probabilistic Timing Analysis
abstract
Measurement-Based Probabilistic Timing Analysis (MBPTA) responds to the challenge of analysing the timing behaviour of real-time software running on hardware deploying high-performance features (e.g., data caches). MBPTA provides a WCET estimate that upper-bounds the execution time of the set of paths exercised with the data input vectors provided by the user. However, in several scenarios, the user is unaware of the input vector leading to the worst-case path. In this paper we present PUB, a new method that makes the WCET estimates obtained with MBPTA a trustworthy upper-bound of the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. This significantly reduces the requirements imposed on the user to apply MBPTA. For Malardarlen and EEMBC respectively, PUB provides WCET estimates 5% and 11% higher than the WCET estimates computed with MBPTA.
Leonidas Kosmidis, Jaume Abella 0001, Franck Wartel, Eduardo Quiñones, Antoine Colin, Francisco J. Cazorla
ECRTS1
2014 Efficient Cache Designs for Probabilistically Analysable Real-Time Systems
abstract
The increasing performance demand in the critical real-time embedded systems (CRTES) domain calls for high-performance features such as cache memories. Unfortunately, the cost to provide trustworthy and tight Worst-Case Execution Time (WCET) estimates in the presence of caches is high with current practice WCET analysis tools, because they need detailed knowledge of program’s cache accesses to provide tight WCET estimates. The advent of Probabilistic timing analysis (PTA) opens the door to economically viable timing analysis in the presence of caches, but it imposes new requirements on hardware design. At cache level, so far only fully associative random-replacement caches have been proven to fulfill the needs of PTA, but their energy, delay, and area cost are unaffordable for CRTES. In this paper, we propose the first PTA-compliant cache design based on set-associative and direct-mapped arrangements, as those are the most common arrangements. In particular, we propose a novel parametric random placement policy suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement.
Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
IEEE Trans. Computers1
2013 A cache design for probabilistically analysable real-time systems
abstract
Caches provide significant performance improvements, though their use in real-time industry is low because current WCET analysis tools require detailed knowledge of program's cache accesses to provide tight WCET estimates. Probabilistic Timing Analysis (PTA) has emerged as a solution to reduce the amount of information needed to provide tight WCET estimates, although it imposes new requirements on hardware design. At cache level, so far only fully-associative random-replacement caches have been proven to fulfill the needs of PTA, but they are expensive in size and energy. In this paper we propose a cache design that allows set-associative and direct-mapped caches to be analysed with PTA techniques. In particular we propose a novel parametric random placement suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement.
Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
DATE1
2013 Probabilistic timing analysis on conventional cache designs
abstract
Probabilistic timing analysis (PTA), a promising alternative to traditional worst-case execution time (WCET) analyses, enables pairing time bounds (named probabilistic WCET or pWCET) with an exceedance probability (e.g., 10−16), resulting in far tighter bounds than conventional analyses. However, the applicability of PTA has been limited because of its dependence on relatively exotic hardware: fully-associative caches using random replacement. This paper extends the applicability of PTA to conventional cache designs via a software-only approach. We show that, by using a combination of compiler techniques and runtime system support to randomise the memory layout of both code and data, conventional caches behave as fully-associative ones with random replacement.
Leonidas Kosmidis, Charlie Curtsinger, Eduardo Quiñones, Jaume Abella 0001, Emery D. Berger, Francisco J. Cazorla
DATE1
2013 DTM: Degraded Test Mode for Fault-Aware Probabilistic Timing Analysis
abstract
Existing timing analysis techniques to derive Worst-Case Execution Time (WCET) estimates assume that hardware in the target platform (e.g., the CPU) is fault-free. Given the performance requirements increase in current Critical Real-Time Embedded Systems (CRTES), the use of high-performance features and smaller transistors in current and future hardware becomes a must. The use of smaller transistors helps providing more performance while maintaining low energy budgets, however, hardware fault rates increase noticeably, affecting the temporal behaviour of the system in general, and WCET in particular. In this paper, we reconcile these two emergent needs of CRTES, namely, tight (and trustworthy) WCET estimates and the use of hardware implemented with smaller transistors. To that end we propose the Degraded Test Mode (DTM) that, in combination with fault-tolerant hardware designs and probabilistic timing analysis techniques, (i) enables the computation of tight and trustworthy WCET estimates in the presence of faults, (ii) provides graceful average and worst-case performance degradation due to faults, and (iii) requires modifications neither in WCET analysis tools nor in applications. Our results show that DTM allows accounting for the effect of faults at analysis time with low impact in WCET estimates and negligible hardware modifications.
Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
ECRTS2
2013 Achieving timing composability with measurement-based probabilistic timing analysis
abstract
Probabilistic Timing Analysis (PTA) allows complex hardware acceleration features, which defeat classic timing analysis, to be used in hard real-time systems. PTA can do that because it drastically reduces intrinsic dependence on execution history. This distinctive feature is a great facilitator to time composability, which is a must for industry needing incremental development and qualification. In this paper we show how time composability is achieved in PTA-conformant systems and how the pessimism of worst-case execution time bounds obtained from PTA is contained within a 5% to 25% range for representative application scenarios.
Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla
ISORC1
2013 Multi-level Unified Caches for Probabilistically Time Analysable Real-Time Systems
abstract
Abstract—Caches are key resources in high-end processor architectures to increase performance. In fact, most high-performance processors come equipped with a multi-level cache hierarchy. In terms of guaranteed performance, however, cache hierarchies severely challenge the computation of tight worst-case execution time (WCET) estimates. On the one hand, the analysis of the timing behaviour of a single level of cache is already challenging, particularly for data accesses. On the other hand, unifying data and instructions in each level, makes the problem of cache analysis significantly more complex requiring tracking simultaneously data and instruction accesses to cache. In this paper we prove that multi-level cache hierarchies can be used in the context of Probabilistic Timing Analysis and tight WCET estimates can be obtained. Our detailed analysis (1) covers unified data and instruction caches, (2) covers different cache-write policies (write-through and write back), write allocation policies (write-allocate and non-write-allocate) and several inclusion mechanisms (inclusive, non-inclusive and exclusive caches), and (3) scales to an arbitrary number of cache levels. Our results show that the probabilistic WCET (pWCET) estimates provided by our analysis technique effectively benefit from having multi-level caches. For a two-level cache configura-tion and for EEMBC benchmarks, pWCET reductions are 55% on average (and up to 90%) with respect to a processor with a single level of cache. I.
Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
RTSS1
2013 PROARTIS: Probabilistically Analyzable Real-Time Systems
abstract
Static timing analysis is the state-of-the-art practice of ascertaining the timing behavior of current-generation real-time embedded systems. The adoption of more complex hardware to respond to the increasing demand for computing power in next-generation systems exacerbates some of the limitations of static timing analysis. In particular, the effort of acquiring (1) detailed information on the hardware to develop an accurate model of its execution latency as well as (2) knowledge of the timing behavior of the program in the presence of varying hardware conditions, such as those dependent on the history of previously executed instructions. We call these problems the timing analysis walls. In this vision-statement article, we present probabilistic timing analysis , a novel approach to the analysis of the timing behavior of next-generation real-time embedded systems. We show how probabilistic timing analysis attacks the timing analysis walls; we then illustrate the mathematical foundations on which this method is based and the challenges we face in the effort of efficiently implementing it. We also present experimental evidence that shows how probabilistic timing analysis reduces the extent of knowledge about the execution platform required to produce probabilistically accurate WCET estimations.
Francisco J. Cazorla, Eduardo Quiñones, Tullio Vardanega, Liliana Cucu-Grosjean, Benoit Triquet, Guillem Bernat, Emery D. Berger, Jaume Abella 0001, Franck Wartel, Michael Houston, Luca Santinelli, Leonidas Kosmidis, Code Lo, Dorin Maxim
ACM Trans. Embed. Comput. Syst.12
2012 Measurement-Based Probabilistic Timing Analysis for Multi-path Programs
abstract
The rigorous application of static timing analysis requires a large and costly amount of detail knowledge on the hardware and software components of the system. Probabilistic Timing Analysis has potential for reducing the weight of that demand. In this paper, we present a sound measurement-based probabilistic timing analysis technique based on Extreme Value Theory. In all the experiments made as part of this work, the timing bounds determined by our technique were less than 15% pessimistic in comparison with the tightest possible bounds obtainable with any probabilistic timing analysis technique. As a point of interest to industrial users, our technique also requires a comparatively low number of measurement runs of the program under analysis, less than 650 runs were needed for the benchmarks presented in this paper.
Liliana Cucu-Grosjean, Luca Santinelli, Michael Houston, Code Lo, Tullio Vardanega, Leonidas Kosmidis, Jaume Abella 0001, Enrico Mezzetti, Eduardo Quiñones, Francisco J. Cazorla
ECRTS6