VLDB 2026 Research / reviewers in the wild / expert
Enrico Mezzetti
dblp:72/8239
· DBLP profile ↗
45ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-1886-2931ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Partner Project: UP2DATE4SDV : Enabling Safe and Secure Modular Updates, Upgrades and Dynamic Task-Reallocation and -Execution for Software Defined VehiclesabstractThe European automotive industry is undergoing a revolution by the upcoming technologies of software-defined vehicles (SDV) and connected cooperative and automated mobility (CCAM). In a globally challenging context, in which Europe lost market share, the local automotive software and electronics market still expects a 11.9% compound annual growth rate from 2025 to 2030 – even more accentuated with the increasing adoption of advanced driver assistance (ADAS) and autonomous driving (AD). Both SDV and CCAM are key technological paradigms for enabling a shift of the European automotive sector towards a regained strategic competitiveness. Targeting the resulting need for faster development, deployment and test cycles the Horizon Europe RIA UP2DATE4SDV aims to develop a comprehensive ecosystem for seamless and efficient, safe and secure software updates, hardware upgrades, and situation-dependent reconfigurations of SDVs. For that goal, the UP2DATE4SDV consortium collaborates on the definition and development of two abstraction layers – the hardware abstraction layer and the operating system & middleware abstraction layer – as well as on researching and prototyping a safe and secure orchestration and reconfiguration plane between vehicle and cloud. Based on the resulting modular architecture concept and the corresponding DevOps process the project develops demonstrators to showcase safe and secure updates, upgrades and dynamic task reallocation for automotive hardware and software components. In this paper, we introduce the project, its objectives and planned results, draft first outcomes by refining our demonstrator definitions, and conclude with an outlook into the automotive future based on a safe and secure adaptive SDV stack. Gregor Nitsche, Patrick Uven, Hannes Fuchs, Enrico Mezzetti, Marcus Hähnel, Mijangos Ane, Kim Grüttner |
DATE | 4 |
| 2026 | ROSBand: A Bandwidth Regulation Approach on ROS2-Based Systems
Jon Altonaga Puente, Enrico Mezzetti, Irune Agirre, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 2 |
| 2026 | Evaluating quantile regression neural networks for optimizing real-time applications on heterogeneous platformsabstractModern cyber-physical systems increasingly rely on computationally demanding applications, particularly at the edge, where Artificial Intelligence-based algorithms are deployed. To meet these demands, industry trends are shifting towards heterogeneous MultiProcessor Systems on Chip (MPSoCs), which must also satisfy strict real-time and functional safety requirements. A major challenge in such systems is memory contention, where multiple processing units compete for shared memory resources, affecting application performance and the accurate estimation of Worst-Case Execution Times (WCETs). Traditional static analysis becomes impractical as system configurations grow in complexity. This work presents the design of an analysis and optimization framework for real-time systems that re-evaluates WCET estimates based on system configurations to reflect the impact of memory contention on heterogeneous platforms. The proposed method estimates new WCETs using Quantile Regression Neural Networks (QRNNs), which infer memory contention from Event Monitor data. Experimental results reveal that QRNN models must be system-specific for accurate predictions and that memory access patterns significantly affect model generalization. Two strategies are proposed: using generic models for simplicity or task-specific models for higher accuracy. Despite some potential underestimations, QRNNs maintain a strong correlation with actual observed contention, enabling effective worst-case scenario identification. Furthermore, a comparative analysis highlights the superior scalability of the estimation-based approach over empirical measurements, especially in large system optimization processes where performance can be easily enhanced by at least two orders of magnitude, making it a practical solution for real-time system design and analysis. Iosu Gomez, David Fonts, Sergi Vilardell, Unai Díaz-de-Cerio, Juan Maria Rivas, Enrico Mezzetti, J. Javier Gutiérrez, Francisco J. Cazorla |
Future Gener. Comput. Syst. | 6 |
| 2026 | Supporting Timing-related Metrics for Autonomous Driving Frameworks in CyberRTabstractThe provision of increasingly advanced autonomous software functionalities builds on cutting-edge autonomous driving frameworks to enable modular interactions among multiple software components. This approach helps to support functional cause-effect chains from multiple sensors to actuators. The complexity of the (software) component interactions makes it more difficult to ascertain the correctness of the timing behavior of the system. This is so because traditional timing-related metrics like worst-case execution and worst-case response time do not capture the inter-dependency in cause-effect chains between the input sampling time and the time at which computation based on those inputs is performed. Complementary timing-related metrics, such as maximum reaction time and maximum data age have been considered to capture timing requirements, typically with an end-to-end scope, in cause-effect chains. These metrics have been formalized and demonstrated in ROS2-based automotive and autonomous driving setups [ 44 , 46 ]. However, the formalization of those metrics, which is necessary for deriving analytical lower and upper bounds and monitoring them at run-time, largely depends on the execution model and semantics offered by the run-time. Any concrete application of those metrics need to be tailored and adapted to the system at hand. Apollo auto is a popular, industrial-quality, open-source autonomous driving framework that is seeing increasing adoption both for industrial and academic projects. Apollo builds on CyberRT , an ad-hoc run-time that is similar in mechanism and intent to ROS2 but differentiates from it with respect to execution model and supported semantics. In contrast to ROS2, CyberRT is highly specialized to support the Apollo AD framework, is neither extensively documented or thoroughly analysed in the literature, especially in relation to execution model and instantiation of timing-related metrics. In this work, for the first time, we provide an insightful analysis and discussion on CyberRT execution model and semantics, starting from its raw and non-extensively documented codebase. Based on the identified semantics, we elaborate a formalization of timing-related metrics on CyberRT , across different granularity scopes, namely end-to-end and node levels. In particular, we develop on the importance of node-level timing properties to intercept any latent timing misbehavior before it is too late, and it severely impacts end-to-end execution. We provide a concrete mapping of a comprehensive set of timing-related metrics to the CyberRT execution model, both at end-to-end and node level, and develop a monitoring library that allows to intercept them on the specific software stack. We exploit the proposed library on a set of Apollo autonomous driving scenarios to demonstrate its effectiveness in monitoring the considered timing metrics and to promptly intercept a subtle timing misbehavior beyond end-to-end execution scope in a representative autonomous driving stack. Miguel Alcon, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | SAFEXPLAIN: a Complete Approach Towards Trustworthy AI-Based Safety-Critical SystemsabstractAI becomes increasingly important in safetycritical systems, especially in the case of autonomous systems, since navigation relies on AI for object detection and collision avoidance. However, safety-critical systems must adhere to functional safety standards that enforce software to be correct-by-construction, component decomposition to simplify design and validation, and the use of data only for testing purposes not to design the system itself. AI in general, and Deep Learning (DL) in particular have opposed characteristics since they have error rates (e.g., due to mispredictions), AI/DL modules can only be designed and validated monolithically, and they build on data for their design (i.e. for training purposes). Hence, DL solutions are at odds with the development process of safetycritical systems. A number of standards have recently emerged in different domains to reconcile the requirements of safety-critical systems with the characteristics of DL solutions, such as ISO 21448, ISO/IEC TR 5469, and ISO 8800, among others. However, there is a lack of realistic practice to design a DL-based safety-critical system in accordance with those regulations, and existing solutions only cover some aspects in isolation, and are often incompatible among them. SAFEXPLAIN is a 3-year Horizon Europe project addressing this challenge. SAFEXPLAIN, which finishes in September 2025, has already reached its main goals providing specific and complementary solutions to all those challenges so that AIbased safety-critical systems can be designed, implemented and validated adhering to the relevant functional safety standards in domains such as automotive, space and railway. In particular, SAFEXPLAIN provides the concepts, processes, tools and frameworks addressing the challenge end-to-end, from concept to solution. This is proven by the successful application of the SAFEXPLAIN approach in three case studies from the automotive, space and railway domains, whose results will see the light very soon. Jaume Abella 0001, Irune Agirre, Thanh Hai Bui, Frank Geujen, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Axel Brando, Javier Fernández 0004, Irune Yarza, Joanes Plazaola, Maria Ulan, Rob Lavreysen, Lucas Tosi, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Stefano Lodico, William Guarienti, Giuseppe Nicosia, Valeria Dallara |
DSD | 8 |
| 2025 | Impact of Contention-Aware Placement in Heterogeneous Edge DevicesabstractTime predictability is an increasing concern in functionally-rich mixed-criticality applications at the Edge, which often carry different timing requirements. Edge devices, in turn, are increasingly complex to sustain the increasing computational requirements, which hinders providing predictable performance without seriously affecting performance. One of the main threats to predictable performance is the impact of timing interference arising from contention in an increasing number of shared hardware resources. The impact of software to hardware mapping on performance is a well-studied topic, seeking optimal memory mappings to reduce average and worst-case performance, and, more recently, to control and limit timing interference. These methods normally focus on code and data placement, especially in relation to specific properties of the memory hierarchy, either architectural (e.g. heterogeneous memory modules) or obtained through partitioning techniques. In practice, however, these works build on a uniform memory hierarchy model, where the source of a memory request, namely, where a task accessing a given memory is eventually executed, is not directly relevant. In this work, we consider a large class of systems (e.g., TriCore families) where memory hierarchies are non-uniform, and access latency depends on the computing element issuing the request. In those architectures, the impact of code and data placement on timing interference cannot be addressed without considering architectural constraints and task locality. Through empirical exploration, we show that code, data, and locality collectively have a substantial impact on contention bounds, leading to a significantly expanded optimization space compared to approaches considering only code and data placement under the uniform memory assumption. Our results motivate the need for novel, efficient optimization approaches that integrate task mapping and architectural constraints to reduce timing interference. Jeremy Giesen, Ibai Irigoyen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2025 | Detecting Low-Density Mixtures in High-Quantile Tails for pWCET Estimation
Blau Manau, Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 4 |
| 2025 | EMR: Removing Multicollinear Event Monitors to Improve Timing Modelling of Real-Time SystemsabstractMulticollinearity of Event Monitors (EMs) negatively impacts the modeling of non-functional critical metrics in real-time systems like worst-case timing and energy usage since some EMs are over-represented and can reduce model accuracy. To address this challenge, we propose Event Monitor Reduction (EMR), a method to select a reduced set of non-related (independent) features (EMs), hence eliminating multicollinearity. In particular, EMR finds linear relations between the EMs and removes dependent ones without data loss. EMR does not create new features like Principal Component Analysis does, simplifying interpretability. Results on synthetic data and data collected from the execution of representative benchmarks on an avionicsgrade processor show the benefits of our method in removing multicollinear EMs. We further illustrate the benefits of EMR on two different multicore timing contention models, showing how its application helps to reduce execution time requirements and increase the accuracy of the models. David Fonts, Diego Palacios, Sergi Vilardell, Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 6 |
| 2024 | TAP: Task-Aware Profiling on Integrated SystemsabstractHardware Performance Monitors (HPM) are increasingly exploited for timing verification and validation of time-critical embedded systems (TECS). HPMs are typically collected at the lowest software level, which makes it difficult to unequivocally account events to specific run-time entities, a prerequisite for any form of analysis, without relying on ad-hoc support from the run-time or operating system layer. The latter, however, is either unavailable or not fully adequate for verification requirements. Moreover, timing-related concerns in the analysis of embedded systems are typically addressed in the final stages of the software development process where multiple tasks are fully or partially integrated on the platform and it is therefore hard, if not impossible, to enforce controlled testing scenarios where contributions to event counts can be dissected. In this work, we present TAP a generic concept for allowing Task-Aware Profiling of individual tasks in an already integrated system on MPSoCs with on-core and off-core HPM support. The proposed approach combines a lightweight user-level configurable API and minimally intrusive extensions to the operating system layer to enforce separation of contexts when collecting HPM. We implement and assess TAP on top of an Infineon AURIX MPSoC and the OSEK-compliant ERIKA Enterpise RTOS, offering a consistent and intuitive interface for governing and filtering the different sources of events. Our results on synthetic and automotive benchmarks show that TAP can transparently gather and filter the events of interest while incurring negligible overheads. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 2 |
| 2024 | Event Monitor Validation in High-Integrity SystemsabstractPlatforms for modern embedded systems equip an increasing number of high-performance features to provide the required levels of performance. Timing analysis solutions handle the complexity of these platforms by relying on hardware event monitors (HEMs) that provide insightful information about resource utilization and, hence, contention among tasks. As a result, HEMs have become a key element to warrant a safe timing behavior of a system, for which reason they must be validated. While some initial works target HEMs validation, they consider one HEM at a time and focus on those HEMs for which an expert can establish an expected value for relatively small code snippets. In this paper, we propose a methodology for the validation of those HEMs for which a specific expected value cannot be established a priori even for simple cases and, instead, needs to be validated in conjunction with other HEMs. Our method also deals with the natural variability of the HEMs' values in high-performance platforms when collected in different experiments. We illustrate the effectiveness of our proposed technique for validating HEMs related to cache coherence in a relevant platform in the avionics domain. Roger Pujol, Sergi Vilardell, Enrico Mezzetti, Mohamed Hassan 0002, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2024 | Achieving Flexible Performance Isolation on the AMD Xilinx Zynq UltraScale+abstractCo-hosting different tasks on the same MPSoC contributes to increasing average performance by allowing them to share MPSoC's resources that, otherwise, could be underutilized. However, resource sharing challenges performance isolation among tasks, as required in time-sensitive embedded critical systems like automotive and avionics. On the other hand, resource isolation through segregation (the reference solution for preventing the propagation of time-related safety issues) is detrimental to average performance. In this work, we show that the built-in QoS support in modern MPSoCs can be smartly leveraged to adapt to the timing and performance requirements of the running applications. In particular, we develop specific configurations of the complex QoS support in the Zynq UltraScale+ MPSoC that deliver performance isolation for time-sensitive tasks (TSTs) and ensure that non-time-sensitive tasks (NTSTs) maximize their average performance by exploiting the resources not used by TSTs. Our results on the Xilinx UtltraScale+ show that the TSTs with the most stringent constraints achieve high degrees of isolation, 96.0% of their solo performance on average, while NTSTs exploit the resources not used by TSTs achieving performance ranging from 72% to 4% depending on the resource left by TSTs. Alejandro Serrano-Cases, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 2 |
| 2023 | SAFEXPLAIN: Safe and Explainable Critical Embedded Systems Based on AIabstractDeep Learning (DL) techniques are at the heart of most future advanced software functions in Critical Autonomous AI-based Systems (CAIS), where they also represent a major competitive factor. Hence, the economic success of CAIS industries (e.g., automotive, space, railway) depends on their ability to design, implement, qualify, and certify DL-based software products under bounded effort/cost. However, there is a fundamental gap between Functional Safety (FUSA) requirements on CAIS and the nature of DL solutions. This gap stems from the development process of DL libraries and affects high-level safety concepts such as (1) explainability and traceability, (2) suitability for varying safety requirements, (3) FUSA-compliant implementations, and (4) real-time constraints. As a matter of fact, the data-dependent and stochastic nature of DL algorithms clashes with current FUSA practice, which instead builds on deterministic, verifiable, and pass/fail test-based software. The SAFEXPLAIN project tackles these challenges and targets by providing a flexible approach to allow the certification - hence adoption - of DL-based solutions in CAIS building on: (1) DL solutions that provide end-to-end traceability, with specific approaches to explain whether predictions can be trusted and strategies to reach (and prove) correct operation, in accordance to certification standards; (2) alternative and increasingly sophisticated design safety patterns for DL with varying criticality and fault tolerance requirements; (3) DL library implementations that adhere to safety requirements; and (4) computing platform configurations, to regain determinism, and probabilistic timing analyses, to handle the remaining non-determinism. Jaume Abella 0001, Jon Pérez 0001, Cristofer Englund, Bahram Zonooz, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Isabel Serra, Axel Brando, Irune Agirre, Fernando Eizaguirre, Thanh Hai Bui, Elahe Arani, Fahad Sarfraz, Ajay Balasubramaniam, Ahmed Badar, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Davide Brighenti, Davide Cunial |
DATE | 8 |
| 2023 | Quasi Isolation QoS Setups to Control MPSoC Contention in Integrated Software Architectures
Sergio Garcia-Esteban, Alejandro Serrano-Cases, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 4 |
| 2023 | Improving Timing-Related Guarantees for Main Memory in Multicore Critical Embedded SystemsabstractMain memory is one of the most complex resources to analyze in multicore-based embedded real-time systems, with contention in the memory controller and the timing constraints of the main memory device as the main contributors to that complexity. One of the main challenges in multicore real-time systems is producing the required evidence on the management of contention delay for the certification. This stems from the fact that current MPSoCs barely provide any event monitors on how tasks interact and delay each other in memory. Besides, even if hardware and software mechanisms are in place to mitigate contention in the memory system, it is hard - if at all possible - to provide evidence about their correctness. In this work, we cover this gap by proposing a lightweight hardware mechanism that tightly tracks inter-core contention in memory. The proposed hardware mechanism, which we evaluate in detail, improves the quality of timing-related evidence that must be provided on how contention in main memory of multicore real-time systems is handled in adherence to applicable safety standards. Asier Fernández de Lecea, Mohamed Hassan 0002, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 3 |
| 2023 | Main sources of variability and non-determinism in AD software: taxonomy and prospects to handle them
Miguel Alcon, Axel Brando, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
Real Time Syst. | 3 |
| 2023 | Vector Extensions in COTS Processors to Increase Guaranteed Performance in Real-Time SystemsabstractThe need for increased application performance in high-integrity systems such as those in avionics is on the rise as software continues to implement more complex functionalities. The prevalent computing solution for future high-integrity embedded products is multi-processor systems-on-chip (MPSoC) processors. MPSoCs include central processing unit (CPU) multicores that enable improving performance via thread-level parallelism. MPSoCs also include generic accelerators (graphics processing units [GPUs]) and application-specific accelerators. However, the data processing approach (DPA) required to exploit each of these underlying parallel hardware blocks carries several open challenges to enable the safe deployment in high-integrity domains. The main challenges include the qualification of its associated runtime system and the difficulties in analyzing programs deploying the DPA with out-of-the-box timing analysis and code coverage tools. In this work, we perform a thorough analysis of vector extensions (VExts) in current commercial off-the-shelf (COTS) processors for high-integrity systems. We show that VExts prevent many of the challenges arising with parallel programming models and GPUs. Unlike other DPAs, VExts require no runtime support, prevent design race conditions that might arise with parallel programming models, and have minimum impact on the software ecosystem, enabling the use of existing code coverage and timing analysis tools. We develop vectorized versions of neural network kernels and show that the NVIDIA Xavier VExts provide a reasonable increase in guaranteed application performance of up to 2.7x. Our analysis contends that VExts are the DPA approach with arguably the fastest path for adoption in high-integrity systems. Roger Pujol, Josep Jorba 0002, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | Accurately Measuring Contention in Mesh NoCs in Time-Sensitive Embedded SystemsabstractThe computing capacity demanded by embedded systems is on the rise as software implements more functionalities, ranging from best-effort entertainment functions to performance-guaranteed safety-related functions. Heterogeneous manycore processors, using wormhole mesh (wmesh) Network-on-Chips (NoCs) as the main communication means, and contention block among applications, are increasingly considered to deliver the required computing performance. Most research efforts on software timing analysis have focused on deriving bounds (estimates) to the contention that tasks can suffer when accessing wmesh NoCs. However, less effort has been devoted to an equally important problem, namely,accuratelymeasuring the actual contention tasks generate each other on the wmesh which is instrumental during system validation to diagnose any software timing misbehavior and determine which tasks are particularly affected by contention on specific wmesh routers. In this article, we work on the foundations ofcontention measuringin wmesh NoCs and propose and explain the rationale of agolden metric, called taskPairWise Contention(PWC). PWC allows ascribing the actual share of the contention a given task suffers in the wmesh to each of its co-runner tasks at packet level. We also introduce and formalize aGolden Reference Value(GRV) for PWC that specifically defines a criterion to fairly break down the contention suffered by a task among its co-runner tasks in the wmesh. Our evaluation shows that GRV effectively captures how contention occurs by identifying the actual core (task) causing contention and whether contention is caused by local or remote interference in the wmesh. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2022 | Using Quantile Regression in Neural Networks for Contention Prediction in Multicore ProcessorsabstractMachine learning has enabled significant benefits in diverse fields, but, with a few exceptions, has had limited impact on computer architecture. Recent work, however, has explored broader applicability for design, optimization, and simulation. Notably, machine learning based strategies often surpass prior state-of-the-art analytical, heuristic, and human-expert approaches. This paper reviews machine learning applied system-wide to simulation and run-time optimization, and in many individual components, including memory systems, branch predictors, networks-on-chip, and GPUs. The paper further analyzes current practice to highlight useful design strategies and identify areas for future work, based on optimized implementation strategies, opportune extensions to existing work, and ambitious long term possibilities. Taken together, these strategies and techniques present a promising future for increasingly automated architectural design. Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 3 |
| 2022 | Using Markov's Inequality with Power-Of-k Function for Probabilistic WCET Estimation
Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Joan del Castillo |
ECRTS | 3 |
| 2021 | Leveraging Hardware QoS to Control Contention in the Xilinx Zynq UltraScale+ MPSoCabstractThe interference co-running tasks generate on each other’s timing behavior continues to be one of the main challenges to be addressed before Multi-Processor System-on-Chip (MPSoCs) are fully embraced in critical systems like those deployed in avionics and automotive domains. Modern MPSoCs like the Xilinx Zynq UltraScale+ incorporate hardware Quality of Service (QoS) mechanisms that can help controlling contention among tasks. Given the distributed nature of modern MPSoCs, the route a request follows from its source (usually a compute element like a CPU) to its target (usually a memory) crosses several QoS points, each one potentially implementing a different QoS mechanism. Mastering QoS mechanisms individually, as well as their combined operation, is pivotal to obtain the expected benefits from the QoS support. In this work, we perform, to our knowledge, the first qualitative and quantitative analysis of the distributed QoS mechanisms in the Xilinx UltraScale+ MPSoC. We empirically derive QoS information not covered by the technical documentation, and show limitations and benefits of the available QoS support. To that end, we use a case study building on neural network kernels commonly used in autonomous systems in different real-time domains. Alejandro Serrano-Cases, Juan M. Reina, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 4 |
| 2021 | PRL: Standardizing Performance Monitoring Library for High-Integrity Real-Time SystemsabstractThe use of complex processors is becoming ubiquitous in High-Integrity Systems (HIS). To deal with processor’s increased complexity, Performance Monitoring Counters (PMCs) are increasingly used to reason on software behavior and provide the necessary evidence to support software certification. However, the use of PMCs in HIS is relatively recent and hence far from being standardized. As a result, software engineers are forced to resort to highly-customized, low-level programming of platform-specific PMC control registers, which is both error prone and time consuming. To cover this gap, we propose building on the PAPI library, a standardized performance monitoring solution in the mainstream domain, and develop a PMC Reading Library (PRL) for configuring and collecting traceable events while capturing HIS specific requirements and peculiarities. We instantiate PRL in a reference automotive configuration to show that PRL meets key HIS requirements: negligible footprint, limited and predictable overhead, and accuracy collecting hardware events by filtering out the impact of interrupts and context switches. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ICCD | 2 |
| 2020 | Tracing Hardware Monitors in the GR712RC Multicore Platform: Challenges and Lessons Learnt from a Space Case StudyabstractThe demand for increased computing performance is driving industry in critical-embedded systems (CES) domains, e.g. space, towards the use of multicores processors. Multicores, however, pose several challenges that must be addressed before their safe adoption in critical embedded domains. One of the prominent challenges is software timing analysis, a fundamental step in the verification and validation process. Monitoring and profiling solutions, traditionally used for debugging and optimization, are increasingly exploited for software timing in multicores. In particular, hardware event monitors related to requests to shared hardware resources are building block to assess and restraining multicore interference. Modern timing analysis techniques build on event monitors to track and control the contention tasks can generate each other in a multicore platform. In this paper we look into the hardware profiling problem from an industrial perspective and address both methodological and practical problems when monitoring a multicore application. We assess pros and cons of several profiling and tracing solutions, showing that several aspects need to be taken into account while considering the appropriate mechanism to collect and extract the profiling information from a multicore COTS platform. We address the profiling problem on a representative COTS platform for the aerospace domain to find that the availability of directly-accessible hardware counters is not a given, and it may be necessary to the develop specific tools that capture the needs of both the user’s and the timing analysis technique requirements. We report challenges in developing an event monitor tracing tool that works for bare-metal and RTEMS configurations and show the accuracy of the developed tool-set in profiling a real aerospace application. We also show how the profiling tools can be exploited, together with handcrafted benchmarks, to characterize the application behavior in terms of multicore timing interference. Xavier Palomo, Mikel Fernández, Sylvain Girbal, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Laurent Rioux |
ECRTS | 4 |
| 2020 | Timing of Autonomous Driving Software: Problem Analysis and Prospects for Future SolutionsabstractThe software used to implement advanced functionalities in critical domains (e.g. autonomous operation) impairs software timing. This is not only due to the complexity of the underlying high-performance hardware deployed to provide the required levels of computing performance, but also due to the complexity, non-deterministic nature, and huge input space of the artificial intelligence (AI) algorithms used. In this paper, we focus on Apollo, an industrial-quality Autonomous Driving (AD) software framework: we statistically characterize its observed execution time variability and reason on the sources behind it. We discuss the main challenges and limitations in finding a satisfactory software timing analysis solution for Apollo and also show the main traits for the acceptability of statistical timing analysis techniques as a feasible path. While providing a consolidated solution for the software timing analysis of Apollo is a huge effort far beyond the scope of a single research paper, our work aims to set the basis for future and more elaborated techniques for the timing analysis of AD software. Miguel Alcon, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 4 |
| 2020 | Modeling Contention Interference in Crossbar-based Systems via Sequence-Aware Pairing (SeAP)abstractThe Infineon AURIX TriCore family of microcontrollers has consolidated as the reference multicore computing platform for safety-critical systems in the automotive domain. As a distinctive trait, AURIX microcontrollers are designed to promote high timing predictability as witnessed by the presence of large scratchpad memories and a crossbar interconnect. The latter has been introduced to reduce inter-core interference in accessing the memory system and peripherals. Nonetheless, the crossbar does not prevent requests from different cores to the same target resource to suffer contention. Applications are, therefore, inherently exposed to inter-core timing interference, which needs to be taken into account in the determination of reliable execution time bounds. In this paper we propose a contention modeling technique for crossbar-based systems, and hence suitable for bounding contention effects in the AURIX family. Unlike state of the art techniques that build on total request counts, we exploit the sequence of requests to the different target resources produced by each core to produce tighter bounds by discarding contention scenarios that cannot occur in practice. To that end, we adapt existing techniques from the pattern matching domain to derive the worst-case contention effects from the sequences of requests each core sends over the crossbar. Results on a wide set of synthetic and real scenarios and benchmark on an AURIX TC297TX show that our technique outperforms other contention modeling approaches. Jeremy Giesen, Pedro Benedicte, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 3 |
| 2020 | HRM: Merging Hardware Event Monitors for Improved Timing Analysis of Complex MPSoCsabstractThe performance monitoring unit (PMU) in multiprocessor system-on-chips (MPSoCs) is at the heart of the latest measurement-based timing analysis techniques in critical embedded systems. In particular, hardware event monitors (HEMs) in the PMU are used as building blocks in the process of budgeting and verifying software timing by tracking and controlling access counts to shared resources. While the number of HEMs in current MPSoCs reaches hundreds, they are read via performance monitoring counters whose number is usually limited to 4-8, thus requiring multiple runs of each experiment in order to collect all desired HEMs. Despite the effort of engineers in controlling the execution conditions of each experiment, the complexity of current MPSoCs makes it arguably impossible to completely remove the noise affecting each run. As a result, HEMs read in different runs are subject to different variability, and hence, those HEMs captured in different runs cannot be “blindly” merged. In this work, we focus on the NXP T2080 platform where we observed up to 59% variability across different runs of the same experiment for some relevant HEMs (e.g., processor cycles). We develop a HEM reading and merging (HRM) approach to join reliably HEMs across different runs as a fundamental element of any measurement-based timing budgeting and verification technique. Our method builds on order statistics and the selection of an anchor HEM read in all runs to derive the most plausible combination of HEM readings that keep the distribution of each HEM and their relationship with the anchor HEM intact. Sergi Vilardell, Isabel Serra, Roberto Santalla, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | Towards limiting the impact of timing anomalies in complex real-time processorsabstractTiming verification of embedded critical real-time systems is hindered by complex designs. Timing anomalies, deeply analyzed in static timing analysis, require specific solutions to bound their impact. For the first time, we study the concept and impact of timing anomalies in measurement-based timing analysis, the most used in industry, showing that they require to be considered and handled differently. In addition, we analyze anomalies in the context of Measurement-Based Probabilistic Timing Analysis, which simplifies quantifying their impact. Pedro Benedicte, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Francisco J. Cazorla |
ASP-DAC | 4 |
| 2019 | AURIX TC277 Multicore Contention Model Integration for Automotive ApplicationsabstractEmbedded systems industry needs reliable and tight worst-case execution time (WCET) estimates for critical applications running on multicores, as a prerequisite to their adoption. While industry already uses reliable tools for single-core WCET estimation and several multicore contention models (MCMs) have been proposed, their combination have not been shown to be fully compatible with the automotive industrial practice yet. This paper reduces this gap by presenting a framework for the integration of MCMs into industrial WCET estimation practice. We illustrate such integration for a Magneti Marelli powertrain control unit on an Infineon AURIX TC277 multicore platform. Enrico Mezzetti, Luca Barbina, Jaume Abella 0001, Stefania Botta, Francisco J. Cazorla |
DATE | 1 |
| 2019 | Generating and Exploiting Deep Learning Variants to Increase Heterogeneous Resource Utilization in the NVIDIA XavierabstractDeep learning-based solutions and, in particular, deep neural networks (DNNs) are at the heart of several functionalities in critical-real time embedded systems (CRTES) from vision-based perception (object detection and tracking) systems to trajectory planning. As a result, several DNN instances simultaneously run at any time on the same computing platform. However, while modern GPUs offer a variety of computing elements (e.g. CPUs, GPUs, and specific accelerators) in which those DNN tasks can be executed depending on their computational requirements and temporal constraints, current DNNs are mainly programmed to exploit one of them, namely, regular cores in the GPU. This creates resource imbalance and under-utilization of GPU resources when executing several DNN instances, causing an increase in DNN tasks' execution time requirements. In this paper, (a) we develop different variants (implementations) of well-known DNN libraries used in the Apollo Autonomous Driving (AD) software for each of the computing elements of the latest NVIDIA Xavier SoC. Each variant can be configured to balance resource requirements and performance: the regular CPU core implementation that can run on 2, 4, and 6 cores; the GPU regular and Tensor core variants that can run in 4 or 8 GPU’s Streaming Multiprocessors (SM); and 1 or 2 NVIDIA’s Deep Learning Accelerators (NVDLA); (b) we show that each particular variant/configuration offers a different resource utilization/performance point; finally, (c) we show how those heterogeneous computing elements can be exploited by a static scheduler to sustain the execution of multiple and diverse DNN variants on the same platform. Roger Pujol, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 4 |
| 2019 | Accurate ILP-Based Contention Modeling on Statically Scheduled Multicore SystemsabstractCommercially available Off The Shelf (COTS) multicores have been assessed as the baseline computing platform even in the most conservative real-time domains. Multicore contention arising on shared hardware resources, with its circular dependence with scheduling, is among the most challenging issues that require urgent attention before multicores can be fully embraced for real-time computing. In the context of static scheduling, still the most used scheduling approach in real-time industries, we propose an ILP formulation for computing the worst-case contention delay suffered by a task due to interference on a shared bus. Our model provides accurate contention delay bounds that avoid unnecessary over-accounting of conflicts between bus requests, by considering contention effects at system-level (i.e., across tasks) rather than at task-level only. This allows precisely capturing the interdependence between timing interference of conflicting requests, issued in parallel by other cores (tasks), and the identification of the particular set of tasks co-running on those cores. We assess our technique both analytically and empirically on a real COTS multicore platform. We show, via extensive evaluation, that jointly accounting for worst-case task overlapping and request distribution scenarios always provides tighter contention bounds when compared to state-of-the-art solutions. Xavier Palomo, Enrico Mezzetti, Jaume Abella 0001, Reinder J. Bril, Francisco J. Cazorla |
RTAS | 2 |
| 2019 | Increasing the Reliability of Software Timing Analysis for Cache-Based ProcessorsabstractReal-time systems are witnessing a significant increase in critical software's size, complexity, and performance needs, which can only be satisfied with high-performance hardware features. Cache memories, pervasively used to improve average performance, complicate Worst-Case Execution Time analysis: cache placement (i.e., how software objects are mapped to cache) during the testing phase does not only critically affect the observed performance, but also proves to be arduous to control and preserve up to operation. The probabilistic variant of Measurement-Based Timing Analysis (MBPTA) responds to this challenge by deploying time-randomized caches that naturally explore a different random cache placement in each run, relieving the user from producing tests that intercept relevant Cache Conflict Placements (CCP). Yet, to meet an adequate probabilistic CCP coverage, the user is required to collect a minimum number of measurements. We present two mechanisms, CCP-RM and CCP-HRP, to identify CCP with relevant probability of occurrence and large impact on execution-time, for the random modulo (RM) and hash-based random placement (HRP) policies. CCP-RM and CCP-HRP enable a reliable application of MBPTA by computing the number of runsR' necessary to meet the desired CCP coverage. We exhaustively evaluate CCP-RM and CCP-HRP, showing their effectiveness on well-known benchmarks and a railway case study, on top of an accurate simulator and a concrete RTL implementation. Suzana Milutinovic, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Computers | 2 |
| 2018 | Modelling multicore contention on the AURIXTM TC27xabstractMulticores are becoming ubiquitous in automotive. Yet, the expected benefits on integration are challenged by multicore contention concerns on timing V&V. Worst-case execution time (WCET) estimates are required as early as possible in the software development, to enable prompt detection of timing misbehavior. Factoring in multicore contention necessarily builds on conservative assumptions on interference, independent of co-runners load on shared hardware. We propose a contention model for automotive multicores that balances time-composability with tightness by exploiting available information on contenders. We tailor the model to the AURIX TC27x and provide tight WCET estimates using information from performance monitors and software configurations. Enrique Díaz, Enrico Mezzetti, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 2 |
| 2018 | Measurement-based cache representativeness on multipath programsabstractAutonomous vehicles in embedded real-time systems increase critical-software size and complexity whose performance needs are covered with high-performance hardware features like caches, which however hampers obtaining WCET estimates that hold valid for all program execution paths. This requires assessing that all cache layouts have been properly factored in the WCET process. For measurement-based timing analysis, the most common analysis method, we provide a solution to achieve cache representativeness and full path coverage: we create a modified program for analysis purposes where cache impact is upper-bounded across any path, and derive the minimum number of runs required to capture in the test campaign cache layouts resulting in high execution times. Suzana Milutinovic, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
DAC | 3 |
| 2018 | NoCo: ILP-Based Worst-Case Contention Estimation for Mesh Real-Time ManycoresabstractManycores are capable of providing the computational demands required by functionally-advanced critical applications in domains such as automotive and avionics. In manycores a network-on-chip (NoC) provides access to shared caches and memories and hence concentrates most of the contention that tasks suffer, with effects on the worst-case contention delay (WCD) of packets and tasks' WCET. While several proposals minimize the impact of individual NoC parameters on WCD, e.g. mapping and routing, there are strong dependences among these NoC parameters. Hence, finding the optimal NoC configurations requires optimizing all parameters simultaneously, which represents a multidimensional optimization problem. In this paper we propose NoCo, a novel approach that combines ILP and stochastic optimization to find NoC configurations in terms of packet routing, application mapping, and arbitration weight allocation. Our results show that NoCo improves other techniques that optimize a subset of NoC parameters. Jordi Cardona, Carles Hernández 0001, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 3 |
| 2018 | Fitting Software Execution-Time Exceedance into a Residual Random Fault in ISO-26262abstractCar manufacturers relentlessly replace or augment the functionality of mechanical subsystems with electronic components. Most such subsystems (e.g., steer-by-wire) are safety related, hence, subject to regulation. ISO-26262, the dominant standard for road vehicles, regards software faults as systematic, while differentiating hardware faults between systematic and random. The analysis of systematic faults entails rigorous processes and qualitative considerations. The increasing complexity of modern on-board computers, however, questions the very notion of treating the violation of execution-time envelopes for software programs as a systematic fault. Modern hardware in fact reduces the user's ability to delve deep enough into the fabric of hardware-software interaction to gage its extent of contribution to the worst-case execution time (WCET). Changing the nature of the WCET-analysis problem may help address that challenge effectively. To this end, we propose a solution that should allow ISO-26262 to quantify the likelihood of execution-time exceedance events, relating it to target failure metrics employed in support of certification arguments, similarly to random faults in hardware. To this end, we inject randomization in the timing behavior of the computer hardware to relieve the user from the need to control hard-to-reach low-level parts, and use measurement-based probabilistic timing analysis to quantify, constructively, the failure rates resulting from the likelihood of execution-time exceedance events. Irune Agirre, Francisco J. Cazorla, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Mikel Azkarate-askatsua, Tullio Vardanega |
IEEE Trans. Reliab. | 5 |
| 2017 | EPC Enacted: Integration in an Industrial Toolbox and Use against a Railway ApplicationabstractMeasurement-based timing analysis approaches are increasingly making their way into several industrial domains on account of their good cost-benefit ratio. The trustworthiness of those methods, however, suffers from the limitation that their results are only valid for the particular paths and execution conditions that the user is able to explore with the available input vectors. It is generally not possible to guarantee that the collected measurements are fully representative of the worst-case timing behaviour. In the context of measurement-based probabilistic timing analysis, the Extended Path Coverage (EPC) approach has been recently proposed as a means to extend the representativeness of measurement observations, to obtain the same effect of full path coverage. At the time of its first publication, EPC had not reached an implementation maturity that could be trialled industrially. In this work we analyze the practical implications of using EPC with real-world applications, and discuss the challenges in integrating it in an industrial-quality toolchain. We show that we were able to meet EPC requirements and successfully evaluate the technique on a real Railway application, on top of a commercial toolchain and full execution stack. Enrico Mezzetti, Mikel Fernández, Alen Bardizbanyan, Irune Agirre, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla |
RTAS | 1 |
| 2017 | Work-in-Progress Paper: An Analysis of the Impact of Dependencies on Probabilistic Timing Analysis and Task SchedulingabstractRecently there has been a renewed interest for probabilistic timing analysis (PTA) and probabilistic task scheduling (PTS). Despite the number of works in both fields, the link between them is weak: works on the latter build upon a series of assumptions on the probabilistic behavior of each task – or instances (jobs) of it – that have not been shown how to be fulfilled by PTA. This paper makes a first step towards covering this gap with emphasis on providing the right meaning of pWCET estimate as understood by both PTA and PTS. We show that the main issue related to ensuring that PTS assumptions on pWCET estimates are captured by PTA relates to the dependencies among tasks, and even jobs of a given task. Both change the scope of applicability of pWCET estimates provided by PTA and hence, their use by PTS. Enrico Mezzetti, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla |
RTSS | 1 |
| 2016 | PROXIMA: Improving Measurement-Based Timing Analysis through Randomisation and Probabilistic AnalysisabstractThe use of increasingly complex hardware and software platforms in response to the ever rising performance demands of modern real-time systems complicates the verification and validation of their timing behaviour, which form a time-and-effort-intensive step of system qualification or certification. In this paper we relate the current state of practice in measurement-based timing analysis, the predominant choice for industrial developers, to the proceedings of the PROXIMA (Probabilistic real-time control of mixed-criticality multicore systems) project in that very field. We recall the difficulties that the shift towards more complex computing platforms causes in that regard. Then we discuss the probabilistic approach proposed by PROXIMA to overcome some of those limitations. We present the main principles behind the PROXIMA approach as well as the changes it requires at hardware or software level underneath the application. We also present the current status of the project against its overall goals, and highlight some of the principal confidence-building results achieved so far. Francisco J. Cazorla, Jaume Abella 0001, Jan Andersson, Tullio Vardanega, Francis Vatrinet, Iain Bate, Ian Broster, Mikel Azkarate-askatsua, Franck Wartel, Liliana Cucu-Grosjean, Fabrice Cros, Glenn Farrall, Adriana Gogonel, Andrea Gianarro, Benoit Triquet, Carles Hernández 0001, Code Lo, Cristian Maxim, David Morales, Eduardo Quiñones, Enrico Mezzetti, Leonidas Kosmidis, Irune Agirre, Mikel Fernández, Mladen Slijepcevic, Philippa Conmy, Walid Talaboulma |
DSD | 21 |
| 2015 | Timing analysis of an avionics case study on complex hardware/software platforms
Franck Wartel, Leonidas Kosmidis, Adriana Gogonel, Andrea Baldovin, Zoë Stephenson, Benoit Triquet, Eduardo Quiñones, Code Lo, Enrico Mezzetti, Ian Broster, Jaume Abella 0001, Liliana Cucu-Grosjean, Tullio Vardanega, Francisco J. Cazorla |
DATE | 9 |
| 2015 | Experimental Evaluation of Optimal Schedulers Based on Partitioned Proportionate FairnessabstractThe Quasi-Partitioning Scheduling algorithm optimally solves the problem of scheduling a feasible set of independent implicit-deadline sporadic tasks on a symmetric multiprocessor. It iteratively combines bin-packing solutions to determine a feasible task-to-processor allocation, splitting task loads as needed along the way so that the excess computation on one processor is assigned to a paired processor. Though different in formulation, QPS belongs in the same family of schedulers as RUN, which achieve optimality using a relaxed (partitioned) version of proportionate fairness. Unlike RUN, QPS departs from the dual schedule equivalence, thus yielding a simpler implementation with less use of global data structures. One might therefore expect that QPS should outperform RUN in the general case. Surprisingly instead, our implementation of QPS on LITMUS^RT invalidates this conjecture, showing that the QPS offline decisions may have an important influence on run-time performance. In this work, we present an extensive comparison between RUN and QPS, looking at both the offline and the online phases, to highlight their relative strengths and weaknesses. Davide Compagnin, Enrico Mezzetti, Tullio Vardanega |
ECRTS | 2 |
| 2015 | EPC: Extended Path Coverage for Measurement-Based Probabilistic Timing AnalysisabstractMeasurement-based probabilistic timing analysis (MBPTA) computes trustworthy upper bounds to the execution time of software programs. MBPTA has the connotation, typical of measurement-based techniques, that the bounds computed with it only relate to what is observed in actual program traversals, which may not include the effective worst-case phenomena. To overcome this limitation, we propose Extended Path Coverage (EPC), a novel technique that allows extending the representativeness of the bounds computed by MBPTA. We make the observation data probabilistically path-independent by modifying the probability distribution of the observed timing behaviour so as to negatively compensate for any benefits that a basic block may draw from a path leading to it. This enables the derivation of trustworthy upper bounds to the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. Our results confirm that using MBPTA with EPC produces fully trustworthy upper bounds with competitively small overestimation in comparison to state-of-the-art MBPTA techniques. Marco Ziccardi, Enrico Mezzetti, Tullio Vardanega, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 2 |
| 2014 | Putting RUN into Practice: Implementation and EvaluationabstractThe Reduction to UNiprocessor (RUN) algorithm represents an original approach to multiprocessor scheduling that exhibits the prerogatives of both global and partitioned algorithms, without incurring the respective drawbacks. As an interesting trait, RUN promises to reduce the amount of migration interference. However, RUN has also raised some concerns on the complexity and specialization of its run-time support. To the best of our knowledge, no practical implementation and empirical evaluation of RUN have been presented yet, which is rather surprising, given its potential. In this paper we present the first solid implementation of RUN and extensively evaluate its performance against P-EDF and G-EDF, with respect to observed utilization cap, kernel overheads and inter-core interference. Our results show that RUN can be efficiently implemented on top of standard operating system primitives incurring modest overhead and interference, also supporting much higher schedulable utilization than its partitioned and global counterparts. Davide Compagnin, Enrico Mezzetti, Tullio Vardanega |
ECRTS | 2 |
| 2013 | Limited preemptive scheduling of non-independent task setsabstractPreemption is a key factor against architectural coupling in concurrent systems. The whole verification process of real-time systems postulates composability in multiple dimensions, including time. As coupling wrecks composability, the design of real-time systems really needs preemption. However preemption effects complicate feasibility analysis or make it more pessimistic. Hence methods that limit preemptions without affecting feasibility are attractive. State-of-the-art approaches to limited preemption, however, do not treat resource sharing with the importance that it deserves. The placement of non-preemptive regions - and their interactions with shared resources - should not become a design problem, but rather stay as an implementation level feature that does not backtrack to the design space. In this paper we present a refinement to the state-of-the-art limited preemption model that addresses the interaction with resource sharing, and discuss a kernel implementation that uses run-time knowledge to warrant safe and efficient overlaps between critical sections and non-preemptive regions. Experimental results prove the effectiveness of the proposed solution. Andrea Baldovin, Enrico Mezzetti, Tullio Vardanega |
EMSOFT | 2 |
| 2013 | A rapid cache-aware procedure positioning optimization to favor incremental developmentabstractTruly incremental development is a holy grail of verification-intensive software industry. All factors that threaten it should be removed. Cache memories have an intrinsically jittery timing behavior. The WCET variability that this causes wrecks incrementality. This hazard occurs as the WCET bounds of a software system can only be safely determined when its final memory map is known, which only happens at the end of development. Interestingly, the memory layout optimization techniques, originally devised to optimize average- or worst-case cache response time, open some avenue to control the innate dependence of cache behavior on memory layout. The state-of-the-art approaches, though effective to their own goal, are onerous to use and intrinsically iterative, hence arch-enemy of incrementality. As such they do not lend themselves to effective application in real-world industrial development. In this paper, looking at instruction caches, we describe a novel procedure positioning technique that makes it possible to control the memory layout across incremental software releases. Experimental evidence confirms that our approach facilitates early reasoning on the timing behaviour of system increments and also improves cache performance. Enrico Mezzetti, Tullio Vardanega |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2012 | Measurement-Based Probabilistic Timing Analysis for Multi-path ProgramsabstractThe rigorous application of static timing analysis requires a large and costly amount of detail knowledge on the hardware and software components of the system. Probabilistic Timing Analysis has potential for reducing the weight of that demand. In this paper, we present a sound measurement-based probabilistic timing analysis technique based on Extreme Value Theory. In all the experiments made as part of this work, the timing bounds determined by our technique were less than 15% pessimistic in comparison with the tightest possible bounds obtainable with any probabilistic timing analysis technique. As a point of interest to industrial users, our technique also requires a comparatively low number of measurement runs of the program under analysis, less than 650 runs were needed for the benchmarks presented in this paper. Liliana Cucu-Grosjean, Luca Santinelli, Michael Houston, Code Lo, Tullio Vardanega, Leonidas Kosmidis, Jaume Abella 0001, Enrico Mezzetti, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 8 |
| 2010 | Towards a Cache-Aware Development of High Integrity Real-Time SystemsabstractThe job description of caches is to speed up memory accesses in the average case. Their intrinsic unpredictability however can seriously hamper the practicality and trustworthiness of system analysis and validation. In effect, this conflict asks system designers to take side between best average-case performance and maximum assurance, since both can't be had. In this paper we study the I-cache predictability problem from a system-level perspective. We identify some sources of cache-related variability that can be addressed whilst considering the architectural specification of the system and thus at an early stage of development. We discuss an example of what we call a "cache-aware" software architecture and experimentally evaluate its effectiveness on a representative application. Enrico Mezzetti, Tullio Vardanega |
RTCSA | 1 |