VLDB 2026 Research / reviewers in the wild / expert
Francisco J. Cazorla
dblp:52/1629 · also Francisco Javier Cazorla-Almeida
· DBLP profile ↗
172ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0002-3344-376XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 133 · 9 first-author · 25 since 2021Software engineering, systems software and programming languages · 27 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ROSBand: A Bandwidth Regulation Approach on ROS2-Based Systems
Jon Altonaga Puente, Enrico Mezzetti, Irune Agirre, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 5 |
| 2026 | Evaluating quantile regression neural networks for optimizing real-time applications on heterogeneous platformsabstractModern cyber-physical systems increasingly rely on computationally demanding applications, particularly at the edge, where Artificial Intelligence-based algorithms are deployed. To meet these demands, industry trends are shifting towards heterogeneous MultiProcessor Systems on Chip (MPSoCs), which must also satisfy strict real-time and functional safety requirements. A major challenge in such systems is memory contention, where multiple processing units compete for shared memory resources, affecting application performance and the accurate estimation of Worst-Case Execution Times (WCETs). Traditional static analysis becomes impractical as system configurations grow in complexity. This work presents the design of an analysis and optimization framework for real-time systems that re-evaluates WCET estimates based on system configurations to reflect the impact of memory contention on heterogeneous platforms. The proposed method estimates new WCETs using Quantile Regression Neural Networks (QRNNs), which infer memory contention from Event Monitor data. Experimental results reveal that QRNN models must be system-specific for accurate predictions and that memory access patterns significantly affect model generalization. Two strategies are proposed: using generic models for simplicity or task-specific models for higher accuracy. Despite some potential underestimations, QRNNs maintain a strong correlation with actual observed contention, enabling effective worst-case scenario identification. Furthermore, a comparative analysis highlights the superior scalability of the estimation-based approach over empirical measurements, especially in large system optimization processes where performance can be easily enhanced by at least two orders of magnitude, making it a practical solution for real-time system design and analysis. Iosu Gomez, David Fonts, Sergi Vilardell, Unai Díaz-de-Cerio, Juan Maria Rivas, Enrico Mezzetti, J. Javier Gutiérrez, Francisco J. Cazorla |
Future Gener. Comput. Syst. | 8 |
| 2026 | Supporting Timing-related Metrics for Autonomous Driving Frameworks in CyberRTabstractThe provision of increasingly advanced autonomous software functionalities builds on cutting-edge autonomous driving frameworks to enable modular interactions among multiple software components. This approach helps to support functional cause-effect chains from multiple sensors to actuators. The complexity of the (software) component interactions makes it more difficult to ascertain the correctness of the timing behavior of the system. This is so because traditional timing-related metrics like worst-case execution and worst-case response time do not capture the inter-dependency in cause-effect chains between the input sampling time and the time at which computation based on those inputs is performed. Complementary timing-related metrics, such as maximum reaction time and maximum data age have been considered to capture timing requirements, typically with an end-to-end scope, in cause-effect chains. These metrics have been formalized and demonstrated in ROS2-based automotive and autonomous driving setups [ 44 , 46 ]. However, the formalization of those metrics, which is necessary for deriving analytical lower and upper bounds and monitoring them at run-time, largely depends on the execution model and semantics offered by the run-time. Any concrete application of those metrics need to be tailored and adapted to the system at hand. Apollo auto is a popular, industrial-quality, open-source autonomous driving framework that is seeing increasing adoption both for industrial and academic projects. Apollo builds on CyberRT , an ad-hoc run-time that is similar in mechanism and intent to ROS2 but differentiates from it with respect to execution model and supported semantics. In contrast to ROS2, CyberRT is highly specialized to support the Apollo AD framework, is neither extensively documented or thoroughly analysed in the literature, especially in relation to execution model and instantiation of timing-related metrics. In this work, for the first time, we provide an insightful analysis and discussion on CyberRT execution model and semantics, starting from its raw and non-extensively documented codebase. Based on the identified semantics, we elaborate a formalization of timing-related metrics on CyberRT , across different granularity scopes, namely end-to-end and node levels. In particular, we develop on the importance of node-level timing properties to intercept any latent timing misbehavior before it is too late, and it severely impacts end-to-end execution. We provide a concrete mapping of a comprehensive set of timing-related metrics to the CyberRT execution model, both at end-to-end and node level, and develop a monitoring library that allows to intercept them on the specific software stack. We exploit the proposed library on a set of Apollo autonomous driving scenarios to demonstrate its effectiveness in monitoring the considered timing metrics and to promptly intercept a subtle timing misbehavior beyond end-to-end execution scope in a representative autonomous driving stack. Miguel Alcon, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | SAFEXPLAIN: a Complete Approach Towards Trustworthy AI-Based Safety-Critical SystemsabstractAI becomes increasingly important in safetycritical systems, especially in the case of autonomous systems, since navigation relies on AI for object detection and collision avoidance. However, safety-critical systems must adhere to functional safety standards that enforce software to be correct-by-construction, component decomposition to simplify design and validation, and the use of data only for testing purposes not to design the system itself. AI in general, and Deep Learning (DL) in particular have opposed characteristics since they have error rates (e.g., due to mispredictions), AI/DL modules can only be designed and validated monolithically, and they build on data for their design (i.e. for training purposes). Hence, DL solutions are at odds with the development process of safetycritical systems. A number of standards have recently emerged in different domains to reconcile the requirements of safety-critical systems with the characteristics of DL solutions, such as ISO 21448, ISO/IEC TR 5469, and ISO 8800, among others. However, there is a lack of realistic practice to design a DL-based safety-critical system in accordance with those regulations, and existing solutions only cover some aspects in isolation, and are often incompatible among them. SAFEXPLAIN is a 3-year Horizon Europe project addressing this challenge. SAFEXPLAIN, which finishes in September 2025, has already reached its main goals providing specific and complementary solutions to all those challenges so that AIbased safety-critical systems can be designed, implemented and validated adhering to the relevant functional safety standards in domains such as automotive, space and railway. In particular, SAFEXPLAIN provides the concepts, processes, tools and frameworks addressing the challenge end-to-end, from concept to solution. This is proven by the successful application of the SAFEXPLAIN approach in three case studies from the automotive, space and railway domains, whose results will see the light very soon. Jaume Abella 0001, Irune Agirre, Thanh Hai Bui, Frank Geujen, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Axel Brando, Javier Fernández 0004, Irune Yarza, Joanes Plazaola, Maria Ulan, Rob Lavreysen, Lucas Tosi, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Stefano Lodico, William Guarienti, Giuseppe Nicosia, Valeria Dallara |
DSD | 7 |
| 2025 | Impact of Contention-Aware Placement in Heterogeneous Edge DevicesabstractTime predictability is an increasing concern in functionally-rich mixed-criticality applications at the Edge, which often carry different timing requirements. Edge devices, in turn, are increasingly complex to sustain the increasing computational requirements, which hinders providing predictable performance without seriously affecting performance. One of the main threats to predictable performance is the impact of timing interference arising from contention in an increasing number of shared hardware resources. The impact of software to hardware mapping on performance is a well-studied topic, seeking optimal memory mappings to reduce average and worst-case performance, and, more recently, to control and limit timing interference. These methods normally focus on code and data placement, especially in relation to specific properties of the memory hierarchy, either architectural (e.g. heterogeneous memory modules) or obtained through partitioning techniques. In practice, however, these works build on a uniform memory hierarchy model, where the source of a memory request, namely, where a task accessing a given memory is eventually executed, is not directly relevant. In this work, we consider a large class of systems (e.g., TriCore families) where memory hierarchies are non-uniform, and access latency depends on the computing element issuing the request. In those architectures, the impact of code and data placement on timing interference cannot be addressed without considering architectural constraints and task locality. Through empirical exploration, we show that code, data, and locality collectively have a substantial impact on contention bounds, leading to a significantly expanded optimization space compared to approaches considering only code and data placement under the uniform memory assumption. Our results motivate the need for novel, efficient optimization approaches that integrate task mapping and architectural constraints to reduce timing interference. Jeremy Giesen, Ibai Irigoyen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 5 |
| 2025 | Detecting Low-Density Mixtures in High-Quantile Tails for pWCET Estimation
Blau Manau, Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 6 |
| 2025 | Practical do-Shapley Explanations with Estimand-Agnostic Causal InferenceabstractAmong explainability techniques, SHAP stands out as one of the most popular, but often overlooks the causal structure of the problem. In response, do-SHAP employs interventional queries, but its reliance on estimands hinders its practical application. To address this problem, we propose the use of estimand-agnostic approaches, which allow for the estimation of any identifiable query from a single model, making do-SHAP feasible on complex graphs. We also develop a novel algorithm to significantly accelerate its computation at a negligible cost, as well as a method to explain inaccessible Data Generating Processes. We demonstrate the estimation and computational performance of our approach, and validate it on two real-world datasets, highlighting its potential in obtaining reliable explanations. Álvaro Parafita, Tomas Garriga, Axel Brando, Francisco J. Cazorla |
NeurIPS | 4 |
| 2025 | EMR: Removing Multicollinear Event Monitors to Improve Timing Modelling of Real-Time SystemsabstractMulticollinearity of Event Monitors (EMs) negatively impacts the modeling of non-functional critical metrics in real-time systems like worst-case timing and energy usage since some EMs are over-represented and can reduce model accuracy. To address this challenge, we propose Event Monitor Reduction (EMR), a method to select a reduced set of non-related (independent) features (EMs), hence eliminating multicollinearity. In particular, EMR finds linear relations between the EMs and removes dependent ones without data loss. EMR does not create new features like Principal Component Analysis does, simplifying interpretability. Results on synthetic data and data collected from the execution of representative benchmarks on an avionicsgrade processor show the benefits of our method in removing multicollinear EMs. We further illustrate the benefits of EMR on two different multicore timing contention models, showing how its application helps to reduce execution time requirements and increase the accuracy of the models. David Fonts, Diego Palacios, Sergi Vilardell, Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 8 |
| 2025 | Hardware support for contention tracking in CPU and GPU last-level cacheabstractModern MPSoCs increasingly rely on resource utilization to improve application performance with different computation needs. The last-level cache (LLC) is one of the main shared resources, contributing to the improvement of aggregated performance. However, LLC sharing also increases individual application performance variability, which is undesirable in scenarios where performance guarantees are required. While deploying cache partitioning mechanisms allows regaining predictability, they negatively affect aggregated performance. This confronts system designers with the dire conundrum of choosing between aggregated performance and predictability. We contend that adding hardware support to track contention among tasks (kernels) in the LLC enables it to be shared, removing shortcomings brought by partitioning while providing a clear view of how tasks (kernels) affect each other in the LLC of the CPUs and GPUs. This approach enables achieving the desired balance between performance and predictability. Thus, we propose a low-overhead hardware mechanism, called demotion counters (DC), that tightly estimates the contention tasks (kernels) generate on each other in the shared LLC, outperforming other solutions that build on existing hardware contention-tracking proposals which suffer an average workload breakdown deviation (wbd) over 0.13. Our results also show that DC introduces 0.66% area overhead. 1 Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 5 |
| 2024 | TAP: Task-Aware Profiling on Integrated SystemsabstractHardware Performance Monitors (HPM) are increasingly exploited for timing verification and validation of time-critical embedded systems (TECS). HPMs are typically collected at the lowest software level, which makes it difficult to unequivocally account events to specific run-time entities, a prerequisite for any form of analysis, without relying on ad-hoc support from the run-time or operating system layer. The latter, however, is either unavailable or not fully adequate for verification requirements. Moreover, timing-related concerns in the analysis of embedded systems are typically addressed in the final stages of the software development process where multiple tasks are fully or partially integrated on the platform and it is therefore hard, if not impossible, to enforce controlled testing scenarios where contributions to event counts can be dissected. In this work, we present TAP a generic concept for allowing Task-Aware Profiling of individual tasks in an already integrated system on MPSoCs with on-core and off-core HPM support. The proposed approach combines a lightweight user-level configurable API and minimally intrusive extensions to the operating system layer to enforce separation of contexts when collecting HPM. We implement and assess TAP on top of an Infineon AURIX MPSoC and the OSEK-compliant ERIKA Enterpise RTOS, offering a consistent and intuitive interface for governing and filtering the different sources of events. Our results on synthetic and automotive benchmarks show that TAP can transparently gather and filter the events of interest while incurring negligible overheads. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2024 | Event Monitor Validation in High-Integrity SystemsabstractPlatforms for modern embedded systems equip an increasing number of high-performance features to provide the required levels of performance. Timing analysis solutions handle the complexity of these platforms by relying on hardware event monitors (HEMs) that provide insightful information about resource utilization and, hence, contention among tasks. As a result, HEMs have become a key element to warrant a safe timing behavior of a system, for which reason they must be validated. While some initial works target HEMs validation, they consider one HEM at a time and focus on those HEMs for which an expert can establish an expected value for relatively small code snippets. In this paper, we propose a methodology for the validation of those HEMs for which a specific expected value cannot be established a priori even for simple cases and, instead, needs to be validated in conjunction with other HEMs. Our method also deals with the natural variability of the HEMs' values in high-performance platforms when collected in different experiments. We illustrate the effectiveness of our proposed technique for validating HEMs related to cache coherence in a relevant platform in the avionics domain. Roger Pujol, Sergi Vilardell, Enrico Mezzetti, Mohamed Hassan 0002, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 6 |
| 2024 | Achieving Flexible Performance Isolation on the AMD Xilinx Zynq UltraScale+abstractCo-hosting different tasks on the same MPSoC contributes to increasing average performance by allowing them to share MPSoC's resources that, otherwise, could be underutilized. However, resource sharing challenges performance isolation among tasks, as required in time-sensitive embedded critical systems like automotive and avionics. On the other hand, resource isolation through segregation (the reference solution for preventing the propagation of time-related safety issues) is detrimental to average performance. In this work, we show that the built-in QoS support in modern MPSoCs can be smartly leveraged to adapt to the timing and performance requirements of the running applications. In particular, we develop specific configurations of the complex QoS support in the Zynq UltraScale+ MPSoC that deliver performance isolation for time-sensitive tasks (TSTs) and ensure that non-time-sensitive tasks (NTSTs) maximize their average performance by exploiting the resources not used by TSTs. Our results on the Xilinx UtltraScale+ show that the TSTs with the most stringent constraints achieve high degrees of isolation, 96.0% of their solo performance on average, while NTSTs exploit the resources not used by TSTs achieving performance ranging from 72% to 4% depending on the resource left by TSTs. Alejandro Serrano-Cases, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2024 | Safety-Relevant AI-Based System Robustification with Neural Network EnsemblesabstractFunctional safety requirements of AI-based safety critical applications challenge AI models, whose accuracy can be limited. In this paper, we show how using several cooperative deep learning (DL) models helps to raise global accuracy and reject making a prediction when confidence is below a pre-established threshold. Adrià Aldomà, Axel Brando, Francisco J. Cazorla, Jaume Abella 0001 |
IOLTS | 3 |
| 2023 | SAFEXPLAIN: Safe and Explainable Critical Embedded Systems Based on AIabstractDeep Learning (DL) techniques are at the heart of most future advanced software functions in Critical Autonomous AI-based Systems (CAIS), where they also represent a major competitive factor. Hence, the economic success of CAIS industries (e.g., automotive, space, railway) depends on their ability to design, implement, qualify, and certify DL-based software products under bounded effort/cost. However, there is a fundamental gap between Functional Safety (FUSA) requirements on CAIS and the nature of DL solutions. This gap stems from the development process of DL libraries and affects high-level safety concepts such as (1) explainability and traceability, (2) suitability for varying safety requirements, (3) FUSA-compliant implementations, and (4) real-time constraints. As a matter of fact, the data-dependent and stochastic nature of DL algorithms clashes with current FUSA practice, which instead builds on deterministic, verifiable, and pass/fail test-based software. The SAFEXPLAIN project tackles these challenges and targets by providing a flexible approach to allow the certification - hence adoption - of DL-based solutions in CAIS building on: (1) DL solutions that provide end-to-end traceability, with specific approaches to explain whether predictions can be trusted and strategies to reach (and prove) correct operation, in accordance to certification standards; (2) alternative and increasingly sophisticated design safety patterns for DL with varying criticality and fault tolerance requirements; (3) DL library implementations that adhere to safety requirements; and (4) computing platform configurations, to regain determinism, and probabilistic timing analyses, to handle the remaining non-determinism. Jaume Abella 0001, Jon Pérez 0001, Cristofer Englund, Bahram Zonooz, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Isabel Serra, Axel Brando, Irune Agirre, Fernando Eizaguirre, Thanh Hai Bui, Elahe Arani, Fahad Sarfraz, Ajay Balasubramaniam, Ahmed Badar, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Davide Brighenti, Davide Cunial |
DATE | 7 |
| 2023 | Quasi Isolation QoS Setups to Control MPSoC Contention in Integrated Software Architectures
Sergio Garcia-Esteban, Alejandro Serrano-Cases, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 5 |
| 2023 | EUROPULS: NEUROmorphic energy-efficient secure accelerators based on Phase change materials aUgmented siLicon photonicSabstractThis special session paper introduces the Horizon Europe NEUROPULS project, which targets the development of secure and energy-efficient RISC-V interfaced neuromorphic accelerators using augmented silicon photonics technology. Our approach aims to develop an augmented silicon photonics platform, an FPGA-powered RISC-V-connected computing platform, and a complete simulation platform to demonstrate the neuromorphic accelerator capabilities. In particular, their main advantages and limitations will be addressed concerning the underpinning technology for each platform. Then, we will discuss three targeted use cases for edge-computing applications: Global National Satellite System (GNSS) anti-jamming, autonomous driving, and anomaly detection in edge devices. Finally, we will address the reliability and security aspects of the stand-alone accelerator implementation and the project use cases. Fabio Pavanello, Cédric Marchand 0002, Ian O'Connor, Régis Orobtchouk, Fabien Mandorlo, Xavier Letartre, Sébastien Cueff, Elena I. Vatajelu, Giorgio Di Natale, Benoit Cluzel, Aurelien Coillet, Benoît Charbonnier, Pierre Noe, Frantisek Kavan, Martin Zoldak, Michal Szaj, Peter Bienstman, Thomas Van Vaerenbergh, Ulrich Rührmair, Paulo F. Flores, Luís Guerra e Silva, Ricardo Chaves, Luís Miguel Silveira, Mariano Ceccato, Dimitris Gizopoulos, George Papadimitriou 0001, Vasileios Karakostas, Axel Brando, Francisco J. Cazorla, Ramon Canal, Pau Closas, Adria Gusi-Amigo, Paolo Crovetti, Alessio Carpegna, Tzamn Melendez Carmona, Stefano Di Carlo, Alessandro Savino 0001 |
ETS | 29 |
| 2023 | SafeLS: An Open Source Implementation of a Lockstep NOEL-V RISC-V CoreabstractMicrocontrollers running safety-critical applications with high integrity requirements must provide appropriate safety measures to manage random hardware faults. For instance, automotive safety regulations (e.g., ISO26262) impose the use of diverse redundancy for items at the highest automotive safety integrity level (ASIL), ASIL-D. In the case of computing cores, this is realized with dual core lockstep (DCLS). The advent of the RISC-VISA has made open source hardware gain popularity. However, there are few industrial open source SoCs meeting the requirements of safety-critical systems, and, to our knowledge, none of them provides lockstep cores. This paper presents the realization of a RISC-V open source lockstep core based on Gaisler's NOEL-V core for the space domain, as well as its integration in the SELENE SoC that provides a complete microcontroller synthesizable on FPGA successfully assessed against space, automotive and railway safety-critical applications in the past. Marcel Sarraseca, Sergi Alcaide, Francisco Fuentes, Juan Carlos Rodriguez, Feng Chang, Ilham Lasfar, Ramon Canal, Francisco J. Cazorla, Jaume Abella 0001 |
IOLTS | 8 |
| 2023 | Improving Timing-Related Guarantees for Main Memory in Multicore Critical Embedded SystemsabstractMain memory is one of the most complex resources to analyze in multicore-based embedded real-time systems, with contention in the memory controller and the timing constraints of the main memory device as the main contributors to that complexity. One of the main challenges in multicore real-time systems is producing the required evidence on the management of contention delay for the certification. This stems from the fact that current MPSoCs barely provide any event monitors on how tasks interact and delay each other in memory. Besides, even if hardware and software mechanisms are in place to mitigate contention in the memory system, it is hard - if at all possible - to provide evidence about their correctness. In this work, we cover this gap by proposing a lightweight hardware mechanism that tightly tracks inter-core contention in memory. The proposed hardware mechanism, which we evaluate in detail, improves the quality of timing-related evidence that must be provided on how contention in main memory of multicore real-time systems is handled in adherence to applicable safety standards. Asier Fernández de Lecea, Mohamed Hassan 0002, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 5 |
| 2023 | An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt |
J. Syst. Archit. | 9 |
| 2023 | Main sources of variability and non-determinism in AD software: taxonomy and prospects to handle them
Miguel Alcon, Axel Brando, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
Real Time Syst. | 5 |
| 2023 | Dynamic and execution views to improve validation, testing, and optimization of autonomous driving software
Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
Softw. Qual. J. | 4 |
| 2023 | Vector Extensions in COTS Processors to Increase Guaranteed Performance in Real-Time SystemsabstractThe need for increased application performance in high-integrity systems such as those in avionics is on the rise as software continues to implement more complex functionalities. The prevalent computing solution for future high-integrity embedded products is multi-processor systems-on-chip (MPSoC) processors. MPSoCs include central processing unit (CPU) multicores that enable improving performance via thread-level parallelism. MPSoCs also include generic accelerators (graphics processing units [GPUs]) and application-specific accelerators. However, the data processing approach (DPA) required to exploit each of these underlying parallel hardware blocks carries several open challenges to enable the safe deployment in high-integrity domains. The main challenges include the qualification of its associated runtime system and the difficulties in analyzing programs deploying the DPA with out-of-the-box timing analysis and code coverage tools. In this work, we perform a thorough analysis of vector extensions (VExts) in current commercial off-the-shelf (COTS) processors for high-integrity systems. We show that VExts prevent many of the challenges arising with parallel programming models and GPUs. Unlike other DPAs, VExts require no runtime support, prevent design race conditions that might arise with parallel programming models, and have minimum impact on the software ecosystem, enabling the use of existing code coverage and timing analysis tools. We develop vectorized versions of neural network kernels and show that the NVIDIA Xavier VExts provide a reasonable increase in guaranteed application performance of up to 2.7x. Our analysis contends that VExts are the DPA approach with arguably the fastest path for adoption in high-integrity systems. Roger Pujol, Josep Jorba 0002, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2023 | Accurately Measuring Contention in Mesh NoCs in Time-Sensitive Embedded SystemsabstractThe computing capacity demanded by embedded systems is on the rise as software implements more functionalities, ranging from best-effort entertainment functions to performance-guaranteed safety-related functions. Heterogeneous manycore processors, using wormhole mesh (wmesh) Network-on-Chips (NoCs) as the main communication means, and contention block among applications, are increasingly considered to deliver the required computing performance. Most research efforts on software timing analysis have focused on deriving bounds (estimates) to the contention that tasks can suffer when accessing wmesh NoCs. However, less effort has been devoted to an equally important problem, namely,accuratelymeasuring the actual contention tasks generate each other on the wmesh which is instrumental during system validation to diagnose any software timing misbehavior and determine which tasks are particularly affected by contention on specific wmesh routers. In this article, we work on the foundations ofcontention measuringin wmesh NoCs and propose and explain the rationale of agolden metric, called taskPairWise Contention(PWC). PWC allows ascribing the actual share of the contention a given task suffers in the wmesh to each of its co-runner tasks at packet level. We also introduce and formalize aGolden Reference Value(GRV) for PWC that specifically defines a criterion to fairly break down the contention suffered by a task among its co-runner tasks in the wmesh. Our evaluation shows that GRV effectively captures how contention occurs by identifying the actual core (task) causing contention and whether contention is caused by local or remote interference in the wmesh. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | Using Quantile Regression in Neural Networks for Contention Prediction in Multicore ProcessorsabstractMachine learning has enabled significant benefits in diverse fields, but, with a few exceptions, has had limited impact on computer architecture. Recent work, however, has explored broader applicability for design, optimization, and simulation. Notably, machine learning based strategies often surpass prior state-of-the-art analytical, heuristic, and human-expert approaches. This paper reviews machine learning applied system-wide to simulation and run-time optimization, and in many individual components, including memory systems, branch predictors, networks-on-chip, and GPUs. The paper further analyzes current practice to highlight useful design strategies and identify areas for future work, based on optimized implementation strategies, opportune extensions to existing work, and ambitious long term possibilities. Taken together, these strategies and techniques present a promising future for increasingly automated architectural design. Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 5 |
| 2022 | Using Markov's Inequality with Power-Of-k Function for Probabilistic WCET Estimation
Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Joan del Castillo |
ECRTS | 5 |
| 2022 | Contention Tracking in GPU Last-Level CacheabstractThe Last-level cache (LLC) is one of the main GPU’s shared resources that contributes to improve performance but also increases individual kernel’s performance variability. This is detrimental in scenarios in which some level of performance predictability is required. While predictability can be regained by deploying cache partitioning (isolation) mechanisms, isolation negatively affects performance efficiency. This work shows that not partitioning the LLC and providing the ability to track the contention that kernels generate on each other allows them to share LLC space, hence increasing efficiency, while the system designer obtains a clear view of how each kernel affects each other in the LLC so as to balance performance and predictability goals. In this line, we propose GPU demotion counters (GDC), a low-overhead hardware mechanism to track contention that kernels generate on each other in the shared LLC. Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
ICCD | 5 |
| 2021 | Empirical Evidence for MPSoCs in Critical Systems: The Case of NXP's T2080 Cache CoherenceabstractThe adoption of complex MPSoCs in critical realtime embedded systems mandates a detailed analysis of their architecture to facilitate certification. This analysis is hindered by the lack of a thorough understanding of the MPSoC system due to the unobvious and/or insufficiently documented behavior of some key hardware features. Confidence in those features can only be regained by building specific tests to both, assess whether their behavior matches specifications and unveil their behavior when it is not fully known a priori. In this line, in this work we develop a thorough understanding of the cache coherence protocol in the avionics-relevant NXP T2080 architecture. Roger Pujol, Hamid Tabani, Jaume Abella 0001, Mohamed Hassan 0002, Francisco J. Cazorla |
DATE | 5 |
| 2021 | Enabling Unit Testing of Already-Integrated AI Software Systems: The Case of Apollo for Autonomous DrivingabstractThe advanced AI-based software used for autonomous driving comprises multiple highly-coupled modules that are data and control dependent. Deploying those already-integrated software frameworks makes unit testing, a fundamental step in the validation process of critical software, very challenging in safety-critical systems. To tackle this issue, in this paper, we show the steps we followed to develop standalone versions of the modules in an industry-level autonomous driving framework (Apollo) by applying several modifications to its architectural design. We show how the standalone modules have the same functional behavior as their integrated counterpart modules. We exemplify the benefits of standalone modules by performing incremental analysis of the software timing requirements of each module running on a heterogeneous System on Chip (SoC). This is a mandatory step to consolidate and integrate software modules guaranteeing timing constraints (e.g. related to freedom from interference) while maximizing SoC utilization. Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2021 | Leveraging Hardware QoS to Control Contention in the Xilinx Zynq UltraScale+ MPSoCabstractThe interference co-running tasks generate on each other’s timing behavior continues to be one of the main challenges to be addressed before Multi-Processor System-on-Chip (MPSoCs) are fully embraced in critical systems like those deployed in avionics and automotive domains. Modern MPSoCs like the Xilinx Zynq UltraScale+ incorporate hardware Quality of Service (QoS) mechanisms that can help controlling contention among tasks. Given the distributed nature of modern MPSoCs, the route a request follows from its source (usually a compute element like a CPU) to its target (usually a memory) crosses several QoS points, each one potentially implementing a different QoS mechanism. Mastering QoS mechanisms individually, as well as their combined operation, is pivotal to obtain the expected benefits from the QoS support. In this work, we perform, to our knowledge, the first qualitative and quantitative analysis of the distributed QoS mechanisms in the Xilinx UltraScale+ MPSoC. We empirically derive QoS information not covered by the technical documentation, and show limitations and benefits of the available QoS support. To that end, we use a case study building on neural network kernels commonly used in autonomous systems in different real-time domains. Alejandro Serrano-Cases, Juan M. Reina, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 5 |
| 2021 | Demystifying the Characteristics of High Bandwidth Memory for Real-Time SystemsabstractThe number of functionalities controlled by software on every critical real-time product is on the rise in domains like automotive, avionics and space. To implement these advanced functionalities, software applications increasingly adopt artificial intelligence algorithms that manage massive amounts of data transmitted from various sensors. This translates into unprecedented memory performance requirements in critical systems that the commonly used DRAM memories struggle to provide. High-Bandwidth Memory (HBM) can satisfy these requirements offering high bandwidth, low power and high-integration capacity features. However, it remains unclear whether the predictability and isolation properties of HBM are compatible with the requirements of critical embedded systems. In this work, we perform to our knowledge the first timing analysis of HBM. We show the unique structural and timing characteristics of HBM with respect to DRAM memories and how they can be exploited for better time predictability, with emphasis on increased isolation among tasks and reduced worst-case memory latency. Kazi Asifuzzaman, Mohamed Abuelala, Mohamed Hassan 0002, Francisco J. Cazorla |
ICCAD | 4 |
| 2021 | PRL: Standardizing Performance Monitoring Library for High-Integrity Real-Time SystemsabstractThe use of complex processors is becoming ubiquitous in High-Integrity Systems (HIS). To deal with processor’s increased complexity, Performance Monitoring Counters (PMCs) are increasingly used to reason on software behavior and provide the necessary evidence to support software certification. However, the use of PMCs in HIS is relatively recent and hence far from being standardized. As a result, software engineers are forced to resort to highly-customized, low-level programming of platform-specific PMC control registers, which is both error prone and time consuming. To cover this gap, we propose building on the PAPI library, a standardized performance monitoring solution in the mainstream domain, and develop a PMC Reading Library (PRL) for configuring and collecting traceable events while capturing HIS specific requirements and peculiarities. We instantiate PRL in a reference automotive configuration to show that PRL meets key HIS requirements: negligible footprint, limited and predictable overhead, and accuracy collecting hardware events by filtering out the impact of interrupts and context switches. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ICCD | 4 |
| 2021 | Performance Analysis and Optimization Opportunities for NVIDIA Automotive GPUs
Hamid Tabani, Fabio Mazzocchetti, Pedro Benedicte, Jaume Abella 0001, Francisco J. Cazorla |
J. Parallel Distributed Comput. | 5 |
| 2021 | Towards functional safety compliance of matrix-matrix multiplication for machine learning-based autonomous systems
Javier Fernández 0004, Jon Pérez 0001, Irune Agirre, Imanol Allende, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 6 |
| 2021 | Worst-Case Energy Consumption: A New Challenge for Battery-Powered Critical DevicesabstractThe number of (edge) devices connected to the IoT is on the rise, reaching hundreds of billions in the next years. Many devices will implement some type of critical functionality, for instance in the medical market this includes infusion pumps and implantable defibrillators. Energy awareness is mandatory in the design of IoT devices given their huge impact on worldwide energy consumption and the fact that many of them are battery powered. Critical IoT devices further require addressing new energy-related challenges. On the one hand, factoring in the impact of energy-solutions on device's performance, providing evidence of adherence to domain-specific safety standards. On the other hand, deriving safe worst-case energy consumption (WCEC) estimates is fundamental to ensure the system can continuously operate under a pre-established set of power/energy caps, safely delivering its critical functionality. In this line, we analyze for the first time the impact that different hardware physical parameters have on both model-based and measurement-based WCEC modeling, for which we also show the main challenges they face compared to chip manufacturers' current practice for energy modeling and validation. Under the set of constraints that emanate from how certain physical parameters can be actually modeled, we show that measurement-based WCEC is a promising way forward for WCEC estimation. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Sustain. Comput. | 4 |
| 2020 | An On-board Algorithm Implementation on an Embedded GPU: A Space Case StudyabstractOn-board processing requirements of future space missions are constantly increasing, calling for new hardware than the traditional ones used in space. Embedded GPUs are an attractive candidate offering both high performance capabilities and low power consumption, but there are no complex industrial case studies from the space domain demonstrating these advantages. In this paper we present the GPU parallelisation of an on-board algorithm, as well as its performance on a promising embedded GPU COTS platform targeting critical systems. Iván Rodriguez, Leonidas Kosmidis, Olivier Notebaert, Francisco J. Cazorla, David Steenari |
DATE | 4 |
| 2020 | UP2DATE: Safe and secure over-the-air software updates on high-performance mixed-criticality systemsabstractFollowing the same trend of consumer electronics, safety-critical industries are starting to adopt Over-The-Air Software Updates (OTASU) on their embedded systems. The motivation behind this trend is twofold. On the one hand, OTASU offer several benefits to the product makers and users by improving or adding new functionality and services to the product without a complete redesign. On the other hand, the increasing connectivity trend makes OTASU a crucial cyber-security demand to download latest security patches. However, the application of OTASU in the safety-critical domain is not free of challenges, specially when considering the dramatic increase of software complexity and the resulting high computing performance demands. This is the mission of UP2DATE, a recently launched project funded within the European H2020 programme focused on new software update architectures for heterogeneous high-performance mixed-criticality systems. This paper gives an overview of UP2DATE and its foundations, which seeks to improve existing OTASU solutions by considering safety, security and availability from the ground up in an architecture that builds around composability and modularity. Irune Agirre, Peio Onaindia, Tomaso Poggi, Irune Yarza, Francisco J. Cazorla, Leonidas Kosmidis, Kim Grüttner, Mohammed Abuteir, Jan Loewe, Juan M. Orbegozo, Stefania Botta |
DSD | 5 |
| 2020 | Tracing Hardware Monitors in the GR712RC Multicore Platform: Challenges and Lessons Learnt from a Space Case StudyabstractThe demand for increased computing performance is driving industry in critical-embedded systems (CES) domains, e.g. space, towards the use of multicores processors. Multicores, however, pose several challenges that must be addressed before their safe adoption in critical embedded domains. One of the prominent challenges is software timing analysis, a fundamental step in the verification and validation process. Monitoring and profiling solutions, traditionally used for debugging and optimization, are increasingly exploited for software timing in multicores. In particular, hardware event monitors related to requests to shared hardware resources are building block to assess and restraining multicore interference. Modern timing analysis techniques build on event monitors to track and control the contention tasks can generate each other in a multicore platform. In this paper we look into the hardware profiling problem from an industrial perspective and address both methodological and practical problems when monitoring a multicore application. We assess pros and cons of several profiling and tracing solutions, showing that several aspects need to be taken into account while considering the appropriate mechanism to collect and extract the profiling information from a multicore COTS platform. We address the profiling problem on a representative COTS platform for the aerospace domain to find that the availability of directly-accessible hardware counters is not a given, and it may be necessary to the develop specific tools that capture the needs of both the user’s and the timing analysis technique requirements. We report challenges in developing an event monitor tracing tool that works for bare-metal and RTEMS configurations and show the accuracy of the developed tool-set in profiling a real aerospace application. We also show how the profiling tools can be exploited, together with handcrafted benchmarks, to characterize the application behavior in terms of multicore timing interference. Xavier Palomo, Mikel Fernández, Sylvain Girbal, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Laurent Rioux |
ECRTS | 6 |
| 2020 | A Cross-Layer Review of Deep Learning Frameworks to Ease Their Optimization and ReuseabstractMachine learning and especially Deep Learning (DL) approaches are at the heart of many domains, from computer vision and speech processing to predicting trajectories in autonomous driving and data science. Those approaches mainly build upon Neural Networks (NNs), which are compute-intensive in nature. A plethora of frameworks, libraries and platforms have been deployed for the implementation of those NNs, but end users often lack guidance on what frameworks, platforms and libraries to use to obtain the best implementation for their particular needs. This paper analyzes the DL ecosystem providing a structured view of some of the main frameworks, platforms and libraries for DL implementation. We show how those DL applications build ultimately on some form of linear algebra operations such as matrix multiplication, vector addition, dot product and the like. This analysis allows understanding how optimizations of specific linear algebra functions for specific platforms can be effectively leveraged to maximize specific targets (e.g. performance or power-efficiency) at application level reusing components across frameworks and domains. Hamid Tabani, Roger Pujol, Jaume Abella 0001, Francisco J. Cazorla |
ISORC | 4 |
| 2020 | Timing of Autonomous Driving Software: Problem Analysis and Prospects for Future SolutionsabstractThe software used to implement advanced functionalities in critical domains (e.g. autonomous operation) impairs software timing. This is not only due to the complexity of the underlying high-performance hardware deployed to provide the required levels of computing performance, but also due to the complexity, non-deterministic nature, and huge input space of the artificial intelligence (AI) algorithms used. In this paper, we focus on Apollo, an industrial-quality Autonomous Driving (AD) software framework: we statistically characterize its observed execution time variability and reason on the sources behind it. We discuss the main challenges and limitations in finding a satisfactory software timing analysis solution for Apollo and also show the main traits for the acceptability of statistical timing analysis techniques as a feasible path. While providing a consolidated solution for the software timing analysis of Apollo is a huge effort far beyond the scope of a single research paper, our work aims to set the basis for future and more elaborated techniques for the timing analysis of AD software. Miguel Alcon, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 6 |
| 2020 | Modeling Contention Interference in Crossbar-based Systems via Sequence-Aware Pairing (SeAP)abstractThe Infineon AURIX TriCore family of microcontrollers has consolidated as the reference multicore computing platform for safety-critical systems in the automotive domain. As a distinctive trait, AURIX microcontrollers are designed to promote high timing predictability as witnessed by the presence of large scratchpad memories and a crossbar interconnect. The latter has been introduced to reduce inter-core interference in accessing the memory system and peripherals. Nonetheless, the crossbar does not prevent requests from different cores to the same target resource to suffer contention. Applications are, therefore, inherently exposed to inter-core timing interference, which needs to be taken into account in the determination of reliable execution time bounds. In this paper we propose a contention modeling technique for crossbar-based systems, and hence suitable for bounding contention effects in the AURIX family. Unlike state of the art techniques that build on total request counts, we exploit the sequence of requests to the different target resources produced by each core to produce tighter bounds by discarding contention scenarios that cannot occur in practice. To that end, we adapt existing techniques from the pattern matching domain to derive the worst-case contention effects from the sequences of requests each core sends over the crossbar. Results on a wide set of synthetic and real scenarios and benchmark on an AURIX TC297TX show that our technique outperforms other contention modeling approaches. Jeremy Giesen, Pedro Benedicte, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 5 |
| 2020 | HRM: Merging Hardware Event Monitors for Improved Timing Analysis of Complex MPSoCsabstractThe performance monitoring unit (PMU) in multiprocessor system-on-chips (MPSoCs) is at the heart of the latest measurement-based timing analysis techniques in critical embedded systems. In particular, hardware event monitors (HEMs) in the PMU are used as building blocks in the process of budgeting and verifying software timing by tracking and controlling access counts to shared resources. While the number of HEMs in current MPSoCs reaches hundreds, they are read via performance monitoring counters whose number is usually limited to 4-8, thus requiring multiple runs of each experiment in order to collect all desired HEMs. Despite the effort of engineers in controlling the execution conditions of each experiment, the complexity of current MPSoCs makes it arguably impossible to completely remove the noise affecting each run. As a result, HEMs read in different runs are subject to different variability, and hence, those HEMs captured in different runs cannot be “blindly” merged. In this work, we focus on the NXP T2080 platform where we observed up to 59% variability across different runs of the same experiment for some relevant HEMs (e.g., processor cycles). We develop a HEM reading and merging (HRM) approach to join reliably HEMs across different runs as a fundamental element of any measurement-based timing budgeting and verification technique. Our method builds on order statistics and the selection of an anchor HEM read in all runs to derive the most plausible combination of HEM readings that keep the distribution of each HEM and their relationship with the anchor HEM intact. Sergi Vilardell, Isabel Serra, Roberto Santalla, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | GMAI: Understanding and Exploiting the Internals of GPU Resource Allocation in Critical SystemsabstractCritical real-time systems require strict resource provisioning in terms of memory and timing. The constant need for higher performance in these systems has led industry to recently include GPUs. However, GPU software ecosystems are by their nature closed source, forcing system engineers to consider them as black boxes, complicating resource provisioning. In this work, we reverse engineer the internal operations of the GPU system software to increase the understanding of their observed behaviour and how resources are internally managed. We present our methodology that is incorporated in GMAI (GPU Memory Allocation Inspector), a tool that allows system engineers to accurately determine the exact amount of resources required by their critical systems, avoiding underprovisioning. We first apply our methodology on a wide range of GPU hardware from different vendors showing its generality in obtaining the properties of the GPU memory allocators. Next, we demonstrate the benefits of such knowledge in resource provisioning of two case studies from the automotive domain, where the actual memory consumption is up to 5.6× more than the memory requested by the application. Alejandro J. Calderón, Leonidas Kosmidis, Carlos F. Nicolás, Francisco J. Cazorla, Peio Onaindia |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | Towards limiting the impact of timing anomalies in complex real-time processorsabstractTiming verification of embedded critical real-time systems is hindered by complex designs. Timing anomalies, deeply analyzed in static timing analysis, require specific solutions to bound their impact. For the first time, we study the concept and impact of timing anomalies in measurement-based timing analysis, the most used in industry, showing that they require to be considered and handled differently. In addition, we analyze anomalies in the context of Measurement-Based Probabilistic Timing Analysis, which simplifies quantifying their impact. Pedro Benedicte, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Francisco J. Cazorla |
ASP-DAC | 5 |
| 2019 | Assessing the Adherence of an Industrial Autonomous Driving Framework to ISO 26262 Software GuidelinesabstractThe complexity and size of Autonomous Driving (AD) software are comparably higher than that of software implementing other (standard) functionalities in the car. To make things worse, a big fraction of AD software is not specifically designed for the automotive (or any other critical) domain, but the mainstream market. This brings uncertainty on to which extent AD software adheres to guidelines in safety standards. In this paper, we present our experience in applying ISO 26262 -- the applicable functional safety standard for road vehicles -- software safety guidelines to industrial AD software, in particular, Apollo, a heterogeneous Autonomous Driving framework used extensively in industry. We provide quantitative and qualitative metrics of compliance for many ISO 26262 recommendations on software design, implementation, and testing. Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla, Guillem Bernat |
DAC | 4 |
| 2019 | LAEC: Look-Ahead Error Correction Codes in Embedded Processors L1 Data CacheabstractAs implementation technology shrinks, the presence of errors in cache memories is becoming an increasing issue in all computing domains. Critical systems, e.g. space and automotive, are specially exposed and susceptible to reliability issues. Furthermore, hardware designs in these systems are migrating to multilevel cache multicore systems, in which write-through first level data (DL1) caches have been shown to heavily harm average and guaranteed performance. While write-back DL1 caches solve this problem they come with their own challenges: they need Error Correction Codes (ECC) to tolerate soft errors, but implementing DL1 ECC in simple embedded micro-controllers requires either complex hardware to squash instructions consuming erroneous data, or delayed delivery of data to correct potential errors, which impacts performance even if such process is pipelined. In this paper we present a low-complexity hardware mechanism to anticipate data fetch and error correction in DL1 so that both (1) correct data is always delivered, but (2) avoiding additional delays in most of the cases. This achieves both high guaranteed performance and an effective solutions against errors. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 4 |
| 2019 | Maximum-Contention Control Unit (MCCU): Resource Access Count and Contention Time EnforcementabstractIn real-time systems, the techniques to derive bounds to the contention tasks can suffer in multicore build on resource quota monitoring and enforcement. Existing techniques track and bound the number of requests to hardware shared resources that each core (task) is allowed to perform. In this paper we show that current software-only solutions work well when there is a single resource and type of request to track and bound, but do not scale to the more general case of several shared resources that accept different request types, each with a different associated latency. To handle this (more general) case, we propose low-overhead hardware support called Maximum-Contention Control Unit (MCCU). The MCCU performs fine-grain tracking of different types of requests, preventing a core to cause more interference on its contenders than budgeted. In this process, the MCCU also helps verifying that individual requests duration does not exceed their theoretical bounds, hence dealing with scenarios in which requests can have an arbitrarily large duration. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 4 |
| 2019 | Multicore Early Design Stage Guaranteed Performance Estimates for the Space DomainabstractThe ability to produce early guaranteed performance (worst-case execution time) estimates for multicores, i.e. before software from different providers gets integrated onto the same critical system, is pivotal. This helps reducing lately-detected costly-to-handle timing violations. An existing methodology creates `copy' (surrogate) applications from the execution in isolation of each target application. Surrogate applications can be used to upperbound multicore contention delay, and hence WCET estimates in multicores. However, this methodology has only been shown to work on a simulation environment. In this paper we show the work we have carried out to adapt this technology to a real multicore processor for the space domain. Mikel Fernández, Gabriel Fernandez 0002, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 4 |
| 2019 | AURIX TC277 Multicore Contention Model Integration for Automotive ApplicationsabstractEmbedded systems industry needs reliable and tight worst-case execution time (WCET) estimates for critical applications running on multicores, as a prerequisite to their adoption. While industry already uses reliable tools for single-core WCET estimation and several multicore contention models (MCMs) have been proposed, their combination have not been shown to be fully compatible with the automotive industrial practice yet. This paper reduces this gap by presenting a framework for the integration of MCMs into industrial WCET estimation practice. We illustrate such integration for a Magneti Marelli powertrain control unit on an Infineon AURIX TC277 multicore platform. Enrico Mezzetti, Luca Barbina, Jaume Abella 0001, Stefania Botta, Francisco J. Cazorla |
DATE | 5 |
| 2019 | GPU4S: Embedded GPUs in SpaceabstractFollowing the same trend of automotive and avionics, the space domain is witnessing an increase in the on-board computing performance demands. This raise in performance needs comes from both control and payload parts of the spacecraft and calls for advanced electronics able to provide high computational power under the constraints of the harsh space environment. On the non-technical side, for strategic reasons it is mandatory to get European independence on the used computing technology. In this project, which is still in its early phases, we study the applicability of embedded GPUs in space, which have shown a dramatic improvement of their performance per-watt ratio coming from their proliferation in consumer markets based on competitive European technology. To that end, we perform an analysis of the existing space application domains to identify which software domains can benefit from their use. Moreover, we survey the embedded GPU domain in order to assess whether embedded GPUs can provide the required computational power and identify the challenges which need to be addressed for their adoption in space. In this paper, we describe the steps to be followed in the project, as well as the results of our preliminary analyses in the first months of the project. Leonidas Kosmidis, Jérôme Lachaize, Jaume Abella 0001, Olivier Notebaert, Francisco J. Cazorla, David Steenari |
DSD | 5 |
| 2019 | An Approach for Detecting Power Peaks During Testing and Breaking Systematic Pathological BehaviorabstractThe verification and validation process of embedded critical systems requires providing evidence of their functional correctness and also that their non-functional behavior stays within limits. In this work, we focus on power peaks, which may cause voltage droops and thus, challenge performance to preserve correct operation upon droops. In this line, the use of complex software and hardware in critical embedded systems jeopardizes the confidence that can be placed on the tests carried out during the campaigns performed at analysis. This is so because it is unknown whether tests have triggered the highest power peaks that can occur during operation and whether any such peak can occur systematically. In this paper we propose the use of randomization, already used for timing analysis of real-time systems, as an enabler to guarantee that (1) tests expose those peaks that can arise during operation and (2) peaks cannot occur systematically inadvertently. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2019 | Modeling the Impact of Process Variations in Worst-Case Energy Consumption EstimationabstractThe advent of autonomous power-limited systems poses a new challenge for system verification. Powerful processors needed to enable autonomous operation, are typically power-hungry, jeopardizing battery duration. Therefore, guaranteeing a given battery duration requires worst-case energy consumption (WCEC) estimation for tasks running on those systems. Unfortunately, processor energy and power can suffer significant variation across different units due to process variation (PV), i.e. variability in the electrical properties of transistors and wires due to imperfect manufacturing, which challenges existing WCEC estimation methods for applications. In this paper, we propose a statistical modeling approach to capture PV impact on applications energy and a methodology to compute their WCEC capturing PV, as required to deploy portable critical devices. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2019 | Generating and Exploiting Deep Learning Variants to Increase Heterogeneous Resource Utilization in the NVIDIA XavierabstractDeep learning-based solutions and, in particular, deep neural networks (DNNs) are at the heart of several functionalities in critical-real time embedded systems (CRTES) from vision-based perception (object detection and tracking) systems to trajectory planning. As a result, several DNN instances simultaneously run at any time on the same computing platform. However, while modern GPUs offer a variety of computing elements (e.g. CPUs, GPUs, and specific accelerators) in which those DNN tasks can be executed depending on their computational requirements and temporal constraints, current DNNs are mainly programmed to exploit one of them, namely, regular cores in the GPU. This creates resource imbalance and under-utilization of GPU resources when executing several DNN instances, causing an increase in DNN tasks' execution time requirements. In this paper, (a) we develop different variants (implementations) of well-known DNN libraries used in the Apollo Autonomous Driving (AD) software for each of the computing elements of the latest NVIDIA Xavier SoC. Each variant can be configured to balance resource requirements and performance: the regular CPU core implementation that can run on 2, 4, and 6 cores; the GPU regular and Tensor core variants that can run in 4 or 8 GPU’s Streaming Multiprocessors (SM); and 1 or 2 NVIDIA’s Deep Learning Accelerators (NVDLA); (b) we show that each particular variant/configuration offers a different resource utilization/performance point; finally, (c) we show how those heterogeneous computing elements can be exploited by a static scheduler to sustain the execution of multiple and diverse DNN variants on the same platform. Roger Pujol, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 6 |
| 2019 | Understanding and Exploiting the Internals of GPU Resource Allocation for Critical SystemsabstractCritical real-time systems require strict resource provisioning in terms of memory and timing. The constant need for higher performance in these systems has led industry to recently include GPUs. However, GPU software ecosystems are by their nature closed source, forcing system engineers to consider them as black boxes, complicating resource provisioning. In this work we reverse engineer the internal operations of the GPU system software to increase the understanding of their observed behaviour and how resources are internally managed. This way, we allow system engineers to accurately determine the exact amount of resources required by their critical systems, avoiding underprovisioning. We first apply our methodology on a wide range of GPU hardware showing its generality in obtaining the properties of the GPU memory allocators. Next, we demonstrate the benefits of such knowledge in resource provisioning of two case studies from the automotive domain, where the actual memory consumption is up to 5.6 × more than the memory requested by the application. Alejandro J. Calderón, Leonidas Kosmidis, Carlos F. Nicolás, Francisco J. Cazorla, Peio Onaindia |
ICCAD | 4 |
| 2019 | Software Timing Analysis for Complex Hardware with Survivability and Risk AnalysisabstractThe increasing automation of safety-critical real-time systems, such as those in cars and planes, leads, to more complex and performance-demanding on-board software and the subsequent adoption of multicores and accelerators. This causes software's execution time dispersion to increase due to variable-latency resources such as caches, NoCs, advanced memory controllers and the like. Statistical analysis has been proposed to model the Worst-Case Execution Time (WCET) of software running such complex systems by providing reliable probabilistic WCET (pWCET) estimates. However, statistical models used so far, which are based on risk analysis, are overly pessimistic by construction. In this paper we prove that statistical survivability and risk analyses are equivalent in terms of tail analysis and, building upon survivability analysis theory, we show that Weibull tail models can be used to estimate pWCET distributions reliably and tightly. In particular, our methodology proves the correctness-by-construction of the approach, and our evaluation provides evidence about the tightness of the pWCET estimates obtained, which allow decreasing them reliably by 40% for a railway case study w.r.t. state-of-the-art exponential tails. Sergi Vilardell, Isabel Serra, Jaume Abella 0001, Joan del Castillo, Francisco J. Cazorla |
ICCD | 5 |
| 2019 | Accurate ILP-Based Contention Modeling on Statically Scheduled Multicore SystemsabstractCommercially available Off The Shelf (COTS) multicores have been assessed as the baseline computing platform even in the most conservative real-time domains. Multicore contention arising on shared hardware resources, with its circular dependence with scheduling, is among the most challenging issues that require urgent attention before multicores can be fully embraced for real-time computing. In the context of static scheduling, still the most used scheduling approach in real-time industries, we propose an ILP formulation for computing the worst-case contention delay suffered by a task due to interference on a shared bus. Our model provides accurate contention delay bounds that avoid unnecessary over-accounting of conflicts between bus requests, by considering contention effects at system-level (i.e., across tasks) rather than at task-level only. This allows precisely capturing the interdependence between timing interference of conflicting requests, issued in parallel by other cores (tasks), and the identification of the particular set of tasks co-running on those cores. We assess our technique both analytically and empirically on a real COTS multicore platform. We show, via extensive evaluation, that jointly accounting for worst-case task overlapping and request distribution scenarios always provides tighter contention bounds when compared to state-of-the-art solutions. Xavier Palomo, Enrico Mezzetti, Jaume Abella 0001, Reinder J. Bril, Francisco J. Cazorla |
RTAS | 5 |
| 2019 | Performance Analysis and Optimization of Automotive GPUsabstractAdvanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) have drastically increased the performance demands of automotive systems. Suitable high-performance platforms building upon Graphic Processing Units (GPUs) have been developed to respond to this demand, being NVIDIA Jetson TX2 a relevant representative. However, whether high-performance GPU configurations are appropriate for automotive setups remains as an open question. This paper aims at providing light on this question by modelling an automotive GPU (Jetson TX2), analyzing its microarchitectural parameters against relevant benchmarks, and identifying specific configurations able to meaningfully increase performance within similar cost envelopes, or to decrease costs preserving original performance levels. Overall, our analysis opens the door to the optimization of automotive GPUs for further system efficiency. Fabio Mazzocchetti, Pedro Benedicte, Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla |
SBAC-PAD | 6 |
| 2019 | Time-Randomized Wormhole NoCs for Critical ApplicationsabstractWormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this article, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g., wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance w.r.t. conventional time-deterministic setups. Our results show that performance guarantees of applications running on top of probabilistic wNoC designs improve by 40% and 93% on average for 4 × 4 and 6 × 6 wNoC setups, respectively. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2019 | Locality-aware cache random replacement policies
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 4 |
| 2019 | Increasing the Reliability of Software Timing Analysis for Cache-Based ProcessorsabstractReal-time systems are witnessing a significant increase in critical software's size, complexity, and performance needs, which can only be satisfied with high-performance hardware features. Cache memories, pervasively used to improve average performance, complicate Worst-Case Execution Time analysis: cache placement (i.e., how software objects are mapped to cache) during the testing phase does not only critically affect the observed performance, but also proves to be arduous to control and preserve up to operation. The probabilistic variant of Measurement-Based Timing Analysis (MBPTA) responds to this challenge by deploying time-randomized caches that naturally explore a different random cache placement in each run, relieving the user from producing tests that intercept relevant Cache Conflict Placements (CCP). Yet, to meet an adequate probabilistic CCP coverage, the user is required to collect a minimum number of measurements. We present two mechanisms, CCP-RM and CCP-HRP, to identify CCP with relevant probability of occurrence and large impact on execution-time, for the random modulo (RM) and hash-based random placement (HRP) policies. CCP-RM and CCP-HRP enable a reliable application of MBPTA by computing the number of runsR' necessary to meet the desired CCP coverage. We exhaustively evaluate CCP-RM and CCP-HRP, showing their effectiveness on well-known benchmarks and a railway case study, on top of an accurate simulator and a concrete RTL implementation. Suzana Milutinovic, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Computers | 4 |
| 2018 | Modelling multicore contention on the AURIXTM TC27xabstractMulticores are becoming ubiquitous in automotive. Yet, the expected benefits on integration are challenged by multicore contention concerns on timing V&V. Worst-case execution time (WCET) estimates are required as early as possible in the software development, to enable prompt detection of timing misbehavior. Factoring in multicore contention necessarily builds on conservative assumptions on interference, independent of co-runners load on shared hardware. We propose a contention model for automotive multicores that balances time-composability with tightness by exploiting available information on contenders. We tailor the model to the AURIX TC27x and provide tight WCET estimates using information from performance monitors and software configurations. Enrique Díaz, Enrico Mezzetti, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 5 |
| 2018 | Measurement-based cache representativeness on multipath programsabstractAutonomous vehicles in embedded real-time systems increase critical-software size and complexity whose performance needs are covered with high-performance hardware features like caches, which however hampers obtaining WCET estimates that hold valid for all program execution paths. This requires assessing that all cache layouts have been properly factored in the WCET process. For measurement-based timing analysis, the most common analysis method, we provide a solution to achieve cache representativeness and full path coverage: we create a modified program for analysis purposes where cache impact is upper-bounded across any path, and derive the minimum number of runs required to capture in the test campaign cache layouts resulting in high execution times. Suzana Milutinovic, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
DAC | 4 |
| 2018 | Cache side-channel attacks and time-predictability in high-performance critical real-time systemsabstractEmbedded computers control an increasing number of systems directly interacting with humans, while also manage more and more personal or sensitive information. As a result, both safety and security are becoming ubiquitous requirements in embedded computers, and automotive is not an exception to that. In this paper we analyze time-predictability (as an example of safety concern) and side-channel attacks (as an example of security issue) in cache memories. While injecting randomization in cache timing-behavior addresses each of those concerns separately, we show that randomization solutions for time-predictability do not protect against side-channel attacks and vice-versa. We then propose a randomization solution to achieve both safety and security goals. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 4 |
| 2018 | Design and integration of hierarchical-placement multi-level caches for real-time systemsabstractEnabling timing analysis in the presence of caches has been pursued by the real-Time embedded systems (RTES) community for years due to cache's huge potential to reduce software's worst-case execution time (WCET). However, caches heavily complicate timing analysis due to hard-To-predict access patterns, with few works dealing with time analyzability of multi-level cache hierarchies. For measurement-based timing analysis (MBTA) techniques-widely used in domains such as avionics, automotive, and rail-we propose several cache hierarchies amenable to MBTA. We focus on a probabilistic variant of MBTA (or MBPTA) that requires caches with time-randomized behavior whose execution time variability can be captured in the measurements taken during system's test runs. For this type of caches, we explore and propose different multi-level cache setups. From those, we choose a cost-effective cache hierarchy that we implement and integrate in a 4-core LEON3 RTL processor model and prototype in a FPGA. Our results show that our proposed setup implemented in RTL results in better (reduced) WCET estimates with similar implementation cost and no impact on average performance w.r.t. other MBPTA-Amenable setups. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 4 |
| 2018 | HWP: Hardware Support to Reconcile Cache Energy, Complexity, Performance and WCET Estimates in Multicore Real-Time SystemsabstractHigh-performance processors have deployed multilevel cache (MLC) systems for decades. In the embedded real-time market, the use of MLC is also on the rise, with processors for future systems in space, railway, avionics and automotive already featuring two or more cache levels. One of the most critical elements for MLC is the write policy that not only affects several key metrics such as performance, WCET estimates, energy/power, and reliability, but also the design of complexity-prone cache coherence protocol and cache reliability solutions. In this paper we make an extensive analysis of existing write policies, namely write-through (WT) and write-back (WB). In the context of the real-time domain, we show that no write policy is superior for all metrics: WT simplifies the design of the coherence and reliability solutions at the cost of performance, WCET, and energy; while WB improves performance and energy results, but complicates cache design. To take the best of each policy, we propose Hybrid Write Policy (HWP) a low-complexity hardware mechanism that reconciles the benefits of WT in terms of simplifying the cache design (e.g. coherence solution) and the benefits of WB in improved average performance and WCET estimates as the pressure on the interconnection network increases. Guaranteed performance results show that HWP scales with core count similar to WB. Likewise, HWP reduces cache energy usage of WT, to levels similar to those of WB. These benefits are obtained while retaining the reduced coherence complexity of WT, in contrast to high coherence costs under WB. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 4 |
| 2018 | Industrial experiences with resource management under software randomization in ARINC653 avionics environmentsabstractInjecting randomization in different layers of the computing platform has been shown beneficial for security, resilience to software bugs and timing analysis. In this paper, with focus on the latter, we show our experience regarding memory and timing resource management when software randomization techniques are applied to one of the most stringent industrial environments, ARINC653-based avionics. We describe the challenges in this task, we propose a set of solutions and present the results obtained for two commercial avionics applications, executed on COTS hardware and RTOS. Leonidas Kosmidis, Cristian Maxim, Victor Jégu, Francis Vatrinet, Francisco J. Cazorla |
ICCAD | 5 |
| 2018 | NoCo: ILP-Based Worst-Case Contention Estimation for Mesh Real-Time ManycoresabstractManycores are capable of providing the computational demands required by functionally-advanced critical applications in domains such as automotive and avionics. In manycores a network-on-chip (NoC) provides access to shared caches and memories and hence concentrates most of the contention that tasks suffer, with effects on the worst-case contention delay (WCD) of packets and tasks' WCET. While several proposals minimize the impact of individual NoC parameters on WCD, e.g. mapping and routing, there are strong dependences among these NoC parameters. Hence, finding the optimal NoC configurations requires optimizing all parameters simultaneously, which represents a multidimensional optimization problem. In this paper we propose NoCo, a novel approach that combines ILP and stochastic optimization to find NoC configurations in terms of packet routing, application mapping, and arbitration weight allocation. Our results show that NoCo improves other techniques that optimize a subset of NoC parameters. Jordi Cardona, Carles Hernández 0001, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 5 |
| 2018 | Assessing Time Predictability Features of ARM Big. LITTLE MulticoresabstractThe increasing performance needs in critical realtime embedded systems (CRTES), such as for instance the automotive domain, push for the adoption of high-performance hardware from the consumer electronics domain. However, their time-predictability features are quite unexplored. The ARM big. LITTLE architecture is a good candidate for adoption in the CRTES market (i.e. in the automotive market it has already started being used). In this paper we study ARM big. LITTLE's capabilities to meet CRTES requirements. In particular, we perform a qualitative and quantitative assessment of its timing characteristics, focusing on shared multicore resources, and how this architecture can be reliably used in CRTES. Gabriel Fernandez 0002, Francisco J. Cazorla, Jaume Abella 0001, Sylvain Girbal |
SBAC-PAD | 2 |
| 2018 | EOmesh: Combined Flow Balancing and Deterministic Routing for Reduced WCET Estimates in Embedded Real-Time SystemsabstractThe increasing performance needs in critical real-time embedded systems (CRTESs) can only be satisfied with the use of high-performance manycore processors. While NoC-based manycore systems are popular in the high-performance domain due to their high average performance, they challenge deriving tight worst-case execution time (WCET) estimates, as needed in CRTES. Weighted meshes have been proposed to alleviate NoCs pathological behavior-caused by large bandwidth imbalance-by making locally unbalanced arbitration decisions to reach globally balanced bandwidth. In this paper, we show that existing weighted mesh solutions do not completely remove unwanted imbalance, in particular for nodes subject to high congestion. We propose even/odd mesh (EOmesh), an approach that combines heterogeneous predictable routing and weight allocations that delivers near-optimal bandwidth allocation across cores without increasing NoC complexity. EOmesh, which can be implemented either by hardware means or by software means on top of regular weighted meshes, improves the average performance and WCET results of the reference weighted mesh design. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Fitting Software Execution-Time Exceedance into a Residual Random Fault in ISO-26262abstractCar manufacturers relentlessly replace or augment the functionality of mechanical subsystems with electronic components. Most such subsystems (e.g., steer-by-wire) are safety related, hence, subject to regulation. ISO-26262, the dominant standard for road vehicles, regards software faults as systematic, while differentiating hardware faults between systematic and random. The analysis of systematic faults entails rigorous processes and qualitative considerations. The increasing complexity of modern on-board computers, however, questions the very notion of treating the violation of execution-time envelopes for software programs as a systematic fault. Modern hardware in fact reduces the user's ability to delve deep enough into the fabric of hardware-software interaction to gage its extent of contribution to the worst-case execution time (WCET). Changing the nature of the WCET-analysis problem may help address that challenge effectively. To this end, we propose a solution that should allow ISO-26262 to quantify the likelihood of execution-time exceedance events, relating it to target failure metrics employed in support of certification arguments, similarly to random faults in hardware. To this end, we inject randomization in the timing behavior of the computer hardware to relieve the user from the need to control hard-to-reach low-level parts, and use measurement-based probabilistic timing analysis to quantify, constructively, the failure rates resulting from the likelihood of execution-time exceedance events. Irune Agirre, Francisco J. Cazorla, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Mikel Azkarate-askatsua, Tullio Vardanega |
IEEE Trans. Reliab. | 2 |
| 2017 | Dynamic software randomisation: Lessons learnec from an aerospace case studyabstractTiming Validation and Verification (V&V) is an important step in real-time system design, in which a system's timing behaviour is assessed via Worst Case Execution Time (WCET) estimation and scheduling analysis. For WCET estimation, measurement-based timing analysis (MBTA) techniques are widely-used and well-established in industrial environments. However, the advent of complex processors makes it more difficult for the user to provide evidence that the software is tested under stress conditions representative of those at system operation. Measurement-Based Probabilistic Timing Analysis (MBPTA) is a variant of MBTA followed by the PROXIMA European Project that facilitates formulating this representativeness argument. MBPTA requires certain properties to be applicable, which can be obtained by selectively injecting randomisation in platform's timing behaviour via hardware or software means. In this paper, we assess the effectiveness of the PROXIMA's dynamic software randomisation (DSR) with a space industrial case study executed on a real unmodified hardware platform and an industrial operating system. We present the challenges faced in its development, in order to achieve MBPTA compliance and the lessons learned from this process. Our results, obtained using a commercial timing analysis tool, indicate that DSR does not impact the average performance of the application, while it enables the use of MBPTA. This results in tighter pWCET estimates compared to current industrial practice. Fabrice Cros, Leonidas Kosmidis, Franck Wartel, David Morales, Jaume Abella 0001, Ian Broster, Francisco J. Cazorla |
DATE | 7 |
| 2017 | Probabilistic timing analysis on time-randomized platforms for the space domainabstractTiming Verification is a fundamental step in real-time embedded systems, with measurement-based timing analysis (MBTA) being the most common approach used to that end. We present a Space case study on a real platform that has been modified to support a probabilistic variant of MBTA called MBPTA. Our platform provides the properties required by MBPTA with the predicted WCET estimates with MBPTA being competitive to those with current MBTA practice while providing more solid evidence on their correctness for certification. Mikel Fernández, David Morales, Leonidas Kosmidis, Alen Bardizbanyan, Ian Broster, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla, Paulo Machado, Luca Fossati |
DATE | 9 |
| 2017 | Design and implementation of a fair credit-based bandwidth sharing scheme for busesabstractFair arbitration in the access to hardware shared resources is fundamental to obtain low worst-case execution time (WCET) estimates in the context of critical real-time systems, for which performance guarantees are essential. Several hardware mechanisms exist for managing arbitration in those resources (buses, memory controllers, etc.). They typically attain fairness in terms of the number of slots each contender (e.g., core) gets granted access to the shared resource. However, those policies may lead to unfair bandwidth allocations for workloads with contenders issuing short requests and contenders issuing long requests. We propose a Credit-Based Arbitration (CBA) mechanism that achieves fairness in the cycles each core is granted access to the resource rather than in the number of granted slots. Furthermore, we implement CBA as part of a LEON3 4-core processor for the Space domain in an FPGA proving the feasibility and good performance characteristics of the design by comparing it against other arbitration schemes. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 4 |
| 2017 | Boosting Guaranteed Performance in Wormhole NoCs with Probabilistic Timing AnalysisabstractWormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this paper, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g. wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance. Our results show that WCET estimates of applications running on top of probabilistic wNoCs are reduced by 40% and 75% on average for 4x4 and 6x6 wNoC setups respectively when compared against deterministic wNoCs. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2017 | Design and Implementation of a Time Predictable Processor: Evaluation With a Space Case StudyabstractEmbedded real-time systems like those found in automotive, rail and aerospace, steadily require higher levels of guaranteed computing performance (and hence time predictability) motivated by the increasing number of functionalities provided by software. However, high-performance processor design is driven by the average-performance needs of mainstream market. To make things worse, changing those designs is hard since the embedded real-time market is comparatively a small market. A path to address this mismatch is designing low-complexity hardware features that favor time predictability and can be enabled/disabled not to affect average performance when performance guarantees are not required. In this line, we present the lessons learned designing and implementing LEOPARD, a four-core processor facilitating measurement-based timing analysis (widely used in most domains). LEOPARD has been designed adding low-overhead hardware mechanisms to a LEON3 processor baseline that allow capturing the impact of jittery resources (i.e. with variable latency) in the measurements performed at analysis time. In particular, at core level we handle the jitter of caches, TLBs and variable-latency floating point units; and at the chip level, we deal with contention so that time-composable timing guarantees can be obtained. The result of our applied study with a Space application shows how per-resource jitter is controlled facilitating the computation of high-quality WCET estimates. Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla, Alen Bardizbanyan, Jan Andersson, Fabrice Cros, Franck Wartel |
ECRTS | 3 |
| 2017 | EPC Enacted: Integration in an Industrial Toolbox and Use against a Railway ApplicationabstractMeasurement-based timing analysis approaches are increasingly making their way into several industrial domains on account of their good cost-benefit ratio. The trustworthiness of those methods, however, suffers from the limitation that their results are only valid for the particular paths and execution conditions that the user is able to explore with the available input vectors. It is generally not possible to guarantee that the collected measurements are fully representative of the worst-case timing behaviour. In the context of measurement-based probabilistic timing analysis, the Extended Path Coverage (EPC) approach has been recently proposed as a means to extend the representativeness of measurement observations, to obtain the same effect of full path coverage. At the time of its first publication, EPC had not reached an implementation maturity that could be trialled industrially. In this work we analyze the practical implications of using EPC with real-world applications, and discuss the challenges in integrating it in an industrial-quality toolchain. We show that we were able to meet EPC requirements and successfully evaluate the technique on a real Railway application, on top of a commercial toolchain and full execution stack. Enrico Mezzetti, Mikel Fernández, Alen Bardizbanyan, Irune Agirre, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla |
RTAS | 7 |
| 2017 | Work-in-Progress Paper: An Analysis of the Impact of Dependencies on Probabilistic Timing Analysis and Task SchedulingabstractRecently there has been a renewed interest for probabilistic timing analysis (PTA) and probabilistic task scheduling (PTS). Despite the number of works in both fields, the link between them is weak: works on the latter build upon a series of assumptions on the probabilistic behavior of each task – or instances (jobs) of it – that have not been shown how to be fulfilled by PTA. This paper makes a first step towards covering this gap with emphasis on providing the right meaning of pWCET estimate as understood by both PTA and PTS. We show that the main issue related to ensuring that PTS assumptions on pWCET estimates are captured by PTA relates to the dependencies among tasks, and even jobs of a given task. Both change the scope of applicability of pWCET estimates provided by PTA and hence, their use by PTS. Enrico Mezzetti, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla |
RTSS | 4 |
| 2017 | SEDEA: A Sensible Approach to Account DRAM Energy in Multicore SystemsabstractAs the energy cost in todays computing systems keeps increasing, measuring the energy becomes crucial in many scenarios. For instance, due to the fact that the operational cost of datacenters largely depends on the energy consumed by the applications executed, end users should be charged for the energy consumed, which requires a fair and consistent energy measuring approach. However, the use of multicore system complicates per-task energy measurement as the increased Thread Level Parallelism (TLP) allows several tasks to run simultaneously sharing resources. Therefore, the energy usage of each task is hard to determine due to interleaved activities and mutual interferences. To this end, Per-Task Energy Metering (PTEM) has been proposed to measure the actual energy of each task based on their resource utilization in a workload. However, the measured energy depends on the interferences from co-running tasks sharing the resources, and thus fails to provide the consistency across executions. Therefore, Sensible Energy Accounting (SEA) has been proposed to deliver an abstraction of the energy consumption based on a particular allocation of resources to a task. In this work we provide a realization of SEA for the DRAM memory system, SEDEA, where we account a task for the DRAM energy it would have consumed when running in isolation with a fraction of the on-chip shared cache. SEDEA is a mechanism to sensibly account for the DRAM energy of a task based on predicting its memory behavior. Our results show that SEDEA provides accurate estimates, yet with low-cost, beating existing per-task energy models, which do not target accounting energy in multicore system. We also provide a use case showing that SEDEA can be used to guide shared cache and memory bank partition schemes to save energy. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
SBAC-PAD | 4 |
| 2017 | Computing Safe Contention Bounds for Multicore Resources with Round-Robin and FIFO ArbitrationabstractNumerous researchers have studied the contention that arises among tasks running in parallel on a multicore processor. Most of those studies seek to derive a tight and sound upper-bound for the worst-case delay with which a processor resource may serve an incoming request, when its access is arbitrated using time-predictable policies such as round-robin or FIFO. We call this value upper-bound delay (ubd). Deriving trustworthy ubd statically is possible when sufficient public information exists on the timing latency incurred on access to the resource of interest. Unfortunately however, that is rarely granted for commercial-of-the-shelf (COTS) processors. Therefore, the users resort to measurement observations on the target processor and thus compute a “measured” ubdm. However, using ubdm to compute worst-case execution time values for programs running on COTS multicore processors requires qualification on the soundness of the result. In this paper, we present a measurement-based methodology to derive a ubdm under round-robin (RoRo) and first-in-first-out (FIFO) arbitration, which accurately approximates ubd from above, without needing latency information from the hardware provider. Experimental results, obtained on multiple processor configurations, demonstrate the robustness of the proposed methodology. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
IEEE Trans. Computers | 6 |
| 2017 | Measurement-Based Worst-Case Execution Time Estimation Using the Coefficient of VariationabstractExtreme Value Theory (EVT) has been historically used in domains such as finance and hydrology to model worst-case events (e.g., major stock market incidences). EVT takes as input a sample of the distribution of the variable to model and fits the tail of that sample to either the Generalised Extreme Value (GEV) or the Generalised Pareto Distribution (GPD). Recently, EVT has become popular in real-time systems to derive worst-case execution time (WCET) estimates of programs. However, the application of EVT is not straightforward and requires a detailed analysis of, and customisation for, the particular problem at hand. In this article, we tailor the application of EVT to timing analysis. To that end, (1) we analyse the response time of different hardware resources (e.g., cache memories) and identify those that may lead to radically different types of execution time distributions. (2) We show that one of these distributions, known as mixture distribution, causes problems in the use of EVT. In particular, mixture distributions challenge not only properly selecting GEV/GPD parameters (i.e., location, scale and shape) but also determining the size of the sample to ensure that enough tail values are passed to EVT and that only tail values are used by EVT to fit GEV/GPD. Failing to select these parameters has a negative impact on the quality of the derived WCET estimates. We tackle these problems, by (3) proposing Measurement-Based Probabilistic Timing Analysis using the Coefficient of Variation (MBPTA-CV), a new mixture-distribution aware, WCET-suited MBPTA method that builds on recent EVT developments in other fields (e.g., finance) to automatically select the distribution parameters that best fit the maxima of the observed execution times. Our results on a simulation environment and a real board show that MBPTA-CV produces high-quality WCET estimates. Jaume Abella 0001, Maria Padilla, Joan del Castillo, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2016 | Random modulo: a new processor cache design for real-time critical systemsabstractCache memories have a huge impact on software's worst-case execution time (WCET). While enabling the seamless use of caches is key to provide the increasing levels of (guaranteed) performance required by automotive software, caches complicate timing analysis. In the context of Measurement-Based Probabilistic Timing Analysis (MBPTA) -- a promising technique to ease timing analyis of complex hardware -- we propose Random Modulo (RM), a new cache design that provides the probabilistic behavior required by MBPTA and with the following advantages over existing MBPTA-compliant cache designs: (i) an outstanding reduction in WCET estimates, (ii) lower latency and area overhead, and (iii) competitive average performance w.r.t conventional caches. Carles Hernández 0001, Jaume Abella 0001, Andrea Gianarro, Jan Andersson, Francisco J. Cazorla |
DAC | 5 |
| 2016 | Supertask: Maximizing runnable-level parallelism in AUTOSAR applications
Sebastian Kehr, Milos Panic, Eduardo Quiñones, Bert Böddeker, Jorge Becerril Sandoval, Jaume Abella 0001, Francisco J. Cazorla, Günter Schäfer |
DATE | 7 |
| 2016 | Improving performance guarantees in wormhole mesh NoC designs
Milos Panic, Carles Hernández 0001, Jaume Abella 0001, Antoni Roca 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 6 |
| 2016 | PROXIMA: Improving Measurement-Based Timing Analysis through Randomisation and Probabilistic AnalysisabstractThe use of increasingly complex hardware and software platforms in response to the ever rising performance demands of modern real-time systems complicates the verification and validation of their timing behaviour, which form a time-and-effort-intensive step of system qualification or certification. In this paper we relate the current state of practice in measurement-based timing analysis, the predominant choice for industrial developers, to the proceedings of the PROXIMA (Probabilistic real-time control of mixed-criticality multicore systems) project in that very field. We recall the difficulties that the shift towards more complex computing platforms causes in that regard. Then we discuss the probabilistic approach proposed by PROXIMA to overcome some of those limitations. We present the main principles behind the PROXIMA approach as well as the changes it requires at hardware or software level underneath the application. We also present the current status of the project against its overall goals, and highlight some of the principal confidence-building results achieved so far. Francisco J. Cazorla, Jaume Abella 0001, Jan Andersson, Tullio Vardanega, Francis Vatrinet, Iain Bate, Ian Broster, Mikel Azkarate-askatsua, Franck Wartel, Liliana Cucu-Grosjean, Fabrice Cros, Glenn Farrall, Adriana Gogonel, Andrea Gianarro, Benoit Triquet, Carles Hernández 0001, Code Lo, Cristian Maxim, David Morales, Eduardo Quiñones, Enrico Mezzetti, Leonidas Kosmidis, Irune Agirre, Mikel Fernández, Mladen Slijepcevic, Philippa Conmy, Walid Talaboulma |
DSD | 1 |
| 2016 | pTNoC: Probabilistically Time-Analyzable Tree-Based NoC for Mixed-Criticality SystemsabstractThe use of networks-on-chip (NoC) in real-time safety-critical multicore systems challenges deriving tight worst-case execution time (WCET) estimates. This is due to the complexities in tightly upper-bounding the contention in the access to the NoC among running tasks. Probabilistic Timing Analysis (PTA) is a powerful approach to derive WCET estimates on relatively complex processors. However, so far it has only been tested on small multicores comprising an on-chip bus as communication means, which intrinsically does not scale to high core counts. In this paper we propose pTNoC, a new tree-based NoC design compatible with PTA requirements and delivering scalability towards medium/large core counts. pTNoC provides tight WCET estimates by means of asymmetric bandwidth guarantees for mixed-criticality systems with negligible impact on average performance. Finally, our implementation results show the reduced area and power costs of the pTNoC. Mladen Slijepcevic, Mikel Fernández, Carles Hernández 0001, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DSD | 6 |
| 2016 | TASA: toolchain-agnostic static software randomisation for critical real-time systemsabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) derives WCET estimates for tasks running on processors comprising high-performance features such as caches. MBPTA's correct application requires the system to exhibit certain timing properties, which can be achieved by injecting randomisation in the timing behaviour of the task under analysis. However, existing software-randomisation techniques require costly modifications in the industrial production toolchain (compiler, linker, runtime or hardware) in terms of development and certification. In this paper we present TASA, a new software randomisation tool that relies on source-code transformations of the application (i) requiring no changes in existing toolchains, which heavily reduces tool qualification and implementation costs; and (ii) achieving competitive WCET estimates that we assess on a gcc- and a llvm-based compilation toolchain on a real board. Leonidas Kosmidis, Roberto Vargas, David Morales, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
ICCAD | 6 |
| 2016 | A confidence assessment of WCET estimates for software time randomized cachesabstractObtaining Worst-Case Execution Time (WCET) estimates is a required step in real-time embedded systems during software verification. Measurement-Based Probabilistic Timing Analysis (MBPTA) aims at obtaining WCET estimates for industrial-size software running upon hardware platforms comprising high-performance features. MBPTA relies on the randomization of timing behavior (functional behavior is left unchanged) of hard-to-predict events like the location of objects in memory - and hence their associated cache behavior - that significantly impact software's WCET estimates. Software time-randomized caches (sTRc) have been recently proposed to enable MBPTA on top of Commercial off-the-shelf (COTS) caches (e.g. modulo placement). However, some random events may challenge MBPTA reliability on top of sTRc. In this paper, for sTRc and programs with homogeneously accessed addresses, we determine whether the number of observations taken at analysis, as part of the normal MBPTA application process, captures the cache events significantly impacting execution time and WCET. If this is not the case, our techniques provide the user with the number of extra runs to perform to guarantee that cache events are captured for a reliable application of MBPTA. Our techniques are evaluated with synthetic benchmarks and an avionics application. Pedro Benedicte, Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
INDIN | 5 |
| 2016 | Resilient random modulo cache memories for probabilistically-analyzable real-time systemsabstractFault tolerance has often been assessed separately in safety-related real-time systems, which may lead to inefficient solutions. Recently, Measurement-Based Probabilistic Timing Analysis (MBPTA) has been proposed to estimate Worst-Case Execution Time (WCET) on high performance hardware. The intrinsic probabilistic nature of MBPTA-commpliant hardware matches perfectly with the random nature of hardware faults. Joint WCET analysis and reliability assessment has been done so far for some MBPTA-compliant designs, but not for the most promising cache design: random modulo. In this paper we perform, for the first time, an assessment of the aging-robustness of random modulo and propose new implementations preserving the key properties of random modulo, a.k.a. low critical path impact, low miss rates and MBPTA compliance, while enhancing reliability in front of aging by achieving a better - yet random - activity distribution across cache sets. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IOLTS | 4 |
| 2016 | Modelling Probabilistic Cache Representativeness in the Presence of Arbitrary Access PatternsabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) is a promising powerful industry-friendly method to derive worst-case execution time (WCET) estimates as needed for critical real-time embedded systems. MBPTA performs several (R) runs of the program on the target platform collecting the execution times in each run. MBPTA builds a probabilistic representativeness argument on whether those events with high impact on execution time, such as cache misses, arise on the runs made at analysis time so that their impact on execution time is captured. So far only events occurring in cache memories have been shown to challenge providing such representativeness argument. In this context, this paper introduces a representativeness validation method (RVS) to assess the probabilistic representativeness of MBPTA's execution time observations in terms of cache behaviour. RVS resorts to cache simulation to predict worst-case miss scenarios that can appear during the deployment phase. RVS also constructs a probabilistic Worst-Case Miss Count curve based on the miss-counts captured in the R runs. If that curve upperbounds the impact of the predicted cache worst-case scenarios, R is deemed as a sufficient number of runs for which pWCET estimates can be reliably derived. Otherwise, the user is requested to perform more runs until all cache scenarios of interest are captured. Suzana Milutinovic, Jaume Abella 0001, Francisco J. Cazorla |
ISORC | 3 |
| 2016 | Modeling High-Performance Wormhole NoCs for Critical Real-Time Embedded SystemsabstractManycore chips are a promising computing platform to cope with the increasing performance needs of critical real-time embedded systems (CRTES). However, manycores adoption by CRTES industry requires understanding task's timing behavior when their requests use manycore's network-on-chip (NoC) to access hardware shared resources. This paper analyzes the contention in wormhole-based NoC (wNoC) designs - widely implemented in the high-performance domain - for which we introduce a new metric: worst-contention delay (WCD) that captures wNoC impact on worst-case execution time (WCET) in a tighter manner than the existing metric, worst-case traversal time (WCTT). Moreover, we provide an analytical model of the WCD that requests can suffer in a wNoC and we validate it against wNoC designs resembling those in the Tilera-Gx36 and the Intel-SCC 48-core processors. Building on top of our WCD analytical model, we analyze the impact on WCD that different design parameters such as the number of virtual channels, and we make a set of recommendations on what wNoC setups to use in the context of CRTES. Milos Panic, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 5 |
| 2016 | Improving Early Design Stage Timing Modeling in Multicore Based Real-Time SystemsabstractThis paper presents a modelling approach for the timing behavior of real-time embedded systems (RTES) in early design phases. The model focuses on multicore processors - accepted as the next computing platform for RTES - and in particular it predicts the contention tasks suffer in the access to multicore on-chip shared resources. The model presents the key properties of not requiring the application's source code or binary and having high-accuracy and low overhead. The former is of paramount importance in those common scenarios in which several software suppliers work in parallel implementing different applications for a system integrator, subject to different intellectual property (IP) constraints. Our model helps reducing the risk of exceeding the assigned budgets for each application in late design stages and its associated costs. David Trilla, Javier Jalle, Mikel Fernández, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 5 |
| 2016 | Sensible Energy Accounting with Abstract Metering for Multicore SystemsabstractChip multicore processors (CMPs) are the preferred processing platform across different domains such as data centers, real-time systems, and mobile devices. In all those domains, energy is arguably the most expensive resource in a computing system. Accurately quantifying energy usage in a multicore environment presents a challenge as well as an opportunity for optimization. Standard metering approaches are not capable of delivering consistent results with shared resources, since the same task with the same inputs may have different energy consumption based on the mix of co-running tasks. However, it is reasonable for data-center operators to charge on the basis of estimated energy usage rather than time since energy is more correlated with their actual cost. This article introduces the concept of Sensible Energy Accounting (SEA). For a task running in a multicore system, SEA accurately estimates the energy the task would have consumed running in isolation with a given fraction of the CMP shared resources. We explain the potential benefits of SEA in different domains and describe two hardware techniques to implement it for a shared last-level cache and on-core resources in SMT processors. Moreover, with SEA, an energy-aware scheduler can find a highly efficient on-chip resource assignment, reducing by up to 39% the total processor energy for a 4-core system. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Daniel A. Jiménez, Mateo Valero |
ACM Trans. Archit. Code Optim. | 4 |
| 2016 | Thread Assignment in Multicore/Multithreaded Processors: A Statistical ApproachabstractThe introduction of multicore/multithreaded processors, comprised of a large number of hardware contexts (virtual CPUs) that share resources at multiple levels, has made process scheduling, in particular assignment of running threads to available hardware contexts, an important aspect of system performance. Nevertheless, thread assignment of applications running on state-of-the art processors is an NP-complete problem. Over the years, numerous studies have proposed heuristic-based algorithms for thread assignment. Since the thread assignment problem is intractable, it is in general impossible to know the performance of the optimal assignment, so the room for improvement of a given algorithm is also unknown. It is therefore hard to decide whether to invest more effort and time to improve an algorithm that may already be close to optimal. In this paper, we present a statistical approach to the thread assignment problem. First, we present a method that predicts the performance of the optimal thread assignment, based on the observed performance of each thread assignment in a random sample. The method is based on Extreme Value Theory (EVT), a branch of statistics that analyses extreme deviations from the population mean. We also propose sample pruning, a method that significantly reduces the time required to apply the statistical method by reducing the number of candidate solutions that need to be measured. Finally, we show that, if no suitable heuristic-based algorithm is available, a sample of several thousand random thread assignments is enough to obtain, with high confidence, an assignment with performance close to optimal. The presented approach is architecture and application independent, and it can be used to address the thread assignment problem in various domains. It is especially well suited for systems in which the workload seldom changes. An example is network systems, which typically provide a constant set of services that are known in advance, with network applications performing a similar processing algorithm for each packet in the system. In this paper, we validate our methods with an industrial case study for a set of multithreaded network applications on an UltraSPARC T2 processor. This article is an extension of our previous work [44], which was published in Proceedings of 17th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS-2012). Petar Radojkovic, Paul M. Carpenter, Miquel Moretó, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
IEEE Trans. Computers | 7 |
| 2016 | Parallelizing Industrial Hard Real-Time Applications for the parMERASA MulticoreabstractThe EC project parMERASA (Multicore Execution of Parallelized Hard Real-Time Applications Supporting Analyzability) investigated timing-analyzable parallel hard real-time applications running on a predictable multicore processor. A pattern-supported parallelization approach was developed to ease sequential to parallel program transformation based on parallel design patterns that are timing analyzable. The parallelization approach was applied to parallelize the following industrial hard real-time programs: 3D path planning and stereo navigation algorithms (Honeywell International s.r.o.), control algorithm for a dynamic compaction machine (BAUER Maschinen GmbH), and a diesel engine management system (DENSO AUTOMOTIVE Deutschland GmbH). This article focuses on the parallelization approach, experiences during parallelization with the applications, and quantitative results reached by simulation, by static WCET analysis with the OTAWA tool, and by measurement-based WCET analysis with the RapiTime tool. Theo Ungerer, Christian Bradatsch, Martin Frieb, Florian Kluge, Jörg Mische, Alexander Stegmeier, Ralf Jahr, Mike Gerdes 0001, Pavel G. Zaykov, Lucie Matusova, Zai Jian Jia Li, Zlatko Petrov, Bert Böddeker, Sebastian Kehr, Hans Regler, Andreas Hugl, Christine Rochange, Haluk Ozaktas, Hugues Cassé, Armelle Bonenfant, Pascal Sainrat, Nick Lay, Ian Broster, Eduardo Quiñones, Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla, Sascha Uhrig, Mathias Rohde, Arthur Pyka |
ACM Trans. Embed. Comput. Syst. | 29 |
| 2016 | DReAM: An Approach to Estimate per-Task DRAM Energy in Multicore SystemsabstractAccurate per-task energy estimation in multicore systems would allow performing per-task energy-aware task scheduling and energy-aware billing in data centers, among other applications. Per-task energy estimation is challenged by the interaction between tasks in shared resources, which impacts tasks’ energy consumption in uncontrolled ways. Some accurate mechanisms have been devised recently to estimate per-task energy consumed on-chip in multicores, but there is a lack of such mechanisms for DRAM memories. This article makes the case for accurate per-task DRAM energy metering in multicores, which opens new paths to energy/performance optimizations. In particular, the contributions of this article are (i) an ideal per-task energy metering model for DRAM memories; (ii) DReAM, an accurate yet low cost implementation of the ideal model (less than 5% accuracy error when 16 tasks share memory); and (iii) a comparison with standard methods (even distribution and access-count based) proving that DReAM is much more accurate than these other methods. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2015 | Increasing confidence on measurement-based contention bounds for real-time round-robin busesabstractContention among tasks concurrently running in a multicore has been deeply studied in the literature specially for on-chip buses. Most of the works so far focus on deriving exact upper-bounds to the longest delay it takes a bus request to be serviced (ubd), when its access is arbitrated using a time-predictable policy such as round-robin. Deriving ubd for a bus can be done accurately when enough timing information is available, which is not often the case for commercial-of-the-shelf (COTS) processors. Hence, ubd is approximated (ubdm) by directly experimenting on the target processor, i.e by measurements. However, using ubdm makes the timing analysis technique to resort on the accuracy of ubdm to derive trustworthy worst-case execution time estimates. Therefore, accurately estimating ubd by means of ubdm is of paramount importance. In this paper, we propose a systematic measurement-based methodology to accurately approximate ubd without knowing the bus latency or any other latency information, being only required that the underlying bus policy is round-robin. Our experimental results prove the robustness of the proposed methodology by testing it on different bus and processor setups. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
DAC | 6 |
| 2015 | Resource usage templates and signatures for COTS multicore processorsabstractUpper bounding the execution time of tasks running on multicore processors is a hard challenge. This is especially so with commercial-off-the-shelf (COTS) hardware that conceals its internal operation. The main difficulty stems from the contention effects on access to hardware shared resources (e.g., buses) which cause task's timing behavior to depend on the load that co-runner tasks place on them. This dependence reduces time composability and constrains incremental verification. In this paper we introduce the concepts of resource-usage signatures and templates, to abstract the potential contention caused and incurred by tasks running on a multicore. We propose an approach that employs resource-usage signatures and templates to enable the analysis of individual tasks largely in isolation, with low integration costs, producing execution time estimates per task that are easily composable throughout the whole system integration process. We evaluate the proposal on a 4-core NGMP-like multicore architecture. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
DAC | 6 |
| 2015 | PACO: fast average-performance estimation for time-randomized cachesabstractProbabilistic timing analysis is a powerful approach to derive worst-case execution time (WCET) estimates, as needed in safety-critical systems, in the presence of high-performance hardware features (e.g., caches). To that end, the timing behavior of certain hardware resources, such as caches, is randomized. Time-randomized (TR) caches allow deriving hit/miss probabilities for each access and probabilistic WCET estimates for the overall program. Suzana Milutinovic, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 4 |
| 2015 | Timing analysis of an avionics case study on complex hardware/software platforms
Franck Wartel, Leonidas Kosmidis, Adriana Gogonel, Andrea Baldovin, Zoë Stephenson, Benoit Triquet, Eduardo Quiñones, Code Lo, Enrico Mezzetti, Ian Broster, Jaume Abella 0001, Liliana Cucu-Grosjean, Tullio Vardanega, Francisco J. Cazorla |
DATE | 14 |
| 2015 | IEC-61508 SIL 3 Compliant Pseudo-Random Number Generators for Probabilistic Timing AnalysisabstractProbabilistic Timing Analysis (PTA), especially its measurement based variant (MBPTA), has shown to be competitive with state-of-the-art timing analysis techniques. The use of MBPTA to analyse the timing behaviour of safety-critical systems rests on its ability to derive trustworthy WCET bounds. This ability depends on the soundness of the MBPTA method per se, as well as on the satisfaction of safety requirements placed on the pseudo-random number generator (prng) that plays a key role in the platform-level randomisation needed by MBPTA. This paper presents the design of a low-area, low-power prng that meets IEC-61508 SIL 3 safety requirements and allows for seamless integration in a real-world multicore architecture. This work enables the development and the IEC-61508 certification of mixed-criticality systems that use MBPTA for deriving timing bounds for mixed-criticality software programs running on multicore processors. Irune Agirre, Mikel Azkarate-askatsua, Carles Hernández 0001, Jaume Abella 0001, Jon Pérez 0001, Tullio Vardanega, Francisco J. Cazorla |
DSD | 7 |
| 2015 | Enabling TDMA Arbitration in the Context of MBPTAabstractCurrent timing analysis techniques can be broadly classified into two families: deterministic timing analysis (DTA) and probabilistic timing analysis (PTA). Each family defines a set of properties to be provided (enforced) by the hardware and software platform so that valid Worst-Case Execution Time (WCET) estimates can be derived for programs running on that platform. However, the fact that each family relies on each own set of hardware designs limits their applicability and reduces the chances of those designs being adopted by hardware vendors. In this paper we show that Time Division Multiple Access (TDMA), one of the main DTA-compliant arbitration policies, can be made PTA-compliant. To that end, we analyze TDMA in the context of measurement-based PTA (MBPTA) and show that padding execution time observations conveniently leads to trustworthy and tight WCET estimates with MBPTA without introducing any hardware change. In fact, TDMA outperforms round-robin and time-randomized policies in terms of WCET in the context of MBPTA. Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Eduardo Quiñones, Theo Ungerer, Francisco J. Cazorla |
DSD | 6 |
| 2015 | CAP: Communication-Aware Allocation Algorithm for Real-Time Parallel Applications on Many-CoresabstractCritical Real-Time Embedded Systems (CRTES) require additional computing power to match the performance requirements of increasingly complex critical functions. Many-core processors are a key solution to reach the required performance. They allow simultaneous execution of multiple critical functions comprising a number of (parallel) applications. These applications, each implementing some functionality of the system, need to communicate among them in order to make the system work. In many-cores the impact of this communication on the timing behavior of applications depends on the allocation of applications across the chip. In this paper we propose CAP: an allocation algorithm that takes into account communication among applications and tries to reduce its impact on Worst Case Execution Time estimates of applications. We show that using CAP increases the number of applications that can be scheduled on the many-core platform thus facilitating system integration. Milos Panic, Eduardo Quiñones, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 5 |
| 2015 | Increasing multicore system efficiency through intelligent bandwidth shiftingabstractMemory bandwidth is a crucial resource in computing systems. Current CMP/SMT processors have a significant number of cores and they can run many threads concurrently. This large thread count adds high pressure to the memory bus, which demands high bandwidth to service memory requests from the cores. Hardware data prefetching is a well-known technique for hiding memory latency. Due to its speculative nature, however, in some situations prefetching does not effectively work, wasting memory bandwidth and polluting the caches. Data prefetching efficiency depends on the prefetching algorithm. It also depends on the characteristics of the applications running on the system. In this paper we propose an online bandwidth shifting mechanism that dynamically assigns bandwidth to applications according to their prefetch efficiency. This mechanism maximizes the utilization of memory bandwidth, thereby improving system performance and/or reducing memory power consumption. To the best of our knowledge, this solution is the first to not require hardware support. We evaluate the benefits of using our bandwidth shifting mechanism on a real system - the IBM POWER7. We obtain speedups in the order of 10-20% (in one instance, speedup exceeds 1.6X). Our mechanism does not generate a significant degree of unfairness among the applications. In many cases individual thread performance increases by 10-35%, while virtually no thread experiences a slowdown larger than 5%. Víctor Jiménez, Alper Buyuktosunoglu, Pradip Bose, Francis P. O'Connell, Francisco J. Cazorla, Mateo Valero |
HPCA | 5 |
| 2015 | Seeking Time-Composable Partitions of Tasks for COTS Multicore ProcessorsabstractThe timing verification of real-time single core systems involves a timing analysis step that yields an Execution Time Bound (ETB) for each task, followed by a schedulability analysis step, where the scheduling attributes of the individual tasks, including the ETB, are studied from the system level perspective. The transition between those two steps involves accounting for the interference effects that arise when tasks contend for access to shared resource. The advent of multicore processors challenges the viability of this two-step approach because several complex contention effects at the processor level arise that cause tasks to be unable to make progress while actually holding the CPU, which are very difficult to tightly capture by simply inflating the tasks' ETB. In this paper we show how contention on access to hardware shared resources creates a circular dependence between the determination of tasks' ETB and their scheduling at runtime. To help loosen this knot we present an approach that acknowledges different flavors of time compos ability, examining in detail the variant intended for partitioned scheduling, which we evaluate on two real processor boards used in the space domain. Gabriel Fernandez 0002, Jaume Abella 0001, Eduardo Quiñones, Luca Fossati, Marco Zulianello, Tullio Vardanega, Francisco J. Cazorla |
ISORC | 7 |
| 2015 | EPC: Extended Path Coverage for Measurement-Based Probabilistic Timing AnalysisabstractMeasurement-based probabilistic timing analysis (MBPTA) computes trustworthy upper bounds to the execution time of software programs. MBPTA has the connotation, typical of measurement-based techniques, that the bounds computed with it only relate to what is observed in actual program traversals, which may not include the effective worst-case phenomena. To overcome this limitation, we propose Extended Path Coverage (EPC), a novel technique that allows extending the representativeness of the bounds computed by MBPTA. We make the observation data probabilistically path-independent by modifying the probability distribution of the observed timing behaviour so as to negatively compensate for any benefits that a basic block may draw from a path leading to it. This enables the derivation of trustworthy upper bounds to the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. Our results confirm that using MBPTA with EPC produces fully trustworthy upper bounds with competitively small overestimation in comparison to state-of-the-art MBPTA techniques. Marco Ziccardi, Enrico Mezzetti, Tullio Vardanega, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 5 |
| 2014 | Containing Timing-Related Certification Cost in Automotive Systems Deploying Complex HardwareabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) techniques simplify deriving tight and trustworthy WCET estimates for industrial-size programs running on complex processors. MBPTA poses some requirements on the timing behaviour of the hardware/software platform: execution times of end-to-end runs have to be independent and identically distributed (i.i.d.). Hardware and software solutions have been deployed to accomplish MBPTA requirements. The latter has achieved the i.i.d. properties running on some commercial off-the-shelf (COTS) processor designs. Unfortunately, software randomisation challenges functional verification needed for certification since it introduces indirections through pointers in the code. In this paper we propose a new approach to software randomisation able to contain its functional verification costs. Our approach performs software randomisation statically, as opposed to current dynamic approaches. We carefully review the requirements of the new approach and prove its feasibility. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Glenn Farrall, Franck Wartel, Francisco J. Cazorla |
DAC | 6 |
| 2014 | Time-Analysable Non-Partitioned Shared Caches for Real-Time Multicore SystemsabstractShared caches in multicores challenge Worst-Case Execution Time (WCET) estimation due to inter-task interferences. Hardware and software cache partitioning address this issue although they complicate data sharing among tasks and the Operating System (OS) task scheduling and migration. In the context of Probabilistic Timing Analysis (PTA) time-randomised caches are used. We propose a new hardware mechanism to control inter-task interferences in shared time-randomised caches without the need of any hardware or software partitioning. Our proposed mechanism effectively bounds inter-task interferences by limiting the cache eviction frequency of each task, while providing tighter WCET estimates than cache partitioning algorithms. In a 4-core multicore processor setup our proposal improves cache partitioning by 56% in terms of guaranteed performance and 16% in terms of average performance. Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DAC | 5 |
| 2014 | Bus designs for time-probabilistic multicore processorsabstractProbabilistic Timing Analysis (PTA) reduces the amount of information needed to provide tight WCET estimates in real-time systems with respect to classic timing analysis. PTA imposes new requirements on hardware design that have been shown implementable for single-core architectures. However, no support has been proposed for multicores so far. In this paper, we propose several probabilistically-analysable bus designs for multicore processors ranging from 4 cores connected with a single bus, to 16 cores deploying a hierarchical bus design. We derive analytical models of the probabilistic timing behaviour for the different bus designs, show their suitability for PTA and evaluate their hardware cost. Our results show that the proposed bus designs (i) fulfil PTA requirements, (ii) allow deriving WCET estimates with the same cost and complexity as in single-core processors, and (iii) provide higher guaranteed performance than single-core processors, 3.4x and 6.6x on average for an 8-core and a 16-core setup respectively. Javier Jalle, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 5 |
| 2014 | Measurement-Based Probabilistic Timing Analysis and Its Impact on Processor ArchitectureabstractCritical Real-Time Embedded Systems (CRTES) industry needs increasingly complex hardware to attain the performance/cost ratio required to keep competitive edge in the market. Worst-case execution time (WCET) analysis is central to CRTES development. Whereas current timing analysis techniques are sound, their viability is hampered by the soaring cost of acquiring detailed knowledge of the internal operation and state of the system, at both software and hardware level. This is a major hurdle to using them for increasingly complex hardware platforms. Measurement-Based Probabilistic Timing Analysis (PTA) reduces the cost of acquiring the knowledge needed for computing trustworthy WCET bounds. This paper presents the changes required to hardware design to facilitate the use of the PTA techniques. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Ian Broster, Francisco J. Cazorla |
DSD | 6 |
| 2014 | On the Comparison of Deterministic and Probabilistic WCET Estimation TechniquesabstractTiming validation is a critical step in the design of real-time systems, that requires the estimation of Worst-Case Execution Times (WCET) for tasks. A number of different methods have been proposed, such as Static Deterministic Timing Analysis (SDTA). The advent of Probabilistic Timing Analysis, both Measurement-Based (MBPTA) and Static Probabilistic Timing Analyses (SPTA), offers different design points between the tightness of WCET estimates, hardware that can be analyzed and the information needed from the user to carry out the analysis. The lack of comparison among those techniques makes complex the selection of the most appropriate one for a given system. This paper makes a first attempt towards comparing comprehensively SDTA, SPTA and MBPTA, qualitatively and quantitatively, under different cache configurations implementing LRU and random replacement. We identify strengths and limitations of each technique depending on the characteristics of the program under analysis and the hardware platform, thus providing users with guidance on which approach to choose depending on their target application and hardware platform. Jaume Abella 0001, Damien Hardy, Isabelle Puaut, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 5 |
| 2014 | Heart of Gold: Making the Improbable Happen to Increase Confidence in MBPTAabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) has been recently proposed as a viable method to compute probabilistic worst-case execution time (pWCET) bounds for programs with hard real-time constraints. As a key trait, MBPTA needs a comparatively small number of observation runs, made on execution platforms to which MBPTA can be applied, to project the tail of the probability of occurrence of worst-case execution time durations of individual programs. In order for the use of MBPTA to fit the bill of industrial-quality development, it is imperative to understand what factors might threaten the trustworthiness of the pWCET computation. This paper addresses that important question by: (i) identifying the combined characteristics of applications and hardware resources that might lead to optimistic pWCET bounds, (ii) describing why this may occur, and (iii) providing the user with means to detect those cases so that trustworthiness is restored. In particular, we present a method for detecting risk scenarios for time-randomised caches, based on principles that apply to any other time-randomised resource which may challenge the application of MBPTA. Jaume Abella 0001, Eduardo Quiñones, Franck Wartel, Tullio Vardanega, Francisco J. Cazorla |
ECRTS | 5 |
| 2014 | PUB: Path Upper-Bounding for Measurement-Based Probabilistic Timing AnalysisabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) responds to the challenge of analysing the timing behaviour of real-time software running on hardware deploying high-performance features (e.g., data caches). MBPTA provides a WCET estimate that upper-bounds the execution time of the set of paths exercised with the data input vectors provided by the user. However, in several scenarios, the user is unaware of the input vector leading to the worst-case path. In this paper we present PUB, a new method that makes the WCET estimates obtained with MBPTA a trustworthy upper-bound of the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. This significantly reduces the requirements imposed on the user to apply MBPTA. For Malardarlen and EEMBC respectively, PUB provides WCET estimates 5% and 11% higher than the WCET estimates computed with MBPTA. Leonidas Kosmidis, Jaume Abella 0001, Franck Wartel, Eduardo Quiñones, Antoine Colin, Francisco J. Cazorla |
ECRTS | 6 |
| 2014 | Parallel many-core avionics systemsabstractIntegrated Modular Avionics (IMA) enables incremental qualification by encapsulating avionics applications into software partitions (SWPs), as defined by the ARINC 653 standard. SWPs, when running on top of single-core processors, provide robust time partitioning as a means to isolate SWPs timing behavior from each other. However, when moving towards parallel execution in many-core processors, the simultaneous accesses to shared hardware and software resources influence the timing behavior of SWPs, defying the purpose of time partitioning to provide isolation among applications. In this paper, we extend the concept of SWP by introducing parallel software partitions (pSWP) specification that describes the behavior of SWPs required when running in a many-core to enable incremental qualification. pSWP are supported by a new hardware feature called guaranteed resource partition (GRP) that defines an execution environment in which SWPs run and that controls interferences in the accesses to shared hardware resources among SWPs such that time composability can be guaranteed. Milos Panic, Eduardo Quiñones, Pavel G. Zaykov, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
EMSOFT | 6 |
| 2014 | DReAM: Per-Task DRAM Energy Metering in Multicore Systems
Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
Euro-Par | 4 |
| 2014 | AHRB: A high-performance time-composable AMBA AHB busabstractHard real-time systems are moving toward complex systems comprising chips with different IP components connected with standard buses. AMBA is one of the most used bus interfaces and has already been included in processors in the real-time domain. However, AMBA was not designed to provide time composable Worst Case Execution Time (WCET) estimates, which are desirable to reduce timing validation and verification costs. This paper analyzes and extends the AMBA Advanced High-performance Bus (AHB) specification to enable time-composable WCET estimates by design. Concretely, (1) we analyze in detail the AMBA AHB in the context of hard real-time systems proving that it fails to provide time composability; (2) we define a restricted subset of AMBA AHB features, named restricted AHB (resAHB), that allows deriving time-composable, yet not tight, WCET estimates; and (3) we define an extension of resAHB, named Advanced High-performance Real-time Bus (AHRB), that includes the timing constraints in the specification. This allows deriving time-composable and tight WCET estimates. Our results show that AHRB can provide 3.5x tighter estimates than resAHB on average for EEMBC benchmarks. Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Luca Fossati, Marco Zulianello, Francisco J. Cazorla |
RTAS | 6 |
| 2014 | A Dual-Criticality Memory Controller (DCmc): Proposal and Evaluation of a Space Case StudyabstractMulticore Dual-Criticality systems comprise two types of applications, each with a different criticality level. In the space domain these types are referred as payload and control applications, which have high-performance and real time requirements respectively. In order to control the interaction (contention) among payload and control applications in the access to the main memory, reaching the goals of high bandwidth for the former and guaranteed timing bounds for the latter, we propose a Dual-Criticality memory controller (DCmc). DCmc virtually divides memory banks into real-time and high-performance banks, deploying a different request scheduler policy to each bank type, which facilitates achieving both goals. Our evaluation with a multicore cycle-accurate simulator and a real space case study shows that DCmc enables deriving tight WCET estimates, regardless of the co-running payload applications, hence effectively isolating the effect of contention in the access to memory. DCmc also enables payload applications exploiting memory locality, which is needed for high performance. Javier Jalle, Eduardo Quiñones, Jaume Abella 0001, Luca Fossati, Marco Zulianello, Francisco J. Cazorla |
RTSS | 6 |
| 2014 | Efficient Cache Designs for Probabilistically Analysable Real-Time SystemsabstractThe increasing performance demand in the critical real-time embedded systems (CRTES) domain calls for high-performance features such as cache memories. Unfortunately, the cost to provide trustworthy and tight Worst-Case Execution Time (WCET) estimates in the presence of caches is high with current practice WCET analysis tools, because they need detailed knowledge of program’s cache accesses to provide tight WCET estimates. The advent of Probabilistic timing analysis (PTA) opens the door to economically viable timing analysis in the presence of caches, but it imposes new requirements on hardware design. At cache level, so far only fully associative random-replacement caches have been proven to fulfill the needs of PTA, but their energy, delay, and area cost are unaffordable for CRTES. In this paper, we propose the first PTA-compliant cache design based on set-associative and direct-mapped arrangements, as those are the most common arrangements. In particular, we propose a novel parametric random placement policy suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
IEEE Trans. Computers | 4 |
| 2014 | Hybrid Cache Designs for Reliable Hybrid High and Ultra-Low Voltage OperationabstractGeometry scaling of semiconductor devices enables the design of ultra-low-cost (e.g., below 1 USD) battery-powered resource-constrained ubiquitous devices for environment, urban life, and body monitoring. These sensor-based devices require high performance to react in front of infrequent particular events as well as extreme energy efficiency in order to extend battery lifetime during most of the time when low performance is required. In addition, they require real-time guarantees. The most suitable technological solution for these devices consists of using hybrid processors able to operate at: (i) high voltage to provide high performance and (ii) near-/subthreshold voltage to provide ultra-low energy consumption. However, the most efficient SRAM memories for each voltage level differ and trading off different SRAM designs is mandatory. This is particularly true for cache memories, which occupy most of the processor's area. In this article, we propose new, simple, single-Vcc-domain hybrid L1 cache architectures suitable for reliable hybrid high and ultra-low voltage operation. In particular, the cache is designed by combining heterogeneous SRAM cell types: some of the cache ways are optimized to satisfy high-performance requirements during high voltage operation, whereas the rest of the ways provide ultra-low energy consumption and reliability during near-/subthreshold voltage operation. We analyze the performance, energy, and power impact of the proposed cache designs when using them to implement L1 caches in a processor. Experimental results show that our hybrid caches can efficiently and reliably operate across a wide range of voltages, consuming little energy at near-/subthreshold voltage as well as providing high performance at high voltage without decreasing reliability levels to provide strong performance guarantees, as required for our target market. Bojan Maric, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2013 | On the convergence of mainstream and mission-critical marketsabstractThe computing market has been dominated during the last two decades by the well-known convergence of the high-performance computing market and the mobile market. In this paper we witness a new type of convergence between the mission-critical market (such as avionic or automotive) and the mainstream consumer electronics market. Such convergence is fuelled by the common needs of both markets for more reliability, support for mission-critical functionalities and the challenge of harnessing the unsustainable increases in safety margins to guarantee either correctness or timing. In this position paper, we present a description of this new convergence, as well as the main challenges and opportunities that it brings to computing industry. Sylvain Girbal, Miquel Moretó, Arnaud Grasset, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Sami Yehia |
DAC | 6 |
| 2013 | A cache design for probabilistically analysable real-time systemsabstractCaches provide significant performance improvements, though their use in real-time industry is low because current WCET analysis tools require detailed knowledge of program's cache accesses to provide tight WCET estimates. Probabilistic Timing Analysis (PTA) has emerged as a solution to reduce the amount of information needed to provide tight WCET estimates, although it imposes new requirements on hardware design. At cache level, so far only fully-associative random-replacement caches have been proven to fulfill the needs of PTA, but they are expensive in size and energy. In this paper we propose a cache design that allows set-associative and direct-mapped caches to be analysed with PTA techniques. In particular we propose a novel parametric random placement suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 4 |
| 2013 | Probabilistic timing analysis on conventional cache designsabstractProbabilistic timing analysis (PTA), a promising alternative to traditional worst-case execution time (WCET) analyses, enables pairing time bounds (named probabilistic WCET or pWCET) with an exceedance probability (e.g., 10−16), resulting in far tighter bounds than conventional analyses. However, the applicability of PTA has been limited because of its dependence on relatively exotic hardware: fully-associative caches using random replacement. This paper extends the applicability of PTA to conventional cache designs via a software-only approach. We show that, by using a combination of compiler techniques and runtime system support to randomise the memory layout of both code and data, conventional caches behave as fully-associative ones with random replacement. Leonidas Kosmidis, Charlie Curtsinger, Eduardo Quiñones, Jaume Abella 0001, Emery D. Berger, Francisco J. Cazorla |
DATE | 6 |
| 2013 | parMERASA - Multi-core Execution of Parallelised Hard Real-Time Applications Supporting AnalysabilityabstractEngineers who design hard real-time embedded systems express a need for several times the performance available today while keeping safety as major criterion. A breakthrough in performance is expected by parallelizing hard real-time applications and running them on an embedded multi-core processor, which enables combining the requirements for high-performance with timing-predictable execution. parMERASA will provide a timing analyzable system of parallel hard real-time applications running on a scalable multicore processor. parMERASA goes one step beyond mixed criticality demands: It targets future complex control algorithms by parallelizing hard real-time programs to run on predictable multi-/many-core processors. We aim to achieve a breakthrough in techniques for parallelization of industrial hard real-time programs, provide hard real-time support in system software, WCET analysis and verification tools for multi-cores, and techniques for predictable multi-core designs with up to 64 cores. Theo Ungerer, Christian Bradatsch, Mike Gerdes 0001, Florian Kluge, Ralf Jahr, Jörg Mische, Pavel G. Zaykov, Zlatko Petrov, Bert Böddeker, Sebastian Kehr, Hans Regler, Andreas Hugl, Christine Rochange, Haluk Ozaktas, Hugues Cassé, Armelle Bonenfant, Pascal Sainrat, Ian Broster, Nick Lay, Eduardo Quiñones, Milos Panic, Jaume Abella 0001, Francisco J. Cazorla, Sascha Uhrig, Mathias Rohde, Arthur Pyka |
DSD | 25 |
| 2013 | DTM: Degraded Test Mode for Fault-Aware Probabilistic Timing AnalysisabstractExisting timing analysis techniques to derive Worst-Case Execution Time (WCET) estimates assume that hardware in the target platform (e.g., the CPU) is fault-free. Given the performance requirements increase in current Critical Real-Time Embedded Systems (CRTES), the use of high-performance features and smaller transistors in current and future hardware becomes a must. The use of smaller transistors helps providing more performance while maintaining low energy budgets, however, hardware fault rates increase noticeably, affecting the temporal behaviour of the system in general, and WCET in particular. In this paper, we reconcile these two emergent needs of CRTES, namely, tight (and trustworthy) WCET estimates and the use of hardware implemented with smaller transistors. To that end we propose the Degraded Test Mode (DTM) that, in combination with fault-tolerant hardware designs and probabilistic timing analysis techniques, (i) enables the computation of tight and trustworthy WCET estimates in the presence of faults, (ii) provides graceful average and worst-case performance degradation due to faults, and (iii) requires modifications neither in WCET analysis tools nor in applications. Our results show that DTM allows accounting for the effect of faults at analysis time with low impact in WCET estimates and negligible hardware modifications. Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 5 |
| 2013 | Achieving timing composability with measurement-based probabilistic timing analysisabstractProbabilistic Timing Analysis (PTA) allows complex hardware acceleration features, which defeat classic timing analysis, to be used in hard real-time systems. PTA can do that because it drastically reduces intrinsic dependence on execution history. This distinctive feature is a great facilitator to time composability, which is a must for industry needing incremental development and qualification. In this paper we show how time composability is achieved in PTA-conformant systems and how the pessimism of worst-case execution time bounds obtained from PTA is contained within a 5% to 25% range for representative application scenarios. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla |
ISORC | 5 |
| 2013 | Multi-level Unified Caches for Probabilistically Time Analysable Real-Time SystemsabstractAbstract—Caches are key resources in high-end processor architectures to increase performance. In fact, most high-performance processors come equipped with a multi-level cache hierarchy. In terms of guaranteed performance, however, cache hierarchies severely challenge the computation of tight worst-case execution time (WCET) estimates. On the one hand, the analysis of the timing behaviour of a single level of cache is already challenging, particularly for data accesses. On the other hand, unifying data and instructions in each level, makes the problem of cache analysis significantly more complex requiring tracking simultaneously data and instruction accesses to cache. In this paper we prove that multi-level cache hierarchies can be used in the context of Probabilistic Timing Analysis and tight WCET estimates can be obtained. Our detailed analysis (1) covers unified data and instruction caches, (2) covers different cache-write policies (write-through and write back), write allocation policies (write-allocate and non-write-allocate) and several inclusion mechanisms (inclusive, non-inclusive and exclusive caches), and (3) scales to an arbitrary number of cache levels. Our results show that the probabilistic WCET (pWCET) estimates provided by our analysis technique effectively benefit from having multi-level caches. For a two-level cache configura-tion and for EEMBC benchmarks, pWCET reductions are 55% on average (and up to 90%) with respect to a processor with a single level of cache. I. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
RTSS | 4 |
| 2013 | Hardware support for accurate per-task energy metering in multicore systemsabstractAccurately determining the energy consumed by each task in a system will become of prominent importance in future multicore-based systems because it offers several benefits, including (i) better application energy/performance optimizations, (ii) improved energy-aware task scheduling, and (iii) energy-aware billing in data centers. Unfortunately, existing methods for energy metering in multicores fail to provide accurate energy estimates for each task when several tasks run simultaneously. This article makes a case for accurate Per-Task Energy Metering (PTEM) based on tracking the resource utilization and occupancy of each task. Different hardware implementations with different trade-offs between energy prediction accuracy and hardware-implementation complexity are proposed. Our evaluation shows that the energy consumed in a multicore by each task can be accurately measured. For a 32-core, 2-way, simultaneous multithreaded core setup, PTEM reduces the average accuracy error from more than 12% when our hardware support is not used to less than 4% when it is used. The maximum observed error for any task in the workload we used reduces from 58% down to 9% when our hardware support is used. Qixiao Liu, Miquel Moretó, Víctor Jiménez, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Archit. Code Optim. | 5 |
| 2013 | Fair CPU time accounting in CMP+SMT processorsabstractProcessor architectures combining several paradigms of Thread-Level Parallelism (TLP), such as CMP processors in which each core is SMT, are becoming more and more popular as a way to improve performance at a moderate cost. However, the complex interaction between running tasks in hardware shared resources in multi-TLP architectures introduces complexities when accounting CPU time (or CPU utilization) to tasks. The CPU utilization accounted to a task depends on both the time it runs in the processor and the amount of processor hardware resources it receives. Deploying systems with accurate CPU accounting mechanisms is necessary to increase fairness. Moreover, it will allow users to be fairly charged on a shared data center, facilitating server consolidation in future systems. In this article we analyze the accuracy and hardware cost of previous CPU accounting mechanisms for pure-CMP and pure-SMT processors and we show that they are not adequate for CMP+SMT processors. Consequently, we propose a new accounting mechanism for CMP+SMT processors which: (1) increases the accuracy of accounted CPU utilization; (2) provides much more stable results over a wide range of processor setups; and (3) does not require tracking all hardware shared resources, significantly reducing its implementation cost. In particular, previous proposals lead to inaccuracies between 21% and 79% when measuring CPU utilization in an 8-core 2-way SMT processor, while our proposal reduces this inaccuracy to less than 5.0%. Carlos Luque, Miquel Moretó, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Archit. Code Optim. | 3 |
| 2013 | SMT Malleability in IBM POWER5 and POWER6 ProcessorsabstractWhile several hardware mechanisms have been proposed to control the interaction between hardware threads in an SMT processor, few have addressed the issue of software-controllable SMT performance. The IBM POWER5 and POWER6 are the first high-performance processors implementing a software-controllable hardware-thread prioritization mechanism that controls the rate at which each hardware-thread decodes instructions. This paper shows the potential of this basic mechanism to improve several target metrics for various applications on POWER5 and POWER6 processors. Our results show that although the software interface is exactly the same, the software-controlled priority mechanism has a different effect on POWER5 and POWER6. For instance, hardware threads in POWER6 are less sensitive to priorities than in POWER5 due to the in order design. We study the SMT thread malleability to enable user-level optimizations that leverage software-controlled thread priorities. We also show how to achieve various system objectives such as parallel application load balancing, in order to reduce execution time. Finally, we characterize user-level transparent execution on POWER5 and POWER6, and identify the workload mix that best benefits from it. Alessandro Morari, Carlos Boneti, Francisco J. Cazorla, Roberto Gioiosa, Chen-Yong Cher, Alper Buyuktosunoglu, Pradip Bose, Mateo Valero |
IEEE Trans. Computers | 3 |
| 2013 | PROARTIS: Probabilistically Analyzable Real-Time SystemsabstractStatic timing analysis is the state-of-the-art practice of ascertaining the timing behavior of current-generation real-time embedded systems. The adoption of more complex hardware to respond to the increasing demand for computing power in next-generation systems exacerbates some of the limitations of static timing analysis. In particular, the effort of acquiring (1) detailed information on the hardware to develop an accurate model of its execution latency as well as (2) knowledge of the timing behavior of the program in the presence of varying hardware conditions, such as those dependent on the history of previously executed instructions. We call these problems the timing analysis walls. In this vision-statement article, we present probabilistic timing analysis , a novel approach to the analysis of the timing behavior of next-generation real-time embedded systems. We show how probabilistic timing analysis attacks the timing analysis walls; we then illustrate the mathematical foundations on which this method is based and the challenges we face in the effort of efficiently implementing it. We also present experimental evidence that shows how probabilistic timing analysis reduces the extent of knowledge about the execution platform required to produce probabilistically accurate WCET estimations. Francisco J. Cazorla, Eduardo Quiñones, Tullio Vardanega, Liliana Cucu-Grosjean, Benoit Triquet, Guillem Bernat, Emery D. Berger, Jaume Abella 0001, Franck Wartel, Michael Houston, Luca Santinelli, Leonidas Kosmidis, Code Lo, Dorin Maxim |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | A hard real-time capable multi-core SMT processorabstractHard real-time applications in safety critical domains require high performance and time analyzability. Multi-core processors are an answer to these demands, however task interferences make multi-cores more difficult to analyze from a worst-case execution time point of view than single-core processors. We propose a multi-core SMT processor that ensures a bounded maximum delay a task can suffer due to inter-task interferences. Multiple hard real-time tasks can be executed on different cores together with additional non real-time tasks. Our evaluation shows that the proposed MERASA multi-core provides predictability for hard real-time tasks and also high performance for non hard real-time tasks. Marco Paolieri, Jörg Mische, Stefan Metzlaff, Mike Gerdes 0001, Eduardo Quiñones, Sascha Uhrig, Theo Ungerer, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2013 | Timing effects of DDR memory systems in hard real-time multicore architectures: Issues and solutionsabstractMulticore processors are an effective solution to cope with the performance requirements of real-time embedded systems due to their good performance-per-watt ratio and high performance capabilities. Unfortunately, their use in integrated architectures such as IMA or AUTOSAR is limited by the fact that multicores do not guarantee a time composable behavior for the applications: the WCET of a task depends on inter-task interferences introduced by other tasks running simultaneously. This article focuses on the off-chip memory system: the hardware shared resource with the highest impact on the WCET and hence the main impediment for the use of multicores in integrated architectures. We present an analytical model that computes the worst-case delay, also known as Upper Bound Delay (UBD), that a memory request can suffer due to memory interferences generated by other co-running tasks. By considering the UBD in the WCET analysis, the resulting WCET estimation is independent from the other tasks, hence ensuring the time composability property and enabling the use of multicores in integrated architectures. We propose a memory controller for hard real-time multicores compliant with the analytical model that implements extra hardware features to deal with refresh operations and interferences generated by co-running non hard real-time tasks. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2013 | Thread Assignment of Multithreaded Network Applications in Multicore/Multithreaded ProcessorsabstractThe introduction of multithreaded processors comprised of a large number of cores with many shared resources makes thread scheduling, and in particular optimal assignment of running threads to processor hardware contexts to become one of the most promising ways to improve the system performance. However, finding optimal thread assignments for workloads running in state-of-the-art multicore/multithreaded processors is an NP-complete problem. In this paper, we propose BlackBox scheduler, a systematic method for thread assignment of multithreaded network applications running on multicore/multithreaded processors. The method requires minimum information about the target processor architecture and no data about the hardware requirements of the applications under study. The proposed method is evaluated with an industrial case study for a set of multithreaded network applications running on the UltraSPARC T2 processor. In most of the experiments, the proposed thread assignment method detected the best actual thread assignment in the evaluation sample. The method improved the system performance from 5 to 48 percent with respect to load balancing algorithms used in state-of-the-art OSs, and up to 60 percent with respect to a naive thread assignment. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2012 | Making data prefetch smarter: adaptive prefetching on POWER7abstractHardware data prefetch engines are integral parts of many general purpose server-class microprocessors in the field today. Some prefetch engines allow the user to change some of their parameters. The prefetcher, however, is usually enabled in a default configuration during system bring-up and dynamic reconfiguration of the prefetch engine is not an autonomic feature of current machines. Conceptually, however, it is easy to infer that commonly used prefetch algorithms, when applied in a fixed mode will not help performance in many cases. In fact, they may actually degrade performance due to useless bus bandwidth consumption and cache pollution. In this paper, we present an adaptive prefetch scheme that dynamically modifies the prefetch settings in order to adapt to the workload requirements. We implement and evaluate adaptive prefetching in the context of an existing, commercial processor, namely the IBM POWER7. Our adaptive prefetch mechanism improves performance with respect to the default prefetch setting up to 2.7X and 30% for single-threaded and multiprogrammed workloads, respectively. Víctor Jiménez, Roberto Gioiosa, Francisco J. Cazorla, Alper Buyuktosunoglu, Pradip Bose, Francis P. O'Connell |
PACT | 3 |
| 2012 | Optimal task assignment in multithreaded processors: a statistical approachabstractThe introduction of massively multithreaded (MMT) processors, comprised of a large number of cores with many shared resources, has made task scheduling, in particular task to hardware thread assignment, one of the most promising ways to improve system performance. However, finding an optimal task assignment for a workload running on MMT processors is an NP-complete problem. Due to the fact that the performance of the best possible task assignment is unknown, the room for improvement of current task-assignment algorithms cannot be determined. This is a major problem for the industry because it could lead to: (1)~A waste of resources if excessive effort is devoted to improving a task assignment algorithm that already provides a performance that is close to the optimal one, or (2)~significant performance loss if insufficient effort is devoted to improving poorly-performing task assignment algorithms. Petar Radojkovic, Vladimir Cakarevic, Miquel Moretó, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
ASPLOS | 6 |
| 2012 | Measurement-Based Probabilistic Timing Analysis for Multi-path ProgramsabstractThe rigorous application of static timing analysis requires a large and costly amount of detail knowledge on the hardware and software components of the system. Probabilistic Timing Analysis has potential for reducing the weight of that demand. In this paper, we present a sound measurement-based probabilistic timing analysis technique based on Extreme Value Theory. In all the experiments made as part of this work, the timing bounds determined by our technique were less than 15% pessimistic in comparison with the tightest possible bounds obtainable with any probabilistic timing analysis technique. As a point of interest to industrial users, our technique also requires a comparatively low number of measurement runs of the program under analysis, less than 650 runs were needed for the benchmarks presented in this paper. Liliana Cucu-Grosjean, Luca Santinelli, Michael Houston, Code Lo, Tullio Vardanega, Leonidas Kosmidis, Jaume Abella 0001, Enrico Mezzetti, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 10 |
| 2012 | Assessing the suitability of the NGMP multi-core processor in the space domainabstractMulti-core processors are increasingly being considered as a means to provide the performance required by future safety-critical embedded systems. In this line, Aeroflex Gaisler has developed, in conjunction with the European Space Agency, the NGMP, a quad-core processor to be used in the future space missions of the Agency. Unfortunately, the use of multi-core processors in industrial domains is not straightforward since it poses various challenges on the timing behavior of the system. This is mainly due to the interferences tasks suffer when accessing hardware shared resources and which can affect their WCET. Although the effect of inter-task interferences in multi-core shared resources on real-time applications has received attention from academia, most of the solutions proposed require hardware changes. The lack of quantitative studies of the slowdown on applications' performance caused by inter-task interferences on real COTS multi-core processors, limit their use by industry. Mikel Fernández, Roberto Gioiosa, Eduardo Quiñones, Luca Fossati, Marco Zulianello, Francisco J. Cazorla |
EMSOFT | 6 |
| 2012 | Kernel Partitioning of Streaming Applications: A Statistical Approach to an NP-complete ProblemabstractOne of the greatest challenges in computer architecture is how to write efficient, portable, and correct software for multi-core processors. A promising approach is to expose more parallelism to the compiler, through the use of domain-specific languages. The compiler can then perform complex transformations that the programmer would otherwise have had to do. Many important applications related to audio and video encoding, software radio and signal processing have regular behavior that can be represented using a stream programming language. When written in such a language, a portable stream program can be automatically mapped by the stream compiler onto multicore hardware. One of the most difficult tasks of the stream compiler is partitioning the stream program into software threads. The choice of partition significantly affects performance, but finding the optimal partition is an NP-complete problem. This paper presents a method, based on Extreme Value theory (EVT), that statistically estimates the performance of the optimal partition. Knowing the optimal performance improves the evaluation of any partitioning algorithm, and it is the most important piece of information when deciding whether an existing algorithm should be enhanced. We use the method to evaluate a recently-published partitioning algorithm based on a heuristic. We further analyze how the statistical method is affected by the choice of sampling method, and we recommend how sampling should be done. Finally, since a heuristic-based algorithm may not always be available, the user may try to find a good partition by picking the best from a random sample. We analyze whether this approach is likely to find a good partition. To the best of our knowledge, this study is the first application of EVT to a graph partitioning problem. Petar Radojkovic, Paul M. Carpenter, Miquel Moretó, Alex Ramírez, Francisco J. Cazorla |
MICRO | 5 |
| 2012 | On the evaluation of the impact of shared resources in multithreaded COTS processors in time-critical environmentsabstractCommercial Off-The-Shelf (COTS) processors are now commonly used in real-time embedded systems. The characteristics of these processors fulfill system requirements in terms of time-to-market, low cost, and high performance-per-watt ratio. However, multithreaded (MT) processors are still not widely used in real-time systems because the timing analysis is too complex. In MT processors, simultaneously-running tasks share and compete for processor resources, so the timing analysis has to estimate the possible impact that the inter-task interferences have on the execution time of the applications. In this paper, we propose a method that quantifies the slowdown that simultaneously-running tasks may experience due to collision in shared processor resources. To that end, we designed benchmarks that stress specific processor resources and we used them to (1) estimate the upper limit of a slowdown that simultaneously-running tasks may experience because of collision in different shared processor resources, and (2) quantify the sensitivity of time-critical applications to collision in these resources. We used the presented method to determine if a given MT processor is a good candidate for systems with timing requirements. We also present a case study in which the method is used to analyze three multithreaded architectures exhibiting different configurations of resource sharing. Finally, we show that measuring the slowdown that real applications experience when simultaneously-running with resource-stressing benchmarks is an important step in measurement-based timing analysis. This information is a base for incremental verification of MT COTS architectures. Petar Radojkovic, Sylvain Girbal, Arnaud Grasset, Eduardo Quiñones, Sami Yehia, Francisco J. Cazorla |
ACM Trans. Archit. Code Optim. | 6 |
| 2012 | CPU Accounting for Multicore ProcessorsabstractIn single-threaded processors and Symmetric Multiprocessors the execution time of a task depends on the other tasks it runs with (the workload), since the Operating System (OS) time shares the CPU(s) between tasks in the workload. However, the time accounted to a task is roughly the same regardless of the workload in which the task runs in, since the OS takes into account those periods in which the task is not scheduled onto a CPU. Chip Multiprocessors (CMPs) introduce complexities when accounting CPU utilization, since the CPU time to account to a task not only depends on the time that the task is scheduled onto a CPU, but also on the amount of hardware resources it receives during that period. And given that in a CMP hardware resources are dynamically shared between tasks, the CPU time accounted to a task in a CMP depends on the workload it executes in. This is undesirable because the same task with the same input data set may be accounted differently depending on the workload it executes. In this paper, we identify how an inaccurate measurement of the CPU utilization affects several key aspects of the system such as OS statistics or the charging mechanism in data centers. We propose a new hardware CPU accounting mechanism to improve the accuracy when measuring the CPU utilization in CMPs and compare it with the previous accounting mechanisms. Our results show that currently known mechanisms lead to a 16 percent average error when it comes to CPU utilization accounting. Our proposal reduces this error to less than 2.8 percent in a modeled 8-core processor system. Carlos Luque, Miquel Moretó, Francisco J. Cazorla, Roberto Gioiosa, Alper Buyuktosunoglu, Mateo Valero |
IEEE Trans. Computers | 3 |
| 2011 | RVC: a mechanism for time-analyzable real-time processors with faulty cachesabstractGeometry scaling due to technology evolution as well as Vcc scaling lead to failures in large SRAM arrays such as caches. Faulty bits can be tolerated from the average performance perspective, but make critical realtime embedded systems non time-analyzable or worstcase execution time (WCET) estimations unacceptably large. Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Yiannakis Sazeides, Mateo Valero |
HiPEAC | 3 |
| 2011 | Towards improved survivability in safety-critical systemsabstractPerformance demand of Critical Real-Time Embedded (CRTE) systems implementing safety-related system features grows at an exponential rate. Only modern semiconductor technologies can satisfy CRTE systems performance needs efficiently. However, those technologies lead to high failure rates, thus lowering survivability of chips to unacceptable levels for CRTE systems. This paper presents SESACS architecture (Surviving Errors in SAfety-Critical Systems), a paradigm shift in the design of CRTE systems. SESACS is a new system design methodology consisting of three main components: (i) a multicore hardware/firmware platform capable of detecting and diagnosing hardware faults of any type with minimal impact on the worst-case execution time (WCET), recovering quickly from errors, and properly reconfiguring the system so that the resulting system exhibits a predictable and analyzable degradation in WCET; (ii) a set of analysis methods and tools to prove the timing correctness of the reconfigured system; and (iii) a white-box methodology and tools to prove the functional safety of the system and compliance with industry standards. This new design paradigm will deliver huge benefits to the embedded systems industry for several decades by enabling the use of more cost-effective multicore hardware platforms built on top of modern semiconductor technologies, thereby enabling higher performance, and reducing weight and power dissipation. This new paradigm will further extend the life of embedded systems, therefore, reducing warranty and early replacement costs. Jaume Abella 0001, Francisco J. Cazorla, Eduardo Quiñones, Arnaud Grasset, Sami Yehia, Philippe Bonnot 0001, Dimitris Gizopoulos, Riccardo Mariani, Guillem Bernat |
IOLTS | 2 |
| 2011 | RVC-based time-predictable faulty caches for safety-critical systemsabstractTechnology and Vcc scaling lead to significant faulty bit rates in caches. Mechanisms based on disabling faulty parts show to be effective for average performance but are unacceptable in safety critical systems where worst-case execution time (WCET) estimations must be safe and tight. The Reliable Victim Cache (RVC) deals with this issue for a large fraction of the cache bits. However, replacement bits are not protected, thus keeping the probability of failure still high. This paper proposes two mechanisms to tolerate faulty bits in replacement bits and keep time-predictability by extending the RVC. Our solutions offer different tradeoffs between cost and complexity. In particular, the Extended RVC (ERVC) has low energy and area overheads while keeping complexity at a minimum. The Reliable Replacement Bits (RRB) solution has even lower overheads at the expense of some more wiring complexity. Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Mateo Valero, Yiannakis Sazeides |
IOLTS | 3 |
| 2011 | A Quantitative Analysis of OS NoiseabstractOperating system noise is a well-known problem that may limit application scalability on large-scale machines, significantly reducing their performance. Though the problem is well studied, much of the previous work has been qualitative. We have developed a technique to provide a quantitative descriptive analysis for each OS event that contributes to OS noise. The mechanism allows us to detail all sources of OS noise through precise kernel instrumentation and provides frequency and duration analysis for each event. Such a description gives OS developers better guidance for reducing OS noise. We integrated this data with a trace visualizer allowing quicker and more intuitive understanding of the data. Specifically, the contributions of this paper are three-fold. First, we describe a methodology whereby detailed quantitative information may be obtained for each OS noise event. Though not the thrust of the paper, we show how we implemented that methodology by augmenting LTTng. We validate our approach by comparing it to other well-known standard techniques to analyze OS noise. Second, we provide a case study in which we use our methodology to analyze the OS noise when running benchmarks from the LLNL Sequoia applications. Our experiments enrich and expand previous results with our quantitative characterization. Third, we describe how a detailed characterization permits to disambiguate noise signatures of qualitatively similar events, allowing developers to address the true cause of each noise event. Alessandro Morari, Roberto Gioiosa, Robert W. Wisniewski, Francisco J. Cazorla, Mateo Valero |
IPDPS | 4 |
| 2011 | A Software-Pipelined Approach to Multicore Execution of Timing Predictable Multi-threaded Hard Real-Time TasksabstractMulticore processors can deliver higher performance than single-core processors by exploiting thread level parallelism (TLP): applications are split into independent threads, each of which is mapped into a different core, reducing the execution time and potentially its worst-case execution time (WCET). Unfortunately, inter-thread interferences generated by simultaneous accesses to shared resources from different threads may completely destroy the performance benefits brought by TLP. This paper proposes a software/hardware cache partitioning approach that reduces the inter-thread memory interferences generated in hard real-time software-pipelined parallel applications. Our results show that our approach effectively reduces memory interferences, while still guaranteeing a predictable timing behaviour, achieving a WCET estimation reduction of 28% for a software pipelined version of the LU decomposition application with respect to the single-threaded version. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Julian Wolf 0002, Theo Ungerer, Sascha Uhrig, Zlatko Petrov |
ISORC | 3 |
| 2011 | IA^3: An Interference Aware Allocation Algorithm for Multicore Hard Real-Time SystemsabstractIn multicore processors, the execution environment is defined as the environment in which tasks run and it is determined by the hardware resources they get and the workload with which they are executed. Thus, different execution environments lead to different inter-task interferences accessing shared hardware resources due to conflicts with the other corunning tasks, making the WCET estimation of a task dependent on the execution environment in which it runs. Despite such dependency, current partitioned scheduling approaches use a single WCET estimation per task: typically the highest for all execution environments in which a task runs. In this paper we introduce IA3: an interference-aware allocation algorithm that considers not a single WCET estimation but a set of WCET estimations per task. IA3 is based on two novel concepts: the WCET-matrix and the WCET-sensitivity. The former associates every WCET estimation with its corresponding execution environment. The latter measures the impact of changing the execution environment on the WCET estimation. This allows IA3 to reduce the number of resources required to schedule a given taskset. In particular, our results show that in a four-core processor considering tasksets with a total utilization of 2.9, IA3 is able to schedule 70% of the tasksets using 3-cores while a classical partitioned approach with a First-Fit Decreasing heuristic is able to schedule only 5% of the tasksets using 3-cores. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Robert I. Davis 0001, Mateo Valero |
IEEE Real-Time and Embedded Technology and Applications Symposium | 3 |
| 2010 | Power and thermal characterization of POWER6 systemabstractControlling power consumption and temperature is of major concern for modern computing systems. In this work we characterize thermal behavior and power consumption of an IBM POWER6-based system. We perform the characterization at several levels: application, operating system, and hardware level, both when the system is idle, and under load. At hardware level, we report a 25% reduction in total system power consumption by using the processor low power mode. We also study the effect of the hardware thread prioritization mechanism provided by POWER6 on different workloads and how this mechanism can be used to limit power consumption. At OS level, we analyze the power reduction techniques implemented in the Linux kernel, such as the tickless kernel and the CPU idle power manager. At application level, we characterize the power consumption and the temperature of two sets of benchmarks (METbench and SPEC CPU2006) and we study the effect of workload characteristics on power consumption and core temperature. Víctor Jiménez, Francisco J. Cazorla, Roberto Gioiosa, Mateo Valero, Carlos Boneti, Eren Kursun, Chen-Yong Cher, Canturk Isci, Alper Buyuktosunoglu, Pradip Bose |
PACT | 2 |
| 2010 | Adapting cache partitioning algorithms to pseudo-LRU replacement policiesabstractRecent studies have shown that cache partitioning is an efficient technique to improve throughput, fairness and Quality of Service (QoS) in CMP processors. The cache partitioning algorithms proposed so far assume Least Recently Used (LRU) as the underlying replacement policy. However, it has been shown that the true LRU imposes extraordinary complexity and area overheads when implemented on high associativity caches, such as last level caches. As a consequence, current processors available on the market use pseudo-LRU replacement policies, which provide similar behavior as LRU, while reducing the hardware complexity. Thus, the presented so far LRU-based cache partitioning solutions cannot be applied to real CMP architectures. This paper proposes a complete partitioning system for caches using the pseudo-LRU replacement policy. In particular, the paper focuses on the pseudo-LRU implementations proposed by Sun Microsystems and IBM, called Not Recently Used (NRU) and Binary Tree (BT), respectively. We propose a high accuracy profiling logic and a cache partitioning hardware for both schemes. We evaluate our proposals' hardware costs in terms of area and power, and compare them against the LRU partitioning algorithm. Overall, this paper presents two hardware techniques to adapt the existing cache partitioning algorithms to real replacement policies. The results show that our solutions impose negligible performance degradation with respect to the LRU. Kamil Kedzierski, Miquel Moretó, Francisco J. Cazorla, Mateo Valero |
IPDPS | 3 |
| 2010 | Thread to strand binding of parallel network applications in massive multi-threaded systemsabstractIn processors with several levels of hardware resource sharing,like CMPs in which each core is an SMT, the scheduling process becomes more complex than in processors with a single level of resource sharing, such as pure-SMT or pure-CMP processors. Once the operating system selects the set of applications to simultaneously schedule on the processor (workload), each application/thread must be assigned to one of the hardware contexts(strands). We call this last scheduling step the Thread to Strand Binding or TSB. In this paper, we show that the TSB impact on the performance of processors with several levels of shared resources is high. We measure a variation of up to 59% between different TSBs of real multithreaded network applications running on the UltraSPARC T2 processor which has three levels of resource sharing. In our view, this problem is going to be more acute in future multithreaded architectures comprising more cores, more contexts per core, and more levels of resource sharing. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
PPoPP | 5 |
| 2010 | Trends and techniques for energy efficient architecturesabstractMicroprocessor architectures have become increasingly power limited in recent years. Currently power and thermal envelopes dictate peak performance limits more than any other design constraint. As voltage scaling has slowed down, innovative techniques have been pursued to improve the power efficiency of the increasingly demanding multi-core architectures. In this paper we look at recent trends in multi-cores with a special focus on trends and techniques to address these challenges. Víctor Jiménez, Roberto Gioiosa, Eren Kursun, Francisco J. Cazorla, Chen-Yong Cher, Alper Buyuktosunoglu, Pradip Bose, Mateo Valero |
VLSI-SoC | 4 |
| 2010 | On the Problem of Evaluating the Performance of Multiprogrammed WorkloadsabstractMultithreaded architectures are becoming more and more popular. In order to evaluate their behavior, several methodologies and metrics have been proposed. A methodology defines when the measurements for a given workload execution are taken. A metric combines those measurements to obtain a final evaluation result. However, since current evaluation methodologies do not provide representative measurements for these metrics, the analysis and evaluation of novel ideas could be either unfair or misleading. Given the potential impact of multithreaded architectures on current and future processor designs, it is crucial to develop an accurate evaluation methodology for them. This paper presents FAME, a new evaluation methodology aimed to fairly measure the performance of multithreaded processors executing multiprogrammed workloads. FAME reexecutes all programs in the workload until all of them are fairly represented in the final measurements taken. We compare FAME with previously used methodologies showing that it provides more accurate measurements, becoming an ideal evaluation methodology to analyze proposals for multithreaded architectures. Francisco J. Cazorla, Alex Pajuelo, Oliverio J. Santana, Enrique Fernández, Mateo Valero |
IEEE Trans. Computers | 1 |
| 2009 | ITCA: Inter-task Conflict-Aware CPU Accounting for CMPsabstractChip-multiprocessor (CMP) architectures are becoming more and more popular as an alternative to the traditional processors that only extract instruction-level parallelism from an application. CMPs introduce complexities when accounting CPU utilization. This is due to the fact that the progress done by an application during an interval of time highly depends on the activity of the other applications it is co-scheduled with. In this paper, we identify how an inaccurate measurement of the CPU utilization affects several key aspects of the system like the application scheduling or the charging mechanism in data centers. We propose a new hardware CPU accounting mechanism to improve the accuracy when measuring the CPU utilization in CMPs and compare it with the previous accounting mechanisms. Our results show that currently known mechanisms lead to a 19% average error when it comes to CPU utilization accounting. Our proposal reduces this error to less than 1% in a modeled 4-core processor system. Carlos Luque, Miquel Moretó, Francisco J. Cazorla, Roberto Gioiosa, Alper Buyuktosunoglu, Mateo Valero |
PACT | 3 |
| 2009 | Using Randomized Caches in Probabilistic Real-Time SystemsabstractWhile hardware caches are generally effective at improving application performance, they greatly complicate performance prediction. Slight changes in memory layout or data access patterns can lead to large and systematic increases in cache misses, degrading performance. In the worst case, these misses can effectively render the cache useless. These pathological cases, or ``cache risk patterns'', are difficult to predict, test or debug, and their presence limits the usefulness of caches in safety critical real-time systems, especially in hard real-time environments.In this paper, we explore the effect of randomized cache replacement policies in real-time systems with stringent timing constrains. We present simulation-based results on representative examples that illustrate the problem of performance anomalies with standard cache replacement policies. We show that, by eliminating dependencies on access history, randomized replacement greatly reduces the risk of these cache-based performance anomalies, enabling probabilistic worst-case execution time analysis. Eduardo Quiñones, Emery D. Berger, Guillem Bernat, Francisco J. Cazorla |
ECRTS | 4 |
| 2009 | Hardware support for WCET analysis of hard real-time multicore systemsabstractThe increasing demand for new functionalities in current and future hard real-time embedded systems like automotive, avionics and space industries is driving an increase in the performance required in embedded processors. Multicore processors represent a good design solution for such systems due to their high performance, low cost and power consumption characteristics. However, hard real-time embedded systems require time analyzability and current multicore processors are less analyzable than single-core processors due to the interferences between different tasks when accessing shared hardware resources. In this paper we propose a multicore architecture with shared resources that allows the execution of applications with hard real-time and non hard real-time constraints at the same time, providing time analizability for the hard real-time tasks so that they can meet their deadlines. Moreover our architecture proposal provides high-performance for the non hard real-time tasks. Marco Paolieri, Eduardo Quiñones, Francisco J. Cazorla, Guillem Bernat, Mateo Valero |
ISCA | 3 |
| 2009 | Characterizing the resource-sharing levels in the UltraSPARC T2 processorabstractThread level parallelism (TLP) has become a popular trend to improve processor performance, overcoming the limitations of extracting instruction level parallelism. Each TLP paradigm, such as Simultaneous Multithreading or Chip-Multiprocessors, provides di erent bene ts, which has motivated processor vendors to combine several TLP paradigms in each chip design. Even if most of these combined-TLP designs are homogeneous, they present di erent levels of hardware resource sharing, which introduces complexities on the operating system scheduling and load balancing.\nCommonly, processor designs provide two levels of resource sharing: Inter-core in which only the highest levels of the cache hierarchy are shared, and Intracore in which\nmost of the hardware resources of the core are shared . Recently, Sun Microsystems has released the UltraSPARC T2, a processor with three levels of hardware resource sharing:\nInterCore, IntraCore, and IntraPipe. In this work, we provide the rst characterization of a three-level resource sharing processor, the UltraSPARC T2, and we show how\nmulti-level resource sharing a ects the operating system design. We further identify the most critical hardware resources in the T2 and the characteristics of applications that are not sensitive to resource sharing. Finally, we present a case study in which we run a real multithreaded network application, showing that a resource sharing aware scheduler can improve the system throughput up to 55%. Vladimir Cakarevic, Petar Radojkovic, Javier Verdú, Alex Pajuelo, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
MICRO | 5 |
| 2009 | Thread to Core Assignment in SMT On-Chip MultiprocessorsabstractState-of-the-art high-performance processors like the IBM POWER5 and Intel i7 show a trend in industry towards on-chip Multiprocessors (CMP) involving Simultaneous Multithreading (SMT) in each core. In these processors, the way in which applications are assigned to cores plays a key role in the performance of each application and the overall system performance. In this paper we show that the system throughput highly depends on the Thread to Core Assignment (TCA), regardless the SMT Instruction Fetch (IFetch) Policy implemented in the cores. Our results indicate that a good TCA can improve the results of any underlying IFetch Policy, yielding speedups of up to 28%. Given the relevance of TCA, we propose an algorithm to manage it in CMP+SMT processors. The proposed throughput-oriented TCA Algorithm takes into account the workload characteristics and the underlying SMT IFetch Policy. Our results show that the TCA Algorithm obtains thread-to-core assignments 3% close to the optimal assignation for each case, yielding system throughput improvements up to 21%. Carmelo Acosta, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
SBAC-PAD | 2 |
| 2008 | Evolutionary system for prediction and optimization of hardware architecture performanceabstractThe design of computer architectures is a very complex problem. The multiple parameters make the number of possible combinations extremely high.Many researchers have used simulation, although it is a slow solution since evaluating a single point of the search space can take hours. In this work we propose using evolutionary multilayer perceptron (MLP) to compute the performance of an architecture parameter settings. Instead of exploring the search space, simulating many configurations, our method randomly selects some architecture configurations; those are simulated to obtain their performance, and then an artificial neural network is trained to predict the remaining configurations performance. Results obtained show a high accuracy of the estimations using a simple method to select the configurations we have to simulate to optimize the MLP. In order to explore the search space, we have designed a genetic algorithm that uses the MLP as fitness function to find the niche where the best architecture configurations (those with higher performance) are located. Our models need only a small fraction of the design space, obtaining small errors and reducing required simulation by two orders of magnitude. Pedro A. Castillo, Juan Julián Merelo Guervós, Miquel Moretó, Francisco J. Cazorla, Mateo Valero, Antonio Mora García, Juan Luis Jiménez Laredo, Sally A. McKee |
IEEE Congress on Evolutionary Computation | 4 |
| 2008 | MLP-Aware Dynamic Cache Partitioning
Miquel Moretó, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
HiPEAC | 2 |
| 2008 | MFLUSH: Handling Long-Latency Loads in SMT On-Chip MultiprocessorsabstractNowadays, there is a clear trend in industry towards employing the growing amount of transistors on chip in replicating execution cores (CMP), where each core is simultaneous multithreading (SMT). State-of-the-art high-performance processors like the IBM POWER5 and POWER6 corroborate this CMP+SMT trend. Within each SMT core any of the well-known SMT mechanisms may be applied to face SMT related challenges. Among them, probably the most important issue in an SMT execution pipeline concerns the instruction fetch (IFetch) Policy. The FLUSH IFetch Policy represents a choice for throughput-oriented scenarios. It handles L2 cache misses in order to avoid hardware resource monopolization by any given execution thread; involving an additional energy cost via instruction refetching. However, the new constraints imposed by the CMP+SMT scenario may affect well-known SMT mechanisms, like the FLUSH mechanism. In this paper we revisit the FLUSH mechanism and analyze its application in the emerging CMP+SMT scenario. The included analysis points out the new difficulties to be faced by the FLUSH mechanism in the emerging CMP+SMT scenario. Then we propose a novel IFetch Policy designed to cope with the CMP+SMT scenario: the MFLUSH. We also include a complete evaluation of the MFLUSH policy, both in terms of throughput and energy consumption. Our results indicate that the MFLUSH, specifically designed for the emerging CMP+SMT scenario, succeeds not only in overcoming the specific CMP+SMT constraints but also allowing a 20% energy consumption reduction without a significant system throughput loss. Carmelo Acosta, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
ICPP | 2 |
| 2008 | Balancing HPC applications through smart allocation of resources in MT processorsabstractMany studies have shown that load imbalancing causes significant performance degradation in high performance computing (HPC) applications. Nowadays, multi-threaded (MT1) processors are widely used in HPC for their good performance/energy consumption and performance/cost ratios achieved sharing internal resources, like the instruction window or the physical register. Some of these processors provide the software hardware mechanisms for controlling the allocation of processor's internal resources. In this paper, we show, for the first time, that by appropriately using these mechanisms, we are able to control the tasks speed, reducing the imbalance in parallel applications transparently to the user and, hence, reducing the total execution time. Our results show that our proposal leads to a performance improvement up to 18% for one of the NAS benchmark. For a real HPC application (much more dynamic than the benchmark) the performance improvement is 8.1%. Our results also show that, if resource allocation is not used properly, the imbalance of applications is worsened causing performance loss. Carlos Boneti, Roberto Gioiosa, Francisco J. Cazorla, Julita Corbalán, Jesús Labarta, Mateo Valero |
IPDPS | 3 |
| 2008 | Software-Controlled Priority Characterization of POWER5 ProcessorabstractDue to the limitations of instruction-level parallelism, thread-level parallelism has become a popular way to improve processor performance. One example is the IBM POWER5TM processor, a two-context simultaneous-multithreaded dual-core chip. In each SMT core, the IBM POWER5 features two levels of thread resource balancing and prioritization. The first level provides automatic in-hardware resource balancing, while the second level is a software-controlled priority mechanism that presents eight levels of thread priorities. Currently, software-controlled prioritization is only used in limited number of cases in the software platforms due to lack of performance characterization of the effects of this mechanism. In this work, we characterize the effects of the software-based prioritization on several different workloads. We show that the impact of the prioritization significantly depends on the workloads coscheduled on a core. By prioritizing the right task, it is possible to obtain more than two times of throughput improvement for synthetic workloads compared to the baseline. We also present two application case studies targeting two different performance metrics: the first case study improves overall throughput by 23.7% and the second case study reduces the total execution time by 9.3%. In addition, we show the circumstances when a background thread can be run transparently without affecting the performance of the foreground thread. Carlos Boneti, Francisco J. Cazorla, Roberto Gioiosa, Alper Buyuktosunoglu, Chen-Yong Cher, Mateo Valero |
ISCA | 2 |
| 2008 | A Two-Level Load/Store Queue Based on Execution LocalityabstractMulticore processors have emerged as a powerful platform on which to efficiently exploit thread-level parallelism (TLP). However, due to Amdahl’s Law, such designs will be increasingly limited by the remaining sequential components of applications. To overcome this limitation it is necessary to design processors with many lower–performance cores for TLP and some high-performance cores designed to execute sequential algorithms. Such cores will need to address the memory-wall by implementing kilo-instruction windows. Large window processors require large Load/Store Queues that would be too slow if implemented using current CAMbased designs. This paper proposes an Epoch-based Load Store Queue (ELSQ), a new design based on Execution Locality. It is integrated into a large-window processor that has a fast, out-of-order core operating only on L1/L2 cache hits and N slower cores that process L2 misses and their dependent instructions. The large LSQ is coupled with the slow cores and is partitioned into N small and local LSQs, one per core. We evaluate ELSQ in a large-window environment, finding that it enables high performance at low power. By exploiting locality among loads and stores, ELSQ outperforms even an idealized central LSQ when implemented on top of a decoupled processor design. Miquel Pericàs, Adrián Cristal, Francisco J. Cazorla, Rubén González 0001, Alexander V. Veidenbaum, Daniel A. Jiménez, Mateo Valero |
ISCA | 3 |
| 2008 | Selection of the Register File Size and the Resource Allocation Policy on SMT ProcessorsabstractThe performance impact of the Physical Register File(PRF) size on Simultaneous Multithreading processors has not been extensively studied in spite of being a critical shared resource. In this paper we analyze the effect on performance of the PRF size for a broad set of resource allocation policies (Icount, Stall, Flush, Flush++, Static,Dcra and Hill-climbing) and evaluate them under two metrics: instructions per second (IPS) for throughput and harmonic mean of weighted IPCs (Hmean-wIPC) for fairness. We have found that resource allocation policy and PRF size should be considered together in order to obtain the best score in the proposed metrics. For instance, for the analyzed 2 and 4-threaded SPEC CPU2000 workloads,small PRFs are best managed by Flush, whereas for larger PRFs, Hill-climbing and Static lead to the best values for the throughput and fairness metrics, respectively.The second contribution of this work is a simple procedure that, for a given resource allocation policy, selects the PRF size that maximizes IPS and obtains for Hmean-wIPC a value close to its maximum. According to our results, Hill-climbing with a 320-entry PRF achieves the best figures for 2-threaded workloads. When executing 4-threaded workloads, Hill-Climbing with a 384-entry PRF achieves the best throughput whereas Static obtains the best throughput-fairness balance. Jesús Alastruey-Benedé, Teresa Monreal Arnal, Francisco J. Cazorla, Víctor Viñals, Mateo Valero |
SBAC-PAD | 3 |
| 2008 | Measuring Operating System Overhead on CMT ProcessorsabstractNumerous studies have shown that Operating System (OS)noise is one of the reasons for significant performance degradation in clustered architectures. Although many studies examine the OS noise for High Performance Computing (HPC),especially in multi-processor/core systems, most of them focus on 2- or 4-core systems.In this paper, we analyze the major sources of OS noise on a massive multithreading processor, the Sun UltraSPARC T1,running Linux and Solaris. Since a real system is too complex to analyze, we compare those results with a low-overhead runtime environment: the Netra Data Plane Software Suite (Netra DPS).Our results show that the overhead introduced by the OStimer interrupt in Linux and Solaris depends on the particular core and hardware context in which the application is running. This overhead is up to 30% when the application is executed on the same hardware context of the timer interrupt handler and up to 10% when the application and the timer interrupt handler run on different contexts but on the same core.We detect no overhead when the benchmark and the timer interrupt handler run on different cores of the processor. Petar Radojkovic, Vladimir Cakarevic, Javier Verdú, Alex Pajuelo, Roberto Gioiosa, Francisco J. Cazorla, Mario Nemirovsky, Mateo Valero |
SBAC-PAD | 6 |
| 2008 | A dynamic scheduler for balancing HPC applicationsabstractLoad imbalance cause significant performance degradation in High Performance Computing applications. In our previous work we showed that load imbalance can be alleviated by modern MT processors that provide mechanisms for controlling the allocation of processors internal resources. In that work, we applied static, hand-tuned resource allocations to balance HPC applications, providing improvements for benchmarks and real applications. In this paper we propose a dynamic process scheduler for the Linux kernel that automatically and transparently balances HPC applications according to their behavior. We tested our new scheduler on an IBM POWER5 machine, which provides a software-controlled prioritization mechanism that allows us to bias the processor resource allocation. Our experiments show that the scheduler reduces the imbalance of HPC applications, achieving results similar to the ones obtained by hand-tuning the applications (up to 16%). Moreover, our solution reduces the application's execution time combining effect of load balance and high responsive scheduling. Carlos Boneti, Roberto Gioiosa, Francisco J. Cazorla, Mateo Valero |
SC | 3 |
| 2007 | MLP-Aware Dynamic Cache Partitioning
Miquel Moretó, Francisco J. Cazorla, Alex Ramírez, Mateo Valero |
PACT | 2 |
| 2007 | A Flexible Heterogeneous Multi-Core Architecture
Miquel Pericàs, Adrián Cristal, Francisco J. Cazorla, Rubén González 0001, Daniel A. Jiménez, Mateo Valero |
PACT | 3 |
| 2007 | FAME: FAirly MEasuring Multithreaded Architectures
Javier Vera, Francisco J. Cazorla, Alex Pajuelo, Oliverio J. Santana, Enrique Fernández, Mateo Valero |
PACT | 2 |
| 2006 | Predictable Performance in SMT Processors: Synergy between the OS and SMTsabstractCurrent operating systems (OS) perceive the different contexts of simultaneous multithreaded (SMT) processors as multiple independent processing units, although, in reality, threads executed in these units compete for the same hardware resources. Furthermore, hardware resources are assigned to threads implicitly as determined by the SMT instruction fetch (Ifetch) policy, without the control of the OS. Both factors cause a lack of control over how individual threads are executed, which can frustrate the work of the job scheduler. This presents a problem for general purpose systems, where the OS job scheduler cannot enforce priorities, and also for embedded systems, where it would be difficult to guarantee worst-case execution times. In this paper, we propose a novel strategy that enables a two-way interaction between the OS and the SMT processor and allows the OS to run jobs at a certain percentage of their maximum speed, regardless of the workload in which these jobs are executed. In contrast to previous approaches, our approach enables the OS to run time-critical jobs without dedicating all internal resources to them so that non-time-critical jobs can make significant progress as well and without significantly compromising overall throughput. In fact, our mechanism, in addition to fulfilling OS requirements, achieves 90 percent of the throughput of one of the best currently known fetch policies for SMTs. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
IEEE Trans. Computers | 1 |
| 2005 | Architectural support for real-time task scheduling in SMT processorsabstractIn Simultaneous Multithreaded (SMT) architectures most hardware resources are shared between threads. This provides a good cost/performance trade-off which renders these architectures suitable for use in embedded systems. However, since threads share many resources, they also interfere with each other. As a result, execution times of applications become highly unpredictable and dependent on the context in which an application is executed. Obviously, this poses problems if an SMT is to be used in a real-time system.In this paper, we propose two novel hardware mechanisms that can be used to reduce this performance variability. In contrast to previous approaches, our proposed mechanisms do not need any information beyond the information already known by traditional job schedulers. Nor do they require extensive profiling of workloads to determine optimal schedules. Our mechanisms are based on dynamic resource partitioning. The OS level job scheduler needs to be slightly adapted in order to provide the hardware resource allocator some information on how this resource partitioning needs to be done. We show that our mechanisms provide high stability for SMT architectures to be used in real-time systems: the real time benchmarks we used meet their deadlines in more than 98% of the cases considered while the other thread in the workload still achieves high throughput. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
CASES | 1 |
| 2004 | Implicit vs. Explicit Resource Allocation in SMT ProcessorsabstractIn a simultaneous multithreaded (SMT) architecture, the front end of a superscalar is adapted in order to be able to fetch from several threads while the back end is shared among the threads. In this paper, we describe different resource sharing models in SMT processors. We show that explicit resource allocation can improve SMT performance. In addition, it enables SMTs to solve other QoS requirements, not realizable before. Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
DSD | 1 |
| 2004 | Feasibility of QoS for SMT
Francisco J. Cazorla, Peter M. W. Knijnenburg, Rizos Sakellariou, Enrique Fernández, Alex Ramírez, Mateo Valero |
Euro-Par | 1 |
| 2004 | DCache Warn: An I-Fetch Policy to Increase SMT EfficiencyabstractSummary form only given. Simultaneous multithreading (SMT) processors increase performance by executing instructions from multiple threads simultaneously. These threads share the processor's resources, but also compete for them. In this environment, a thread missing in the L2 cache may allocate a large number of resources for a long time, causing other threads to run much slower than they could. To prevent this problem we should know in advance if a thread is going to miss in the L2 cache. L1 misses are a clear indicator of a possible L2 miss. However, to stall a thread on every L1 miss is too severe, because not all L1 misses lead to an L2 miss, and this would cause an unnecessary stall and resource under-use. Also, to wait until an L2 miss is declared and squash the thread to free up the allocated resources is too expensive in terms of complexity and reexecuted instructions. We propose a novel fetch policy, which we call DWarn. DWarn uses L1 misses as indicators of L2 misses, giving higher priority to threads with no outstanding L1 misses. DWarn acts on L1 misses, before L2 misses happen in a controlled manner to reduce resource under-use and to avoid harming a thread when L1 misses do not lead to L2 misses. Our results show that DWarn outperforms previously proposed policies, in both throughput and fairness, while requiring fewer resources and avoiding instruction reexecution. Francisco J. Cazorla, Alex Ramírez, Mateo Valero, Enrique Fernández |
IPDPS | 1 |
| 2004 | Dynamically Controlled Resource Allocation in SMT ProcessorsabstractSMT processors increase performance by executing instructions from several threads simultaneously. These threads use the resources of the processor better by sharing them but, at the same time, threads are competing for these resources. The way critical resources are distributed among threads determines the final performance. Currently, processor resources are distributed among threads as determined by the fetch policy that decides which threads enter the processor to compete for resources. However, current fetch policies only use indirect indicators of resource usage in their decision, which can lead to resource monopolization by a single thread or to resource waste when no thread can use them. Both situations can harm performance and happen, for example, after an L2 cache miss. In this paper, we introduce the concept of dynamic resource control in SMT processors. Using this concept, we propose a novel resource allocation policy for SMT processors. This policy directly monitors the usage of resources by each thread and guarantees that all threads get their fair share of the critical shared resources, avoiding monopolization. We also define a mechanism to allow a thread to borrow resources from another thread if that thread does not require them, thereby reducing resource under-use. Simulation results show that our dynamic resource allocation policy outperforms a static resource allocation policy by 8%, on average. It also improves the best dynamic resource-conscious fetch policies like FLUSH++ by 4%, on average, using the harmonic mean as a metric. This indicates that our policy does not obtain the ILP boost by unfairly running high ILP threads over slow memory-bounded threads. Instead, it achieves a better throughput-fairness balance. Francisco J. Cazorla, Alex Ramírez, Mateo Valero, Enrique Fernández |
MICRO | 1 |