VLDB 2026 Research / reviewers in the wild / expert
Jaume Abella 0001
dblp:97/5544
· DBLP profile ↗
183ranked-venue papers
22as first author
48since 2021 · last 2026
0000-0001-7951-4028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 147 · 20 first-author · 36 since 2021Software engineering, systems software and programming languages · 45 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ROSBand: A Bandwidth Regulation Approach on ROS2-Based Systems
Jon Altonaga Puente, Enrico Mezzetti, Irune Agirre, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 4 |
| 2026 | Supporting Timing-related Metrics for Autonomous Driving Frameworks in CyberRTabstractThe provision of increasingly advanced autonomous software functionalities builds on cutting-edge autonomous driving frameworks to enable modular interactions among multiple software components. This approach helps to support functional cause-effect chains from multiple sensors to actuators. The complexity of the (software) component interactions makes it more difficult to ascertain the correctness of the timing behavior of the system. This is so because traditional timing-related metrics like worst-case execution and worst-case response time do not capture the inter-dependency in cause-effect chains between the input sampling time and the time at which computation based on those inputs is performed. Complementary timing-related metrics, such as maximum reaction time and maximum data age have been considered to capture timing requirements, typically with an end-to-end scope, in cause-effect chains. These metrics have been formalized and demonstrated in ROS2-based automotive and autonomous driving setups [ 44 , 46 ]. However, the formalization of those metrics, which is necessary for deriving analytical lower and upper bounds and monitoring them at run-time, largely depends on the execution model and semantics offered by the run-time. Any concrete application of those metrics need to be tailored and adapted to the system at hand. Apollo auto is a popular, industrial-quality, open-source autonomous driving framework that is seeing increasing adoption both for industrial and academic projects. Apollo builds on CyberRT , an ad-hoc run-time that is similar in mechanism and intent to ROS2 but differentiates from it with respect to execution model and supported semantics. In contrast to ROS2, CyberRT is highly specialized to support the Apollo AD framework, is neither extensively documented or thoroughly analysed in the literature, especially in relation to execution model and instantiation of timing-related metrics. In this work, for the first time, we provide an insightful analysis and discussion on CyberRT execution model and semantics, starting from its raw and non-extensively documented codebase. Based on the identified semantics, we elaborate a formalization of timing-related metrics on CyberRT , across different granularity scopes, namely end-to-end and node levels. In particular, we develop on the importance of node-level timing properties to intercept any latent timing misbehavior before it is too late, and it severely impacts end-to-end execution. We provide a concrete mapping of a comprehensive set of timing-related metrics to the CyberRT execution model, both at end-to-end and node level, and develop a monitoring library that allows to intercept them on the specific software stack. We exploit the proposed library on a set of Apollo autonomous driving scenarios to demonstrate its effectiveness in monitoring the considered timing metrics and to promptly intercept a subtle timing misbehavior beyond end-to-end execution scope in a representative autonomous driving stack. Miguel Alcon, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | SAFEXPLAIN: a Complete Approach Towards Trustworthy AI-Based Safety-Critical SystemsabstractAI becomes increasingly important in safetycritical systems, especially in the case of autonomous systems, since navigation relies on AI for object detection and collision avoidance. However, safety-critical systems must adhere to functional safety standards that enforce software to be correct-by-construction, component decomposition to simplify design and validation, and the use of data only for testing purposes not to design the system itself. AI in general, and Deep Learning (DL) in particular have opposed characteristics since they have error rates (e.g., due to mispredictions), AI/DL modules can only be designed and validated monolithically, and they build on data for their design (i.e. for training purposes). Hence, DL solutions are at odds with the development process of safetycritical systems. A number of standards have recently emerged in different domains to reconcile the requirements of safety-critical systems with the characteristics of DL solutions, such as ISO 21448, ISO/IEC TR 5469, and ISO 8800, among others. However, there is a lack of realistic practice to design a DL-based safety-critical system in accordance with those regulations, and existing solutions only cover some aspects in isolation, and are often incompatible among them. SAFEXPLAIN is a 3-year Horizon Europe project addressing this challenge. SAFEXPLAIN, which finishes in September 2025, has already reached its main goals providing specific and complementary solutions to all those challenges so that AIbased safety-critical systems can be designed, implemented and validated adhering to the relevant functional safety standards in domains such as automotive, space and railway. In particular, SAFEXPLAIN provides the concepts, processes, tools and frameworks addressing the challenge end-to-end, from concept to solution. This is proven by the successful application of the SAFEXPLAIN approach in three case studies from the automotive, space and railway domains, whose results will see the light very soon. Jaume Abella 0001, Irune Agirre, Thanh Hai Bui, Frank Geujen, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Axel Brando, Javier Fernández 0004, Irune Yarza, Joanes Plazaola, Maria Ulan, Rob Lavreysen, Lucas Tosi, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Stefano Lodico, William Guarienti, Giuseppe Nicosia, Valeria Dallara |
DSD | 1 |
| 2025 | Impact of Contention-Aware Placement in Heterogeneous Edge DevicesabstractTime predictability is an increasing concern in functionally-rich mixed-criticality applications at the Edge, which often carry different timing requirements. Edge devices, in turn, are increasingly complex to sustain the increasing computational requirements, which hinders providing predictable performance without seriously affecting performance. One of the main threats to predictable performance is the impact of timing interference arising from contention in an increasing number of shared hardware resources. The impact of software to hardware mapping on performance is a well-studied topic, seeking optimal memory mappings to reduce average and worst-case performance, and, more recently, to control and limit timing interference. These methods normally focus on code and data placement, especially in relation to specific properties of the memory hierarchy, either architectural (e.g. heterogeneous memory modules) or obtained through partitioning techniques. In practice, however, these works build on a uniform memory hierarchy model, where the source of a memory request, namely, where a task accessing a given memory is eventually executed, is not directly relevant. In this work, we consider a large class of systems (e.g., TriCore families) where memory hierarchies are non-uniform, and access latency depends on the computing element issuing the request. In those architectures, the impact of code and data placement on timing interference cannot be addressed without considering architectural constraints and task locality. Through empirical exploration, we show that code, data, and locality collectively have a substantial impact on contention bounds, leading to a significantly expanded optimization space compared to approaches considering only code and data placement under the uniform memory assumption. Our results motivate the need for novel, efficient optimization approaches that integrate task mapping and architectural constraints to reduce timing interference. Jeremy Giesen, Ibai Irigoyen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2025 | Detecting Low-Density Mixtures in High-Quantile Tails for pWCET Estimation
Blau Manau, Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 5 |
| 2025 | Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2PixabstractEnhancing the robustness of object detection systems under adverse weather conditions is crucial for the advancement of autonomous driving technology. This study presents a novel approach leveraging the diffusion model Instruct Pix2Pix to develop prompting methodologies that generate realistic datasets with weather-based augmentations aiming to mitigate the impact of adverse weather on the perception capabilities of state-of-the-art object detection models, including Faster R-CNN and YOLOv10. Experiments were conducted in two environments, in the CARLA simulator where an initial evaluation of the proposed data augmentation was provided, and then on the real-world image data sets BDD100K and ACDC demonstrating the effectiveness of the approach in real environments.The key contributions of this work are twofold: (1) identifying and quantifying the performance gap in object detection models under challenging weather conditions, and (2) demonstrating how tailored data augmentation strategies can significantly enhance the robustness of these models. This research establishes a solid foundation for improving the reliability of perception systems in demanding environmental scenarios, and provides a pathway for future advancements in autonomous driving. Unai Gurbindo, Axel Brando, Jaume Abella 0001, Caroline König |
IJCNN | 3 |
| 2025 | Leveraging Image-Based Transformations to Mitigate Adversarial Attacks in AI-Based Safety-Critical SystemsabstractDual (DMR) and Triple Modular Redundancy (TMR) are widely used techniques to provide fault detection and/or tolerance capabilities in safety-critical systems through – often diverse – redundancy. However, these systems remain vulnerable to adversarial attacks, which can mislead the AI models and lead to severe consequences. In this paper, we propose enhanced DMR and TMR implementations for image-based object detection leveraging image transformations during inference to mitigate the impact of adversarial attacks, hence addressing safety and security concerns simultaneously. Our approach achieves up to 12.9% and 12.2% higher accuracy in adversarial scenarios compared to state-of-the-art solutions in DMR and TMR configurations, respectively. Martí Caro, Axel Brando, Jaume Abella 0001 |
IOLTS | 3 |
| 2025 | GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN accelerationabstractVoltage overscaling, or undervolting, is an enticing approximate technique in the context of energy-efficient Deep Neural Network (DNN) acceleration, given the quadratic relationship between power and voltage. Nevertheless, its very high error rate has thwarted its general adoption. Moreover, recent undervolting accelerators rely on 8-bit arithmetic and cannot compete with state-of-the-art low-precision (<8b) architectures. To overcome these issues, we propose a new technique called Guarded Aggressive underVolting (GAV), which combines the ideas of undervolting and bit-serial computation to create a flexible approximation method based on aggressively lowering the supply voltage on a select number of least significant bit combinations. Based on this idea, we implement GAVINA (GAV mIxed-Precision Accelerator), a novel architecture that supports arbitrary mixed precision and flexible undervolting, with an energy efficiency of up to 89 TOP/sW in its most aggressive configuration. By developing an error model of GAVINA, we show that GAV can achieve an energy efficiency boost of 20% via undervolting, with negligible accuracy degradation on ResNet-18. Jordi Fornt, Pau Fontova, Adrian Gras, Omar Lahyani, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet |
ISLPED | 6 |
| 2025 | EMR: Removing Multicollinear Event Monitors to Improve Timing Modelling of Real-Time SystemsabstractMulticollinearity of Event Monitors (EMs) negatively impacts the modeling of non-functional critical metrics in real-time systems like worst-case timing and energy usage since some EMs are over-represented and can reduce model accuracy. To address this challenge, we propose Event Monitor Reduction (EMR), a method to select a reduced set of non-related (independent) features (EMs), hence eliminating multicollinearity. In particular, EMR finds linear relations between the EMs and removes dependent ones without data loss. EMR does not create new features like Principal Component Analysis does, simplifying interpretability. Results on synthetic data and data collected from the execution of representative benchmarks on an avionicsgrade processor show the benefits of our method in removing multicollinear EMs. We further illustrate the benefits of EMR on two different multicore timing contention models, showing how its application helps to reduce execution time requirements and increase the accuracy of the models. David Fonts, Diego Palacios, Sergi Vilardell, Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 7 |
| 2025 | Expanding SafeSU capabilities by leveraging security frameworks for contention monitoring in complex SoCsabstractThe increased performance requirements of applications running on safety-critical systems have led to the use of complex platforms with several CPUs, GPUs, and AI accelerators. However, higher platform and system complexity challenge performance verification and validation since timing interference across tasks occurs in unobvious ways, hence defeating attempts to optimize application consolidation informedly during design phases and validating that mutual interference across tasks is within bounds during test phases. In that respect, the SafeSU has been proposed to extend inter-task interference monitoring capabilities in simple systems. However, modern mixed-criticality systems are complex, with multilayered interconnects, shared caches, and hardware accelerators. To that end, this paper proposes a non-intrusive add-on approach for monitoring interference across tasks in multilayer heterogeneous systems implemented by leveraging existing security frameworks and the SafeSU infrastructure. The feasibility of the proposed approach has been validated in an RTL RISC-V-based multicore SoC with support for AI hardware acceleration. Our results show that our approach can safely track contention and properly break down contention cycles across the different sources of interference, hence guiding optimization and validation processes. • Enables inter-core interference monitoring in complex SoCs. • Tool to check inter-task timing independence at all NoC levels. • Proposes unified initiator naming to solve initiator dilution across SoC layers. Pablo Andreu, Sergi Alcaide, Pedro López 0001, Jaume Abella 0001, Carles Hernández 0001 |
Future Gener. Comput. Syst. | 4 |
| 2025 | Hardware support for contention tracking in CPU and GPU last-level cacheabstractModern MPSoCs increasingly rely on resource utilization to improve application performance with different computation needs. The last-level cache (LLC) is one of the main shared resources, contributing to the improvement of aggregated performance. However, LLC sharing also increases individual application performance variability, which is undesirable in scenarios where performance guarantees are required. While deploying cache partitioning mechanisms allows regaining predictability, they negatively affect aggregated performance. This confronts system designers with the dire conundrum of choosing between aggregated performance and predictability. We contend that adding hardware support to track contention among tasks (kernels) in the LLC enables it to be shared, removing shortcomings brought by partitioning while providing a clear view of how tasks (kernels) affect each other in the LLC of the CPUs and GPUs. This approach enables achieving the desired balance between performance and predictability. Thus, we propose a low-overhead hardware mechanism, called demotion counters (DC), that tightly estimates the contention tasks (kernels) generate on each other in the shared LLC, outperforming other solutions that build on existing hardware contention-tracking proposals which suffer an average workload breakdown deviation (wbd) over 0.13. Our results also show that DC introduces 0.66% area overhead. 1 Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 4 |
| 2025 | Mix-GEMM: Extending RISC-V CPUs for Energy-Efficient Mixed-Precision DNN Inference Using Binary SegmentationabstractEfficiently computing Deep Neural Networks (DNNs) has become a primary challenge in today's computers, especially on devices targeting mobile or edge applications. Recent progress on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) has shown that the key to high energy efficiency lies in executing deep learning models with low- (8- to 5-bit) or ultra-low-precision (4- to 2-bit). Unfortunately, current Central Processing Unit (CPU) architectures and Instruction Set Architectures (ISAs) present severe limitations on the range of data sizes supported to compute DNN kernels. In this work, we presentMix-GEMM, a hardware-software co-designed architecture that enables RISC-V processors to efficiently compute arbitrary mixed-precision DNN kernels, supporting all data size combinations from 8- to 2-bit. By applyingbinary segmentation, our architecture can scale its throughput by decreasing the data size of the operands, resulting in a flexible approach capable of leveraging state-of-the-art QAT and PTQ to achieve high energy efficiency at a very low cost. Evaluating ourMix-GEMMarchitecture in a dual-issue in-order RISC-V processor shows that we are able to boost its performance and energy efficiency by up to$44\times$and$11\times$with respect to the baseline processor, with an area overhead of only 2%. This allows our extended processor to execute state-of-the-art DNNs with significantly higher performance and energy efficiency than the standard FP32 precision, while retaining almost the same model accuracy. Jordi Fornt, Enrico Reggiani, Pau Fontova, Narcís Rodas, Alessandro Pappalardo, Osman S. Unsal, Adrián Cristal, Josep Altet, Francesc Moll, Jaume Abella 0001 |
IEEE Trans. Computers | 10 |
| 2025 | Semantic Diverse DMR and TMR for High-Integrity AI-Based Function EfficiencyabstractDual Modular Redundancy (DMR) and Triple Modular Redundancy (TMR), often with some form of diversity, are used in safety-critical systems to realize those functionalities at the highest integrity level providing fault detection and/or tolerance capabilities. Redundant executions are intended to provide bit-level identical results, and, upon any mismatch, an error is assumed and recovery actions taken as needed. In this article, we note that many emerging AI-based functionalities are intrinsically stochastic (e.g., camera-based object detection), and hence, their correctness must be judged semantically, with room for variations across correct outcomes (e.g., confidence must be above a given threshold, but how much it exceeds the threshold is irrelevant). Building on this observation, we propose strategies to create DMR and TMR implementations of AI-based functionalities that bring not only fault tolerance against random hardware faults but also against AI model inaccuracies. Those strategies, which can be realized with software-only means and ported to virtually any computing platform, build on input data modifications affecting the inference computations, but not the expected semantic output (e.g., introducing some controlled changes in the input data). Moreover, we provide our solution in the form of an open source tool for image and video processing aimed at facilitating the reproducibility of our evaluation results, and enabling others to use it and conduct further research on input transformations. Martí Caro, Axel Brando, Jaume Abella 0001 |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2024 | TAP: Task-Aware Profiling on Integrated SystemsabstractHardware Performance Monitors (HPM) are increasingly exploited for timing verification and validation of time-critical embedded systems (TECS). HPMs are typically collected at the lowest software level, which makes it difficult to unequivocally account events to specific run-time entities, a prerequisite for any form of analysis, without relying on ad-hoc support from the run-time or operating system layer. The latter, however, is either unavailable or not fully adequate for verification requirements. Moreover, timing-related concerns in the analysis of embedded systems are typically addressed in the final stages of the software development process where multiple tasks are fully or partially integrated on the platform and it is therefore hard, if not impossible, to enforce controlled testing scenarios where contributions to event counts can be dissected. In this work, we present TAP a generic concept for allowing Task-Aware Profiling of individual tasks in an already integrated system on MPSoCs with on-core and off-core HPM support. The proposed approach combines a lightweight user-level configurable API and minimally intrusive extensions to the operating system layer to enforce separation of contexts when collecting HPM. We implement and assess TAP on top of an Infineon AURIX MPSoC and the OSEK-compliant ERIKA Enterpise RTOS, offering a consistent and intuitive interface for governing and filtering the different sources of events. Our results on synthetic and automotive benchmarks show that TAP can transparently gather and filter the events of interest while incurring negligible overheads. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2024 | Event Monitor Validation in High-Integrity SystemsabstractPlatforms for modern embedded systems equip an increasing number of high-performance features to provide the required levels of performance. Timing analysis solutions handle the complexity of these platforms by relying on hardware event monitors (HEMs) that provide insightful information about resource utilization and, hence, contention among tasks. As a result, HEMs have become a key element to warrant a safe timing behavior of a system, for which reason they must be validated. While some initial works target HEMs validation, they consider one HEM at a time and focus on those HEMs for which an expert can establish an expected value for relatively small code snippets. In this paper, we propose a methodology for the validation of those HEMs for which a specific expected value cannot be established a priori even for simple cases and, instead, needs to be validated in conjunction with other HEMs. Our method also deals with the natural variability of the HEMs' values in high-performance platforms when collected in different experiments. We illustrate the effectiveness of our proposed technique for validating HEMs related to cache coherence in a relevant platform in the avionics domain. Roger Pujol, Sergi Vilardell, Enrico Mezzetti, Mohamed Hassan 0002, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 5 |
| 2024 | Achieving Flexible Performance Isolation on the AMD Xilinx Zynq UltraScale+abstractCo-hosting different tasks on the same MPSoC contributes to increasing average performance by allowing them to share MPSoC's resources that, otherwise, could be underutilized. However, resource sharing challenges performance isolation among tasks, as required in time-sensitive embedded critical systems like automotive and avionics. On the other hand, resource isolation through segregation (the reference solution for preventing the propagation of time-related safety issues) is detrimental to average performance. In this work, we show that the built-in QoS support in modern MPSoCs can be smartly leveraged to adapt to the timing and performance requirements of the running applications. In particular, we develop specific configurations of the complex QoS support in the Zynq UltraScale+ MPSoC that deliver performance isolation for time-sensitive tasks (TSTs) and ensure that non-time-sensitive tasks (NTSTs) maximize their average performance by exploiting the resources not used by TSTs. Our results on the Xilinx UtltraScale+ show that the TSTs with the most stringent constraints achieve high degrees of isolation, 96.0% of their solo performance on average, while NTSTs exploit the resources not used by TSTs achieving performance ranging from 72% to 4% depending on the resource left by TSTs. Alejandro Serrano-Cases, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2024 | Safety-Relevant AI-Based System Robustification with Neural Network EnsemblesabstractFunctional safety requirements of AI-based safety critical applications challenge AI models, whose accuracy can be limited. In this paper, we show how using several cooperative deep learning (DL) models helps to raise global accuracy and reject making a prediction when confidence is below a pre-established threshold. Adrià Aldomà, Axel Brando, Francisco J. Cazorla, Jaume Abella 0001 |
IOLTS | 4 |
| 2023 | Efficient Diverse Redundant DNNs for Autonomous DrivingabstractAutomotive applications with safety requirements must adhere to specific regulations such as ISO 26262, which imposes the use of diverse redundancy for the highest integrity levels (i.e., ASIL D). While this has been often achieved by means of Dual-Core LockStep (DCLS) for microcontrollers, it remains an open challenge how to realize diverse redundancy efficiently, i.e., without full duplication and preserving performance, for DNN-based safety-related tasks, such as object detection, needing accelerators for performance reasons.This paper proposes an architecture where the accelerator performing DNN inference is replicated, as in the case of DCLS for cores, but using a cheaper implementation for the replica. In particular, we build on the stochastic nature of DNN-based object detection to realize two redundant accelerators where the secondary accelerator uses smartly chosen lower precision arithmetic (e.g., dropping some bits of the original data) so that it provides diverse redundancy, it can keep the performance of the primary accelerator, does not require as much cost as full- precision replication, and can build on the very same data stream from memory used by the primary accelerator. With a simple heuristic, we show that such a diverse redundancy scheme is able to cope with faults restricting false positives and negatives to a few relatively small objects. Martí Caro, Jordi Fornt, Jaume Abella 0001 |
COMPSAC | 3 |
| 2023 | SAFEXPLAIN: Safe and Explainable Critical Embedded Systems Based on AIabstractDeep Learning (DL) techniques are at the heart of most future advanced software functions in Critical Autonomous AI-based Systems (CAIS), where they also represent a major competitive factor. Hence, the economic success of CAIS industries (e.g., automotive, space, railway) depends on their ability to design, implement, qualify, and certify DL-based software products under bounded effort/cost. However, there is a fundamental gap between Functional Safety (FUSA) requirements on CAIS and the nature of DL solutions. This gap stems from the development process of DL libraries and affects high-level safety concepts such as (1) explainability and traceability, (2) suitability for varying safety requirements, (3) FUSA-compliant implementations, and (4) real-time constraints. As a matter of fact, the data-dependent and stochastic nature of DL algorithms clashes with current FUSA practice, which instead builds on deterministic, verifiable, and pass/fail test-based software. The SAFEXPLAIN project tackles these challenges and targets by providing a flexible approach to allow the certification - hence adoption - of DL-based solutions in CAIS building on: (1) DL solutions that provide end-to-end traceability, with specific approaches to explain whether predictions can be trusted and strategies to reach (and prove) correct operation, in accordance to certification standards; (2) alternative and increasingly sophisticated design safety patterns for DL with varying criticality and fault tolerance requirements; (3) DL library implementations that adhere to safety requirements; and (4) computing platform configurations, to regain determinism, and probabilistic timing analyses, to handle the remaining non-determinism. Jaume Abella 0001, Jon Pérez 0001, Cristofer Englund, Bahram Zonooz, Gabriele Giordana, Carlo Donzella, Francisco J. Cazorla, Enrico Mezzetti, Isabel Serra, Axel Brando, Irune Agirre, Fernando Eizaguirre, Thanh Hai Bui, Elahe Arani, Fahad Sarfraz, Ajay Balasubramaniam, Ahmed Badar, Ilaria Bloise, Lorenzo Feruglio, Ilaria Cinelli, Davide Brighenti, Davide Cunial |
DATE | 1 |
| 2023 | NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated ChipsabstractThe NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2. Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov |
DATE | 3 |
| 2023 | Quasi Isolation QoS Setups to Control MPSoC Contention in Integrated Software Architectures
Sergio Garcia-Esteban, Alejandro Serrano-Cases, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 3 |
| 2023 | SafeLS: An Open Source Implementation of a Lockstep NOEL-V RISC-V CoreabstractMicrocontrollers running safety-critical applications with high integrity requirements must provide appropriate safety measures to manage random hardware faults. For instance, automotive safety regulations (e.g., ISO26262) impose the use of diverse redundancy for items at the highest automotive safety integrity level (ASIL), ASIL-D. In the case of computing cores, this is realized with dual core lockstep (DCLS). The advent of the RISC-VISA has made open source hardware gain popularity. However, there are few industrial open source SoCs meeting the requirements of safety-critical systems, and, to our knowledge, none of them provides lockstep cores. This paper presents the realization of a RISC-V open source lockstep core based on Gaisler's NOEL-V core for the space domain, as well as its integration in the SELENE SoC that provides a complete microcontroller synthesizable on FPGA successfully assessed against space, automotive and railway safety-critical applications in the past. Marcel Sarraseca, Sergi Alcaide, Francisco Fuentes, Juan Carlos Rodriguez, Feng Chang, Ilham Lasfar, Ramon Canal, Francisco J. Cazorla, Jaume Abella 0001 |
IOLTS | 9 |
| 2023 | Improving Timing-Related Guarantees for Main Memory in Multicore Critical Embedded SystemsabstractMain memory is one of the most complex resources to analyze in multicore-based embedded real-time systems, with contention in the memory controller and the timing constraints of the main memory device as the main contributors to that complexity. One of the main challenges in multicore real-time systems is producing the required evidence on the management of contention delay for the certification. This stems from the fact that current MPSoCs barely provide any event monitors on how tasks interact and delay each other in memory. Besides, even if hardware and software mechanisms are in place to mitigate contention in the memory system, it is hard - if at all possible - to provide evidence about their correctness. In this work, we cover this gap by proposing a lightweight hardware mechanism that tightly tracks inter-core contention in memory. The proposed hardware mechanism, which we evaluate in detail, improves the quality of timing-related evidence that must be provided on how contention in main memory of multicore real-time systems is handled in adherence to applicable safety standards. Asier Fernández de Lecea, Mohamed Hassan 0002, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 4 |
| 2023 | An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt |
J. Syst. Archit. | 3 |
| 2023 | Main sources of variability and non-determinism in AD software: taxonomy and prospects to handle them
Miguel Alcon, Axel Brando, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
Real Time Syst. | 4 |
| 2023 | Dynamic and execution views to improve validation, testing, and optimization of autonomous driving software
Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
Softw. Qual. J. | 3 |
| 2023 | Vector Extensions in COTS Processors to Increase Guaranteed Performance in Real-Time SystemsabstractThe need for increased application performance in high-integrity systems such as those in avionics is on the rise as software continues to implement more complex functionalities. The prevalent computing solution for future high-integrity embedded products is multi-processor systems-on-chip (MPSoC) processors. MPSoCs include central processing unit (CPU) multicores that enable improving performance via thread-level parallelism. MPSoCs also include generic accelerators (graphics processing units [GPUs]) and application-specific accelerators. However, the data processing approach (DPA) required to exploit each of these underlying parallel hardware blocks carries several open challenges to enable the safe deployment in high-integrity domains. The main challenges include the qualification of its associated runtime system and the difficulties in analyzing programs deploying the DPA with out-of-the-box timing analysis and code coverage tools. In this work, we perform a thorough analysis of vector extensions (VExts) in current commercial off-the-shelf (COTS) processors for high-integrity systems. We show that VExts prevent many of the challenges arising with parallel programming models and GPUs. Unlike other DPAs, VExts require no runtime support, prevent design race conditions that might arise with parallel programming models, and have minimum impact on the software ecosystem, enabling the use of existing code coverage and timing analysis tools. We develop vectorized versions of neural network kernels and show that the NVIDIA Xavier VExts provide a reasonable increase in guaranteed application performance of up to 2.7x. Our analysis contends that VExts are the DPA approach with arguably the fastest path for adoption in high-integrity systems. Roger Pujol, Josep Jorba 0002, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2023 | Accurately Measuring Contention in Mesh NoCs in Time-Sensitive Embedded SystemsabstractThe computing capacity demanded by embedded systems is on the rise as software implements more functionalities, ranging from best-effort entertainment functions to performance-guaranteed safety-related functions. Heterogeneous manycore processors, using wormhole mesh (wmesh) Network-on-Chips (NoCs) as the main communication means, and contention block among applications, are increasingly considered to deliver the required computing performance. Most research efforts on software timing analysis have focused on deriving bounds (estimates) to the contention that tasks can suffer when accessing wmesh NoCs. However, less effort has been devoted to an equally important problem, namely,accuratelymeasuring the actual contention tasks generate each other on the wmesh which is instrumental during system validation to diagnose any software timing misbehavior and determine which tasks are particularly affected by contention on specific wmesh routers. In this article, we work on the foundations ofcontention measuringin wmesh NoCs and propose and explain the rationale of agolden metric, called taskPairWise Contention(PWC). PWC allows ascribing the actual share of the contention a given task suffers in the wmesh to each of its co-runner tasks at packet level. We also introduce and formalize aGolden Reference Value(GRV) for PWC that specifically defines a criterion to fairly break down the contention suffered by a task among its co-runner tasks in the wmesh. Our evaluation shows that GRV effectively captures how contention occurs by identifying the actual core (task) causing contention and whether contention is caused by local or remote interference in the wmesh. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | An Energy-Efficient GeMM-Based Convolution Accelerator With On-the-Fly im2colabstractSystolic array architectures have recently emerged as successful accelerators for deep convolutional neural network (CNN) inference. Such architectures can be used to efficiently execute general matrix–matrix multiplications (GeMMs), but computing convolutions with this primitive involves transforming the 3-D input tensor into an equivalent matrix, which can lead to an inflation of the input data, increasing the OFF-chip memory traffic which is critical for energy efficiency. In this work, we propose a GeMM-based systolic array accelerator that uses a novel data feeder architecture to perform ON-chip, on-the-fly convolution lowering (also known as im2col), supporting arbitrary tensor and kernel sizes as well as strided and dilated (or atrous) convolutions. By using our data feeder, we reduce memory transactions and required bandwidth on state-of-the-art CNNs by a factor of two, while only adding an area and power overhead of 4% and 7%, respectively. Application specific integrated circuit (ASIC) implementation of our accelerator in 22-nm technology fits in less than 1.1 mm 2 and reaches an energy efficiency of 1.10 TFLOP/sW with 16-bit floating-point arithmetic. Jordi Fornt, Pau Fontova, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet, Christoph Studer |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2022 | SafeDM: a Hardware Diversity Monitor for Redundant Execution on Non-Lockstepped CoresabstractComputing systems in the safety domain, such as those in avionics or space, require specific safety measures related to the criticality of the deployment. A problem these systems face is that of transient failures in hardware. A solution commonly used to tackle potential failures is to introduce redundancy in these systems, for example 2 cores that execute the same program at the same time. However, redundancy does not solve all potential failures, such as Common Cause Failures (CCF), where a single fault affects both cores identically (e.g. a voltage droop). If both redundant cores have identical state when the fault occurs, then there may be a CCF since the fault can affect both cores in the same way. To avoid CCF it is critical to know that there is diversity in the execution amongst the redundant cores. In this paper we introduce SafeDM, a hardware Diversity Monitor that quantifies the diversity of each redundant processor to guarantee that CCF will not go unnoticed, and without needing to deploy lockstepped cores. SafeDM computes data and instruction diversity separately, using different techniques appropriate for each case. We integrate SafeDM in a RISC-V FPGA space MPSoC from Cobham Gaisler where SafeDM is proven effective with a large benchmark suite, incurring low area and power overheads. Overall, SafeDM is an effective hardware solution to quantify diversity in cores performing redundant execution. Francisco Bas, Pedro Benedicte, Sergi Alcaide, Guillem Cabo, Fabio Mazzocchetti, Jaume Abella 0001 |
DATE | 6 |
| 2022 | SafeSU-2: a Safe Statistics Unit for Space MPSoCsabstractAdvanced statistics units (SUs) have been proven effective for the verification, validation and implementation of safety measures as part of safety-related MPSoCs. This is the case, for instance, of the RISC-V MPSoC by CAES Gaisler based on NOEL-V cores that will become commercially ready on FPGAs by the end of 2022. However, while those SUs support safety in the rest of the SoC, they must be built to be safe to be part of commercial products. This paper presents the SafeSU-2, the safety-compliant version of the SafeSU. In particular, we perform a Failure Mode and Effect Analysis (FMEA) for the SafeSU for relevant fault models, and implement fault detection and tolerance features needed to make it compliant with the requirements of safety-related devices in general, and of space MPSoCs in particular. Guillem Cabo, Sergi Alcaide, Carles Hernández 0001, Pedro Benedicte, Francisco Bas, Fabio Mazzocchetti, Jaume Abella 0001 |
DATE | 7 |
| 2022 | De-RISC: A Complete RISC-V Based Space-Grade PlatformabstractThe H2020 EIC-FTI De-RISC project develops a RISC-V space-grade platform to jointly respond to several emerging, as well as longstanding needs in the space domain such as: (1) higher performance than that of monocore and basic multicore space-grade processors in the market; (2) access to an increasingly rich software ecosystem rather than sticking to the slowly fading SPARC and PowerPC-based ones; (3) freedom (or drastic reduction) of export and license restrictions imposed by commercial ISAs such as Arm; and (4) improved support for the design and validation of safety-related real-time applications, (5) being the platform with software qualified and hardware designed per established space industry standards. De-RISC partners have set up the different layers of the platform during the first phases of the project. However, they have recently boosted integration and assessment activities. This paper introduces the De-RISC space platform, presents recent progress such as enabling virtualization and software qualification, new MPSoC features, and use case deployment and evaluation, including a comparison against other commercial platforms. Finally, this paper introduces the ongoing activities that will lead to the hardware and fully qualified software platform at TRL8 on FPGA by September 2022. Nils-Johan Wessman, Fabio Malatesta, Stefano Ribes, Jan Andersson, Antonio García-Vilanova, Miguel Masmano, Vicente Nicolau, Paco Gomez, Jimmy Le Rhun, Sergi Alcaide, Guillem Cabo, Francisco Bas, Pedro Benedicte, Fabio Mazzocchetti, Jaume Abella 0001 |
DATE | 15 |
| 2022 | Using Quantile Regression in Neural Networks for Contention Prediction in Multicore ProcessorsabstractMachine learning has enabled significant benefits in diverse fields, but, with a few exceptions, has had limited impact on computer architecture. Recent work, however, has explored broader applicability for design, optimization, and simulation. Notably, machine learning based strategies often surpass prior state-of-the-art analytical, heuristic, and human-expert approaches. This paper reviews machine learning applied system-wide to simulation and run-time optimization, and in many individual components, including memory systems, branch predictors, networks-on-chip, and GPUs. The paper further analyzes current practice to highlight useful design strategies and identify areas for future work, based on optimized implementation strategies, opportune extensions to existing work, and ambitious long term possibilities. Taken together, these strategies and techniques present a promising future for increasingly automated architectural design. Axel Brando, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 4 |
| 2022 | Using Markov's Inequality with Power-Of-k Function for Probabilistic WCET Estimation
Sergi Vilardell, Isabel Serra, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Joan del Castillo |
ECRTS | 4 |
| 2022 | SafeX: Open Source Hardware and Software Components for Safety-Critical SystemsabstractRISC-V Instruction Set Architecture (ISA) emerges as an opportunity to develop open source hardware without being subject to expensive licenses or export restrictions. A plethora of initiatives are nowadays developing systems-on-chip (SoCs) and its components based on RISC-V targeting a wide variety of markets. However, domains with safety requirements, such as avionics, space, and automotive, impose SoCs to include support to meet those requirements.This work introduces the SafeX family of components, a set of components providing SoC controllability, observability and safety measures support. These components, developed by the Barcelona Supercomputing Center with permissive open source licenses, are intended to be the basis to make SoCs meet the needs of domains with safety requirements. In particular, the SafeX components developed so far include the SafeSU (multicore statistics unit), the SafeTI (flexible and programmable traffic injector), the SafeDE and SafeSoftDR (hardware and software modules to enforce lockstep execution), and the SafeDM (module to monitor diversity across cores). Sergi Alcaide, Guillem Cabo, Francisco Bas, Pedro Benedicte, Francisco Fuentes, Feng Chang, Ilham Lasfar, Ramon Canal, Jaume Abella 0001 |
FDL | 9 |
| 2022 | Contention Tracking in GPU Last-Level CacheabstractThe Last-level cache (LLC) is one of the main GPU’s shared resources that contributes to improve performance but also increases individual kernel’s performance variability. This is detrimental in scenarios in which some level of performance predictability is required. While predictability can be regained by deploying cache partitioning (isolation) mechanisms, isolation negatively affects performance efficiency. This work shows that not partitioning the LLC and providing the ability to track the contention that kernels generate on each other allows them to share LLC space, hence increasing efficiency, while the system designer obtains a clear view of how each kernel affects each other in the LLC so as to balance performance and predictability goals. In this line, we propose GPU demotion counters (GDC), a low-overhead hardware mechanism to track contention that kernels generate on each other in the shared LLC. Javier Barrera, Leonidas Kosmidis, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
ICCD | 4 |
| 2022 | At-scale evaluation of weight clustering to enable energy-efficient object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001 |
J. Syst. Archit. | 3 |
| 2021 | Empirical Evidence for MPSoCs in Critical Systems: The Case of NXP's T2080 Cache CoherenceabstractThe adoption of complex MPSoCs in critical realtime embedded systems mandates a detailed analysis of their architecture to facilitate certification. This analysis is hindered by the lack of a thorough understanding of the MPSoC system due to the unobvious and/or insufficiently documented behavior of some key hardware features. Confidence in those features can only be regained by building specific tests to both, assess whether their behavior matches specifications and unveil their behavior when it is not fully known a priori. In this line, in this work we develop a thorough understanding of the cache coherence protocol in the avionics-relevant NXP T2080 architecture. Roger Pujol, Hamid Tabani, Jaume Abella 0001, Mohamed Hassan 0002, Francisco J. Cazorla |
DATE | 3 |
| 2021 | Enabling Unit Testing of Already-Integrated AI Software Systems: The Case of Apollo for Autonomous DrivingabstractThe advanced AI-based software used for autonomous driving comprises multiple highly-coupled modules that are data and control dependent. Deploying those already-integrated software frameworks makes unit testing, a fundamental step in the validation process of critical software, very challenging in safety-critical systems. To tackle this issue, in this paper, we show the steps we followed to develop standalone versions of the modules in an industry-level autonomous driving framework (Apollo) by applying several modifications to its architectural design. We show how the standalone modules have the same functional behavior as their integrated counterpart modules. We exemplify the benefits of standalone modules by performing incremental analysis of the software timing requirements of each module running on a heterogeneous System on Chip (SoC). This is a mandatory step to consolidate and integrate software modules guaranteeing timing constraints (e.g. related to freedom from interference) while maximizing SoC utilization. Miguel Alcon, Hamid Tabani, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2021 | Leveraging Hardware QoS to Control Contention in the Xilinx Zynq UltraScale+ MPSoCabstractThe interference co-running tasks generate on each other’s timing behavior continues to be one of the main challenges to be addressed before Multi-Processor System-on-Chip (MPSoCs) are fully embraced in critical systems like those deployed in avionics and automotive domains. Modern MPSoCs like the Xilinx Zynq UltraScale+ incorporate hardware Quality of Service (QoS) mechanisms that can help controlling contention among tasks. Given the distributed nature of modern MPSoCs, the route a request follows from its source (usually a compute element like a CPU) to its target (usually a memory) crosses several QoS points, each one potentially implementing a different QoS mechanism. Mastering QoS mechanisms individually, as well as their combined operation, is pivotal to obtain the expected benefits from the QoS support. In this work, we perform, to our knowledge, the first qualitative and quantitative analysis of the distributed QoS mechanisms in the Xilinx UltraScale+ MPSoC. We empirically derive QoS information not covered by the technical documentation, and show limitations and benefits of the available QoS support. To that end, we use a case study building on neural network kernels commonly used in autonomous systems in different real-time domains. Alejandro Serrano-Cases, Juan M. Reina, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
ECRTS | 3 |
| 2021 | Security, Reliability and Test Aspects of the RISC-V EcosystemabstractRISC-V has emerged as a viable solution on academia and industry. However, to use open source hardware for safety-critical applications, we need a deep understanding of the way in which well established mechanisms for testing and reliability could be integrated and deployed on the RISC-V ecosystem, and we need a clear knowledge on how such an ecosystem can be leveraged to improve security. This paper includes four contributions presenting the potential of RISC-V in security research, the way in which RISC-V can be hardened against power analysis attacks, how to implement, using RISC-V, software and hardware/software solutions for dual core lock step, and how to perform system-level testing in the RISC-V ecosystem. Jaume Abella 0001, Sergi Alcaide, Jens Anders, Francisco Bas, Steffen Becker 0001, Elke De Mulder, Nourhan Elhamawy, Frank K. Gürkaynak, Helena Handschuh, Carles Hernández 0001, Michael Hutter, Leonidas Kosmidis, Ilia Polian, Matthias Sauer 0002, Stefan Wagner 0001, Francesco Regazzoni 0001 |
ETS | 1 |
| 2021 | SafeSU: an Extended Statistics Unit for Multicore Timing InterferenceabstractStatistics units (SUs) in MPSoCs are becoming increasingly used for the (1) verification and (2) validation of multicore timing interference, as well as for (3) deploying safety measures in safety-related real-time systems. However, existing SU extensions to manage multicore timing interference have neither been integrated together nor deployed in commercial MPSoCs.This paper presents the realization of the Safe Statistics Unit (SafeSU for short), which smartly integrates existing solutions for multicore timing interference verification, validation and monitoring, and is in turn integrated in commercial space-graded RISC-V and SparcV8 MPSoCs. Our evaluation illustrates the operation of the SafeSU, and paves the way for a thorough validation prior to reaching commercialization and being offered as open source IP. Guillem Cabo, Francisco Bas, Ruben Lorenzo, David Trilla, Sergi Alcaide, Miquel Moretó, Carles Hernández 0001, Jaume Abella 0001 |
ETS | 8 |
| 2021 | PRL: Standardizing Performance Monitoring Library for High-Integrity Real-Time SystemsabstractThe use of complex processors is becoming ubiquitous in High-Integrity Systems (HIS). To deal with processor’s increased complexity, Performance Monitoring Counters (PMCs) are increasingly used to reason on software behavior and provide the necessary evidence to support software certification. However, the use of PMCs in HIS is relatively recent and hence far from being standardized. As a result, software engineers are forced to resort to highly-customized, low-level programming of platform-specific PMC control registers, which is both error prone and time consuming. To cover this gap, we propose building on the PAPI library, a standardized performance monitoring solution in the mainstream domain, and develop a PMC Reading Library (PRL) for configuring and collecting traceable events while capturing HIS specific requirements and peculiarities. We instantiate PRL in a reference automotive configuration to show that PRL meets key HIS requirements: negligible footprint, limited and predictable overhead, and accuracy collecting hardware events by filtering out the impact of interrupts and context switches. Jeremy Giesen, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ICCD | 3 |
| 2021 | SafeDE: a flexible Diversity Enforcement hardware module for light-locksteppingabstractSafety-related systems, such as those in automotive, avionics and space, impose the existence of appropriate safety measures to meet the safety requirements of the system. In the case of the highest integrity level functionalities (e.g. ASIL-D in automotive), diverse redundancy must be deployed to avoid unreasonable risk of a single fault leading the system to a failure (e.g. using lockstepped cores). However, existing lockstep solutions are either (1) highly intrusive and inflexible coupling two cores with hardware means, or (2) costly in terms of execution time and monitoring if a software monitor thread checks that cores running redundantly preserve sufficient staggering. This paper presents SafeDE, a Diversity Enforcement hardware module providing light-lockstep support by means of a non-intrusive and flexible hardware module that preserves staggering across cores running redundant threads, thus bringing time diversity. SafeDE reconciles the lightness and flexibility of software-only solutions, even allowing using the cores without any lockstepping, as well as the tighter staggering of hardware-only solutions that allow using staggering values of few cycles, instead of hundreds of microseconds, as for software-only solutions. Our integration of SafeDE in a RISC-V FPGA-based space multicore from Cobham Gaisler shows that staggering is effectively preserved, and SafeDE overheads are negligible in terms of area and performance due to staggering. Francisco Bas, Sergi Alcaide, Ruben Lorenzo, Guillem Cabo, Guillermo Gil, Oriol Sala, Fabio Mazzocchetti, David Trilla, Jaume Abella 0001 |
IOLTS | 9 |
| 2021 | SafeTI: a Hardware Traffic Injector for MPSoC Functional and Timing ValidationabstractFunctional and timing validation of safety-related MPSoCs requires testing specific traffic patterns in the on-chip interconnects. Generally, testing needs to be performed by using software tests whose degree of control on the traffic generated is indirect, and limited to behavior that can be triggered by software, thus often unable to produce traffic generated by peripherals. Therefore, untested traffic scenarios can be abundant and, to a large extent, it is hard to know what traffic scenarios have been effectively tested. This paper presents the safe traffic injector, SafeTI, which allows injecting programmable traffic in AMBA AHB interconnects with high flexibility and degree of control, thus easing achieving high coverage in terms of traffic scenarios tested, and mitigating the uncertainty due to the difficulties to relate software tests with actual traffic scenarios tested. We also integrate successfully the SafeTI in an industrial MPSoC for the space domain proving the effectiveness of the proposed traffic injector. Oriol Sala, Sergi Alcaide, Guillem Cabo, Francisco Bas, Ruben Lorenzo, Pedro Benedicte, David Trilla, Guillermo Gil, Fabio Mazzocchetti, Jaume Abella 0001 |
IOLTS | 10 |
| 2021 | Performance Analysis and Optimization Opportunities for NVIDIA Automotive GPUs
Hamid Tabani, Fabio Mazzocchetti, Pedro Benedicte, Jaume Abella 0001, Francisco J. Cazorla |
J. Parallel Distributed Comput. | 4 |
| 2021 | Towards functional safety compliance of matrix-matrix multiplication for machine learning-based autonomous systems
Javier Fernández 0004, Jon Pérez 0001, Irune Agirre, Imanol Allende, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 5 |
| 2021 | Worst-Case Energy Consumption: A New Challenge for Battery-Powered Critical DevicesabstractThe number of (edge) devices connected to the IoT is on the rise, reaching hundreds of billions in the next years. Many devices will implement some type of critical functionality, for instance in the medical market this includes infusion pumps and implantable defibrillators. Energy awareness is mandatory in the design of IoT devices given their huge impact on worldwide energy consumption and the fact that many of them are battery powered. Critical IoT devices further require addressing new energy-related challenges. On the one hand, factoring in the impact of energy-solutions on device's performance, providing evidence of adherence to domain-specific safety standards. On the other hand, deriving safe worst-case energy consumption (WCEC) estimates is fundamental to ensure the system can continuously operate under a pre-established set of power/energy caps, safely delivering its critical functionality. In this line, we analyze for the first time the impact that different hardware physical parameters have on both model-based and measurement-based WCEC modeling, for which we also show the main challenges they face compared to chip manufacturers' current practice for energy modeling and validation. Under the set of constraints that emanate from how certain physical parameters can be actually modeled, we show that measurement-based WCEC is a promising way forward for WCEC estimation. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Sustain. Comput. | 3 |
| 2020 | SELENE: Self-Monitored Dependable Platform for High-Performance Safety-Critical SystemsabstractExisting HW/SW platforms for safety-critical systems suffer from limited performance and/or from lack of flexibility due to building on specific proprietary components. This jeopardizes their wide deployment across domains. While some research has been done to overcome these limitations, they have had limited success owing to missing flexibility and extensibility. Flexibility and extensibility are the cornerstones of industry adoption: industries dealing in capital goods need technologies on which they can rely on during decades (e.g. avionics, space, automotive). SELENE aims at covering this gap by proposing a new family of safety-critical computing platforms, which builds upon open source components such as the RISC-V instruction set architecture, GNU/Linux, and the Jailhouse hypervisor. SELENE will develop an advanced computing platform that is able to: (1) adapt the system to the specific requirements of different application domains, to changing environmental conditions, and to internal conditions of the system itself; (2) allow the integration of applications of different criticalities and performance demands in the same platform, guaranteeing functional and temporal isolation properties; (3) achieve flexible diverse redundancy by exploiting the inherent redundant capabilities of the multicore; and (4) efficiently execute compute-intensive applications by means of specific accelerators. Carles Hernández 0001, José Flich, Roberto Paredes, Charles-Alexis Lefebvre, Imanol Allende, Jaume Abella 0001, David Trillin, Martin Matschnig, Bernhard Fischer, Konrad Schwarz, Jan Kiszka, Martin Rönnbäck, Johan Klockars, Nicholas Mc Guire, Franz Rammerstorfer, Christian Schwarzl, Franck Wartel, Dierk Lüdemann, Mikel Labayen |
DSD | 6 |
| 2020 | The ECSEL FRACTAL Project: A Cognitive Fractal and Secure edge based on a unique Open-Safe-Reliable-Low Power Hardware PlatformabstractThe objective of the FRACTAL project is to create a new approach to reliable edge computing. The computing node will be the building block of scalable Internet of Things (from Low Computing to High Computing Edge Nodes). The cognitive skill will be given by an internal and external architecture that allows forecasting its internal performance and the state of the surrounding world. The node will have the capability of learning how to improve its performance against the uncertainty of the environment. New industrial functions will flourish through the created space of the cognitive system. Cognitive advantages are brought to a resilient edge and a computing paradigm that lay down between the physical world and the cloud. Aizea Lojo, Leire Rubio, Jesus Miguel Ruano, Tania Di Mascio, Luigi Pomante, Enrico Ferrari, Ignacio Garcìa Vega, Frank K. Gürkaynak, Mikel Labayen, Vanessa Orani, Jaume Abella 0001 |
DSD | 11 |
| 2020 | Tracing Hardware Monitors in the GR712RC Multicore Platform: Challenges and Lessons Learnt from a Space Case StudyabstractThe demand for increased computing performance is driving industry in critical-embedded systems (CES) domains, e.g. space, towards the use of multicores processors. Multicores, however, pose several challenges that must be addressed before their safe adoption in critical embedded domains. One of the prominent challenges is software timing analysis, a fundamental step in the verification and validation process. Monitoring and profiling solutions, traditionally used for debugging and optimization, are increasingly exploited for software timing in multicores. In particular, hardware event monitors related to requests to shared hardware resources are building block to assess and restraining multicore interference. Modern timing analysis techniques build on event monitors to track and control the contention tasks can generate each other in a multicore platform. In this paper we look into the hardware profiling problem from an industrial perspective and address both methodological and practical problems when monitoring a multicore application. We assess pros and cons of several profiling and tracing solutions, showing that several aspects need to be taken into account while considering the appropriate mechanism to collect and extract the profiling information from a multicore COTS platform. We address the profiling problem on a representative COTS platform for the aerospace domain to find that the availability of directly-accessible hardware counters is not a given, and it may be necessary to the develop specific tools that capture the needs of both the user’s and the timing analysis technique requirements. We report challenges in developing an event monitor tracing tool that works for bare-metal and RTEMS configurations and show the accuracy of the developed tool-set in profiling a real aerospace application. We also show how the profiling tools can be exploited, together with handcrafted benchmarks, to characterize the application behavior in terms of multicore timing interference. Xavier Palomo, Mikel Fernández, Sylvain Girbal, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla, Laurent Rioux |
ECRTS | 5 |
| 2020 | A Cross-Layer Review of Deep Learning Frameworks to Ease Their Optimization and ReuseabstractMachine learning and especially Deep Learning (DL) approaches are at the heart of many domains, from computer vision and speech processing to predicting trajectories in autonomous driving and data science. Those approaches mainly build upon Neural Networks (NNs), which are compute-intensive in nature. A plethora of frameworks, libraries and platforms have been deployed for the implementation of those NNs, but end users often lack guidance on what frameworks, platforms and libraries to use to obtain the best implementation for their particular needs. This paper analyzes the DL ecosystem providing a structured view of some of the main frameworks, platforms and libraries for DL implementation. We show how those DL applications build ultimately on some form of linear algebra operations such as matrix multiplication, vector addition, dot product and the like. This analysis allows understanding how optimizations of specific linear algebra functions for specific platforms can be effectively leveraged to maximize specific targets (e.g. performance or power-efficiency) at application level reusing components across frameworks and domains. Hamid Tabani, Roger Pujol, Jaume Abella 0001, Francisco J. Cazorla |
ISORC | 3 |
| 2020 | Timing of Autonomous Driving Software: Problem Analysis and Prospects for Future SolutionsabstractThe software used to implement advanced functionalities in critical domains (e.g. autonomous operation) impairs software timing. This is not only due to the complexity of the underlying high-performance hardware deployed to provide the required levels of computing performance, but also due to the complexity, non-deterministic nature, and huge input space of the artificial intelligence (AI) algorithms used. In this paper, we focus on Apollo, an industrial-quality Autonomous Driving (AD) software framework: we statistically characterize its observed execution time variability and reason on the sources behind it. We discuss the main challenges and limitations in finding a satisfactory software timing analysis solution for Apollo and also show the main traits for the acceptability of statistical timing analysis techniques as a feasible path. While providing a consolidated solution for the software timing analysis of Apollo is a huge effort far beyond the scope of a single research paper, our work aims to set the basis for future and more elaborated techniques for the timing analysis of AD software. Miguel Alcon, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 5 |
| 2020 | Modeling Contention Interference in Crossbar-based Systems via Sequence-Aware Pairing (SeAP)abstractThe Infineon AURIX TriCore family of microcontrollers has consolidated as the reference multicore computing platform for safety-critical systems in the automotive domain. As a distinctive trait, AURIX microcontrollers are designed to promote high timing predictability as witnessed by the presence of large scratchpad memories and a crossbar interconnect. The latter has been introduced to reduce inter-core interference in accessing the memory system and peripherals. Nonetheless, the crossbar does not prevent requests from different cores to the same target resource to suffer contention. Applications are, therefore, inherently exposed to inter-core timing interference, which needs to be taken into account in the determination of reliable execution time bounds. In this paper we propose a contention modeling technique for crossbar-based systems, and hence suitable for bounding contention effects in the AURIX family. Unlike state of the art techniques that build on total request counts, we exploit the sequence of requests to the different target resources produced by each core to produce tighter bounds by discarding contention scenarios that cannot occur in practice. To that end, we adapt existing techniques from the pattern matching domain to derive the worst-case contention effects from the sequences of requests each core sends over the crossbar. Results on a wide set of synthetic and real scenarios and benchmark on an AURIX TC297TX show that our technique outperforms other contention modeling approaches. Jeremy Giesen, Pedro Benedicte, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 4 |
| 2020 | HRM: Merging Hardware Event Monitors for Improved Timing Analysis of Complex MPSoCsabstractThe performance monitoring unit (PMU) in multiprocessor system-on-chips (MPSoCs) is at the heart of the latest measurement-based timing analysis techniques in critical embedded systems. In particular, hardware event monitors (HEMs) in the PMU are used as building blocks in the process of budgeting and verifying software timing by tracking and controlling access counts to shared resources. While the number of HEMs in current MPSoCs reaches hundreds, they are read via performance monitoring counters whose number is usually limited to 4-8, thus requiring multiple runs of each experiment in order to collect all desired HEMs. Despite the effort of engineers in controlling the execution conditions of each experiment, the complexity of current MPSoCs makes it arguably impossible to completely remove the noise affecting each run. As a result, HEMs read in different runs are subject to different variability, and hence, those HEMs captured in different runs cannot be “blindly” merged. In this work, we focus on the NXP T2080 platform where we observed up to 59% variability across different runs of the same experiment for some relevant HEMs (e.g., processor cycles). We develop a HEM reading and merging (HRM) approach to join reliably HEMs across different runs as a fundamental element of any measurement-based timing budgeting and verification technique. Our method builds on order statistics and the selection of an anchor HEM read in all runs to derive the most plausible combination of HEM readings that keep the distribution of each HEM and their relationship with the anchor HEM intact. Sergi Vilardell, Isabel Serra, Roberto Santalla, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | Towards limiting the impact of timing anomalies in complex real-time processorsabstractTiming verification of embedded critical real-time systems is hindered by complex designs. Timing anomalies, deeply analyzed in static timing analysis, require specific solutions to bound their impact. For the first time, we study the concept and impact of timing anomalies in measurement-based timing analysis, the most used in industry, showing that they require to be considered and handled differently. In addition, we analyze anomalies in the context of Measurement-Based Probabilistic Timing Analysis, which simplifies quantifying their impact. Pedro Benedicte, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Francisco J. Cazorla |
ASP-DAC | 2 |
| 2019 | Assessing the Adherence of an Industrial Autonomous Driving Framework to ISO 26262 Software GuidelinesabstractThe complexity and size of Autonomous Driving (AD) software are comparably higher than that of software implementing other (standard) functionalities in the car. To make things worse, a big fraction of AD software is not specifically designed for the automotive (or any other critical) domain, but the mainstream market. This brings uncertainty on to which extent AD software adheres to guidelines in safety standards. In this paper, we present our experience in applying ISO 26262 -- the applicable functional safety standard for road vehicles -- software safety guidelines to industrial AD software, in particular, Apollo, a heterogeneous Autonomous Driving framework used extensively in industry. We provide quantitative and qualitative metrics of compliance for many ISO 26262 recommendations on software design, implementation, and testing. Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla, Guillem Bernat |
DAC | 3 |
| 2019 | High-Integrity GPU Designs for Critical Real-Time Automotive SystemsabstractAutonomous Driving (AD) imposes the use of high-performance hardware, such as GPUs, to perform object recognition and tracking in real-time. However, differently to the consumer electronics market, critical real-time AD functionalities require a high degree of resilience against faults, in line with the automotive ISO26262 functional safety standard requirements. ISO26262 imposes the use of some source of independent redundancy for the most critical functionalities so that a single fault cannot lead to a failure, being dual core lockstep (DCLS) with diversity the preferred choice for computing devices. Unfortunately, GPUs do not support diverse DCLS by construction, thus failing to meet ISO26262 requirements efficiently.In this paper we propose lightweight modifications to GPUs to enable diverse DCLS for critical real-time applications without diminishing their performance for non-critical applications. In particular, we show how enabling specific mechanisms for software-controlled kernel scheduling in the GPU, allows guaranteeing that redundant kernels can be executed in different resources so that a single fault cannot lead to a failure, as imposed by ISO26262. Our results on a GPU simulator and an NVIDIA GPU prove the viability of the approach and its effectiveness on high-performance GPU designs needed for AD systems. Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001 |
DATE | 4 |
| 2019 | LAEC: Look-Ahead Error Correction Codes in Embedded Processors L1 Data CacheabstractAs implementation technology shrinks, the presence of errors in cache memories is becoming an increasing issue in all computing domains. Critical systems, e.g. space and automotive, are specially exposed and susceptible to reliability issues. Furthermore, hardware designs in these systems are migrating to multilevel cache multicore systems, in which write-through first level data (DL1) caches have been shown to heavily harm average and guaranteed performance. While write-back DL1 caches solve this problem they come with their own challenges: they need Error Correction Codes (ECC) to tolerate soft errors, but implementing DL1 ECC in simple embedded micro-controllers requires either complex hardware to squash instructions consuming erroneous data, or delayed delivery of data to correct potential errors, which impacts performance even if such process is pipelined. In this paper we present a low-complexity hardware mechanism to anticipate data fetch and error correction in DL1 so that both (1) correct data is always delivered, but (2) avoiding additional delays in most of the cases. This achieves both high guaranteed performance and an effective solutions against errors. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 3 |
| 2019 | Maximum-Contention Control Unit (MCCU): Resource Access Count and Contention Time EnforcementabstractIn real-time systems, the techniques to derive bounds to the contention tasks can suffer in multicore build on resource quota monitoring and enforcement. Existing techniques track and bound the number of requests to hardware shared resources that each core (task) is allowed to perform. In this paper we show that current software-only solutions work well when there is a single resource and type of request to track and bound, but do not scale to the more general case of several shared resources that accept different request types, each with a different associated latency. To handle this (more general) case, we propose low-overhead hardware support called Maximum-Contention Control Unit (MCCU). The MCCU performs fine-grain tracking of different types of requests, preventing a core to cause more interference on its contenders than budgeted. In this process, the MCCU also helps verifying that individual requests duration does not exceed their theoretical bounds, hence dealing with scenarios in which requests can have an arbitrarily large duration. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 3 |
| 2019 | Multicore Early Design Stage Guaranteed Performance Estimates for the Space DomainabstractThe ability to produce early guaranteed performance (worst-case execution time) estimates for multicores, i.e. before software from different providers gets integrated onto the same critical system, is pivotal. This helps reducing lately-detected costly-to-handle timing violations. An existing methodology creates `copy' (surrogate) applications from the execution in isolation of each target application. Surrogate applications can be used to upperbound multicore contention delay, and hence WCET estimates in multicores. However, this methodology has only been shown to work on a simulation environment. In this paper we show the work we have carried out to adapt this technology to a real multicore processor for the space domain. Mikel Fernández, Gabriel Fernandez 0002, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 3 |
| 2019 | AURIX TC277 Multicore Contention Model Integration for Automotive ApplicationsabstractEmbedded systems industry needs reliable and tight worst-case execution time (WCET) estimates for critical applications running on multicores, as a prerequisite to their adoption. While industry already uses reliable tools for single-core WCET estimation and several multicore contention models (MCMs) have been proposed, their combination have not been shown to be fully compatible with the automotive industrial practice yet. This paper reduces this gap by presenting a framework for the integration of MCMs into industrial WCET estimation practice. We illustrate such integration for a Magneti Marelli powertrain control unit on an Infineon AURIX TC277 multicore platform. Enrico Mezzetti, Luca Barbina, Jaume Abella 0001, Stefania Botta, Francisco J. Cazorla |
DATE | 3 |
| 2019 | GPU4S: Embedded GPUs in SpaceabstractFollowing the same trend of automotive and avionics, the space domain is witnessing an increase in the on-board computing performance demands. This raise in performance needs comes from both control and payload parts of the spacecraft and calls for advanced electronics able to provide high computational power under the constraints of the harsh space environment. On the non-technical side, for strategic reasons it is mandatory to get European independence on the used computing technology. In this project, which is still in its early phases, we study the applicability of embedded GPUs in space, which have shown a dramatic improvement of their performance per-watt ratio coming from their proliferation in consumer markets based on competitive European technology. To that end, we perform an analysis of the existing space application domains to identify which software domains can benefit from their use. Moreover, we survey the embedded GPU domain in order to assess whether embedded GPUs can provide the required computational power and identify the challenges which need to be addressed for their adoption in space. In this paper, we describe the steps to be followed in the project, as well as the results of our preliminary analyses in the first months of the project. Leonidas Kosmidis, Jérôme Lachaize, Jaume Abella 0001, Olivier Notebaert, Francisco J. Cazorla, David Steenari |
DSD | 3 |
| 2019 | An Approach for Detecting Power Peaks During Testing and Breaking Systematic Pathological BehaviorabstractThe verification and validation process of embedded critical systems requires providing evidence of their functional correctness and also that their non-functional behavior stays within limits. In this work, we focus on power peaks, which may cause voltage droops and thus, challenge performance to preserve correct operation upon droops. In this line, the use of complex software and hardware in critical embedded systems jeopardizes the confidence that can be placed on the tests carried out during the campaigns performed at analysis. This is so because it is unknown whether tests have triggered the highest power peaks that can occur during operation and whether any such peak can occur systematically. In this paper we propose the use of randomization, already used for timing analysis of real-time systems, as an enabler to guarantee that (1) tests expose those peaks that can arise during operation and (2) peaks cannot occur systematically inadvertently. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2019 | Modeling the Impact of Process Variations in Worst-Case Energy Consumption EstimationabstractThe advent of autonomous power-limited systems poses a new challenge for system verification. Powerful processors needed to enable autonomous operation, are typically power-hungry, jeopardizing battery duration. Therefore, guaranteeing a given battery duration requires worst-case energy consumption (WCEC) estimation for tasks running on those systems. Unfortunately, processor energy and power can suffer significant variation across different units due to process variation (PV), i.e. variability in the electrical properties of transistors and wires due to imperfect manufacturing, which challenges existing WCEC estimation methods for applications. In this paper, we propose a statistical modeling approach to capture PV impact on applications energy and a methodology to compute their WCEC capturing PV, as required to deploy portable critical devices. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2019 | Generating and Exploiting Deep Learning Variants to Increase Heterogeneous Resource Utilization in the NVIDIA XavierabstractDeep learning-based solutions and, in particular, deep neural networks (DNNs) are at the heart of several functionalities in critical-real time embedded systems (CRTES) from vision-based perception (object detection and tracking) systems to trajectory planning. As a result, several DNN instances simultaneously run at any time on the same computing platform. However, while modern GPUs offer a variety of computing elements (e.g. CPUs, GPUs, and specific accelerators) in which those DNN tasks can be executed depending on their computational requirements and temporal constraints, current DNNs are mainly programmed to exploit one of them, namely, regular cores in the GPU. This creates resource imbalance and under-utilization of GPU resources when executing several DNN instances, causing an increase in DNN tasks' execution time requirements. In this paper, (a) we develop different variants (implementations) of well-known DNN libraries used in the Apollo Autonomous Driving (AD) software for each of the computing elements of the latest NVIDIA Xavier SoC. Each variant can be configured to balance resource requirements and performance: the regular CPU core implementation that can run on 2, 4, and 6 cores; the GPU regular and Tensor core variants that can run in 4 or 8 GPU’s Streaming Multiprocessors (SM); and 1 or 2 NVIDIA’s Deep Learning Accelerators (NVDLA); (b) we show that each particular variant/configuration offers a different resource utilization/performance point; finally, (c) we show how those heterogeneous computing elements can be exploited by a static scheduler to sustain the execution of multiple and diverse DNN variants on the same platform. Roger Pujol, Hamid Tabani, Leonidas Kosmidis, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 5 |
| 2019 | Software Timing Analysis for Complex Hardware with Survivability and Risk AnalysisabstractThe increasing automation of safety-critical real-time systems, such as those in cars and planes, leads, to more complex and performance-demanding on-board software and the subsequent adoption of multicores and accelerators. This causes software's execution time dispersion to increase due to variable-latency resources such as caches, NoCs, advanced memory controllers and the like. Statistical analysis has been proposed to model the Worst-Case Execution Time (WCET) of software running such complex systems by providing reliable probabilistic WCET (pWCET) estimates. However, statistical models used so far, which are based on risk analysis, are overly pessimistic by construction. In this paper we prove that statistical survivability and risk analyses are equivalent in terms of tail analysis and, building upon survivability analysis theory, we show that Weibull tail models can be used to estimate pWCET distributions reliably and tightly. In particular, our methodology proves the correctness-by-construction of the approach, and our evaluation provides evidence about the tightness of the pWCET estimates obtained, which allow decreasing them reliably by 40% for a railway case study w.r.t. state-of-the-art exponential tails. Sergi Vilardell, Isabel Serra, Jaume Abella 0001, Joan del Castillo, Francisco J. Cazorla |
ICCD | 3 |
| 2019 | Software-only Diverse Redundancy on GPUs for Autonomous Driving PlatformsabstractAutonomous driving (AD) builds upon high-performance computing platforms including (1) general purpose CPUs as well as (2) specific accelerators, being GPUs one of the main representatives. Microcontrollers have reached ASIL-D compliance by implementing diverse redundancy with lockstep execution. However, ASIL-D compliant GPUs rely on either fully redundant lockstep GPUs (i.e. 2 GPUs), which doubles hardware costs, or fully redundant systems with a GPU and another accelerator, which virtually doubles design and validation/verification (V&V) costs. In this paper we analyze the degree of diversity achieved when implementing redundancy on a single GPU, showing that diverse redundancy is not achieved in many cases, and propose software strategies that guarantee achieving diverse redundancy for any kernel on systems using commercial off-the-shelf (COTS) GPUs, thus showing how to achieve ASIL-D compliance on a single COTS GPU in controlled scenarios. Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001 |
IOLTS | 4 |
| 2019 | Accurate ILP-Based Contention Modeling on Statically Scheduled Multicore SystemsabstractCommercially available Off The Shelf (COTS) multicores have been assessed as the baseline computing platform even in the most conservative real-time domains. Multicore contention arising on shared hardware resources, with its circular dependence with scheduling, is among the most challenging issues that require urgent attention before multicores can be fully embraced for real-time computing. In the context of static scheduling, still the most used scheduling approach in real-time industries, we propose an ILP formulation for computing the worst-case contention delay suffered by a task due to interference on a shared bus. Our model provides accurate contention delay bounds that avoid unnecessary over-accounting of conflicts between bus requests, by considering contention effects at system-level (i.e., across tasks) rather than at task-level only. This allows precisely capturing the interdependence between timing interference of conflicting requests, issued in parallel by other cores (tasks), and the identification of the particular set of tasks co-running on those cores. We assess our technique both analytically and empirically on a real COTS multicore platform. We show, via extensive evaluation, that jointly accounting for worst-case task overlapping and request distribution scenarios always provides tighter contention bounds when compared to state-of-the-art solutions. Xavier Palomo, Enrico Mezzetti, Jaume Abella 0001, Reinder J. Bril, Francisco J. Cazorla |
RTAS | 3 |
| 2019 | Performance Analysis and Optimization of Automotive GPUsabstractAdvanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) have drastically increased the performance demands of automotive systems. Suitable high-performance platforms building upon Graphic Processing Units (GPUs) have been developed to respond to this demand, being NVIDIA Jetson TX2 a relevant representative. However, whether high-performance GPU configurations are appropriate for automotive setups remains as an open question. This paper aims at providing light on this question by modelling an automotive GPU (Jetson TX2), analyzing its microarchitectural parameters against relevant benchmarks, and identifying specific configurations able to meaningfully increase performance within similar cost envelopes, or to decrease costs preserving original performance levels. Overall, our analysis opens the door to the optimization of automotive GPUs for further system efficiency. Fabio Mazzocchetti, Pedro Benedicte, Hamid Tabani, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla |
SBAC-PAD | 5 |
| 2019 | Time-Randomized Wormhole NoCs for Critical ApplicationsabstractWormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this article, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g., wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance w.r.t. conventional time-deterministic setups. Our results show that performance guarantees of applications running on top of probabilistic wNoC designs improve by 40% and 93% on average for 4 × 4 and 6 × 6 wNoC setups, respectively. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Locality-aware cache random replacement policies
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
J. Syst. Archit. | 3 |
| 2019 | Increasing the Reliability of Software Timing Analysis for Cache-Based ProcessorsabstractReal-time systems are witnessing a significant increase in critical software's size, complexity, and performance needs, which can only be satisfied with high-performance hardware features. Cache memories, pervasively used to improve average performance, complicate Worst-Case Execution Time analysis: cache placement (i.e., how software objects are mapped to cache) during the testing phase does not only critically affect the observed performance, but also proves to be arduous to control and preserve up to operation. The probabilistic variant of Measurement-Based Timing Analysis (MBPTA) responds to this challenge by deploying time-randomized caches that naturally explore a different random cache placement in each run, relieving the user from producing tests that intercept relevant Cache Conflict Placements (CCP). Yet, to meet an adequate probabilistic CCP coverage, the user is required to collect a minimum number of measurements. We present two mechanisms, CCP-RM and CCP-HRP, to identify CCP with relevant probability of occurrence and large impact on execution-time, for the random modulo (RM) and hash-based random placement (HRP) policies. CCP-RM and CCP-HRP enable a reliable application of MBPTA by computing the number of runsR' necessary to meet the desired CCP coverage. We exhaustively evaluate CCP-RM and CCP-HRP, showing their effectiveness on well-known benchmarks and a railway case study, on top of an accurate simulator and a concrete RTL implementation. Suzana Milutinovic, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Computers | 3 |
| 2018 | Modelling multicore contention on the AURIXTM TC27xabstractMulticores are becoming ubiquitous in automotive. Yet, the expected benefits on integration are challenged by multicore contention concerns on timing V&V. Worst-case execution time (WCET) estimates are required as early as possible in the software development, to enable prompt detection of timing misbehavior. Factoring in multicore contention necessarily builds on conservative assumptions on interference, independent of co-runners load on shared hardware. We propose a contention model for automotive multicores that balances time-composability with tightness by exploiting available information on contenders. We tailor the model to the AURIX TC27x and provide tight WCET estimates using information from performance monitors and software configurations. Enrique Díaz, Enrico Mezzetti, Leonidas Kosmidis, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 4 |
| 2018 | Measurement-based cache representativeness on multipath programsabstractAutonomous vehicles in embedded real-time systems increase critical-software size and complexity whose performance needs are covered with high-performance hardware features like caches, which however hampers obtaining WCET estimates that hold valid for all program execution paths. This requires assessing that all cache layouts have been properly factored in the WCET process. For measurement-based timing analysis, the most common analysis method, we provide a solution to achieve cache representativeness and full path coverage: we create a modified program for analysis purposes where cache impact is upper-bounded across any path, and derive the minimum number of runs required to capture in the test campaign cache layouts resulting in high execution times. Suzana Milutinovic, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla |
DAC | 2 |
| 2018 | Cache side-channel attacks and time-predictability in high-performance critical real-time systemsabstractEmbedded computers control an increasing number of systems directly interacting with humans, while also manage more and more personal or sensitive information. As a result, both safety and security are becoming ubiquitous requirements in embedded computers, and automotive is not an exception to that. In this paper we analyze time-predictability (as an example of safety concern) and side-channel attacks (as an example of security issue) in cache memories. While injecting randomization in cache timing-behavior addresses each of those concerns separately, we show that randomization solutions for time-predictability do not protect against side-channel attacks and vice-versa. We then propose a randomization solution to achieve both safety and security goals. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 3 |
| 2018 | Design and integration of hierarchical-placement multi-level caches for real-time systemsabstractEnabling timing analysis in the presence of caches has been pursued by the real-Time embedded systems (RTES) community for years due to cache's huge potential to reduce software's worst-case execution time (WCET). However, caches heavily complicate timing analysis due to hard-To-predict access patterns, with few works dealing with time analyzability of multi-level cache hierarchies. For measurement-based timing analysis (MBTA) techniques-widely used in domains such as avionics, automotive, and rail-we propose several cache hierarchies amenable to MBTA. We focus on a probabilistic variant of MBTA (or MBPTA) that requires caches with time-randomized behavior whose execution time variability can be captured in the measurements taken during system's test runs. For this type of caches, we explore and propose different multi-level cache setups. From those, we choose a cost-effective cache hierarchy that we implement and integrate in a 4-core LEON3 RTL processor model and prototype in a FPGA. Our results show that our proposed setup implemented in RTL results in better (reduced) WCET estimates with similar implementation cost and no impact on average performance w.r.t. other MBPTA-Amenable setups. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 3 |
| 2018 | A Reliable Statistical Analysis of the Best-Fit Distribution for High Execution TimesabstractExtreme Value Theory has been used to model the WCET probabilistically, relying on the assumption that probabilistic WCET (pWCET) estimates can be upper-bounded with exponential distributions, but this is only assessed on execution time samples with pass/fail hypothesis tests. However, the degree of fulfilment of this hypothesis for the execution time sample has a direct impact on the tightness of the pWCET estimate. This paper tackles this limitation of pass/fail tests by applying 3 alternative methods to model the distribution of high execution times through the analysis of the number of finite moments of execution time samples. These methods provide information on the degree of fulfilment of the exponentiality hypothesis, rather than a simple pass/fail response. Hence, whenever the number of finite moments is shown to be low, despite pass/fail tests are passed, these methods indicate that pWCET estimates may be untight. We show that those methods complement each other and the information obtained - number of finite moments proven to exist - can be used to increase the execution time sample size opportunistically to obtain tighter pWCET estimates. Xavier Civit, Joan del Castillo, Jaume Abella 0001 |
DSD | 3 |
| 2018 | HWP: Hardware Support to Reconcile Cache Energy, Complexity, Performance and WCET Estimates in Multicore Real-Time SystemsabstractHigh-performance processors have deployed multilevel cache (MLC) systems for decades. In the embedded real-time market, the use of MLC is also on the rise, with processors for future systems in space, railway, avionics and automotive already featuring two or more cache levels. One of the most critical elements for MLC is the write policy that not only affects several key metrics such as performance, WCET estimates, energy/power, and reliability, but also the design of complexity-prone cache coherence protocol and cache reliability solutions. In this paper we make an extensive analysis of existing write policies, namely write-through (WT) and write-back (WB). In the context of the real-time domain, we show that no write policy is superior for all metrics: WT simplifies the design of the coherence and reliability solutions at the cost of performance, WCET, and energy; while WB improves performance and energy results, but complicates cache design. To take the best of each policy, we propose Hybrid Write Policy (HWP) a low-complexity hardware mechanism that reconciles the benefits of WT in terms of simplifying the cache design (e.g. coherence solution) and the benefits of WB in improved average performance and WCET estimates as the pressure on the interconnection network increases. Guaranteed performance results show that HWP scales with core count similar to WB. Likewise, HWP reduces cache energy usage of WT, to levels similar to those of WB. These benefits are obtained while retaining the reduced coherence complexity of WT, in contrast to high coherence costs under WB. Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
ECRTS | 3 |
| 2018 | NoCo: ILP-Based Worst-Case Contention Estimation for Mesh Real-Time ManycoresabstractManycores are capable of providing the computational demands required by functionally-advanced critical applications in domains such as automotive and avionics. In manycores a network-on-chip (NoC) provides access to shared caches and memories and hence concentrates most of the contention that tasks suffer, with effects on the worst-case contention delay (WCD) of packets and tasks' WCET. While several proposals minimize the impact of individual NoC parameters on WCD, e.g. mapping and routing, there are strong dependences among these NoC parameters. Hence, finding the optimal NoC configurations requires optimizing all parameters simultaneously, which represents a multidimensional optimization problem. In this paper we propose NoCo, a novel approach that combines ILP and stochastic optimization to find NoC configurations in terms of packet routing, application mapping, and arbitration weight allocation. Our results show that NoCo improves other techniques that optimize a subset of NoC parameters. Jordi Cardona, Carles Hernández 0001, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 4 |
| 2018 | Assessing Time Predictability Features of ARM Big. LITTLE MulticoresabstractThe increasing performance needs in critical realtime embedded systems (CRTES), such as for instance the automotive domain, push for the adoption of high-performance hardware from the consumer electronics domain. However, their time-predictability features are quite unexplored. The ARM big. LITTLE architecture is a good candidate for adoption in the CRTES market (i.e. in the automotive market it has already started being used). In this paper we study ARM big. LITTLE's capabilities to meet CRTES requirements. In particular, we perform a qualitative and quantitative assessment of its timing characteristics, focusing on shared multicore resources, and how this architecture can be reliably used in CRTES. Gabriel Fernandez 0002, Francisco J. Cazorla, Jaume Abella 0001, Sylvain Girbal |
SBAC-PAD | 3 |
| 2018 | EOmesh: Combined Flow Balancing and Deterministic Routing for Reduced WCET Estimates in Embedded Real-Time SystemsabstractThe increasing performance needs in critical real-time embedded systems (CRTESs) can only be satisfied with the use of high-performance manycore processors. While NoC-based manycore systems are popular in the high-performance domain due to their high average performance, they challenge deriving tight worst-case execution time (WCET) estimates, as needed in CRTES. Weighted meshes have been proposed to alleviate NoCs pathological behavior-caused by large bandwidth imbalance-by making locally unbalanced arbitration decisions to reach globally balanced bandwidth. In this paper, we show that existing weighted mesh solutions do not completely remove unwanted imbalance, in particular for nodes subject to high congestion. We propose even/odd mesh (EOmesh), an approach that combines heterogeneous predictable routing and weight allocations that delivers near-optimal bandwidth allocation across cores without increasing NoC complexity. EOmesh, which can be implemented either by hardware means or by software means on top of regular weighted meshes, improves the average performance and WCET results of the reference weighted mesh design. Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Fitting Software Execution-Time Exceedance into a Residual Random Fault in ISO-26262abstractCar manufacturers relentlessly replace or augment the functionality of mechanical subsystems with electronic components. Most such subsystems (e.g., steer-by-wire) are safety related, hence, subject to regulation. ISO-26262, the dominant standard for road vehicles, regards software faults as systematic, while differentiating hardware faults between systematic and random. The analysis of systematic faults entails rigorous processes and qualitative considerations. The increasing complexity of modern on-board computers, however, questions the very notion of treating the violation of execution-time envelopes for software programs as a systematic fault. Modern hardware in fact reduces the user's ability to delve deep enough into the fabric of hardware-software interaction to gage its extent of contribution to the worst-case execution time (WCET). Changing the nature of the WCET-analysis problem may help address that challenge effectively. To this end, we propose a solution that should allow ISO-26262 to quantify the likelihood of execution-time exceedance events, relating it to target failure metrics employed in support of certification arguments, similarly to random faults in hardware. To this end, we inject randomization in the timing behavior of the computer hardware to relieve the user from the need to control hard-to-reach low-level parts, and use measurement-based probabilistic timing analysis to quantify, constructively, the failure rates resulting from the likelihood of execution-time exceedance events. Irune Agirre, Francisco J. Cazorla, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Mikel Azkarate-askatsua, Tullio Vardanega |
IEEE Trans. Reliab. | 3 |
| 2017 | DIMP: A Low-Cost Diversity Metric Based on Circuit Path AnalysisabstractDiversity has been regarded as a desirable property of redundant instances, since it allows circuits to behave differently in front of a given fault. However, while qualitatively diversity is a well-understood concept, usable efficient metrics do not exist to quantify diversity in the context of safety-related systems. In this paper we cover this gap by proposing DIMP, a low-cost diversity metric based on analyzing the paths of the redundant circuits. We relate it to the particular case of automotive microcontrollers implementing lockstep cores and show that it can be successfully used providing relevant information for addressing common cause faults. Sergi Alcaide, Carles Hernández 0001, Antoni Roca 0001, Jaume Abella 0001 |
DAC | 4 |
| 2017 | Dynamic software randomisation: Lessons learnec from an aerospace case studyabstractTiming Validation and Verification (V&V) is an important step in real-time system design, in which a system's timing behaviour is assessed via Worst Case Execution Time (WCET) estimation and scheduling analysis. For WCET estimation, measurement-based timing analysis (MBTA) techniques are widely-used and well-established in industrial environments. However, the advent of complex processors makes it more difficult for the user to provide evidence that the software is tested under stress conditions representative of those at system operation. Measurement-Based Probabilistic Timing Analysis (MBPTA) is a variant of MBTA followed by the PROXIMA European Project that facilitates formulating this representativeness argument. MBPTA requires certain properties to be applicable, which can be obtained by selectively injecting randomisation in platform's timing behaviour via hardware or software means. In this paper, we assess the effectiveness of the PROXIMA's dynamic software randomisation (DSR) with a space industrial case study executed on a real unmodified hardware platform and an industrial operating system. We present the challenges faced in its development, in order to achieve MBPTA compliance and the lessons learned from this process. Our results, obtained using a commercial timing analysis tool, indicate that DSR does not impact the average performance of the application, while it enables the use of MBPTA. This results in tighter pWCET estimates compared to current industrial practice. Fabrice Cros, Leonidas Kosmidis, Franck Wartel, David Morales, Jaume Abella 0001, Ian Broster, Francisco J. Cazorla |
DATE | 5 |
| 2017 | Probabilistic timing analysis on time-randomized platforms for the space domainabstractTiming Verification is a fundamental step in real-time embedded systems, with measurement-based timing analysis (MBTA) being the most common approach used to that end. We present a Space case study on a real platform that has been modified to support a probabilistic variant of MBTA called MBPTA. Our platform provides the properties required by MBPTA with the predicted WCET estimates with MBPTA being competitive to those with current MBTA practice while providing more solid evidence on their correctness for certification. Mikel Fernández, David Morales, Leonidas Kosmidis, Alen Bardizbanyan, Ian Broster, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla, Paulo Machado, Luca Fossati |
DATE | 8 |
| 2017 | Design and implementation of a fair credit-based bandwidth sharing scheme for busesabstractFair arbitration in the access to hardware shared resources is fundamental to obtain low worst-case execution time (WCET) estimates in the context of critical real-time systems, for which performance guarantees are essential. Several hardware mechanisms exist for managing arbitration in those resources (buses, memory controllers, etc.). They typically attain fairness in terms of the number of slots each contender (e.g., core) gets granted access to the shared resource. However, those policies may lead to unfair bandwidth allocations for workloads with contenders issuing short requests and contenders issuing long requests. We propose a Credit-Based Arbitration (CBA) mechanism that achieves fairness in the cycles each core is granted access to the resource rather than in the number of granted slots. Furthermore, we implement CBA as part of a LEON3 4-core processor for the Space domain in an FPGA proving the feasibility and good performance characteristics of the design by comparing it against other arbitration schemes. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DATE | 3 |
| 2017 | Boosting Guaranteed Performance in Wormhole NoCs with Probabilistic Timing AnalysisabstractWormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this paper, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g. wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance. Our results show that WCET estimates of applications running on top of probabilistic wNoCs are reduced by 40% and 75% on average for 4x4 and 6x6 wNoC setups respectively when compared against deterministic wNoCs. Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 3 |
| 2017 | Design and Implementation of a Time Predictable Processor: Evaluation With a Space Case StudyabstractEmbedded real-time systems like those found in automotive, rail and aerospace, steadily require higher levels of guaranteed computing performance (and hence time predictability) motivated by the increasing number of functionalities provided by software. However, high-performance processor design is driven by the average-performance needs of mainstream market. To make things worse, changing those designs is hard since the embedded real-time market is comparatively a small market. A path to address this mismatch is designing low-complexity hardware features that favor time predictability and can be enabled/disabled not to affect average performance when performance guarantees are not required. In this line, we present the lessons learned designing and implementing LEOPARD, a four-core processor facilitating measurement-based timing analysis (widely used in most domains). LEOPARD has been designed adding low-overhead hardware mechanisms to a LEON3 processor baseline that allow capturing the impact of jittery resources (i.e. with variable latency) in the measurements performed at analysis time. In particular, at core level we handle the jitter of caches, TLBs and variable-latency floating point units; and at the chip level, we deal with contention so that time-composable timing guarantees can be obtained. The result of our applied study with a Space application shows how per-resource jitter is controlled facilitating the computation of high-quality WCET estimates. Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla, Alen Bardizbanyan, Jan Andersson, Fabrice Cros, Franck Wartel |
ECRTS | 2 |
| 2017 | EPC Enacted: Integration in an Industrial Toolbox and Use against a Railway ApplicationabstractMeasurement-based timing analysis approaches are increasingly making their way into several industrial domains on account of their good cost-benefit ratio. The trustworthiness of those methods, however, suffers from the limitation that their results are only valid for the particular paths and execution conditions that the user is able to explore with the available input vectors. It is generally not possible to guarantee that the collected measurements are fully representative of the worst-case timing behaviour. In the context of measurement-based probabilistic timing analysis, the Extended Path Coverage (EPC) approach has been recently proposed as a means to extend the representativeness of measurement observations, to obtain the same effect of full path coverage. At the time of its first publication, EPC had not reached an implementation maturity that could be trialled industrially. In this work we analyze the practical implications of using EPC with real-world applications, and discuss the challenges in integrating it in an industrial-quality toolchain. We show that we were able to meet EPC requirements and successfully evaluate the technique on a real Railway application, on top of a commercial toolchain and full execution stack. Enrico Mezzetti, Mikel Fernández, Alen Bardizbanyan, Irune Agirre, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla |
RTAS | 5 |
| 2017 | Work-in-Progress Paper: An Analysis of the Impact of Dependencies on Probabilistic Timing Analysis and Task SchedulingabstractRecently there has been a renewed interest for probabilistic timing analysis (PTA) and probabilistic task scheduling (PTS). Despite the number of works in both fields, the link between them is weak: works on the latter build upon a series of assumptions on the probabilistic behavior of each task – or instances (jobs) of it – that have not been shown how to be fulfilled by PTA. This paper makes a first step towards covering this gap with emphasis on providing the right meaning of pWCET estimate as understood by both PTA and PTS. We show that the main issue related to ensuring that PTS assumptions on pWCET estimates are captured by PTA relates to the dependencies among tasks, and even jobs of a given task. Both change the scope of applicability of pWCET estimates provided by PTA and hence, their use by PTS. Enrico Mezzetti, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla |
RTSS | 2 |
| 2017 | SEDEA: A Sensible Approach to Account DRAM Energy in Multicore SystemsabstractAs the energy cost in todays computing systems keeps increasing, measuring the energy becomes crucial in many scenarios. For instance, due to the fact that the operational cost of datacenters largely depends on the energy consumed by the applications executed, end users should be charged for the energy consumed, which requires a fair and consistent energy measuring approach. However, the use of multicore system complicates per-task energy measurement as the increased Thread Level Parallelism (TLP) allows several tasks to run simultaneously sharing resources. Therefore, the energy usage of each task is hard to determine due to interleaved activities and mutual interferences. To this end, Per-Task Energy Metering (PTEM) has been proposed to measure the actual energy of each task based on their resource utilization in a workload. However, the measured energy depends on the interferences from co-running tasks sharing the resources, and thus fails to provide the consistency across executions. Therefore, Sensible Energy Accounting (SEA) has been proposed to deliver an abstraction of the energy consumption based on a particular allocation of resources to a task. In this work we provide a realization of SEA for the DRAM memory system, SEDEA, where we account a task for the DRAM energy it would have consumed when running in isolation with a fraction of the on-chip shared cache. SEDEA is a mechanism to sensibly account for the DRAM energy of a task based on predicting its memory behavior. Our results show that SEDEA provides accurate estimates, yet with low-cost, beating existing per-task energy models, which do not target accounting energy in multicore system. We also provide a use case showing that SEDEA can be used to guide shared cache and memory bank partition schemes to save energy. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
SBAC-PAD | 3 |
| 2017 | Computing Safe Contention Bounds for Multicore Resources with Round-Robin and FIFO ArbitrationabstractNumerous researchers have studied the contention that arises among tasks running in parallel on a multicore processor. Most of those studies seek to derive a tight and sound upper-bound for the worst-case delay with which a processor resource may serve an incoming request, when its access is arbitrated using time-predictable policies such as round-robin or FIFO. We call this value upper-bound delay (ubd). Deriving trustworthy ubd statically is possible when sufficient public information exists on the timing latency incurred on access to the resource of interest. Unfortunately however, that is rarely granted for commercial-of-the-shelf (COTS) processors. Therefore, the users resort to measurement observations on the target processor and thus compute a “measured” ubdm. However, using ubdm to compute worst-case execution time values for programs running on COTS multicore processors requires qualification on the soundness of the result. In this paper, we present a measurement-based methodology to derive a ubdm under round-robin (RoRo) and first-in-first-out (FIFO) arbitration, which accurately approximates ubd from above, without needing latency information from the hardware provider. Experimental results, obtained on multiple processor configurations, demonstrate the robustness of the proposed methodology. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
IEEE Trans. Computers | 3 |
| 2017 | Measurement-Based Worst-Case Execution Time Estimation Using the Coefficient of VariationabstractExtreme Value Theory (EVT) has been historically used in domains such as finance and hydrology to model worst-case events (e.g., major stock market incidences). EVT takes as input a sample of the distribution of the variable to model and fits the tail of that sample to either the Generalised Extreme Value (GEV) or the Generalised Pareto Distribution (GPD). Recently, EVT has become popular in real-time systems to derive worst-case execution time (WCET) estimates of programs. However, the application of EVT is not straightforward and requires a detailed analysis of, and customisation for, the particular problem at hand. In this article, we tailor the application of EVT to timing analysis. To that end, (1) we analyse the response time of different hardware resources (e.g., cache memories) and identify those that may lead to radically different types of execution time distributions. (2) We show that one of these distributions, known as mixture distribution, causes problems in the use of EVT. In particular, mixture distributions challenge not only properly selecting GEV/GPD parameters (i.e., location, scale and shape) but also determining the size of the sample to ensure that enough tail values are passed to EVT and that only tail values are used by EVT to fit GEV/GPD. Failing to select these parameters has a negative impact on the quality of the derived WCET estimates. We tackle these problems, by (3) proposing Measurement-Based Probabilistic Timing Analysis using the Coefficient of Variation (MBPTA-CV), a new mixture-distribution aware, WCET-suited MBPTA method that builds on recent EVT developments in other fields (e.g., finance) to automatically select the distribution parameters that best fit the maxima of the observed execution times. Our results on a simulation environment and a real board show that MBPTA-CV produces high-quality WCET estimates. Jaume Abella 0001, Maria Padilla, Joan del Castillo, Francisco J. Cazorla |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2016 | Random modulo: a new processor cache design for real-time critical systemsabstractCache memories have a huge impact on software's worst-case execution time (WCET). While enabling the seamless use of caches is key to provide the increasing levels of (guaranteed) performance required by automotive software, caches complicate timing analysis. In the context of Measurement-Based Probabilistic Timing Analysis (MBPTA) -- a promising technique to ease timing analyis of complex hardware -- we propose Random Modulo (RM), a new cache design that provides the probabilistic behavior required by MBPTA and with the following advantages over existing MBPTA-compliant cache designs: (i) an outstanding reduction in WCET estimates, (ii) lower latency and area overhead, and (iii) competitive average performance w.r.t conventional caches. Carles Hernández 0001, Jaume Abella 0001, Andrea Gianarro, Jan Andersson, Francisco J. Cazorla |
DAC | 2 |
| 2016 | Supertask: Maximizing runnable-level parallelism in AUTOSAR applications
Sebastian Kehr, Milos Panic, Eduardo Quiñones, Bert Böddeker, Jorge Becerril Sandoval, Jaume Abella 0001, Francisco J. Cazorla, Günter Schäfer |
DATE | 6 |
| 2016 | Improving performance guarantees in wormhole mesh NoC designs
Milos Panic, Carles Hernández 0001, Jaume Abella 0001, Antoni Roca 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 3 |
| 2016 | A detailed methodology to compute Soft Error Rates in advanced technologies
Marc Riera, Ramon Canal, Jaume Abella 0001, Antonio González 0001 |
DATE | 3 |
| 2016 | PROXIMA: Improving Measurement-Based Timing Analysis through Randomisation and Probabilistic AnalysisabstractThe use of increasingly complex hardware and software platforms in response to the ever rising performance demands of modern real-time systems complicates the verification and validation of their timing behaviour, which form a time-and-effort-intensive step of system qualification or certification. In this paper we relate the current state of practice in measurement-based timing analysis, the predominant choice for industrial developers, to the proceedings of the PROXIMA (Probabilistic real-time control of mixed-criticality multicore systems) project in that very field. We recall the difficulties that the shift towards more complex computing platforms causes in that regard. Then we discuss the probabilistic approach proposed by PROXIMA to overcome some of those limitations. We present the main principles behind the PROXIMA approach as well as the changes it requires at hardware or software level underneath the application. We also present the current status of the project against its overall goals, and highlight some of the principal confidence-building results achieved so far. Francisco J. Cazorla, Jaume Abella 0001, Jan Andersson, Tullio Vardanega, Francis Vatrinet, Iain Bate, Ian Broster, Mikel Azkarate-askatsua, Franck Wartel, Liliana Cucu-Grosjean, Fabrice Cros, Glenn Farrall, Adriana Gogonel, Andrea Gianarro, Benoit Triquet, Carles Hernández 0001, Code Lo, Cristian Maxim, David Morales, Eduardo Quiñones, Enrico Mezzetti, Leonidas Kosmidis, Irune Agirre, Mikel Fernández, Mladen Slijepcevic, Philippa Conmy, Walid Talaboulma |
DSD | 2 |
| 2016 | pTNoC: Probabilistically Time-Analyzable Tree-Based NoC for Mixed-Criticality SystemsabstractThe use of networks-on-chip (NoC) in real-time safety-critical multicore systems challenges deriving tight worst-case execution time (WCET) estimates. This is due to the complexities in tightly upper-bounding the contention in the access to the NoC among running tasks. Probabilistic Timing Analysis (PTA) is a powerful approach to derive WCET estimates on relatively complex processors. However, so far it has only been tested on small multicores comprising an on-chip bus as communication means, which intrinsically does not scale to high core counts. In this paper we propose pTNoC, a new tree-based NoC design compatible with PTA requirements and delivering scalability towards medium/large core counts. pTNoC provides tight WCET estimates by means of asymmetric bandwidth guarantees for mixed-criticality systems with negligible impact on average performance. Finally, our implementation results show the reduced area and power costs of the pTNoC. Mladen Slijepcevic, Mikel Fernández, Carles Hernández 0001, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DSD | 4 |
| 2016 | TASA: toolchain-agnostic static software randomisation for critical real-time systemsabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) derives WCET estimates for tasks running on processors comprising high-performance features such as caches. MBPTA's correct application requires the system to exhibit certain timing properties, which can be achieved by injecting randomisation in the timing behaviour of the task under analysis. However, existing software-randomisation techniques require costly modifications in the industrial production toolchain (compiler, linker, runtime or hardware) in terms of development and certification. In this paper we present TASA, a new software randomisation tool that relies on source-code transformations of the application (i) requiring no changes in existing toolchains, which heavily reduces tool qualification and implementation costs; and (ii) achieving competitive WCET estimates that we assess on a gcc- and a llvm-based compilation toolchain on a real board. Leonidas Kosmidis, Roberto Vargas, David Morales, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
ICCAD | 5 |
| 2016 | A confidence assessment of WCET estimates for software time randomized cachesabstractObtaining Worst-Case Execution Time (WCET) estimates is a required step in real-time embedded systems during software verification. Measurement-Based Probabilistic Timing Analysis (MBPTA) aims at obtaining WCET estimates for industrial-size software running upon hardware platforms comprising high-performance features. MBPTA relies on the randomization of timing behavior (functional behavior is left unchanged) of hard-to-predict events like the location of objects in memory - and hence their associated cache behavior - that significantly impact software's WCET estimates. Software time-randomized caches (sTRc) have been recently proposed to enable MBPTA on top of Commercial off-the-shelf (COTS) caches (e.g. modulo placement). However, some random events may challenge MBPTA reliability on top of sTRc. In this paper, for sTRc and programs with homogeneously accessed addresses, we determine whether the number of observations taken at analysis, as part of the normal MBPTA application process, captures the cache events significantly impacting execution time and WCET. If this is not the case, our techniques provide the user with the number of extra runs to perform to guarantee that cache events are captured for a reliable application of MBPTA. Our techniques are evaluated with synthetic benchmarks and an avionics application. Pedro Benedicte, Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
INDIN | 4 |
| 2016 | Modeling RTL fault models behavior to increase the confidence on TSIM-based fault injectionabstractFuture high-performance safety-relevant applications require microcontrollers delivering higher performance than the existing certified ones. However, means for assessing their dependability are needed so that they can be certified against safety critical certification standars (e.g ISO26262). Dependability assessment analyses performed at high level of abstraction inject single faults to investigate the effects these have in the system. In this work we show that single faults do not comprise the whole picture, due to fault multiplicities and reactivations. Later we prove that, by injecting complex fault models that consider multiplicities and reactivations in higher levels of abstraction, results are substantially different, thus indicating that a change in the methodology is needed. Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001 |
IOLTS | 3 |
| 2016 | Resilient random modulo cache memories for probabilistically-analyzable real-time systemsabstractFault tolerance has often been assessed separately in safety-related real-time systems, which may lead to inefficient solutions. Recently, Measurement-Based Probabilistic Timing Analysis (MBPTA) has been proposed to estimate Worst-Case Execution Time (WCET) on high performance hardware. The intrinsic probabilistic nature of MBPTA-commpliant hardware matches perfectly with the random nature of hardware faults. Joint WCET analysis and reliability assessment has been done so far for some MBPTA-compliant designs, but not for the most promising cache design: random modulo. In this paper we perform, for the first time, an assessment of the aging-robustness of random modulo and propose new implementations preserving the key properties of random modulo, a.k.a. low critical path impact, low miss rates and MBPTA compliance, while enhancing reliability in front of aging by achieving a better - yet random - activity distribution across cache sets. David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
IOLTS | 3 |
| 2016 | Modelling Probabilistic Cache Representativeness in the Presence of Arbitrary Access PatternsabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) is a promising powerful industry-friendly method to derive worst-case execution time (WCET) estimates as needed for critical real-time embedded systems. MBPTA performs several (R) runs of the program on the target platform collecting the execution times in each run. MBPTA builds a probabilistic representativeness argument on whether those events with high impact on execution time, such as cache misses, arise on the runs made at analysis time so that their impact on execution time is captured. So far only events occurring in cache memories have been shown to challenge providing such representativeness argument. In this context, this paper introduces a representativeness validation method (RVS) to assess the probabilistic representativeness of MBPTA's execution time observations in terms of cache behaviour. RVS resorts to cache simulation to predict worst-case miss scenarios that can appear during the deployment phase. RVS also constructs a probabilistic Worst-Case Miss Count curve based on the miss-counts captured in the R runs. If that curve upperbounds the impact of the predicted cache worst-case scenarios, R is deemed as a sufficient number of runs for which pWCET estimates can be reliably derived. Otherwise, the user is requested to perform more runs until all cache scenarios of interest are captured. Suzana Milutinovic, Jaume Abella 0001, Francisco J. Cazorla |
ISORC | 2 |
| 2016 | Modeling High-Performance Wormhole NoCs for Critical Real-Time Embedded SystemsabstractManycore chips are a promising computing platform to cope with the increasing performance needs of critical real-time embedded systems (CRTES). However, manycores adoption by CRTES industry requires understanding task's timing behavior when their requests use manycore's network-on-chip (NoC) to access hardware shared resources. This paper analyzes the contention in wormhole-based NoC (wNoC) designs - widely implemented in the high-performance domain - for which we introduce a new metric: worst-contention delay (WCD) that captures wNoC impact on worst-case execution time (WCET) in a tighter manner than the existing metric, worst-case traversal time (WCTT). Moreover, we provide an analytical model of the WCD that requests can suffer in a wNoC and we validate it against wNoC designs resembling those in the Tilera-Gx36 and the Intel-SCC 48-core processors. Building on top of our WCD analytical model, we analyze the impact on WCD that different design parameters such as the number of virtual channels, and we make a set of recommendations on what wNoC setups to use in the context of CRTES. Milos Panic, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 4 |
| 2016 | Improving Early Design Stage Timing Modeling in Multicore Based Real-Time SystemsabstractThis paper presents a modelling approach for the timing behavior of real-time embedded systems (RTES) in early design phases. The model focuses on multicore processors - accepted as the next computing platform for RTES - and in particular it predicts the contention tasks suffer in the access to multicore on-chip shared resources. The model presents the key properties of not requiring the application's source code or binary and having high-accuracy and low overhead. The former is of paramount importance in those common scenarios in which several software suppliers work in parallel implementing different applications for a system integrator, subject to different intellectual property (IP) constraints. Our model helps reducing the risk of exceeding the assigned budgets for each application in late design stages and its associated costs. David Trilla, Javier Jalle, Mikel Fernández, Jaume Abella 0001, Francisco J. Cazorla |
RTAS | 4 |
| 2016 | Sensible Energy Accounting with Abstract Metering for Multicore SystemsabstractChip multicore processors (CMPs) are the preferred processing platform across different domains such as data centers, real-time systems, and mobile devices. In all those domains, energy is arguably the most expensive resource in a computing system. Accurately quantifying energy usage in a multicore environment presents a challenge as well as an opportunity for optimization. Standard metering approaches are not capable of delivering consistent results with shared resources, since the same task with the same inputs may have different energy consumption based on the mix of co-running tasks. However, it is reasonable for data-center operators to charge on the basis of estimated energy usage rather than time since energy is more correlated with their actual cost. This article introduces the concept of Sensible Energy Accounting (SEA). For a task running in a multicore system, SEA accurately estimates the energy the task would have consumed running in isolation with a given fraction of the CMP shared resources. We explain the potential benefits of SEA in different domains and describe two hardware techniques to implement it for a shared last-level cache and on-core resources in SMT processors. Moreover, with SEA, an energy-aware scheduler can find a highly efficient on-chip resource assignment, reducing by up to 39% the total processor energy for a 4-core system. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Daniel A. Jiménez, Mateo Valero |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Parallelizing Industrial Hard Real-Time Applications for the parMERASA MulticoreabstractThe EC project parMERASA (Multicore Execution of Parallelized Hard Real-Time Applications Supporting Analyzability) investigated timing-analyzable parallel hard real-time applications running on a predictable multicore processor. A pattern-supported parallelization approach was developed to ease sequential to parallel program transformation based on parallel design patterns that are timing analyzable. The parallelization approach was applied to parallelize the following industrial hard real-time programs: 3D path planning and stereo navigation algorithms (Honeywell International s.r.o.), control algorithm for a dynamic compaction machine (BAUER Maschinen GmbH), and a diesel engine management system (DENSO AUTOMOTIVE Deutschland GmbH). This article focuses on the parallelization approach, experiences during parallelization with the applications, and quantitative results reached by simulation, by static WCET analysis with the OTAWA tool, and by measurement-based WCET analysis with the RapiTime tool. Theo Ungerer, Christian Bradatsch, Martin Frieb, Florian Kluge, Jörg Mische, Alexander Stegmeier, Ralf Jahr, Mike Gerdes 0001, Pavel G. Zaykov, Lucie Matusova, Zai Jian Jia Li, Zlatko Petrov, Bert Böddeker, Sebastian Kehr, Hans Regler, Andreas Hugl, Christine Rochange, Haluk Ozaktas, Hugues Cassé, Armelle Bonenfant, Pascal Sainrat, Nick Lay, Ian Broster, Eduardo Quiñones, Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla, Sascha Uhrig, Mathias Rohde, Arthur Pyka |
ACM Trans. Embed. Comput. Syst. | 27 |
| 2016 | DReAM: An Approach to Estimate per-Task DRAM Energy in Multicore SystemsabstractAccurate per-task energy estimation in multicore systems would allow performing per-task energy-aware task scheduling and energy-aware billing in data centers, among other applications. Per-task energy estimation is challenged by the interaction between tasks in shared resources, which impacts tasks’ energy consumption in uncontrolled ways. Some accurate mechanisms have been devised recently to estimate per-task energy consumed on-chip in multicores, but there is a lack of such mechanisms for DRAM memories. This article makes the case for accurate per-task DRAM energy metering in multicores, which opens new paths to energy/performance optimizations. In particular, the contributions of this article are (i) an ideal per-task energy metering model for DRAM memories; (ii) DReAM, an accurate yet low cost implementation of the ideal model (less than 5% accuracy error when 16 tasks share memory); and (iii) a comparison with standard methods (even distribution and access-count based) proving that DReAM is much more accurate than these other methods. Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2015 | Analysis and RTL correlation of instruction set simulators for automotive microcontroller robustness verificationabstractIncreasingly complex microcontroller designs for safety-relevant automotive systems require the adoption of new methods and tools to enable a cost-effective verification of their robustness. In particular, costs associated to the certification against the ISO26262 safety standard must be kept low for economical reasons. In this context, simulation-based verification using instruction set simulators (ISS) arises as a promising approach to partially cope with the increasing cost of the verification process as it allows taking design decisions in early design stages when modifications can be performed quickly and with low cost. However, it remains to be proven that verification in those stages provides accurate enough information to be used in the context of automotive microcontrollers. In this paper we analyze the existing correlation between fault injection experiments in an RTL microcontroller description and the information available at the ISS to enable accurate ISS-based fault injection. Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001, David de Andrés, Juan-Carlos Ruiz-Garcia 0001 |
DAC | 3 |
| 2015 | Increasing confidence on measurement-based contention bounds for real-time round-robin busesabstractContention among tasks concurrently running in a multicore has been deeply studied in the literature specially for on-chip buses. Most of the works so far focus on deriving exact upper-bounds to the longest delay it takes a bus request to be serviced (ubd), when its access is arbitrated using a time-predictable policy such as round-robin. Deriving ubd for a bus can be done accurately when enough timing information is available, which is not often the case for commercial-of-the-shelf (COTS) processors. Hence, ubd is approximated (ubdm) by directly experimenting on the target processor, i.e by measurements. However, using ubdm makes the timing analysis technique to resort on the accuracy of ubdm to derive trustworthy worst-case execution time estimates. Therefore, accurately estimating ubd by means of ubdm is of paramount importance. In this paper, we propose a systematic measurement-based methodology to accurately approximate ubd without knowing the bus latency or any other latency information, being only required that the underlying bus policy is round-robin. Our experimental results prove the robustness of the proposed methodology by testing it on different bus and processor setups. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
DAC | 3 |
| 2015 | Resource usage templates and signatures for COTS multicore processorsabstractUpper bounding the execution time of tasks running on multicore processors is a hard challenge. This is especially so with commercial-off-the-shelf (COTS) hardware that conceals its internal operation. The main difficulty stems from the contention effects on access to hardware shared resources (e.g., buses) which cause task's timing behavior to depend on the load that co-runner tasks place on them. This dependence reduces time composability and constrains incremental verification. In this paper we introduce the concepts of resource-usage signatures and templates, to abstract the potential contention caused and incurred by tasks running on a multicore. We propose an approach that employs resource-usage signatures and templates to enable the analysis of individual tasks largely in isolation, with low integration costs, producing execution time estimates per task that are easily composable throughout the whole system integration process. We evaluate the proposal on a 4-core NGMP-like multicore architecture. Gabriel Fernandez 0002, Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Tullio Vardanega, Francisco J. Cazorla |
DAC | 3 |
| 2015 | PACO: fast average-performance estimation for time-randomized cachesabstractProbabilistic timing analysis is a powerful approach to derive worst-case execution time (WCET) estimates, as needed in safety-critical systems, in the presence of high-performance hardware features (e.g., caches). To that end, the timing behavior of certain hardware resources, such as caches, is randomized. Time-randomized (TR) caches allow deriving hit/miss probabilities for each access and probabilistic WCET estimates for the overall program. Suzana Milutinovic, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla |
DAC | 3 |
| 2015 | Low-cost checkpointing in automotive safety-relevant systems
Carles Hernández 0001, Jaume Abella 0001 |
DATE | 2 |
| 2015 | Timing analysis of an avionics case study on complex hardware/software platforms
Franck Wartel, Leonidas Kosmidis, Adriana Gogonel, Andrea Baldovin, Zoë Stephenson, Benoit Triquet, Eduardo Quiñones, Code Lo, Enrico Mezzetti, Ian Broster, Jaume Abella 0001, Liliana Cucu-Grosjean, Tullio Vardanega, Francisco J. Cazorla |
DATE | 11 |
| 2015 | IEC-61508 SIL 3 Compliant Pseudo-Random Number Generators for Probabilistic Timing AnalysisabstractProbabilistic Timing Analysis (PTA), especially its measurement based variant (MBPTA), has shown to be competitive with state-of-the-art timing analysis techniques. The use of MBPTA to analyse the timing behaviour of safety-critical systems rests on its ability to derive trustworthy WCET bounds. This ability depends on the soundness of the MBPTA method per se, as well as on the satisfaction of safety requirements placed on the pseudo-random number generator (prng) that plays a key role in the platform-level randomisation needed by MBPTA. This paper presents the design of a low-area, low-power prng that meets IEC-61508 SIL 3 safety requirements and allows for seamless integration in a real-world multicore architecture. This work enables the development and the IEC-61508 certification of mixed-criticality systems that use MBPTA for deriving timing bounds for mixed-criticality software programs running on multicore processors. Irune Agirre, Mikel Azkarate-askatsua, Carles Hernández 0001, Jaume Abella 0001, Jon Pérez 0001, Tullio Vardanega, Francisco J. Cazorla |
DSD | 4 |
| 2015 | Enabling TDMA Arbitration in the Context of MBPTAabstractCurrent timing analysis techniques can be broadly classified into two families: deterministic timing analysis (DTA) and probabilistic timing analysis (PTA). Each family defines a set of properties to be provided (enforced) by the hardware and software platform so that valid Worst-Case Execution Time (WCET) estimates can be derived for programs running on that platform. However, the fact that each family relies on each own set of hardware designs limits their applicability and reduces the chances of those designs being adopted by hardware vendors. In this paper we show that Time Division Multiple Access (TDMA), one of the main DTA-compliant arbitration policies, can be made PTA-compliant. To that end, we analyze TDMA in the context of measurement-based PTA (MBPTA) and show that padding execution time observations conveniently leads to trustworthy and tight WCET estimates with MBPTA without introducing any hardware change. In fact, TDMA outperforms round-robin and time-randomized policies in terms of WCET in the context of MBPTA. Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Eduardo Quiñones, Theo Ungerer, Francisco J. Cazorla |
DSD | 2 |
| 2015 | CAP: Communication-Aware Allocation Algorithm for Real-Time Parallel Applications on Many-CoresabstractCritical Real-Time Embedded Systems (CRTES) require additional computing power to match the performance requirements of increasingly complex critical functions. Many-core processors are a key solution to reach the required performance. They allow simultaneous execution of multiple critical functions comprising a number of (parallel) applications. These applications, each implementing some functionality of the system, need to communicate among them in order to make the system work. In many-cores the impact of this communication on the timing behavior of applications depends on the allocation of applications across the chip. In this paper we propose CAP: an allocation algorithm that takes into account communication among applications and tries to reduce its impact on Worst Case Execution Time estimates of applications. We show that using CAP increases the number of applications that can be scheduled on the many-core platform thus facilitating system integration. Milos Panic, Eduardo Quiñones, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
DSD | 4 |
| 2015 | Characterizing fault propagation in safety-critical processor designsabstractAchieving reduced time-to-market in modern electronic designs targeting safety critical applications is becoming very challenging, as these designs need to go through a certification step that introduces a non-negligible overhead in the verification and validation process. To cope with this challenge, safety-critical systems industry is demanding new tools and methodologies allowing quick and cost-effective means for robustness verification. Microarchitectural simulators have been widely used to test reliability properties in different domains but their use in the process of robustness verification remains yet to be validated against other accepted methods such as RTL or gate-level simulation. In this paper we perform fault injections in an RTL model of a processor to characterize fault propagation. The results and conclusions of this characterization will serve to devise to what extent fault injection methodologies for robustness verification using microarchitectural simulators can be employed. Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001 |
IOLTS | 3 |
| 2015 | Seeking Time-Composable Partitions of Tasks for COTS Multicore ProcessorsabstractThe timing verification of real-time single core systems involves a timing analysis step that yields an Execution Time Bound (ETB) for each task, followed by a schedulability analysis step, where the scheduling attributes of the individual tasks, including the ETB, are studied from the system level perspective. The transition between those two steps involves accounting for the interference effects that arise when tasks contend for access to shared resource. The advent of multicore processors challenges the viability of this two-step approach because several complex contention effects at the processor level arise that cause tasks to be unable to make progress while actually holding the CPU, which are very difficult to tightly capture by simply inflating the tasks' ETB. In this paper we show how contention on access to hardware shared resources creates a circular dependence between the determination of tasks' ETB and their scheduling at runtime. To help loosen this knot we present an approach that acknowledges different flavors of time compos ability, examining in detail the variant intended for partitioned scheduling, which we evaluate on two real processor boards used in the space domain. Gabriel Fernandez 0002, Jaume Abella 0001, Eduardo Quiñones, Luca Fossati, Marco Zulianello, Tullio Vardanega, Francisco J. Cazorla |
ISORC | 2 |
| 2015 | EPC: Extended Path Coverage for Measurement-Based Probabilistic Timing AnalysisabstractMeasurement-based probabilistic timing analysis (MBPTA) computes trustworthy upper bounds to the execution time of software programs. MBPTA has the connotation, typical of measurement-based techniques, that the bounds computed with it only relate to what is observed in actual program traversals, which may not include the effective worst-case phenomena. To overcome this limitation, we propose Extended Path Coverage (EPC), a novel technique that allows extending the representativeness of the bounds computed by MBPTA. We make the observation data probabilistically path-independent by modifying the probability distribution of the observed timing behaviour so as to negatively compensate for any benefits that a basic block may draw from a path leading to it. This enables the derivation of trustworthy upper bounds to the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. Our results confirm that using MBPTA with EPC produces fully trustworthy upper bounds with competitively small overestimation in comparison to state-of-the-art MBPTA techniques. Marco Ziccardi, Enrico Mezzetti, Tullio Vardanega, Jaume Abella 0001, Francisco J. Cazorla |
RTSS | 4 |
| 2015 | Timely Error Detection for Effective Recovery in Light-Lockstep Automotive SystemsabstractSafety-relevant systems in the automotive domain often implement features such as lockstep execution for error detection, and reset and re-execution for error correction. Light-lockstep has already been adopted in some such systems due to its relatively low-implementation cost given that it does not require deep changes into nonlockstep hardware. Instead, as only off-core activities (i.e., data/addresses sent) need to be compared across different cores, light-lockstep designs are lowly intrusive. This approach has been proven sufficient to guarantee functional correctness of the system in the presence of errors in the cores, in particular in relation with certification against safety standards such as ISO26262 in the automotive domain. However, error detection in light-lockstep systems may occur long after the error actually occurs, thus jeopardizing timing guarantees, which are as critical as functional ones in hard real-time systems. In this paper, we analyze the timing behavior of errors due to transient and permanent faults in light-lockstep systems. Our results show that the time elapsed until an error is detected can be inordinately large, especially for permanent faults. Based on this observation and building upon the specific characteristics of light-lockstep systems, we propose lightly verbose (LiVe), a new mechanism to enforce the early detection of errors, due to both transient and permanent faults, thus enabling the computation of tight error detection timing bounds. We also analyze how existing mechanisms for error recovery in multicore systems increase their effectiveness when light-lockstep operates in LiVe mode in the context of mixed-criticality workloads. Carles Hernández 0001, Jaume Abella 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | LiVe: Timely Error Detection in Light-Lockstep Safety Critical SystemsabstractSafety-critical systems rely on features such as lockstep execution for error detection, and reset and reexecution for error correction. In particular, light lockstep is an attractive choice since it does not require redesigning cores but, instead, comparing off-core activities (i.e. data/addresses sent). While this approach suffices to guarantee functional correctness of the system, as needed for certification against safety standards (e.g., ISO26262), it fails to provide any timing guarantee as the time elapsed since the error occurs until lockstep detects it can be inordinately large. Carles Hernández 0001, Jaume Abella 0001 |
DAC | 2 |
| 2014 | Containing Timing-Related Certification Cost in Automotive Systems Deploying Complex HardwareabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) techniques simplify deriving tight and trustworthy WCET estimates for industrial-size programs running on complex processors. MBPTA poses some requirements on the timing behaviour of the hardware/software platform: execution times of end-to-end runs have to be independent and identically distributed (i.i.d.). Hardware and software solutions have been deployed to accomplish MBPTA requirements. The latter has achieved the i.i.d. properties running on some commercial off-the-shelf (COTS) processor designs. Unfortunately, software randomisation challenges functional verification needed for certification since it introduces indirections through pointers in the code. In this paper we propose a new approach to software randomisation able to contain its functional verification costs. Our approach performs software randomisation statically, as opposed to current dynamic approaches. We carefully review the requirements of the new approach and prove its feasibility. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Glenn Farrall, Franck Wartel, Francisco J. Cazorla |
DAC | 3 |
| 2014 | Time-Analysable Non-Partitioned Shared Caches for Real-Time Multicore SystemsabstractShared caches in multicores challenge Worst-Case Execution Time (WCET) estimation due to inter-task interferences. Hardware and software cache partitioning address this issue although they complicate data sharing among tasks and the Operating System (OS) task scheduling and migration. In the context of Probabilistic Timing Analysis (PTA) time-randomised caches are used. We propose a new hardware mechanism to control inter-task interferences in shared time-randomised caches without the need of any hardware or software partitioning. Our proposed mechanism effectively bounds inter-task interferences by limiting the cache eviction frequency of each task, while providing tighter WCET estimates than cache partitioning algorithms. In a 4-core multicore processor setup our proposal improves cache partitioning by 56% in terms of guaranteed performance and 16% in terms of average performance. Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DAC | 3 |
| 2014 | Bus designs for time-probabilistic multicore processorsabstractProbabilistic Timing Analysis (PTA) reduces the amount of information needed to provide tight WCET estimates in real-time systems with respect to classic timing analysis. PTA imposes new requirements on hardware design that have been shown implementable for single-core architectures. However, no support has been proposed for multicores so far. In this paper, we propose several probabilistically-analysable bus designs for multicore processors ranging from 4 cores connected with a single bus, to 16 cores deploying a hierarchical bus design. We derive analytical models of the probabilistic timing behaviour for the different bus designs, show their suitability for PTA and evaluate their hardware cost. Our results show that the proposed bus designs (i) fulfil PTA requirements, (ii) allow deriving WCET estimates with the same cost and complexity as in single-core processors, and (iii) provide higher guaranteed performance than single-core processors, 3.4x and 6.6x on average for an 8-core and a 16-core setup respectively. Javier Jalle, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 3 |
| 2014 | Measurement-Based Probabilistic Timing Analysis and Its Impact on Processor ArchitectureabstractCritical Real-Time Embedded Systems (CRTES) industry needs increasingly complex hardware to attain the performance/cost ratio required to keep competitive edge in the market. Worst-case execution time (WCET) analysis is central to CRTES development. Whereas current timing analysis techniques are sound, their viability is hampered by the soaring cost of acquiring detailed knowledge of the internal operation and state of the system, at both software and hardware level. This is a major hurdle to using them for increasingly complex hardware platforms. Measurement-Based Probabilistic Timing Analysis (PTA) reduces the cost of acquiring the knowledge needed for computing trustworthy WCET bounds. This paper presents the changes required to hardware design to facilitate the use of the PTA techniques. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Ian Broster, Francisco J. Cazorla |
DSD | 3 |
| 2014 | On the Comparison of Deterministic and Probabilistic WCET Estimation TechniquesabstractTiming validation is a critical step in the design of real-time systems, that requires the estimation of Worst-Case Execution Times (WCET) for tasks. A number of different methods have been proposed, such as Static Deterministic Timing Analysis (SDTA). The advent of Probabilistic Timing Analysis, both Measurement-Based (MBPTA) and Static Probabilistic Timing Analyses (SPTA), offers different design points between the tightness of WCET estimates, hardware that can be analyzed and the information needed from the user to carry out the analysis. The lack of comparison among those techniques makes complex the selection of the most appropriate one for a given system. This paper makes a first attempt towards comparing comprehensively SDTA, SPTA and MBPTA, qualitatively and quantitatively, under different cache configurations implementing LRU and random replacement. We identify strengths and limitations of each technique depending on the characteristics of the program under analysis and the hardware platform, thus providing users with guidance on which approach to choose depending on their target application and hardware platform. Jaume Abella 0001, Damien Hardy, Isabelle Puaut, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 1 |
| 2014 | Heart of Gold: Making the Improbable Happen to Increase Confidence in MBPTAabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) has been recently proposed as a viable method to compute probabilistic worst-case execution time (pWCET) bounds for programs with hard real-time constraints. As a key trait, MBPTA needs a comparatively small number of observation runs, made on execution platforms to which MBPTA can be applied, to project the tail of the probability of occurrence of worst-case execution time durations of individual programs. In order for the use of MBPTA to fit the bill of industrial-quality development, it is imperative to understand what factors might threaten the trustworthiness of the pWCET computation. This paper addresses that important question by: (i) identifying the combined characteristics of applications and hardware resources that might lead to optimistic pWCET bounds, (ii) describing why this may occur, and (iii) providing the user with means to detect those cases so that trustworthiness is restored. In particular, we present a method for detecting risk scenarios for time-randomised caches, based on principles that apply to any other time-randomised resource which may challenge the application of MBPTA. Jaume Abella 0001, Eduardo Quiñones, Franck Wartel, Tullio Vardanega, Francisco J. Cazorla |
ECRTS | 1 |
| 2014 | PUB: Path Upper-Bounding for Measurement-Based Probabilistic Timing AnalysisabstractMeasurement-Based Probabilistic Timing Analysis (MBPTA) responds to the challenge of analysing the timing behaviour of real-time software running on hardware deploying high-performance features (e.g., data caches). MBPTA provides a WCET estimate that upper-bounds the execution time of the set of paths exercised with the data input vectors provided by the user. However, in several scenarios, the user is unaware of the input vector leading to the worst-case path. In this paper we present PUB, a new method that makes the WCET estimates obtained with MBPTA a trustworthy upper-bound of the probabilistic execution time of all paths in the program, even when the user-provided input vectors do not exercise the worst-case path. This significantly reduces the requirements imposed on the user to apply MBPTA. For Malardarlen and EEMBC respectively, PUB provides WCET estimates 5% and 11% higher than the WCET estimates computed with MBPTA. Leonidas Kosmidis, Jaume Abella 0001, Franck Wartel, Eduardo Quiñones, Antoine Colin, Francisco J. Cazorla |
ECRTS | 2 |
| 2014 | Parallel many-core avionics systemsabstractIntegrated Modular Avionics (IMA) enables incremental qualification by encapsulating avionics applications into software partitions (SWPs), as defined by the ARINC 653 standard. SWPs, when running on top of single-core processors, provide robust time partitioning as a means to isolate SWPs timing behavior from each other. However, when moving towards parallel execution in many-core processors, the simultaneous accesses to shared hardware and software resources influence the timing behavior of SWPs, defying the purpose of time partitioning to provide isolation among applications. In this paper, we extend the concept of SWP by introducing parallel software partitions (pSWP) specification that describes the behavior of SWPs required when running in a many-core to enable incremental qualification. pSWP are supported by a new hardware feature called guaranteed resource partition (GRP) that defines an execution environment in which SWPs run and that controls interferences in the accesses to shared hardware resources among SWPs such that time composability can be guaranteed. Milos Panic, Eduardo Quiñones, Pavel G. Zaykov, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla |
EMSOFT | 5 |
| 2014 | DReAM: Per-Task DRAM Energy Metering in Multicore Systems
Qixiao Liu, Miquel Moretó, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
Euro-Par | 3 |
| 2014 | AHRB: A high-performance time-composable AMBA AHB busabstractHard real-time systems are moving toward complex systems comprising chips with different IP components connected with standard buses. AMBA is one of the most used bus interfaces and has already been included in processors in the real-time domain. However, AMBA was not designed to provide time composable Worst Case Execution Time (WCET) estimates, which are desirable to reduce timing validation and verification costs. This paper analyzes and extends the AMBA Advanced High-performance Bus (AHB) specification to enable time-composable WCET estimates by design. Concretely, (1) we analyze in detail the AMBA AHB in the context of hard real-time systems proving that it fails to provide time composability; (2) we define a restricted subset of AMBA AHB features, named restricted AHB (resAHB), that allows deriving time-composable, yet not tight, WCET estimates; and (3) we define an extension of resAHB, named Advanced High-performance Real-time Bus (AHRB), that includes the timing constraints in the specification. This allows deriving time-composable and tight WCET estimates. Our results show that AHRB can provide 3.5x tighter estimates than resAHB on average for EEMBC benchmarks. Javier Jalle, Jaume Abella 0001, Eduardo Quiñones, Luca Fossati, Marco Zulianello, Francisco J. Cazorla |
RTAS | 2 |
| 2014 | A Dual-Criticality Memory Controller (DCmc): Proposal and Evaluation of a Space Case StudyabstractMulticore Dual-Criticality systems comprise two types of applications, each with a different criticality level. In the space domain these types are referred as payload and control applications, which have high-performance and real time requirements respectively. In order to control the interaction (contention) among payload and control applications in the access to the main memory, reaching the goals of high bandwidth for the former and guaranteed timing bounds for the latter, we propose a Dual-Criticality memory controller (DCmc). DCmc virtually divides memory banks into real-time and high-performance banks, deploying a different request scheduler policy to each bank type, which facilitates achieving both goals. Our evaluation with a multicore cycle-accurate simulator and a real space case study shows that DCmc enables deriving tight WCET estimates, regardless of the co-running payload applications, hence effectively isolating the effect of contention in the access to memory. DCmc also enables payload applications exploiting memory locality, which is needed for high performance. Javier Jalle, Eduardo Quiñones, Jaume Abella 0001, Luca Fossati, Marco Zulianello, Francisco J. Cazorla |
RTSS | 3 |
| 2014 | Efficient Cache Designs for Probabilistically Analysable Real-Time SystemsabstractThe increasing performance demand in the critical real-time embedded systems (CRTES) domain calls for high-performance features such as cache memories. Unfortunately, the cost to provide trustworthy and tight Worst-Case Execution Time (WCET) estimates in the presence of caches is high with current practice WCET analysis tools, because they need detailed knowledge of program’s cache accesses to provide tight WCET estimates. The advent of Probabilistic timing analysis (PTA) opens the door to economically viable timing analysis in the presence of caches, but it imposes new requirements on hardware design. At cache level, so far only fully associative random-replacement caches have been proven to fulfill the needs of PTA, but their energy, delay, and area cost are unaffordable for CRTES. In this paper, we propose the first PTA-compliant cache design based on set-associative and direct-mapped arrangements, as those are the most common arrangements. In particular, we propose a novel parametric random placement policy suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
IEEE Trans. Computers | 2 |
| 2014 | Hybrid Cache Designs for Reliable Hybrid High and Ultra-Low Voltage OperationabstractGeometry scaling of semiconductor devices enables the design of ultra-low-cost (e.g., below 1 USD) battery-powered resource-constrained ubiquitous devices for environment, urban life, and body monitoring. These sensor-based devices require high performance to react in front of infrequent particular events as well as extreme energy efficiency in order to extend battery lifetime during most of the time when low performance is required. In addition, they require real-time guarantees. The most suitable technological solution for these devices consists of using hybrid processors able to operate at: (i) high voltage to provide high performance and (ii) near-/subthreshold voltage to provide ultra-low energy consumption. However, the most efficient SRAM memories for each voltage level differ and trading off different SRAM designs is mandatory. This is particularly true for cache memories, which occupy most of the processor's area. In this article, we propose new, simple, single-Vcc-domain hybrid L1 cache architectures suitable for reliable hybrid high and ultra-low voltage operation. In particular, the cache is designed by combining heterogeneous SRAM cell types: some of the cache ways are optimized to satisfy high-performance requirements during high voltage operation, whereas the rest of the ways provide ultra-low energy consumption and reliability during near-/subthreshold voltage operation. We analyze the performance, energy, and power impact of the proposed cache designs when using them to implement L1 caches in a processor. Experimental results show that our hybrid caches can efficiently and reliably operate across a wide range of voltages, consuming little energy at near-/subthreshold voltage as well as providing high performance at high voltage without decreasing reliability levels to provide strong performance guarantees, as required for our target market. Bojan Maric, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Analyzing the Efficiency of L1 Caches for Reliable Hybrid-Voltage Operation Using EDC CodesabstractThe increasing demand for highly miniaturized battery-powered ultralow cost systems (e.g., below 1 dollar) in emerging applications such as body, urban life and environment monitoring, and so on, has introduced many challenges in chip design. Such applications require high performance occasionally and very little energy consumption during most of the time to extend battery lifetime. In addition, they require real-time guarantees. Caches have been shown to be the most critical blocks in these systems due to their high energy/area consumption and hard-to-predict behavior. New, simple, hybrid-voltage operation (high Vccand ultralow Vcc), single-Vccdomain L1 cache architectures based on replacing energy-hungry bitcells (e.g., 10T) by more energy-efficient and smaller cells (e.g., 8T) enhanced with error detection and correction codes have been recently proposed. Such designs provide significant energy and area efficiency without jeopardizing reliability levels to still provide strong performance guarantees. In this brief, we analyze the efficiency of these designs during ultralow voltage operation. We identify the limits of such approaches by finding an energy-optimal voltage region through experimental models. The experimental results show that area efficiency is always achieved in the range 200-400 mV, whereas both energy and area gains occur above 250 mV, i.e., in near-threshold regime. Bojan Maric, Jaume Abella 0001, Mateo Valero |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | On the convergence of mainstream and mission-critical marketsabstractThe computing market has been dominated during the last two decades by the well-known convergence of the high-performance computing market and the mobile market. In this paper we witness a new type of convergence between the mission-critical market (such as avionic or automotive) and the mainstream consumer electronics market. Such convergence is fuelled by the common needs of both markets for more reliability, support for mission-critical functionalities and the challenge of harnessing the unsustainable increases in safety margins to guarantee either correctness or timing. In this position paper, we present a description of this new convergence, as well as the main challenges and opportunities that it brings to computing industry. Sylvain Girbal, Miquel Moretó, Arnaud Grasset, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Sami Yehia |
DAC | 4 |
| 2013 | APPLE: adaptive performance-predictable low-energy caches for reliable hybrid voltage operationabstractSemiconductor technology evolution enables the design of resource-constrained battery-powered ultra-low-cost chips required for new market segments such as environment, urban life and body monitoring. Caches have been shown to be the main energy and area consumer in those chips. Bojan Maric, Jaume Abella 0001, Mateo Valero |
DAC | 2 |
| 2013 | A cache design for probabilistically analysable real-time systemsabstractCaches provide significant performance improvements, though their use in real-time industry is low because current WCET analysis tools require detailed knowledge of program's cache accesses to provide tight WCET estimates. Probabilistic Timing Analysis (PTA) has emerged as a solution to reduce the amount of information needed to provide tight WCET estimates, although it imposes new requirements on hardware design. At cache level, so far only fully-associative random-replacement caches have been proven to fulfill the needs of PTA, but they are expensive in size and energy. In this paper we propose a cache design that allows set-associative and direct-mapped caches to be analysed with PTA techniques. In particular we propose a novel parametric random placement suitable for PTA that is proven to have low hardware complexity and energy consumption while providing comparable performance to that of conventional modulo placement. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
DATE | 2 |
| 2013 | Probabilistic timing analysis on conventional cache designsabstractProbabilistic timing analysis (PTA), a promising alternative to traditional worst-case execution time (WCET) analyses, enables pairing time bounds (named probabilistic WCET or pWCET) with an exceedance probability (e.g., 10−16), resulting in far tighter bounds than conventional analyses. However, the applicability of PTA has been limited because of its dependence on relatively exotic hardware: fully-associative caches using random replacement. This paper extends the applicability of PTA to conventional cache designs via a software-only approach. We show that, by using a combination of compiler techniques and runtime system support to randomise the memory layout of both code and data, conventional caches behave as fully-associative ones with random replacement. Leonidas Kosmidis, Charlie Curtsinger, Eduardo Quiñones, Jaume Abella 0001, Emery D. Berger, Francisco J. Cazorla |
DATE | 4 |
| 2013 | Efficient cache architectures for reliable hybrid voltage operation using EDC codesabstractSemiconductor technology evolution enables the design of sensor-based battery-powered ultra-low-cost chips (e.g., below 1 €) required for new market segments such as body, urban life and environment monitoring. Caches have been shown to be the highest energy and area consumer in those chips. This paper proposes a novel, hybrid-operation (high Vcc, ultra-low Vcc), single-Vccdomain cache architecture based on replacing energy-hungry bitcells (e.g., 10T) by more energy-efficient and smaller cells (e.g., 8T) enhanced with Error Detection and Correction (EDC) features for high reliability and performance predictability. Our architecture is proven to largely outperform existing solutions in terms of energy and area. Bojan Maric, Jaume Abella 0001, Mateo Valero |
DATE | 2 |
| 2013 | parMERASA - Multi-core Execution of Parallelised Hard Real-Time Applications Supporting AnalysabilityabstractEngineers who design hard real-time embedded systems express a need for several times the performance available today while keeping safety as major criterion. A breakthrough in performance is expected by parallelizing hard real-time applications and running them on an embedded multi-core processor, which enables combining the requirements for high-performance with timing-predictable execution. parMERASA will provide a timing analyzable system of parallel hard real-time applications running on a scalable multicore processor. parMERASA goes one step beyond mixed criticality demands: It targets future complex control algorithms by parallelizing hard real-time programs to run on predictable multi-/many-core processors. We aim to achieve a breakthrough in techniques for parallelization of industrial hard real-time programs, provide hard real-time support in system software, WCET analysis and verification tools for multi-cores, and techniques for predictable multi-core designs with up to 64 cores. Theo Ungerer, Christian Bradatsch, Mike Gerdes 0001, Florian Kluge, Ralf Jahr, Jörg Mische, Pavel G. Zaykov, Zlatko Petrov, Bert Böddeker, Sebastian Kehr, Hans Regler, Andreas Hugl, Christine Rochange, Haluk Ozaktas, Hugues Cassé, Armelle Bonenfant, Pascal Sainrat, Ian Broster, Nick Lay, Eduardo Quiñones, Milos Panic, Jaume Abella 0001, Francisco J. Cazorla, Sascha Uhrig, Mathias Rohde, Arthur Pyka |
DSD | 24 |
| 2013 | DTM: Degraded Test Mode for Fault-Aware Probabilistic Timing AnalysisabstractExisting timing analysis techniques to derive Worst-Case Execution Time (WCET) estimates assume that hardware in the target platform (e.g., the CPU) is fault-free. Given the performance requirements increase in current Critical Real-Time Embedded Systems (CRTES), the use of high-performance features and smaller transistors in current and future hardware becomes a must. The use of smaller transistors helps providing more performance while maintaining low energy budgets, however, hardware fault rates increase noticeably, affecting the temporal behaviour of the system in general, and WCET in particular. In this paper, we reconcile these two emergent needs of CRTES, namely, tight (and trustworthy) WCET estimates and the use of hardware implemented with smaller transistors. To that end we propose the Degraded Test Mode (DTM) that, in combination with fault-tolerant hardware designs and probabilistic timing analysis techniques, (i) enables the computation of tight and trustworthy WCET estimates in the presence of faults, (ii) provides graceful average and worst-case performance degradation due to faults, and (iii) requires modifications neither in WCET analysis tools nor in applications. Our results show that DTM allows accounting for the effect of faults at analysis time with low impact in WCET estimates and negligible hardware modifications. Mladen Slijepcevic, Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 3 |
| 2013 | Supporting industrial use of probabilistic timing analysis with explicit argumentationabstractProbabilistic Timing Analysis (PTA) in general and its measurement-based variant called MBPTA in particular have potential for mitigating the problems that impair current worstcase execution time (WCET) analysis techniques whether as in industrial practice or in state-of-the-art research. MBPTA can compute tight upper bounds on the execution time of software programs, which it expresses as probabilistic exceedance functions, without needing much information on the hardware and software internals of the system. To exploit this capability in practice, some reasoned argument must be constructed to explain why the method is suitable. This paper details our experience with the construction of such an argument, and in particular shows how the structure of the argument allows it to be easily configured for the needs of different industries. Zoë Stephenson, Jaume Abella 0001, Tullio Vardanega |
INDIN | 2 |
| 2013 | Achieving timing composability with measurement-based probabilistic timing analysisabstractProbabilistic Timing Analysis (PTA) allows complex hardware acceleration features, which defeat classic timing analysis, to be used in hard real-time systems. PTA can do that because it drastically reduces intrinsic dependence on execution history. This distinctive feature is a great facilitator to time composability, which is a must for industry needing incremental development and qualification. In this paper we show how time composability is achieved in PTA-conformant systems and how the pessimism of worst-case execution time bounds obtained from PTA is contained within a 5% to 25% range for representative application scenarios. Leonidas Kosmidis, Eduardo Quiñones, Jaume Abella 0001, Tullio Vardanega, Francisco J. Cazorla |
ISORC | 3 |
| 2013 | Implicit-storing and redundant-encoding-of-attribute information in error-correction-codesabstractThis paper proposes implicit-storing to extend the logical capacity of a memory array without increasing its physical capacity by leveraging the array's error-correction-codes to infer the implicitly stored bits. Implicit-storing is related to error-code-tagging, a technique that distinguishes between faults in data and invariant attributes of a location when the attributes are not stored in the memory array but are encoded in the error-correction-codes. Both error-code-tagging and implicit-storing cause a code-strength reduction due to their encoding of additional information in the code meant to only protect data. Yiannakis Sazeides, Emre Ozer 0001, Danny Kershaw, Panagiota Nikolaou, Marios Kleanthous, Jaume Abella 0001 |
MICRO | 6 |
| 2013 | Multi-level Unified Caches for Probabilistically Time Analysable Real-Time SystemsabstractAbstract—Caches are key resources in high-end processor architectures to increase performance. In fact, most high-performance processors come equipped with a multi-level cache hierarchy. In terms of guaranteed performance, however, cache hierarchies severely challenge the computation of tight worst-case execution time (WCET) estimates. On the one hand, the analysis of the timing behaviour of a single level of cache is already challenging, particularly for data accesses. On the other hand, unifying data and instructions in each level, makes the problem of cache analysis significantly more complex requiring tracking simultaneously data and instruction accesses to cache. In this paper we prove that multi-level cache hierarchies can be used in the context of Probabilistic Timing Analysis and tight WCET estimates can be obtained. Our detailed analysis (1) covers unified data and instruction caches, (2) covers different cache-write policies (write-through and write back), write allocation policies (write-allocate and non-write-allocate) and several inclusion mechanisms (inclusive, non-inclusive and exclusive caches), and (3) scales to an arbitrary number of cache levels. Our results show that the probabilistic WCET (pWCET) estimates provided by our analysis technique effectively benefit from having multi-level caches. For a two-level cache configura-tion and for EEMBC benchmarks, pWCET reductions are 55% on average (and up to 90%) with respect to a processor with a single level of cache. I. Leonidas Kosmidis, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla |
RTSS | 2 |
| 2013 | Hardware support for accurate per-task energy metering in multicore systemsabstractAccurately determining the energy consumed by each task in a system will become of prominent importance in future multicore-based systems because it offers several benefits, including (i) better application energy/performance optimizations, (ii) improved energy-aware task scheduling, and (iii) energy-aware billing in data centers. Unfortunately, existing methods for energy metering in multicores fail to provide accurate energy estimates for each task when several tasks run simultaneously. This article makes a case for accurate Per-Task Energy Metering (PTEM) based on tracking the resource utilization and occupancy of each task. Different hardware implementations with different trade-offs between energy prediction accuracy and hardware-implementation complexity are proposed. Our evaluation shows that the energy consumed in a multicore by each task can be accurately measured. For a 32-core, 2-way, simultaneous multithreaded core setup, PTEM reduces the average accuracy error from more than 12% when our hardware support is not used to less than 4% when it is used. The maximum observed error for any task in the workload we used reduces from 58% down to 9% when our hardware support is used. Qixiao Liu, Miquel Moretó, Víctor Jiménez, Jaume Abella 0001, Francisco J. Cazorla, Mateo Valero |
ACM Trans. Archit. Code Optim. | 4 |
| 2013 | PROARTIS: Probabilistically Analyzable Real-Time SystemsabstractStatic timing analysis is the state-of-the-art practice of ascertaining the timing behavior of current-generation real-time embedded systems. The adoption of more complex hardware to respond to the increasing demand for computing power in next-generation systems exacerbates some of the limitations of static timing analysis. In particular, the effort of acquiring (1) detailed information on the hardware to develop an accurate model of its execution latency as well as (2) knowledge of the timing behavior of the program in the presence of varying hardware conditions, such as those dependent on the history of previously executed instructions. We call these problems the timing analysis walls. In this vision-statement article, we present probabilistic timing analysis , a novel approach to the analysis of the timing behavior of next-generation real-time embedded systems. We show how probabilistic timing analysis attacks the timing analysis walls; we then illustrate the mathematical foundations on which this method is based and the challenges we face in the effort of efficiently implementing it. We also present experimental evidence that shows how probabilistic timing analysis reduces the extent of knowledge about the execution platform required to produce probabilistically accurate WCET estimations. Francisco J. Cazorla, Eduardo Quiñones, Tullio Vardanega, Liliana Cucu-Grosjean, Benoit Triquet, Guillem Bernat, Emery D. Berger, Jaume Abella 0001, Franck Wartel, Michael Houston, Luca Santinelli, Leonidas Kosmidis, Code Lo, Dorin Maxim |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2012 | Measurement-Based Probabilistic Timing Analysis for Multi-path ProgramsabstractThe rigorous application of static timing analysis requires a large and costly amount of detail knowledge on the hardware and software components of the system. Probabilistic Timing Analysis has potential for reducing the weight of that demand. In this paper, we present a sound measurement-based probabilistic timing analysis technique based on Extreme Value Theory. In all the experiments made as part of this work, the timing bounds determined by our technique were less than 15% pessimistic in comparison with the tightest possible bounds obtainable with any probabilistic timing analysis technique. As a point of interest to industrial users, our technique also requires a comparatively low number of measurement runs of the program under analysis, less than 650 runs were needed for the benchmarks presented in this paper. Liliana Cucu-Grosjean, Luca Santinelli, Michael Houston, Code Lo, Tullio Vardanega, Leonidas Kosmidis, Jaume Abella 0001, Enrico Mezzetti, Eduardo Quiñones, Francisco J. Cazorla |
ECRTS | 7 |
| 2012 | ADAM: an efficient data management mechanism for hybrid high and ultra-low voltage operation cachesabstractSemiconductor technology evolution enables the design of ultra-low-cost chips (e.g., below 1 USD) required for new market segments such as environment, urban life and body monitoring, etc. Recently, hybrid-operation (high Vcc, ultra-low Vcc) single-Vcc-domain cache designs have been proposed to tackle the needs of those chips. However, existing data management policies are far from being optimal during high Vcc operation. Bojan Maric, Jaume Abella 0001, Mateo Valero |
ACM Great Lakes Symposium on VLSI | 2 |
| 2011 | RVC: a mechanism for time-analyzable real-time processors with faulty cachesabstractGeometry scaling due to technology evolution as well as Vcc scaling lead to failures in large SRAM arrays such as caches. Faulty bits can be tolerated from the average performance perspective, but make critical realtime embedded systems non time-analyzable or worstcase execution time (WCET) estimations unacceptably large. Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Yiannakis Sazeides, Mateo Valero |
HiPEAC | 1 |
| 2011 | Hardware/software-based diagnosis of load-store queues using expandable activity logsabstractThe increasing device count and design complexity are posing significant challenges to post-silicon validation. Bug diagnosis is the most difficult step during post-silicon validation. Limited reproducibility and low testing speeds are common limitations in current testing techniques. Moreover, low observability defies full-speed testing approaches. Modern solutions like on-chip trace buffers alleviate these issues, but are unable to store long activity traces. As a consequence, the cost of post-Si validation now represents a large fraction of the total design cost. This work describes a hybrid post-Si approach to validate a modern load-store queue. We use an effective error detection mechanism and an expandable logging mechanism to observe the microarchitectural activity for long periods of time, at processor full-speed. Validation is performed by analyzing the log activity by means of a diagnosis algorithm. Correct memory ordering is checked to root the cause of errors. Javier Carretero, Xavier Vera, Jaume Abella 0001, Tanausú Ramírez, Matteo Monchiero, Antonio González 0001 |
HPCA | 3 |
| 2011 | Towards improved survivability in safety-critical systemsabstractPerformance demand of Critical Real-Time Embedded (CRTE) systems implementing safety-related system features grows at an exponential rate. Only modern semiconductor technologies can satisfy CRTE systems performance needs efficiently. However, those technologies lead to high failure rates, thus lowering survivability of chips to unacceptable levels for CRTE systems. This paper presents SESACS architecture (Surviving Errors in SAfety-Critical Systems), a paradigm shift in the design of CRTE systems. SESACS is a new system design methodology consisting of three main components: (i) a multicore hardware/firmware platform capable of detecting and diagnosing hardware faults of any type with minimal impact on the worst-case execution time (WCET), recovering quickly from errors, and properly reconfiguring the system so that the resulting system exhibits a predictable and analyzable degradation in WCET; (ii) a set of analysis methods and tools to prove the timing correctness of the reconfigured system; and (iii) a white-box methodology and tools to prove the functional safety of the system and compliance with industry standards. This new design paradigm will deliver huge benefits to the embedded systems industry for several decades by enabling the use of more cost-effective multicore hardware platforms built on top of modern semiconductor technologies, thereby enabling higher performance, and reducing weight and power dissipation. This new paradigm will further extend the life of embedded systems, therefore, reducing warranty and early replacement costs. Jaume Abella 0001, Francisco J. Cazorla, Eduardo Quiñones, Arnaud Grasset, Sami Yehia, Philippe Bonnot 0001, Dimitris Gizopoulos, Riccardo Mariani, Guillem Bernat |
IOLTS | 1 |
| 2011 | RVC-based time-predictable faulty caches for safety-critical systemsabstractTechnology and Vcc scaling lead to significant faulty bit rates in caches. Mechanisms based on disabling faulty parts show to be effective for average performance but are unacceptable in safety critical systems where worst-case execution time (WCET) estimations must be safe and tight. The Reliable Victim Cache (RVC) deals with this issue for a large fraction of the cache bits. However, replacement bits are not protected, thus keeping the probability of failure still high. This paper proposes two mechanisms to tolerate faulty bits in replacement bits and keep time-predictability by extending the RVC. Our solutions offer different tradeoffs between cost and complexity. In particular, the Extended RVC (ERVC) has low energy and area overheads while keeping complexity at a minimum. The Reliable Replacement Bits (RRB) solution has even lower overheads at the expense of some more wiring complexity. Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla, Mateo Valero, Yiannakis Sazeides |
IOLTS | 1 |
| 2011 | Implementing End-to-End Register Data-Flow Continuous Self-TestabstractWhile Moore's Law predicts the ability of semiconductor industry to engineer smaller and more efficient transistors and circuits, there are serious issues not contemplated in that law. One concern is the verification effort of modern computing systems, which has grown to dominate the cost of system design. On the other hand, technology scaling leads to burn-in phase out. As a result, in-the-field error rate may increase due to both actual errors and latent defects. Whereas data can be protected with arithmetic codes, there is a lack of cost-effective mechanisms for control logic. This paper presents a light-weight microarchitectural mechanism that ensures that data consumed through registers are correct. The structures protected include the issue queue logic and the data associated (i.e., tags and control signals), input multiplexors, rename data, replay logic, register free-list and release logic, and register file logic. Our results show a coverage around 90 percent for the targeted structures with a cost in power and area of about four percent, and without impact in performance. Javier Carretero, Pedro Chaparro, Xavier Vera, Jaume Abella 0001, Antonio González 0001 |
IEEE Trans. Computers | 4 |
| 2010 | The split register fileabstractTechnology scaling requires lowering Vcc due to power constraints. Unfortunately, permanent faulty bit rates grow due to the higher impact of process variations at low Vcc, especially in the register file whose critical timing limits circuit optimizations. This paper proposes a novel register file design based on splitting registers and discarding faulty blocks to increase the number of registers available. By increasing the number of registers available higher performance can be obtained and yield increases because a larger number of processors reaches the minimum number of registers required to operate. Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Xavier Vera |
DATE | 1 |
| 2010 | High-Performance low-vcc in-order coreabstractPower density grows in new technology nodes, thus requiring Vcc to scale especially in mobile platforms where energy is critical. This paper presents a novel approach to decrease Vcc while keeping operating frequency high. Our mechanism is referred to as immediate read after write (IRAW) avoidance. We propose an implementation of the mechanism for an Intel®SilverthorneTMin-order core. Furthermore, we show that our mechanism can be adapted dynamically to provide the highest performance and lowest energy-delay product (EDP) at each Vcc level. Results show that IRAW avoidance increases operating frequency by 57% at 500mV and 99% at 400mV with negligible area and power overhead (below 1%), which translates into large speedups (48% at 500mV and 90% at 400mV) and EDP reductions (0.61 EDP at 500mV and 0.33 at 400mV). Jaume Abella 0001, Pedro Chaparro, Xavier Vera, Javier Carretero, Antonio González 0001 |
HPCA | 1 |
| 2010 | VCTA: A Via-Configurable Transistor Array regular fabricabstractLayout regularity is introduced progressively by integrated circuit manufacturers to reduce the increasing systematic process variations in the deep sub-micron era. In this paper we focus on a scenario where layout regularity must be pushed to the limit to deal with severe systematic process variations in future technology nodes. With this objective, we propose and evaluate a new regular layout style called Via-Configurable Transistor Array (VCTA) that maximizes regularity at device and interconnect levels. In order to assess VCTA maximum layout regularity tradeoffs, we implement 32-bit adders in the 90 nm technology node for VCTA and compare them with implementations that make use of standard cells. For this purpose we study the impact of photolithography proximity and coma effects on channel length variations, and the impact of shallow trench isolation mechanical stress on threshold voltage variations. We demonstrate that both variations, that are important sources of energy and delay circuit variability, are minimized through VCTA regularity. Marc Pons 0001, Francesc Moll, Antonio Rubio 0001, Jaume Abella 0001, Xavier Vera, Antonio González 0001 |
VLSI-SoC | 4 |
| 2010 | Microarchitectural Online Testing for Failure Detection in Memory Order BuffersabstractTechnology scaling leads to burn-in phase out and higher postsilicon test complexity, which increases in-the-field failure rate due to both latent defects and actual errors, respectively. As a consequence, current reliability qualification methods will likely be infeasible. Microarchitecture knowledge of application runtime behavior offers a possibility to have low-cost continuous online testing techniques detect hard errors in the field. Whereas data can be protected with redundancy (like parity or ECC), there is a lack of mechanism for control logic. This paper proposes a microarchitectural approach for validating that the memory order buffer logic works correctly. Our design relies on a small cache-like structure that keeps track of the last store to each cached address. Each load is checked to have obtained the data from the youngest older producing store. We present three different implementations of this idea, offering different trade-offs for error coverage, performance overhead, and design complexity. Javier Carretero, Xavier Vera, Pedro Chaparro, Jaume Abella 0001 |
IEEE Trans. Computers | 4 |
| 2009 | Online error detection and correction of erratic bits in register filesabstractAggressive voltage scaling needed for low power in each new process generation causes large deviations in the threshold voltage of minimally sized devices of the 6T SRAM cell. Gate oxide scaling can cause large transient gate leakage (a trap in the gate oxide), which is known as the erratic bits phenomena. Register file protection is necessary to prevent errors from quickly spreading to different parts of the system, which may cause applications to crash or silent data corruption. This paper proposes a simple and cost-effective mechanism that increases the resiliency of the register files to erratic bits. Our mechanism detects those registers that have erratic bits, recovers from the error and quarantines the faulty register. After the quarantine period, it is able to detect whether they are fully operational with low overhead. Xavier Vera, Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Antonio González 0001 |
IOLTS | 2 |
| 2009 | End-to-end register data-flow continuous self-testabstractWhile Moore's Law predicts the ability of semi-conductor industry to engineer smaller and more efficient transistors and circuits, there are serious issues not contemplated in that law. One concern is the verification effort of modern computing systems, which has grown to dominate the cost of system design. On the other hand, technology scaling leads to burn-in phase out. As a result, in-the-field error rate may increase due to both actual errors and latent defects. Whereas data can be protected with arithmetic codes (like parity or ECC), there is a lack of cost-effective mechanisms for control logic. Javier Carretero, Pedro Chaparro, Xavier Vera, Jaume Abella 0001, Antonio González 0001 |
ISCA | 4 |
| 2009 | Low Vccmin fault-tolerant cache with highly predictable performanceabstractTransistors per area unit double in every new technology node. However, the electric field density and power demand grow if Vcc is not scaled. Therefore, Vcc must be scaled in pace with new technology nodes to prevent excessive degradation and keep power demand within reasonable limits. Unfortunately, low Vcc operation exacerbates the effect of variations and decreases noise and stability margins, increasing the likelihood of errors in SRAM memories such as caches. Those errors translate into performance loss and performance variation across different cores, which is especially undesirable in a multi-core processor. Jaume Abella 0001, Javier Carretero, Pedro Chaparro, Xavier Vera, Antonio González 0001 |
MICRO | 1 |
| 2009 | Exploring the limits of early register release: Exploiting compiler analysisabstractRegister pressure in modern superscalar processors can be reduced by releasing registers early and by copying their contents to cheap back-up storage. This article quantifies the potential benefits of register occupancy reduction and shows that existing hardware-based schemes typically achieve only a small fraction of this potential. This is because they are unable to accurately determine the last use of a register and must wait until the redefining instruction enters the pipeline. On the other hand, compilers have a global view of the program and, using simple dataflow analysis, can determine the last use. This article evaluates the extent to which compiler analysis can aid early releasing, explores the design space, and introduces commit and issue-based early releasing schemes, quantifying their benefits. Using simple compiler analysis and microarchitecture changes, we achieve 70% of the potential register file occupancy reduction. By adding more hardware support, we can increase this to 94%. Our schemes are compared to state-of-the-art approaches for varying register file sizes and are shown to outperform these existing techniques. Timothy M. Jones 0001, Michael F. P. O'Boyle, Jaume Abella 0001, Antonio González 0001, Oguz Ergin |
ACM Trans. Archit. Code Optim. | 3 |
| 2009 | Energy-efficient register caching with compiler assistanceabstractThe register file is a critical component in a modern superscalar processor. It must be large enough to accommodate the results of all in-flight instructions. It must also have enough ports to allow simultaneous issue and writeback of many values each cycle. However, this makes it one of the most energy-consuming structures within the processor with a high access latency. As technology scales, there comes a point where register accesses are the bottleneck to performance and so must be pipelined over several cycles. This increases the pipeline depth, lowering performance. To overcome these challenges, we propose a novel use of compiler analysis to aid register caching. Adding a register cache allows us to preserve single-cycle register accesses, maintaining performance and reducing energy consumption. We do this by passing information to the processor using free bits in a real ISA, allowing us to cache only the most important registers. Evaluating the register cache over a variety of sizes and associativities and varying the read ports into the cache, our best scheme achieves an energy-delay-squared (EDD) product of 0.81, with a performance increase of 11%. Another configuration saves 13% of register system energy. Using four register cache read ports brings both performance gains and energy savings, consistently outperforming two state-of-the-art hardware approaches. Timothy M. Jones 0001, Michael F. P. O'Boyle, Jaume Abella 0001, Antonio González 0001, Oguz Ergin |
ACM Trans. Archit. Code Optim. | 3 |
| 2009 | Selective replication: A lightweight technique for soft errorsabstractSoft errors are an important challenge in contemporary microprocessors. Modern processors have caches and large memory arrays protected by parity or error detection and correction codes. However, today's failure rate is dominated by flip flops, latches, and the increasing sensitivity of combinational logic to particle strikes. Moreover, as Chip Multi-Processors (CMPs) become ubiquitous, meeting the FIT budget for new designs is becoming a major challenge. Solutions based on replicating threads have been explored deeply; however, their high cost in performance and energy make them unsuitable for current designs. Moreover, our studies based on a typical configuration for a modern processor show that focusing on the top 5 most vulnerable structures can provide up to 70% reduction in FIT rate. Therefore, full replication may overprotect the chip by reducing the FIT much below budget. We propose Selective Replication , a lightweight-reconfigurable mechanism that achieves a high FIT reduction by protecting the most vulnerable instructions with minimal performance and energy impact. Low performance degradation is achieved by not requiring additional issue slots and reissuing instructions only during the time window between when they are retirable and they actually retire. Coverage can be reconfigured online by replicating only a subset of the instructions (the most vulnerable ones). Instructions' vulnerability is estimated based on the area they occupy and the time they spend in the issue queue. By changing the vulnerability threshold, we can adjust the trade-off between coverage and performance loss. Results for an out-of-order processor configured similarly to Intel® Core™ Micro-Architecture show that our scheme can achieve over 65% FIT reduction with less than 4% performance degradation with small area and complexity overhead. Xavier Vera, Jaume Abella 0001, Javier Carretero, Antonio González 0001 |
ACM Trans. Comput. Syst. | 2 |
| 2008 | Issue system protection mechanismsabstractMulti-core microprocessors require reducing the FIT (failures-in-time) rate per core drastically to enable a larger number of cores within a FIT budget. Since large arrays like caches and register flies are typically protected with either ECC or parity, the issue system becomes as one of the largest contributors to the core's FIT rate. Soft-errors are an important concern in contemporary microprocessors. Particle hits on the components of a processor are expected to create an increasing number of transient errors in each new microprocessor generation. In addition, the number of hard-errors in the field is expected to grow as burn-in becomes less effective. Moreover, the continuous device shrinking increases the likelihood of in-the-field failures due to rather small defects exacerbated by degradation. This paper proposes on-line mechanisms to detect and recover to a consistent state, classify and confine in-the-field errors in the issue system of both in-order and out-of-order cores. Such mechanisms provide high coverage at a small cost. Pedro Chaparro, Jaume Abella 0001, Javier Carretero, Xavier Vera |
ICCD | 2 |
| 2008 | On-Line Failure Detection and Confinement in CachesabstractTechnology scaling leads to burn-in phase out and increasing post-silicon test complexity, which increases in-the-field error rate due to both latent defects and actual errors. As a consequence, there is an increasing need for continuous on-line testing techniques to cope with hard errors in the field. Similarly, those techniques are needed for detecting soft errors in logic, whose error rate is expected to raise in future technologies. Cache memories, which occupy most of the area of the chip, are typically protected with parity or ECC, but most of the wires as well as some combinational blocks remain unprotected against both soft and hard errors. This paper presents a set of techniques to detect and confine hard and soft errors in cache memories in combination with parity/ECC at very low cost. By means of hard signatures in data rows and error tracking, faults can be detected, classified properly and confined for hardware reconfiguration. Jaume Abella 0001, Pedro Chaparro, Xavier Vera, Javier Carretero, Antonio González 0001 |
IOLTS | 1 |
| 2008 | On-line Failure Detection in Memory Order BuffersabstractTechnology scaling leads to burn-in phase out and higher post-silicon test complexity, which increases in-the-field error rate due to both latent defects and actual errors respectively. As a consequence, current reliability qualification methods will likely be infeasible. Microarchitecture knowledge of application runtime behavior offers a possibility to have low-cost continuous online testing techniques to cope with hard errors in the field. Whereas data can be protected with redundancy (like parity or ECC), there is a lack of mechanisms for control logic. This paper proposes a microarchitectural approach for validating that the memory order buffer logic works correctly. Javier Carretero, Xavier Vera, Pedro Chaparro, Jaume Abella 0001 |
ITC | 4 |
| 2007 | Fuse: A Technique to Anticipate Failures due to Degradation in ALUsabstractThis paper proposes the fuse, a technique to anticipate failures due to degradation in any ALU (arithmetic logic unit), and particularly in an adder. The fuse consists of a replica of the weakest transistor in the adder and the circuitry required to measure its degradation. By mimicking the behavior of the replicated transistor the fuse anticipates the failure short before the first failure in the adder appears, and hence, data corruption and program crashes can be avoided. Our results show that the fuse anticipates the failure in more than 99.9% of the cases after 96.6% of the lifetime, even for pessimistic random within-die variations. Jaume Abella 0001, Xavier Vera, Osman S. Unsal, Oguz Ergin, Antonio González 0001 |
IOLTS | 1 |
| 2007 | Surviving to Errors in Multi-Core EnvironmentsabstractIn this paper, the authors present a global view of the issues outlined above as well as some directions to address them. First, the most important sources of failure (SOF) are presented as well as their impact on CMOS technology. Then, techniques and key parameters to measure degradation due to different SOF are introduced and microarchitectural approaches to mitigate degradation are outlined. The problem of error detection and anticipation is illustrated as well as pros and cons of different types of mechanisms to perform such detection and anticipation. Finally, we illustrate the whole picture where performance and reliability must be traded carefully. We point out some directions to use the information about the detected errors and the amount of degradation of each component to configure the multi-core in such a way that performance is maximized without compromising reliability. Xavier Vera, Jaume Abella 0001 |
IOLTS | 2 |
| 2007 | Penelope: The NBTI-Aware ProcessorabstractTransistors consist of lower number of atoms with every technology generation. Such atoms may be displaced due to the stress caused by high temperature, frequency and current, leading to failures. NBTI (negative bias temperature instability) is one of the most important sources of failure affecting transistors. NBTI degrades PMOS transistors whenever the voltage at the gate is negative (logic input "0"). The main consequence is a reduction in the maximum operating frequency and an increase in the minimum supply voltage of storage structures to cope for the degradation. Many PMOS transistors affected by NBTI can be found in both combinational and storage blocks since they observe a "0 " at their gates most of the time. This paper proposes and evaluates the design of Penelope, an NBTI-aware processor. We propose (i) generic strategies to mitigate degradation in both combinational and storage blocks, (ii) specific techniques to protect individual blocks by applying the global strategies, and (Hi) a metric to assess the benefits of reduced degradation and the overheads in performance and power. Jaume Abella 0001, Xavier Vera, Antonio González 0001 |
MICRO | 1 |
| 2006 | Heterogeneous way-size cacheabstractSet-associative cache architectures are commonly used. These caches consist of a number of ways, each of the same size. We have observed that the different ways have very different utilization, which motivates the design of caches with heterogeneous way sizes. This can potentially result in higher performance for the same area, better capabilities to implement dynamically adaptive schemes, and more flexibility for choosing the size of the cache.This paper proposes a novel cache architecture, the Heterogeneous Way-Size cache (HWS cache), in which the different cache ways may have different sizes. HWS caches are shown to outperform conventional caches for L1 (data and instruction) and L2 caches. For instance, a HWS cache can achieve up to 20% dynamic and leakage energy savings with respect to its conventional cache counterpart, while the hit ratio is practically the same.We also present a Dynamically Adaptive version of the HWS cache (DAHWS cache). DAHWS caches are shown to be more adaptive than conventional architectures. Using state-of-the-art resizing schemes, we show that DAHWS caches achieve higher energy savings and lower miss rates than conventional caches when using the same resizing schemes, due to their higher flexibility. For a L1 instruction cache the active ratio is reduced 4% more (66% total reduction) than state-of-the-art techniques and the DAHWS cache hit ratio is higher. The active ratio is also reduced up to 55% and 41% for L1 data and L2 caches respectively. Jaume Abella 0001, Antonio González 0001 |
ICS | 1 |
| 2006 | SAMIE-LSQ: set-associative multiple-instruction entry load/store queueabstractThe load/store queue (LSQ) is one of the most complex parts of contemporary processors. Its latency is critical for the processor performance and it is usually one of the processor hotspots. This paper presents a highly banked, set-associative, multiple-instruction entry LSQ (SAMIE-LSQ,) that achieves high performance with small energy requirements. The SAMIE-LSQ classifies the memory instructions (loads and stores) based on the address to be accessed, and groups those instructions accessing the same cache line in the same entry. Our approach relies on the fact that many in-flight memory instructions access the same cache lines. Each SAMIE-LSQ entry has space for several memory instructions accessing the same cache line. This arrangement has a number of advantages. First, it significantly reduces the address comparison activity needed for memory disambiguation since there are less addresses to be compared. It also reduces the activity in the data TLB, the cache tag and cache data arrays. This is achieved by caching the cache line location and address translation in the corresponding SAMIE-LSQ entry once the access of one of the instructions in an entry is performed, so instructions that share an entry can reuse the translation, avoid the tag check and get the data directly from the concrete cache way without checking the others. Besides, the delay of the proposed scheme is lower than that required by a conventional LSQ. We show that the SAMIE-LSQ saves 82% dynamic energy for the load/store queue, 42% for the LI data cache and 73% for the data TLB, with a negligible impact on performance (0.6%) Jaume Abella 0001, Antonio González 0001 |
IPDPS | 1 |
| 2005 | Software Directed Issue Queue Power ReductionabstractThe issue logic of a superscalar processor dissipates a large amount of static and dynamic power. Furthermore, its power density makes it a hot-spot requiring expensive cooling systems and additional packaging. In this paper we present a novel software assisted approach to power reduction where the processor dynamically resizes the issue queue based on compiler analysis. The compiler passes information to the processor about the number of entries needed which limits the number of instructions dispatched and resident in the queue. This saves power without adversely affecting performance. Compared with recently proposed hardware techniques, our approach is faster, simpler and saves more power. Using a simplistic scheme we achieve 47% dynamic and 31% static power savings in the issue queue with only a 2.2% performance loss. We then show that the performance loss can be reduced to less than 1.3% with 45% dynamic and 30% static power savings, outperforming all current approaches. Timothy M. Jones 0001, Michael F. P. O'Boyle, Jaume Abella 0001, Antonio González 0001 |
HPCA | 3 |
| 2005 | IATAC: a smart predictor to turn-off L2 cache linesabstractAs technology evolves, power dissipation increases and cooling systems become more complex and expensive. There are two main sources of power dissipation in a processor: dynamic power and leakage. Dynamic power has been the most significant factor, but leakage will become increasingly significant in future. It is predicted that leakage will shortly be the most significant cost as it grows at about a 5× rate per generation. Thus, reducing leakage is essential for future processor design. Since large caches occupy most of the area, they are one of the leakiest structures in the chip and hence, a main source of energy consumption for future processors.This paper introduces IATAC (inter-access time per access count), a new hardware technique to reduce cache leakage for L2 caches. IATAC dynamically adapts the cache size to the program requirements turning off cache lines whose content is not likely to be reused. Our evaluation shows that this approach outperforms all previous state-of-the-art techniques. IATAC turns off 65% of the cache lines across different L2 cache configurations with a very small performance degradation of around 2%. Jaume Abella 0001, Antonio González 0001, Xavier Vera, Michael F. P. O'Boyle |
ACM Trans. Archit. Code Optim. | 1 |
| 2005 | An accurate cost model for guiding data locality transformationsabstractCaches have become increasingly important with the widening gap between main memory and processor speeds. Small and fast cache memories are designed to bridge this discrepancy. However, they are only effective when programs exhibit sufficient data locality.The performance of the memory hierarchy can be improved by means of data and loop transformations. Tiling is a loop transformation that aims at reducing capacity misses by shortening the reuse distance. Padding is a data layout transformation targeted to reduce conflict misses.This article presents an accurate cost model that describes misses across different hierarchy levels and considers the effects of other hardware components such as branch predictors. The cost model drives the application of tiling and padding transformations. We combine the cost model with a genetic algorithm to compute the tile and pad factors that enhance the program performance.To validate our strategy, we ran experiments for a set of benchmarks on a large set of modern architectures. Our results show that this scheme is useful to optimize programs' performance. When compared to previous approaches, we observe that with a reasonable compile-time overhead, our approach gives significant performance improvements for all studied kernels on all architectures. Xavier Vera, Jaume Abella 0001, Josep Llosa, Antonio González 0001 |
ACM Trans. Program. Lang. Syst. | 2 |
| 2004 | Low-Complexity Distributed Issue QueueabstractAs technology evolves, power density significantly increases and cooling systems become more complex and expensive. The issue logic is one of the processor hotspots and, at the same time, its latency is crucial for the processor performance. We present a low-complexity FP issue logic (MB/spl I.bar/distr) that achieves high performance with small energy requirements. The MB/spl I.bar/distr scheme is based on classifying instructions and dispatching them into a set of queues depending on their data dependences. These instructions are selected for issuing based on an estimation of when their operands will be available, so the conventional wakeup activity is not required. Additionally, the functional units are distributed across the different queues. The energy required by the proposed scheme is substantially lower than that required by a conventional issue design, even if the latter has the ability of waking-up only unready operands. MB/spl I.bar/distr scheme reduces the energy-delay product by 35% and the energy-delay product by 18% with respect to a state-of-the-art approach. Jaume Abella 0001, Antonio González 0001 |
HPCA | 1 |
| 2003 | Power-Aware Adaptive Issue Queue and Register File
Jaume Abella 0001, Antonio González 0001 |
HiPC | 1 |
| 2003 | Power Efficient Data Cache DesignsabstractWe investigate some power efficient data cache designs that try to significantly reduce the cache energy consumption, both static and dynamic, with a minimal impact in performance. The basic idea is to combine different threshold voltages with different cache organizations that provide different levels of performance. Multibanked organizations in combination with different approaches to allocate data to cache banks are explored. Some of the resulting cache architectures are shown to provide a good tradeoff between power and performance. Jaume Abella 0001, Antonio González 0001 |
ICCD | 1 |
| 2003 | On Reducing Register Pressure and Energy in Multiple-Banked Register FilesabstractThe storage for speculative values in superscalar processors is one of the main sources of complexity and power dissipation. We present a novel technique to reduce register requirements as well as their dynamic and static power dissipation that is based on delaying the dispatch of instructions while minimizing its impact on performance. The proposed technique outperforms previous schemes in both performance and power savings. With only 1.77% IPC loss, the mechanism achieves more than 13% dynamic and 15% static extra power savings in the integer rename buffers and more than 9% dynamic and 10% static extra power savings in the FP rename buffers. Significant power savings are also achieved if the processor uses a physical register file for both committed and noncommitted values instead of rename buffers. Additionally the register requirements are reduced by more than 18% and 13% for integer and FP programs respectively. Jaume Abella 0001, Antonio González 0001 |
ICCD | 1 |