Carles Hernández 0001

dblp:77/7434 · also Carles Hernández Luz · DBLP profile ↗
← Back
72ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0001-5393-3195ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 64 · 11 first-author · 20 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TinyATD: Compressing Auxiliary Tag Directories using tag hashing and set-sampling
abstract
Auxiliary Tag Directories (ATD) are hardware structures widely analyzed in academia to estimate the interference that a task experiences when sharing a cache in a multithreaded system. ATDs are used for execution time inflation calculations, reducing power consumption in multiprocessor systems, improving cache replacement, and improving cache prefetching accuracy. However, the commercial adoption of ATD-based approaches seems to be limited due to the area overheads that they impose. We propose TinyATD, a mechanism that aims to improve state-of-the-art ATD area reduction by combining ATD set sampling and tag hashing area reduction techniques to build smaller and more precise ATDs. TinyATD achieves a 41% area reduction for the same error or a 23% error reduction for the same area over the baseline set-sampling when sampling 32 out of 2048 sets. Furthermore, TinyATD offers the stated reduction over a wide array of commonly used LLC replacement policies.
Pablo Andreu, Pedro López 0001, Carles Hernández 0001
J. Syst. Archit.3
2026 HashTAG With CALM: Low-Overhead Hardware Support for Inter-Task Eviction Monitoring
abstract
Multicore processors have emerged as the preferred architecture for safetycritical systems due to their significant performance advantages. However, concurrent access by multiple cores to a shared cache induces intercore evictions that generate nondeterministic interference and compromise timing predictability. Static partitioning of the cache among cores is a wellestablished countermeasure that effectively eliminates such evictions but reduces flexibility and system throughput. To accurately estimate inter-core cache contention, Auxiliary Tag Directories (ATDs) are widely adopted. However, ATDs incur substantial hardware area costs, which often motivates the use of heuristic-based reductions. These reduced ATD designs, while more compact, compromise accuracy and therefore are not suitable for safety-critical domains. This paper extends the proposal of HashTAG, a novel approach to accurately upper-bound inter-core eviction interference. HashTAG introduces a safe and lightweight Auxiliary Tag Directory mechanism that tracks which cores are responsible for evicting cache lines used by others, thus measuring contention. We further refine the proposed HashTAG approach by creating CALM, a custom-made memory allocator that significantly improves HashTAG performance in multicore systems. Our results show that no inter-task interference underprediction is possible with HashTAG, making it suitable for the safety domain. HashTAG provides a 47% reduction in the Auxiliary Tag Directory area, presenting perfect measurements on 80% of cases and only a 1% error on maximum inter-core eviction measurements for a HashTAG tag size of ten bits.
Pablo Andreu, Pedro López 0001, Carles Hernández 0001
IEEE Trans. Parallel Distributed Syst.3
2026 An Open Source Methodology to Emulate Transient Faults in ASIC Storage Cells Using AMD Ultrascale+ FPGAs
abstract
The dependability assessment of critical systems must consider the emulation of transient faults, as they pose an important dependability threat for modern VLSI designs. FPGA-based systems are dominated by bit-flips in configuration memory (CM), which are relatively easy to emulate using partial runtime reconfiguration (RTR) FPGA fault injection (FFI) approaches. However, when FPGA is used as an ASIC prototyping platform, transient fault models representative of ASIC designs must be considered, such as bit-flips in sequential logic cells (Flip-Flops and on-chip RAM blocks). Existing RTR-FFI approaches do not adequately cover these faults, as they require the orchestrated manipulation of multiple CM bits for each target logic cell, and the location of these CM bits is unknown (not documented) for modern FPGA generations. This work experimentally formalises the mapping of the necessary CM bits, proposes an enhanced RTR-FFI methodology to emulate bit-flips in registers and on-chip RAMs of current-generation AMD Ultrascale+ FPGAs, and provides an upgraded publicly available open source FFI tool (BAFFI, https://gitlab.com/selene-riscv-platform/DAVOS ) supporting the proposed methodology. The validity of the proposed FFI approach is demonstrated by comparison with gate-level simulation-based fault injection in a case study of two soft-core processors (MC8051 and NOEL-V).
Ilya Tuzov, David de Andrés, Juan-Carlos Ruiz-Garcia 0001, Carles Hernández 0001
ACM Trans. Reconfigurable Technol. Syst.4
2025 Improving Out-of-Distribution Detection via Test-Time Augmentation
Imanol Allende, Nicholas Mc Guire, Javier del Campo, Carles Hernández 0001
SAFECOMP4
2025 Expanding SafeSU capabilities by leveraging security frameworks for contention monitoring in complex SoCs
abstract
The increased performance requirements of applications running on safety-critical systems have led to the use of complex platforms with several CPUs, GPUs, and AI accelerators. However, higher platform and system complexity challenge performance verification and validation since timing interference across tasks occurs in unobvious ways, hence defeating attempts to optimize application consolidation informedly during design phases and validating that mutual interference across tasks is within bounds during test phases. In that respect, the SafeSU has been proposed to extend inter-task interference monitoring capabilities in simple systems. However, modern mixed-criticality systems are complex, with multilayered interconnects, shared caches, and hardware accelerators. To that end, this paper proposes a non-intrusive add-on approach for monitoring interference across tasks in multilayer heterogeneous systems implemented by leveraging existing security frameworks and the SafeSU infrastructure. The feasibility of the proposed approach has been validated in an RTL RISC-V-based multicore SoC with support for AI hardware acceleration. Our results show that our approach can safely track contention and properly break down contention cycles across the different sources of interference, hence guiding optimization and validation processes. • Enables inter-core interference monitoring in complex SoCs. • Tool to check inter-task timing independence at all NoC levels. • Proposes unified initiator naming to solve initiator dilution across SoC layers.
Pablo Andreu, Sergi Alcaide, Pedro López 0001, Jaume Abella 0001, Carles Hernández 0001
Future Gener. Comput. Syst.5
2025 Leveraging HLS to Design a Versatile & High-Performance Classic McEliece Accelerator
abstract
By harnessing fundamental quantum properties, a large-scale quantum computer could undermine currently deployed public-key algorithms. The post-quantum, code-based cryptosystem Classic McEliece (CM) addresses this security concern. However, its large public key size (up to 1.3 MB) poses various hardware implementation challenges. In this article, we focus on the high memory bandwidth requirements of the CM encoding function, in the context of heterogeneous CPU-FPGA devices. More concretely, we target the acceleration of public-key loading and processing from any globally shared or accelerator-private memory system. We present a novel and constant-time accelerator eEnc that exploits the elevated parallelization potential of FPGA devices to yield high-performance results. Our accelerator implements the encoding and the random error vector generation functions, which comprise the main computational load of Encapsulation. Two accelerator design variants are introduced, providing different hardware tradeoffs. Regarding intra-accelerator data communication, and unlike other state-of-the-art (SOTA) works, we combine a streaming protocol with task-level parallelization to remove the need to store the public key in accelerator-private memories. Our proposed design shows new record execution times over its SOTA counterparts, ranging on average from 3.5× up to 7.7× across the five security level parameter sets. Our end-to-end implementation in a Zynq SoC shows an average speedup of 2.2× compared to a 64-bit vectorized CM software-baseline. The elevated logic resource consumption, characteristic of HLS designs, can be readily adjusted with a performance tradeoff.
Vatistas Kostalabros, Jordi Ribes-González, Xavier Carril, Oriol Farràs, Carles Hernández 0001, Miquel Moretó
ACM Trans. Embed. Comput. Syst.5
2025 Exploiting neural networks bit-level redundancy to mitigate the impact of faults at inference
abstract
Abstract Neural networks are widely used in critical environments such as healthcare, autonomous vehicles, or video surveillance. To ensure the safety of the systems that rely on their functionality, it is essential to validate their correct behaviour in the presence of faults. This paper studies the behaviour of state-of-the-art neural network models with fault injection in their weights. For this purpose, we analyse the sensitivity of these models and identify the impact of bit flips on their accuracy. To mitigate the effects of faults, we introduce two mechanisms that leverage bit-level redundancy for protection. The first mechanism, Fixed Protection , safeguards consecutive sets of bits, while the second, Variable Protection , targets non-consecutive bits. Our findings demonstrate that, on average, random bit flip faults cause the accuracy of the original models to drop by 1.3% to over 3%. However, with our protection mechanisms in place, accuracy reductions are significantly minimised, ranging from only 0.0001% to 0.4%.
Izan Catalán, José Flich, Carles Hernández 0001
J. Supercomput.3
2024 Hardware Acceleration for High-Volume Operations of CRYSTALS-Kyber and CRYSTALS-Dilithium
abstract
Many high-demand digital services need to perform several cryptographic operations, such as key exchange or security credentialing, in a concise amount of time. In turn, the security of some of these cryptographic schemes is threatened by advances in quantum computing, as quantum computer could break their security in the near future. Post-quantum cryptography (PQC) is an emerging field that studies cryptographic algorithms that resist such attacks. The National Institute of Standards and Technology (NIST) has selected the CRYSTALS-Kyber Key Encapsulation Mechanism and the CRYSTALS-Dilithium Digital Signature algorithm as primary PQC standards. In this article, we present field-programmable gate array (FPGA)-based hardware accelerators for high-volume operations of both schemes. We apply high-level synthesis (HLS) for hardware optimization, leveraging a batch processing approach to maximize the memory throughput and applying custom HLS logic to specific algorithmic components. Using reconfigurable FPGAs, we show that our hardware accelerators achieve speedups between 3 \(\times\) and 9 \(\times\) over software baseline implementations, even over ones leveraging CPU vector architectures. Furthermore, the methods used in this study can also be extended to the new CRYSTALS-based NIST FIPS drafts, ML-KEM and ML-DSA, with similar acceleration results.
Xavier Carril, Charalampos Kardaris, Jordi Ribes-González, Oriol Farràs, Carles Hernández 0001, Vatistas Kostalabros, Joel Ulises González-Jiménez, Miquel Moretó
ACM Trans. Reconfigurable Technol. Syst.5
2023 NimbleAI: Towards Neuromorphic Sensing-Processing 3D-integrated Chips
abstract
The NimbleAI Horizon Europe project leverages key principles of energy-efficient visual sensing and processing in biological eyes and brains, and harnesses the latest advances in$\mathbf{33D}$stacked silicon integration, to create an integral sensing-processing neuromorphic architecture that efficiently and accurately runs computer vision algorithms in area-constrained endpoint chips. The rationale behind the NimbleAI architecture is: sense data only with high information value and discard data as soon as they are found not to be useful for the application (in a given context). The NimbleAI sensing-processing architecture is to be specialized after-deployment by tunning system-level trade-offs for each particular computer vision algorithm and deployment environment. The objectives of NimbleAI are: (1)$\mathbf{100x}$performance per mW gains compared to state-of-the-practice solutions (i.e., CPU/GPUs processing frame-based video); (2)$\mathbf{50x}$processing latency reduction compared to CPU/GPUs; (3) energy consumption in the order of tens of mWs; and (4) silicon area of approx. 50 mm2.
Xabier Iturbe, Nassim Abderrahmane, Jaume Abella 0001, Sergi Alcaide, Eric Beyne, Henri-Pierre Charles, Christelle Charpin-Nicolle, Lars Chittka, Angélica Dávila, Arne Erdmann, Carles Estrada, Ander Fernández, Anna Fontanelli, José Flich, Gianluca Furano, Alejandro Hernán Gloriani, Erik Isusquiza, Radu Grosu, Carles Hernández 0001, Daniele Ielmini, Maha Kooli, Nicola Lepri, Bernabé Linares-Barranco, Jean-Loup Lachese, Eric Laurent, Menno Lindwer, Frank Linsenmaier, Mikel Luján, Karel Masarík, Nele Mentens, Orlando Moreira, Chinmay Nawghane, Luca Peres, Jean-Philippe Noël, Arash Pourtaherian, Christoph Posch, Peter Priller, Zdenek Prikryl, Felix Resch, Oliver Rhodes, Todor P. Stefanov, Moritz Storring, Michele Taliercio, Rafael Tornero, Marcel D. van de Burgwal, Geert Van der Plas, Elisa Vianello, Pavel Zaykov
DATE19
2023 BAFFI: a bit-accurate fault injector for improved dependability assessment of FPGA prototypes
abstract
FPGA-based fault injection (FFI) is an indispensable technique for verification and dependability assessment of FPGA designs and prototypes. Existing FFI tools make use of Xilinx essential bits technology to locate the relevant fault targets in FPGA configuration memory (CM). Most FFI tools treat essential bits as black-box, while few of them are able to filter essential bits on the area basis in order to selectively target design components contained within the predefined Pblocks. This approach, however, remains insufficiently precise since the granularity of Pblocks in practice does not reach the smallest design components. This paper proposes an open-source FFI tool that enables much more fine-grained FFI experiments for Xilinx 7-series and Ultrascale+ FPGAs. By mapping the essential bits with the hierarchical netlist, it allows to precisely target any component in the design tree, up to an individual LUT or register, without the need for defining Pblocks (floorplanning). With minimal experimental effort it estimates the contribution of each DUT component into the resulting dependability features, and discovers weak points of the DUT. Through case studies we show how the proposed tool can be applied to different kinds of DUTs: from small-footprint microcontrollers, up to multicore RISC-V SoC. The correctness of FFI results is validated by means of RT-level and gate-level simulation-based fault injection.
Ilya Tuzov, David de Andrés, Juan-Carlos Ruiz-Garcia 0001, Carles Hernández 0001
DATE4
2023 A Survey of Recent Developments in Testability, Safety and Security of RISC-V Processors
abstract
With the continued success of the open RISC-V architecture, practical deployment of RISC-V processors necessitates an in-depth consideration of their testability, safety and security aspects. This survey provides an overview of recent developments in this quickly-evolving field. We start with discussing the application of state-of-the-art functional and system-level test solutions to RISC-V processors. Then, we discuss the use of RISC-V processors for safety-related applications; to this end, we outline the essential techniques necessary to obtain safety both in the functional and in the timing domain and review recent processor designs with safety features. Finally, we survey the different aspects of security with respect to RISC-V implementations and discuss the relationship between cryptographic protocols and primitives on the one hand and the RISC-V processor architecture and hardware implementation on the other. We also comment on the role of a RISC-V processor for system security and its resilience against side-channel attacks.
Jens Anders, Pablo Andreu, Bernd Becker 0001, Steffen Becker 0001, Riccardo Cantoro, Nikolaos Ioannis Deligiannis, Nourhan Elhamawy, Tobias Faller, Carles Hernández 0001, Nele Mentens, Mahnaz Namazi Rizi, Ilia Polian, Abolfazl Sajadi, Matthias Sauer 0002, Denis Schwachhofer, Matteo Sonza Reorda, Todor Stefanov, Ilya Tuzov, Stefan Wagner 0001, Nusa Zidaric
ETS9
2023 Accurately Measuring Contention in Mesh NoCs in Time-Sensitive Embedded Systems
abstract
The computing capacity demanded by embedded systems is on the rise as software implements more functionalities, ranging from best-effort entertainment functions to performance-guaranteed safety-related functions. Heterogeneous manycore processors, using wormhole mesh (wmesh) Network-on-Chips (NoCs) as the main communication means, and contention block among applications, are increasingly considered to deliver the required computing performance. Most research efforts on software timing analysis have focused on deriving bounds (estimates) to the contention that tasks can suffer when accessing wmesh NoCs. However, less effort has been devoted to an equally important problem, namely,accuratelymeasuring the actual contention tasks generate each other on the wmesh which is instrumental during system validation to diagnose any software timing misbehavior and determine which tasks are particularly affected by contention on specific wmesh routers. In this article, we work on the foundations ofcontention measuringin wmesh NoCs and propose and explain the rationale of agolden metric, called taskPairWise Contention(PWC). PWC allows ascribing the actual share of the contention a given task suffers in the wmesh to each of its co-runner tasks at packet level. We also introduce and formalize aGolden Reference Value(GRV) for PWC that specifically defines a criterion to fairly break down the contention suffered by a task among its co-runner tasks in the wmesh. Our evaluation shows that GRV effectively captures how contention occurs by identifying the actual core (task) causing contention and whether contention is caused by local or remote interference in the wmesh.
Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Enrico Mezzetti, Francisco J. Cazorla
ACM Trans. Design Autom. Electr. Syst.2
2022 SafeSU-2: a Safe Statistics Unit for Space MPSoCs
abstract
Advanced statistics units (SUs) have been proven effective for the verification, validation and implementation of safety measures as part of safety-related MPSoCs. This is the case, for instance, of the RISC-V MPSoC by CAES Gaisler based on NOEL-V cores that will become commercially ready on FPGAs by the end of 2022. However, while those SUs support safety in the rest of the SoC, they must be built to be safe to be part of commercial products. This paper presents the SafeSU-2, the safety-compliant version of the SafeSU. In particular, we perform a Failure Mode and Effect Analysis (FMEA) for the SafeSU for relevant fault models, and implement fault detection and tolerance features needed to make it compliant with the requirements of safety-related devices in general, and of space MPSoCs in particular.
Guillem Cabo, Sergi Alcaide, Carles Hernández 0001, Pedro Benedicte, Francisco Bas, Fabio Mazzocchetti, Jaume Abella 0001
DATE3
2022 The SELENE Deep Learning Acceleration Framework for Safety-related Applications
abstract
The goal of the H2020 SELENE project is the development of a flexible computing platform for autonomous applications that includes built-in hardware support for safety. The SELENE computing platform is an open-source RISC-V heterogeneous multicore system-on-chip (SoC) that includes 6 NOEL-V RISC-V cores and artificial intelligence accelerators. In this paper, we describe the approach followed in the SELENE project to accelerate neural network inference processes. Our intermediate results show that both the FPGA and ASIC accel-erators provide real-time inference performance for the analyzed network models at a reasonable implementation cost.
Laura Medina, Salva Carrion, Pablo Andreu, Tomás Picornell, José Flich, Carles Hernández 0001, Michael Sandoval, Markel Sainz, Charles-Alexis Lefebvre, Martin Rönnbäck, Martin Matschnig, Matthias Wess, Herbert Taucher
DATE6
2022 Efficient Inference Of Image-Based Neural Network Models In Reconfigurable Systems With Pruning And Quantization
abstract
Neural networks (NN) for image processing in embedded systems expose two conflicting requirements: increasing computing power needs as models become more complex and constrained resource budget. In order to alleviate this problems, model compression based on quantization and pruning techniques are common. Derived models then need to fit on reconfigurable systems such as FPGAs for the embedded system to work properly. In this paper, we present HLSinf, an open source framework for the development of custom NN accelerators for FPGAs which provides efficient support to quantized and pruned NN models. With HLSinf, significant inference speedups can be obtained for typical medical image-based applications. In particular, we obtain up to 90x speedup factor when we combine quantization/pruning with the flexibility of HLSinf compared to CPU.
José Flich, Laura Medina, Izan Catalán, Carles Hernández 0001, Andrea Bragagnolo, Fabrice Auzanneau, David Briand
ICIP4
2021 From a FPGA Prototyping Platform to a Computing Platform: The MANGO Experience
abstract
In this paper we describe the evolution of the FPGA-based prototype deployed in the MANGO project, from a hardware prototyping platform of HPC architectures to a computing platform targeting HPC and AI applications. Our main goal is to reinvest on the MANGO cluster by providing a duality in its use for both large-scale hardware prototyping and highperformance computation. From our experience we can reach several interesting conclusions about the complexities and hurdles that lay below FPGA technologies, and therefore, shedding some light onto the real complexities that difficult the adoption of FPGAs on either large-scale pure HPC systems or on hybrid systems (HPC + BigData/Ai).
José Flich, Rafael Tornero, Davide Russo, José Maria Martínez, Carles Hernández 0001
DATE6
2021 Security, Reliability and Test Aspects of the RISC-V Ecosystem
abstract
RISC-V has emerged as a viable solution on academia and industry. However, to use open source hardware for safety-critical applications, we need a deep understanding of the way in which well established mechanisms for testing and reliability could be integrated and deployed on the RISC-V ecosystem, and we need a clear knowledge on how such an ecosystem can be leveraged to improve security. This paper includes four contributions presenting the potential of RISC-V in security research, the way in which RISC-V can be hardened against power analysis attacks, how to implement, using RISC-V, software and hardware/software solutions for dual core lock step, and how to perform system-level testing in the RISC-V ecosystem.
Jaume Abella 0001, Sergi Alcaide, Jens Anders, Francisco Bas, Steffen Becker 0001, Elke De Mulder, Nourhan Elhamawy, Frank K. Gürkaynak, Helena Handschuh, Carles Hernández 0001, Michael Hutter, Leonidas Kosmidis, Ilia Polian, Matthias Sauer 0002, Stefan Wagner 0001, Francesco Regazzoni 0001
ETS10
2021 SafeSU: an Extended Statistics Unit for Multicore Timing Interference
abstract
Statistics units (SUs) in MPSoCs are becoming increasingly used for the (1) verification and (2) validation of multicore timing interference, as well as for (3) deploying safety measures in safety-related real-time systems. However, existing SU extensions to manage multicore timing interference have neither been integrated together nor deployed in commercial MPSoCs.This paper presents the realization of the Safe Statistics Unit (SafeSU for short), which smartly integrates existing solutions for multicore timing interference verification, validation and monitoring, and is in turn integrated in commercial space-graded RISC-V and SparcV8 MPSoCs. Our evaluation illustrates the operation of the SafeSU, and paves the way for a thorough validation prior to reaching commercialization and being offered as open source IP.
Guillem Cabo, Francisco Bas, Ruben Lorenzo, David Trilla, Sergi Alcaide, Miquel Moretó, Carles Hernández 0001, Jaume Abella 0001
ETS7
2021 HLS-Based HW/SW Co-Design of the Post-Quantum Classic McEliece Cryptosystem
abstract
While quantum computers are rapidly becoming more powerful, the current cryptographic infrastructure is imminently threatened. In a preventive manner, the U.S. National Institute of Standards and Technology (NIST) has initiated a process to evaluate quantum-resistant cryptosystems, to form the first post-quantum (PQ) cryptographic standard. Classic McEliece (CM) is one of the most prominent cryptosystems considered for standardization in NIST’s PQ cryptography contest. However, its computational cost poses notable challenges to a big fraction of existing computing devices. This work presents an HLS-based, HW/SW co-design acceleration of the CM Key Encapsulation Mechanism (CM KEM). We demonstrate significant maximum speedups of up to 55.2 ×, 3.3 ×, and 8.7 × in the CM KEM algorithms of key generation, encapsulation, and decapsulation respectively, comparing to a SW-only scalar implementation.
Vatistas Kostalabros, Jordi Ribes-González, Oriol Farràs, Miquel Moretó, Carles Hernández 0001
FPL5
2021 Improving the Robustness of Redundant Execution with Register File Randomization
abstract
Staggered Redundant execution (SRE) is a fault-tolerance mechanism that has been widely deployed in the context of safety-critical applications. SRE not only protects the system in the presence of faults but also helps relaxing safety requirements of individual elements. However, in this paper, we show that SRE does not effectively protect the system against a wide range of faults and thus, new mechanisms to increase the diversity of homogeneous cores are needed. In this paper, we propose Register File Randomization (RFR), a low-cost diversity mechanism that significantly increases the robustness of homogeneous multicores in front of common-cause faults (CCFs) and register file wearout. Our results show that RFR completely removes the failure rate for register file CCFs for certain workloads and reduces by a factor of 5X the impact of stress related register file aging for the workloads analysed. Our implementation requires less than 50 RTL lines of code and the area (FPGA logic) overhead of RFR is less than 0.2% of a 64-bit RISC-V core FPGA implementation.
Ilya Tuzov, Pablo Andreu, Laura Medina, Tomás Picornell, Antonio Robles, Pedro López 0001, José Flich, Carles Hernández 0001
ICCAD8
2021 Enforcing Predictability of Many-Cores With DCFNoC
abstract
The ever need for higher performance forces industry to include technology based on multi-processors system on chip (MPSoCs) in their safety-critical embedded systems. MPSoCs include a network-on-chip (NoC) to interconnect the cores between them and with memory and the rest of shared resources. Unfortunately, the inclusion of NoCs compromises guaranteeing time predictability as network-level conflicts may occur. To overcome this problem, in this article we propose DCFNoC, a new time-predictable NoC design paradigm where conflicts within the network are eliminated by design. This new paradigm builds on top of the Channel Dependency Graph (CDG) in order to deterministically avoid network conflicts. The network guarantees predictability to applications and is able to naturally inject messages using a TDM period equal to the optimal theoretical bound without the need of using a computationally demanding offline process. DCFNoC is integrated in a tile-based many-core system and adapted to its memory hierarchy. Our results show that DCFNoC guarantees time predictability avoiding network interference among multiple running applications. DCFNoC always guarantees performance and also improves wormhole performance in a 4 x 4 setting by a factor of 3.7x when interference traffic is injected. For a 8 x 8 network differences are even larger. In addition, DCFNoC obtains a total area saving of 10.79 percent over a standard wormhole implementation.
Tomás Picornell, José Flich, Carles Hernández 0001, José Duato
IEEE Trans. Computers3
2021 Worst-Case Energy Consumption: A New Challenge for Battery-Powered Critical Devices
abstract
The number of (edge) devices connected to the IoT is on the rise, reaching hundreds of billions in the next years. Many devices will implement some type of critical functionality, for instance in the medical market this includes infusion pumps and implantable defibrillators. Energy awareness is mandatory in the design of IoT devices given their huge impact on worldwide energy consumption and the fact that many of them are battery powered. Critical IoT devices further require addressing new energy-related challenges. On the one hand, factoring in the impact of energy-solutions on device's performance, providing evidence of adherence to domain-specific safety standards. On the other hand, deriving safe worst-case energy consumption (WCEC) estimates is fundamental to ensure the system can continuously operate under a pre-established set of power/energy caps, safely delivering its critical functionality. In this line, we analyze for the first time the impact that different hardware physical parameters have on both model-based and measurement-based WCEC modeling, for which we also show the main challenges they face compared to chip manufacturers' current practice for energy modeling and validation. Under the set of constraints that emanate from how certain physical parameters can be actually modeled, we show that measurement-based WCEC is a promising way forward for WCEC estimation.
David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
IEEE Trans. Sustain. Comput.2
2020 SELENE: Self-Monitored Dependable Platform for High-Performance Safety-Critical Systems
abstract
Existing HW/SW platforms for safety-critical systems suffer from limited performance and/or from lack of flexibility due to building on specific proprietary components. This jeopardizes their wide deployment across domains. While some research has been done to overcome these limitations, they have had limited success owing to missing flexibility and extensibility. Flexibility and extensibility are the cornerstones of industry adoption: industries dealing in capital goods need technologies on which they can rely on during decades (e.g. avionics, space, automotive). SELENE aims at covering this gap by proposing a new family of safety-critical computing platforms, which builds upon open source components such as the RISC-V instruction set architecture, GNU/Linux, and the Jailhouse hypervisor. SELENE will develop an advanced computing platform that is able to: (1) adapt the system to the specific requirements of different application domains, to changing environmental conditions, and to internal conditions of the system itself; (2) allow the integration of applications of different criticalities and performance demands in the same platform, guaranteeing functional and temporal isolation properties; (3) achieve flexible diverse redundancy by exploiting the inherent redundant capabilities of the multicore; and (4) efficiently execute compute-intensive applications by means of specific accelerators.
Carles Hernández 0001, José Flich, Roberto Paredes, Charles-Alexis Lefebvre, Imanol Allende, Jaume Abella 0001, David Trillin, Martin Matschnig, Bernhard Fischer, Konrad Schwarz, Jan Kiszka, Martin Rönnbäck, Johan Klockars, Nicholas Mc Guire, Franz Rammerstorfer, Christian Schwarzl, Franck Wartel, Dierk Lüdemann, Mikel Labayen
DSD1
2019 Towards limiting the impact of timing anomalies in complex real-time processors
abstract
Timing verification of embedded critical real-time systems is hindered by complex designs. Timing anomalies, deeply analyzed in static timing analysis, require specific solutions to bound their impact. For the first time, we study the concept and impact of timing anomalies in measurement-based timing analysis, the most used in industry, showing that they require to be considered and handled differently. In addition, we analyze anomalies in the context of Measurement-Based Probabilistic Timing Analysis, which simplifies quantifying their impact.
Pedro Benedicte, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Francisco J. Cazorla
ASP-DAC3
2019 DCFNoC: A Delayed Conflict-Free Time Division Multiplexing Network on Chip
abstract
The adoption of many-cores in safety-critical systems requires real-time capable networks on chip (NoC). In this paper we propose a new time-predictable NoC design paradigm where contention within the network is eliminated. This new paradigm builds on the Channel Dependency Graph (CDG) and guarantees by design the absence of contention. Our delayed conflict-free NoC (DCFNoC) is able to naturally inject messages using a TDM period equal to the optimal theoretical bound and without the need of using a computationally demanding offline process. Results show that DCFNoC guarantees time predictability with very low implementation cost.
Tomás Picornell, José Flich, Carles Hernández 0001, José Duato
DAC3
2019 High-Integrity GPU Designs for Critical Real-Time Automotive Systems
abstract
Autonomous Driving (AD) imposes the use of high-performance hardware, such as GPUs, to perform object recognition and tracking in real-time. However, differently to the consumer electronics market, critical real-time AD functionalities require a high degree of resilience against faults, in line with the automotive ISO26262 functional safety standard requirements. ISO26262 imposes the use of some source of independent redundancy for the most critical functionalities so that a single fault cannot lead to a failure, being dual core lockstep (DCLS) with diversity the preferred choice for computing devices. Unfortunately, GPUs do not support diverse DCLS by construction, thus failing to meet ISO26262 requirements efficiently.In this paper we propose lightweight modifications to GPUs to enable diverse DCLS for critical real-time applications without diminishing their performance for non-critical applications. In particular, we show how enabling specific mechanisms for software-controlled kernel scheduling in the GPU, allows guaranteeing that redundant kernels can be executed in different resources so that a single fault cannot lead to a failure, as imposed by ISO26262. Our results on a GPU simulator and an NVIDIA GPU prove the viability of the approach and its effectiveness on high-performance GPU designs needed for AD systems.
Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001
DATE3
2019 LAEC: Look-Ahead Error Correction Codes in Embedded Processors L1 Data Cache
abstract
As implementation technology shrinks, the presence of errors in cache memories is becoming an increasing issue in all computing domains. Critical systems, e.g. space and automotive, are specially exposed and susceptible to reliability issues. Furthermore, hardware designs in these systems are migrating to multilevel cache multicore systems, in which write-through first level data (DL1) caches have been shown to heavily harm average and guaranteed performance. While write-back DL1 caches solve this problem they come with their own challenges: they need Error Correction Codes (ECC) to tolerate soft errors, but implementing DL1 ECC in simple embedded micro-controllers requires either complex hardware to squash instructions consuming erroneous data, or delayed delivery of data to correct potential errors, which impacts performance even if such process is pipelined. In this paper we present a low-complexity hardware mechanism to anticipate data fetch and error correction in DL1 so that both (1) correct data is always delivered, but (2) avoiding additional delays in most of the cases. This achieves both high guaranteed performance and an effective solutions against errors.
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DATE2
2019 Maximum-Contention Control Unit (MCCU): Resource Access Count and Contention Time Enforcement
abstract
In real-time systems, the techniques to derive bounds to the contention tasks can suffer in multicore build on resource quota monitoring and enforcement. Existing techniques track and bound the number of requests to hardware shared resources that each core (task) is allowed to perform. In this paper we show that current software-only solutions work well when there is a single resource and type of request to track and bound, but do not scale to the more general case of several shared resources that accept different request types, each with a different associated latency. To handle this (more general) case, we propose low-overhead hardware support called Maximum-Contention Control Unit (MCCU). The MCCU performs fine-grain tracking of different types of requests, preventing a core to cause more interference on its contenders than budgeted. In this process, the MCCU also helps verifying that individual requests duration does not exceed their theoretical bounds, hence dealing with scenarios in which requests can have an arbitrarily large duration.
Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DATE2
2019 Challenges in Deeply Heterogeneous High Performance Systems
abstract
RECIPE (REliable power and time-ConstraInts-aware Predictive management of heterogeneous Exascale systems) is a recently started project funded within the H2020 FETHPC programme, which is expressly targeted at exploring new High-Performance Computing (HPC) technologies. RECIPE aims at introducing a hierarchical runtime resource management infrastructure to optimize energy efficiency and minimize the occurrence of thermal hotspots, while enforcing the time constraints imposed by the applications and ensuring reliability for both time-critical and throughput-oriented computation that run on deeply heterogeneous accelerator-based systems. This paper presents a detailed overview of RECIPE, identifying the fundamental challenges as well as the key innovations addressed by the project, which span run-time management, heterogeneous computing architectures, HPC memory/interconnection infrastructures, thermal modelling, reliability, programming models, and timing analysis. For each of these areas, the paper describes the relevant state of the art as well as the specific actions that the project will take to effectively address the identified technological challenges.
Giovanni Agosta, William Fornaciari, David Atienza 0001, Ramon Canal, Alessandro Cilardo, José Flich, Carles Hernández 0001, Michal Kulczewski, Giuseppe Massari, Rafael Tornero, Marina Zapater
DSD7
2019 An Approach for Detecting Power Peaks During Testing and Breaking Systematic Pathological Behavior
abstract
The verification and validation process of embedded critical systems requires providing evidence of their functional correctness and also that their non-functional behavior stays within limits. In this work, we focus on power peaks, which may cause voltage droops and thus, challenge performance to preserve correct operation upon droops. In this line, the use of complex software and hardware in critical embedded systems jeopardizes the confidence that can be placed on the tests carried out during the campaigns performed at analysis. This is so because it is unknown whether tests have triggered the highest power peaks that can occur during operation and whether any such peak can occur systematically. In this paper we propose the use of randomization, already used for timing analysis of real-time systems, as an enabler to guarantee that (1) tests expose those peaks that can arise during operation and (2) peaks cannot occur systematically inadvertently.
David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DSD2
2019 Modeling the Impact of Process Variations in Worst-Case Energy Consumption Estimation
abstract
The advent of autonomous power-limited systems poses a new challenge for system verification. Powerful processors needed to enable autonomous operation, are typically power-hungry, jeopardizing battery duration. Therefore, guaranteeing a given battery duration requires worst-case energy consumption (WCEC) estimation for tasks running on those systems. Unfortunately, processor energy and power can suffer significant variation across different units due to process variation (PV), i.e. variability in the electrical properties of transistors and wires due to imperfect manufacturing, which challenges existing WCEC estimation methods for applications. In this paper, we propose a statistical modeling approach to capture PV impact on applications energy and a methodology to compute their WCEC capturing PV, as required to deploy portable critical devices.
David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DSD2
2019 Software-only Diverse Redundancy on GPUs for Autonomous Driving Platforms
abstract
Autonomous driving (AD) builds upon high-performance computing platforms including (1) general purpose CPUs as well as (2) specific accelerators, being GPUs one of the main representatives. Microcontrollers have reached ASIL-D compliance by implementing diverse redundancy with lockstep execution. However, ASIL-D compliant GPUs rely on either fully redundant lockstep GPUs (i.e. 2 GPUs), which doubles hardware costs, or fully redundant systems with a GPU and another accelerator, which virtually doubles design and validation/verification (V&V) costs. In this paper we analyze the degree of diversity achieved when implementing redundancy on a single GPU, showing that diverse redundancy is not achieved in many cases, and propose software strategies that guarantee achieving diverse redundancy for any kernel on systems using commercial off-the-shelf (COTS) GPUs, thus showing how to achieve ASIL-D compliance on a single COTS GPU in controlled scenarios.
Sergi Alcaide, Leonidas Kosmidis, Carles Hernández 0001, Jaume Abella 0001
IOLTS3
2019 Time-Randomized Wormhole NoCs for Critical Applications
abstract
Wormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this article, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g., wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance w.r.t. conventional time-deterministic setups. Our results show that performance guarantees of applications running on top of probabilistic wNoC designs improve by 40% and 93% on average for 4 × 4 and 6 × 6 wNoC setups, respectively.
Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
ACM J. Emerg. Technol. Comput. Syst.2
2019 Locality-aware cache random replacement policies
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
J. Syst. Archit.2
2018 Cache side-channel attacks and time-predictability in high-performance critical real-time systems
abstract
Embedded computers control an increasing number of systems directly interacting with humans, while also manage more and more personal or sensitive information. As a result, both safety and security are becoming ubiquitous requirements in embedded computers, and automotive is not an exception to that. In this paper we analyze time-predictability (as an example of safety concern) and side-channel attacks (as an example of security issue) in cache memories. While injecting randomization in cache timing-behavior addresses each of those concerns separately, we show that randomization solutions for time-predictability do not protect against side-channel attacks and vice-versa. We then propose a randomization solution to achieve both safety and security goals.
David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DAC2
2018 Design and integration of hierarchical-placement multi-level caches for real-time systems
abstract
Enabling timing analysis in the presence of caches has been pursued by the real-Time embedded systems (RTES) community for years due to cache's huge potential to reduce software's worst-case execution time (WCET). However, caches heavily complicate timing analysis due to hard-To-predict access patterns, with few works dealing with time analyzability of multi-level cache hierarchies. For measurement-based timing analysis (MBTA) techniques-widely used in domains such as avionics, automotive, and rail-we propose several cache hierarchies amenable to MBTA. We focus on a probabilistic variant of MBTA (or MBPTA) that requires caches with time-randomized behavior whose execution time variability can be captured in the measurements taken during system's test runs. For this type of caches, we explore and propose different multi-level cache setups. From those, we choose a cost-effective cache hierarchy that we implement and integrate in a 4-core LEON3 RTL processor model and prototype in a FPGA. Our results show that our proposed setup implemented in RTL results in better (reduced) WCET estimates with similar implementation cost and no impact on average performance w.r.t. other MBPTA-Amenable setups.
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DATE2
2018 HWP: Hardware Support to Reconcile Cache Energy, Complexity, Performance and WCET Estimates in Multicore Real-Time Systems
abstract
High-performance processors have deployed multilevel cache (MLC) systems for decades. In the embedded real-time market, the use of MLC is also on the rise, with processors for future systems in space, railway, avionics and automotive already featuring two or more cache levels. One of the most critical elements for MLC is the write policy that not only affects several key metrics such as performance, WCET estimates, energy/power, and reliability, but also the design of complexity-prone cache coherence protocol and cache reliability solutions. In this paper we make an extensive analysis of existing write policies, namely write-through (WT) and write-back (WB). In the context of the real-time domain, we show that no write policy is superior for all metrics: WT simplifies the design of the coherence and reliability solutions at the cost of performance, WCET, and energy; while WB improves performance and energy results, but complicates cache design. To take the best of each policy, we propose Hybrid Write Policy (HWP) a low-complexity hardware mechanism that reconciles the benefits of WT in terms of simplifying the cache design (e.g. coherence solution) and the benefits of WB in improved average performance and WCET estimates as the pressure on the interconnection network increases. Guaranteed performance results show that HWP scales with core count similar to WB. Likewise, HWP reduces cache energy usage of WT, to levels similar to those of WB. These benefits are obtained while retaining the reduced coherence complexity of WT, in contrast to high coherence costs under WB.
Pedro Benedicte, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
ECRTS2
2018 NoCo: ILP-Based Worst-Case Contention Estimation for Mesh Real-Time Manycores
abstract
Manycores are capable of providing the computational demands required by functionally-advanced critical applications in domains such as automotive and avionics. In manycores a network-on-chip (NoC) provides access to shared caches and memories and hence concentrates most of the contention that tasks suffer, with effects on the worst-case contention delay (WCD) of packets and tasks' WCET. While several proposals minimize the impact of individual NoC parameters on WCD, e.g. mapping and routing, there are strong dependences among these NoC parameters. Hence, finding the optimal NoC configurations requires optimizing all parameters simultaneously, which represents a multidimensional optimization problem. In this paper we propose NoCo, a novel approach that combines ILP and stochastic optimization to find NoC configurations in terms of packet routing, application mapping, and arbitration weight allocation. Our results show that NoCo improves other techniques that optimize a subset of NoC parameters.
Jordi Cardona, Carles Hernández 0001, Enrico Mezzetti, Jaume Abella 0001, Francisco J. Cazorla
RTSS2
2018 EOmesh: Combined Flow Balancing and Deterministic Routing for Reduced WCET Estimates in Embedded Real-Time Systems
abstract
The increasing performance needs in critical real-time embedded systems (CRTESs) can only be satisfied with the use of high-performance manycore processors. While NoC-based manycore systems are popular in the high-performance domain due to their high average performance, they challenge deriving tight worst-case execution time (WCET) estimates, as needed in CRTES. Weighted meshes have been proposed to alleviate NoCs pathological behavior-caused by large bandwidth imbalance-by making locally unbalanced arbitration decisions to reach globally balanced bandwidth. In this paper, we show that existing weighted mesh solutions do not completely remove unwanted imbalance, in particular for nodes subject to high congestion. We propose even/odd mesh (EOmesh), an approach that combines heterogeneous predictable routing and weight allocations that delivers near-optimal bandwidth allocation across cores without increasing NoC complexity. EOmesh, which can be implemented either by hardware means or by software means on top of regular weighted meshes, improves the average performance and WCET results of the reference weighted mesh design.
Jordi Cardona, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Fitting Software Execution-Time Exceedance into a Residual Random Fault in ISO-26262
abstract
Car manufacturers relentlessly replace or augment the functionality of mechanical subsystems with electronic components. Most such subsystems (e.g., steer-by-wire) are safety related, hence, subject to regulation. ISO-26262, the dominant standard for road vehicles, regards software faults as systematic, while differentiating hardware faults between systematic and random. The analysis of systematic faults entails rigorous processes and qualitative considerations. The increasing complexity of modern on-board computers, however, questions the very notion of treating the violation of execution-time envelopes for software programs as a systematic fault. Modern hardware in fact reduces the user's ability to delve deep enough into the fabric of hardware-software interaction to gage its extent of contribution to the worst-case execution time (WCET). Changing the nature of the WCET-analysis problem may help address that challenge effectively. To this end, we propose a solution that should allow ISO-26262 to quantify the likelihood of execution-time exceedance events, relating it to target failure metrics employed in support of certification arguments, similarly to random faults in hardware. To this end, we inject randomization in the timing behavior of the computer hardware to relieve the user from the need to control hard-to-reach low-level parts, and use measurement-based probabilistic timing analysis to quantify, constructively, the failure rates resulting from the likelihood of execution-time exceedance events.
Irune Agirre, Francisco J. Cazorla, Jaume Abella 0001, Carles Hernández 0001, Enrico Mezzetti, Mikel Azkarate-askatsua, Tullio Vardanega
IEEE Trans. Reliab.4
2017 DIMP: A Low-Cost Diversity Metric Based on Circuit Path Analysis
abstract
Diversity has been regarded as a desirable property of redundant instances, since it allows circuits to behave differently in front of a given fault. However, while qualitatively diversity is a well-understood concept, usable efficient metrics do not exist to quantify diversity in the context of safety-related systems. In this paper we cover this gap by proposing DIMP, a low-cost diversity metric based on analyzing the paths of the redundant circuits. We relate it to the particular case of automotive microcontrollers implementing lockstep cores and show that it can be successfully used providing relevant information for addressing common cause faults.
Sergi Alcaide, Carles Hernández 0001, Antoni Roca 0001, Jaume Abella 0001
DAC2
2017 Probabilistic timing analysis on time-randomized platforms for the space domain
abstract
Timing Verification is a fundamental step in real-time embedded systems, with measurement-based timing analysis (MBTA) being the most common approach used to that end. We present a Space case study on a real platform that has been modified to support a probabilistic variant of MBTA called MBPTA. Our platform provides the properties required by MBPTA with the predicted WCET estimates with MBPTA being competitive to those with current MBTA practice while providing more solid evidence on their correctness for certification.
Mikel Fernández, David Morales, Leonidas Kosmidis, Alen Bardizbanyan, Ian Broster, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla, Paulo Machado, Luca Fossati
DATE6
2017 Design and implementation of a fair credit-based bandwidth sharing scheme for buses
abstract
Fair arbitration in the access to hardware shared resources is fundamental to obtain low worst-case execution time (WCET) estimates in the context of critical real-time systems, for which performance guarantees are essential. Several hardware mechanisms exist for managing arbitration in those resources (buses, memory controllers, etc.). They typically attain fairness in terms of the number of slots each contender (e.g., core) gets granted access to the shared resource. However, those policies may lead to unfair bandwidth allocations for workloads with contenders issuing short requests and contenders issuing long requests. We propose a Credit-Based Arbitration (CBA) mechanism that achieves fairness in the cycles each core is granted access to the resource rather than in the number of granted slots. Furthermore, we implement CBA as part of a LEON3 4-core processor for the Space domain in an FPGA proving the feasibility and good performance characteristics of the design by comparing it against other arbitration schemes.
Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DATE2
2017 Boosting Guaranteed Performance in Wormhole NoCs with Probabilistic Timing Analysis
abstract
Wormhole-based NoCs (wNoCs) are widely accepted in high-performance domains as the most appropriate solution to interconnect an increasing number of cores in the chip. However, wNoCs suitability in the context of critical real-time applications has not been demonstrated yet. In this paper, in the context of probabilistic timing analysis (PTA), we propose a PTA-compatible wNoC design that provides tight time-composable contention bounds. The proposed wNoC design builds on PTA ability to reason in probabilistic terms about hardware events impacting execution time (e.g. wNoC contention), discarding those sequences of events occurring with a negligible low probability. This allows our wNoC design to deliver improved guaranteed performance. Our results show that WCET estimates of applications running on top of probabilistic wNoCs are reduced by 40% and 75% on average for 4x4 and 6x6 wNoC setups respectively when compared against deterministic wNoCs.
Mladen Slijepcevic, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DSD2
2017 Design and Implementation of a Time Predictable Processor: Evaluation With a Space Case Study
abstract
Embedded real-time systems like those found in automotive, rail and aerospace, steadily require higher levels of guaranteed computing performance (and hence time predictability) motivated by the increasing number of functionalities provided by software. However, high-performance processor design is driven by the average-performance needs of mainstream market. To make things worse, changing those designs is hard since the embedded real-time market is comparatively a small market. A path to address this mismatch is designing low-complexity hardware features that favor time predictability and can be enabled/disabled not to affect average performance when performance guarantees are not required. In this line, we present the lessons learned designing and implementing LEOPARD, a four-core processor facilitating measurement-based timing analysis (widely used in most domains). LEOPARD has been designed adding low-overhead hardware mechanisms to a LEON3 processor baseline that allow capturing the impact of jittery resources (i.e. with variable latency) in the measurements performed at analysis time. In particular, at core level we handle the jitter of caches, TLBs and variable-latency floating point units; and at the chip level, we deal with contention so that time-composable timing guarantees can be obtained. The result of our applied study with a Space application shows how per-resource jitter is controlled facilitating the computation of high-quality WCET estimates.
Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla, Alen Bardizbanyan, Jan Andersson, Fabrice Cros, Franck Wartel
ECRTS1
2017 Work-in-Progress Paper: An Analysis of the Impact of Dependencies on Probabilistic Timing Analysis and Task Scheduling
abstract
Recently there has been a renewed interest for probabilistic timing analysis (PTA) and probabilistic task scheduling (PTS). Despite the number of works in both fields, the link between them is weak: works on the latter build upon a series of assumptions on the probabilistic behavior of each task – or instances (jobs) of it – that have not been shown how to be fulfilled by PTA. This paper makes a first step towards covering this gap with emphasis on providing the right meaning of pWCET estimate as understood by both PTA and PTS. We show that the main issue related to ensuring that PTS assumptions on pWCET estimates are captured by PTA relates to the dependencies among tasks, and even jobs of a given task. Both change the scope of applicability of pWCET estimates provided by PTA and hence, their use by PTS.
Enrico Mezzetti, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla
RTSS3
2016 Random modulo: a new processor cache design for real-time critical systems
abstract
Cache memories have a huge impact on software's worst-case execution time (WCET). While enabling the seamless use of caches is key to provide the increasing levels of (guaranteed) performance required by automotive software, caches complicate timing analysis. In the context of Measurement-Based Probabilistic Timing Analysis (MBPTA) -- a promising technique to ease timing analyis of complex hardware -- we propose Random Modulo (RM), a new cache design that provides the probabilistic behavior required by MBPTA and with the following advantages over existing MBPTA-compliant cache designs: (i) an outstanding reduction in WCET estimates, (ii) lower latency and area overhead, and (iii) competitive average performance w.r.t conventional caches.
Carles Hernández 0001, Jaume Abella 0001, Andrea Gianarro, Jan Andersson, Francisco J. Cazorla
DAC1
2016 Improving performance guarantees in wormhole mesh NoC designs
Milos Panic, Carles Hernández 0001, Jaume Abella 0001, Antoni Roca 0001, Eduardo Quiñones, Francisco J. Cazorla
DATE2
2016 PROXIMA: Improving Measurement-Based Timing Analysis through Randomisation and Probabilistic Analysis
abstract
The use of increasingly complex hardware and software platforms in response to the ever rising performance demands of modern real-time systems complicates the verification and validation of their timing behaviour, which form a time-and-effort-intensive step of system qualification or certification. In this paper we relate the current state of practice in measurement-based timing analysis, the predominant choice for industrial developers, to the proceedings of the PROXIMA (Probabilistic real-time control of mixed-criticality multicore systems) project in that very field. We recall the difficulties that the shift towards more complex computing platforms causes in that regard. Then we discuss the probabilistic approach proposed by PROXIMA to overcome some of those limitations. We present the main principles behind the PROXIMA approach as well as the changes it requires at hardware or software level underneath the application. We also present the current status of the project against its overall goals, and highlight some of the principal confidence-building results achieved so far.
Francisco J. Cazorla, Jaume Abella 0001, Jan Andersson, Tullio Vardanega, Francis Vatrinet, Iain Bate, Ian Broster, Mikel Azkarate-askatsua, Franck Wartel, Liliana Cucu-Grosjean, Fabrice Cros, Glenn Farrall, Adriana Gogonel, Andrea Gianarro, Benoit Triquet, Carles Hernández 0001, Code Lo, Cristian Maxim, David Morales, Eduardo Quiñones, Enrico Mezzetti, Leonidas Kosmidis, Irune Agirre, Mikel Fernández, Mladen Slijepcevic, Philippa Conmy, Walid Talaboulma
DSD16
2016 pTNoC: Probabilistically Time-Analyzable Tree-Based NoC for Mixed-Criticality Systems
abstract
The use of networks-on-chip (NoC) in real-time safety-critical multicore systems challenges deriving tight worst-case execution time (WCET) estimates. This is due to the complexities in tightly upper-bounding the contention in the access to the NoC among running tasks. Probabilistic Timing Analysis (PTA) is a powerful approach to derive WCET estimates on relatively complex processors. However, so far it has only been tested on small multicores comprising an on-chip bus as communication means, which intrinsically does not scale to high core counts. In this paper we propose pTNoC, a new tree-based NoC design compatible with PTA requirements and delivering scalability towards medium/large core counts. pTNoC provides tight WCET estimates by means of asymmetric bandwidth guarantees for mixed-criticality systems with negligible impact on average performance. Finally, our implementation results show the reduced area and power costs of the pTNoC.
Mladen Slijepcevic, Mikel Fernández, Carles Hernández 0001, Jaume Abella 0001, Eduardo Quiñones, Francisco J. Cazorla
DSD3
2016 Modeling RTL fault models behavior to increase the confidence on TSIM-based fault injection
abstract
Future high-performance safety-relevant applications require microcontrollers delivering higher performance than the existing certified ones. However, means for assessing their dependability are needed so that they can be certified against safety critical certification standars (e.g ISO26262). Dependability assessment analyses performed at high level of abstraction inject single faults to investigate the effects these have in the system. In this work we show that single faults do not comprise the whole picture, due to fault multiplicities and reactivations. Later we prove that, by injecting complex fault models that consider multiplicities and reactivations in higher levels of abstraction, results are substantially different, thus indicating that a change in the methodology is needed.
Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001
IOLTS2
2016 Resilient random modulo cache memories for probabilistically-analyzable real-time systems
abstract
Fault tolerance has often been assessed separately in safety-related real-time systems, which may lead to inefficient solutions. Recently, Measurement-Based Probabilistic Timing Analysis (MBPTA) has been proposed to estimate Worst-Case Execution Time (WCET) on high performance hardware. The intrinsic probabilistic nature of MBPTA-commpliant hardware matches perfectly with the random nature of hardware faults. Joint WCET analysis and reliability assessment has been done so far for some MBPTA-compliant designs, but not for the most promising cache design: random modulo. In this paper we perform, for the first time, an assessment of the aging-robustness of random modulo and propose new implementations preserving the key properties of random modulo, a.k.a. low critical path impact, low miss rates and MBPTA compliance, while enhancing reliability in front of aging by achieving a better - yet random - activity distribution across cache sets.
David Trilla, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
IOLTS2
2016 Modeling High-Performance Wormhole NoCs for Critical Real-Time Embedded Systems
abstract
Manycore chips are a promising computing platform to cope with the increasing performance needs of critical real-time embedded systems (CRTES). However, manycores adoption by CRTES industry requires understanding task's timing behavior when their requests use manycore's network-on-chip (NoC) to access hardware shared resources. This paper analyzes the contention in wormhole-based NoC (wNoC) designs - widely implemented in the high-performance domain - for which we introduce a new metric: worst-contention delay (WCD) that captures wNoC impact on worst-case execution time (WCET) in a tighter manner than the existing metric, worst-case traversal time (WCTT). Moreover, we provide an analytical model of the WCD that requests can suffer in a wNoC and we validate it against wNoC designs resembling those in the Tilera-Gx36 and the Intel-SCC 48-core processors. Building on top of our WCD analytical model, we analyze the impact on WCD that different design parameters such as the number of virtual channels, and we make a set of recommendations on what wNoC setups to use in the context of CRTES.
Milos Panic, Carles Hernández 0001, Eduardo Quiñones, Jaume Abella 0001, Francisco J. Cazorla
RTAS2
2016 Parallelizing Industrial Hard Real-Time Applications for the parMERASA Multicore
abstract
The EC project parMERASA (Multicore Execution of Parallelized Hard Real-Time Applications Supporting Analyzability) investigated timing-analyzable parallel hard real-time applications running on a predictable multicore processor. A pattern-supported parallelization approach was developed to ease sequential to parallel program transformation based on parallel design patterns that are timing analyzable. The parallelization approach was applied to parallelize the following industrial hard real-time programs: 3D path planning and stereo navigation algorithms (Honeywell International s.r.o.), control algorithm for a dynamic compaction machine (BAUER Maschinen GmbH), and a diesel engine management system (DENSO AUTOMOTIVE Deutschland GmbH). This article focuses on the parallelization approach, experiences during parallelization with the applications, and quantitative results reached by simulation, by static WCET analysis with the OTAWA tool, and by measurement-based WCET analysis with the RapiTime tool.
Theo Ungerer, Christian Bradatsch, Martin Frieb, Florian Kluge, Jörg Mische, Alexander Stegmeier, Ralf Jahr, Mike Gerdes 0001, Pavel G. Zaykov, Lucie Matusova, Zai Jian Jia Li, Zlatko Petrov, Bert Böddeker, Sebastian Kehr, Hans Regler, Andreas Hugl, Christine Rochange, Haluk Ozaktas, Hugues Cassé, Armelle Bonenfant, Pascal Sainrat, Nick Lay, Ian Broster, Eduardo Quiñones, Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Francisco J. Cazorla, Sascha Uhrig, Mathias Rohde, Arthur Pyka
ACM Trans. Embed. Comput. Syst.28
2015 Analysis and RTL correlation of instruction set simulators for automotive microcontroller robustness verification
abstract
Increasingly complex microcontroller designs for safety-relevant automotive systems require the adoption of new methods and tools to enable a cost-effective verification of their robustness. In particular, costs associated to the certification against the ISO26262 safety standard must be kept low for economical reasons. In this context, simulation-based verification using instruction set simulators (ISS) arises as a promising approach to partially cope with the increasing cost of the verification process as it allows taking design decisions in early design stages when modifications can be performed quickly and with low cost. However, it remains to be proven that verification in those stages provides accurate enough information to be used in the context of automotive microcontrollers. In this paper we analyze the existing correlation between fault injection experiments in an RTL microcontroller description and the information available at the ISS to enable accurate ISS-based fault injection.
Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001, David de Andrés, Juan-Carlos Ruiz-Garcia 0001
DAC2
2015 Low-cost checkpointing in automotive safety-relevant systems
Carles Hernández 0001, Jaume Abella 0001
DATE1
2015 IEC-61508 SIL 3 Compliant Pseudo-Random Number Generators for Probabilistic Timing Analysis
abstract
Probabilistic Timing Analysis (PTA), especially its measurement based variant (MBPTA), has shown to be competitive with state-of-the-art timing analysis techniques. The use of MBPTA to analyse the timing behaviour of safety-critical systems rests on its ability to derive trustworthy WCET bounds. This ability depends on the soundness of the MBPTA method per se, as well as on the satisfaction of safety requirements placed on the pseudo-random number generator (prng) that plays a key role in the platform-level randomisation needed by MBPTA. This paper presents the design of a low-area, low-power prng that meets IEC-61508 SIL 3 safety requirements and allows for seamless integration in a real-world multicore architecture. This work enables the development and the IEC-61508 certification of mixed-criticality systems that use MBPTA for deriving timing bounds for mixed-criticality software programs running on multicore processors.
Irune Agirre, Mikel Azkarate-askatsua, Carles Hernández 0001, Jaume Abella 0001, Jon Pérez 0001, Tullio Vardanega, Francisco J. Cazorla
DSD3
2015 Enabling TDMA Arbitration in the Context of MBPTA
abstract
Current timing analysis techniques can be broadly classified into two families: deterministic timing analysis (DTA) and probabilistic timing analysis (PTA). Each family defines a set of properties to be provided (enforced) by the hardware and software platform so that valid Worst-Case Execution Time (WCET) estimates can be derived for programs running on that platform. However, the fact that each family relies on each own set of hardware designs limits their applicability and reduces the chances of those designs being adopted by hardware vendors. In this paper we show that Time Division Multiple Access (TDMA), one of the main DTA-compliant arbitration policies, can be made PTA-compliant. To that end, we analyze TDMA in the context of measurement-based PTA (MBPTA) and show that padding execution time observations conveniently leads to trustworthy and tight WCET estimates with MBPTA without introducing any hardware change. In fact, TDMA outperforms round-robin and time-randomized policies in terms of WCET in the context of MBPTA.
Milos Panic, Jaume Abella 0001, Carles Hernández 0001, Eduardo Quiñones, Theo Ungerer, Francisco J. Cazorla
DSD3
2015 CAP: Communication-Aware Allocation Algorithm for Real-Time Parallel Applications on Many-Cores
abstract
Critical Real-Time Embedded Systems (CRTES) require additional computing power to match the performance requirements of increasingly complex critical functions. Many-core processors are a key solution to reach the required performance. They allow simultaneous execution of multiple critical functions comprising a number of (parallel) applications. These applications, each implementing some functionality of the system, need to communicate among them in order to make the system work. In many-cores the impact of this communication on the timing behavior of applications depends on the allocation of applications across the chip. In this paper we propose CAP: an allocation algorithm that takes into account communication among applications and tries to reduce its impact on Worst Case Execution Time estimates of applications. We show that using CAP increases the number of applications that can be scheduled on the many-core platform thus facilitating system integration.
Milos Panic, Eduardo Quiñones, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
DSD3
2015 Characterizing fault propagation in safety-critical processor designs
abstract
Achieving reduced time-to-market in modern electronic designs targeting safety critical applications is becoming very challenging, as these designs need to go through a certification step that introduces a non-negligible overhead in the verification and validation process. To cope with this challenge, safety-critical systems industry is demanding new tools and methodologies allowing quick and cost-effective means for robustness verification. Microarchitectural simulators have been widely used to test reliability properties in different domains but their use in the process of robustness verification remains yet to be validated against other accepted methods such as RTL or gate-level simulation. In this paper we perform fault injections in an RTL model of a processor to characterize fault propagation. The results and conclusions of this characterization will serve to devise to what extent fault injection methodologies for robustness verification using microarchitectural simulators can be employed.
Jaime Espinosa, Carles Hernández 0001, Jaume Abella 0001
IOLTS2
2015 Timely Error Detection for Effective Recovery in Light-Lockstep Automotive Systems
abstract
Safety-relevant systems in the automotive domain often implement features such as lockstep execution for error detection, and reset and re-execution for error correction. Light-lockstep has already been adopted in some such systems due to its relatively low-implementation cost given that it does not require deep changes into nonlockstep hardware. Instead, as only off-core activities (i.e., data/addresses sent) need to be compared across different cores, light-lockstep designs are lowly intrusive. This approach has been proven sufficient to guarantee functional correctness of the system in the presence of errors in the cores, in particular in relation with certification against safety standards such as ISO26262 in the automotive domain. However, error detection in light-lockstep systems may occur long after the error actually occurs, thus jeopardizing timing guarantees, which are as critical as functional ones in hard real-time systems. In this paper, we analyze the timing behavior of errors due to transient and permanent faults in light-lockstep systems. Our results show that the time elapsed until an error is detected can be inordinately large, especially for permanent faults. Based on this observation and building upon the specific characteristics of light-lockstep systems, we propose lightly verbose (LiVe), a new mechanism to enforce the early detection of errors, due to both transient and permanent faults, thus enabling the computation of tight error detection timing bounds. We also analyze how existing mechanisms for error recovery in multicore systems increase their effectiveness when light-lockstep operates in LiVe mode in the context of mixed-criticality workloads.
Carles Hernández 0001, Jaume Abella 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2014 LiVe: Timely Error Detection in Light-Lockstep Safety Critical Systems
abstract
Safety-critical systems rely on features such as lockstep execution for error detection, and reset and reexecution for error correction. In particular, light lockstep is an attractive choice since it does not require redesigning cores but, instead, comparing off-core activities (i.e. data/addresses sent). While this approach suffices to guarantee functional correctness of the system, as needed for certification against safety standards (e.g., ISO26262), it fails to provide any timing guarantee as the time elapsed since the error occurs until lockstep detects it can be inordinately large.
Carles Hernández 0001, Jaume Abella 0001
DAC1
2014 Parallel many-core avionics systems
abstract
Integrated Modular Avionics (IMA) enables incremental qualification by encapsulating avionics applications into software partitions (SWPs), as defined by the ARINC 653 standard. SWPs, when running on top of single-core processors, provide robust time partitioning as a means to isolate SWPs timing behavior from each other. However, when moving towards parallel execution in many-core processors, the simultaneous accesses to shared hardware and software resources influence the timing behavior of SWPs, defying the purpose of time partitioning to provide isolation among applications. In this paper, we extend the concept of SWP by introducing parallel software partitions (pSWP) specification that describes the behavior of SWPs required when running in a many-core to enable incremental qualification. pSWP are supported by a new hardware feature called guaranteed resource partition (GRP) that defines an execution environment in which SWPs run and that controls interferences in the accesses to shared hardware resources among SWPs such that time composability can be guaranteed.
Milos Panic, Eduardo Quiñones, Pavel G. Zaykov, Carles Hernández 0001, Jaume Abella 0001, Francisco J. Cazorla
EMSOFT4
2013 Silicon-aware distributed switch architecture for on-chip networks
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
J. Syst. Archit.2
2012 Enabling High-Performance Crossbars through a Floorplan-Aware Design
abstract
Networks-on-Chip (NoC) with low-radix switches forming a simple and planar topology is typically accepted as the right interconnection infrastructure for current Chip Multi Processor and high-end Multi Processor System-on-Chip. This is mainly due to its simplicity in the physical mapping on the chip. However, as the network diameter increases, latency and power consumption are increased due to the rapidly growing queuing delay in each switch do not scale with system size. In this context, topologies with high-radix switches have been recently proposed in the NoC scenario to keep message latency low when interconnecting a large number of devices. However, the use of high-radix switches present several well-known drawbacks, being the most important the scalability in area and frequency when implemented. In addition, average and maximum wire length is increased. In this paper we present a distributed crossbar NoC architecture that reduces network latency, increases network throughput significantly. For the distributed crossbar implementation trees of 2-to-1 multiplexers with arbitration and buffer capabilities are spread over the chip, avoiding the negative impact of a high radix switch degree on NoC operating frequency, and minimizing the impact of long wires. Results show that in a 64-node NoC our most aggressive distributed crossbar configuration reduces flit latency by 42% and increases throughput by 544% with respect to the low latency flattened butterfly architecture, meanwhile area is increased a 110%. A more conservative distributed crossbar configuration obtains an increment in throughput of 276.1%, latency is decreased a 29.7%, but area is also decreased by 7%.
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
ICPP2
2012 On the Impact of Within-Die Process Variation in GALS-Based NoC Performance
abstract
Current integration scales allow designing chip multiprocessors (CMP), where cores are interconnected by means of a network-on-chip (NoC). Unfortunately, the small feature size of current integration scales causes some unpredictability in manufactured devices because of process variation. In NoCs, variability may affect links and routers causing them not to match the parameters established at design time. In this paper, we first analyze the way that manufacturing deviations affect the components of a NoC by applying a new comprehensive and detailed within-die variability model to 200 instances of an 8×8 mesh NoC synthesized using 45 nm technology. Later, we show that GALS-based NoCs present communication bottlenecks under process variation which cannot be avoided by using just device-level solutions but higher level architectural approaches are required. Therefore, to overcome this performance reduction, we draft a novel architectural approach, called performance domains, intended to reduce the negative impact of variability on application execution time. This mechanism is suitable when several applications are simultaneously running in the CMP chip.
Carles Hernández 0001, Antoni Roca 0001, Federico Silla, José Flich, José Duato
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2011 Energy and Performance Efficient Thread Mapping in NoC-Based CMPs under Process Variations
abstract
Within-die process variation causes cores, memories, and network resources in NoC-based CMPs to present different speeds and leakage power. In this context, thread mapping strategies that consider the effects of process variability on chip resources arise as a suitable choice to maximize performance while energy consumption constraints are satisfied. However, other factors, as the location of memory controllers and the concurrent execution of several applications in the chip, can bound the possible benefits of such mapping strategies. In this paper we propose a mapping strategy, named as uniform regions, that takes variability effects into account when assigning application threads to cores in the chip. More specifically, uniform regions, in terms of operating frequency, that additionally present the highest available frequency, are selected so that the benefits of such a variation-aware mapping strategy in a NoC-based CMP are maximized. We additionally present two different ways of configuring the frequency and voltage of the cores in the selected region. The first one is intended to provide the maximum performance while keeping energy as low as possible, while the second one is much more for energy-aware. The first one reduces the execution time up to a 23% while reducing the energy up to 24% whereas the second one provides smaller speed ups while reduces energy up to 33%.
Carles Hernández 0001, Federico Silla, José Duato
ICPP1
2011 A Distributed Switch Architecture for On-Chip Networks
abstract
It is well-known that current Chip Multiprocessor (CMP) and high-end MultiProcessor System-on-Chip (MPSoC) designs are growing in their number of components. Networks-on-Chip (NoC) provide the required connectivity for such CMP and MPSoC designs at reasonable costs. However, as technology advances, links become the critical component in the NoC. First, because the power consumption of the link is extremely high with respect the power consumption of the rest of components (mainly switches), becoming unacceptable for long global interconnects. Second, the delay of a link does not scale with technology, thus, degrading the performance of the network. To solve both problems, several solutions have been previously proposed. In this paper, we present a new switch architecture that reduces the negative impact of links on the NoC. We call our proposal distributed switch. The distributed switch moves the circuitry of a standard switch onto the links. Then, packets are buffered, routed, and forwarded at the same time they are crossing the link. Distributing a standard switch onto the link improves the trade off between the power consumption and the operating frequency of the entire network. In contrast, area requirements are increased. The distributed switch reduces up to 14.8% the peak power consumption while increases its area up to 22%. Furthermore, the distributed switch is able to increase the maximum achievable frequency with respect to the standard switch. In particular, the maximum operating frequency of the distributed switch can be increased up to 14.3%.
Antoni Roca 0001, Carles Hernández 0001, José Flich, Federico Silla, José Duato
ICPP2
2011 Characterizing the impact of process variation on 45 nm NoC-based CMPs
Carles Hernández 0001, Antoni Roca 0001, José Flich, Federico Silla, José Duato
J. Parallel Distributed Comput.1
2010 A methodology for the characterization of process variation in NoC links
abstract
Associated with the ever growing integration scales is the increase in process variability. In the context of network-on-chip, this variability affects the maximum frequency that could be sustained by each link that interconnects two cores in a chip multiprocessor. In this paper we present a methodology to model delay variations in NoC links. We also show its application to several technologies, namely 45nm, 32nm, 22nm, and 16nm. Simulation results show that conclusions about variability greatly depend on the implementation context.
Carles Hernández 0001, Federico Silla, José Duato
DATE1
2010 Improving the Performance of GALS-Based NoCs in the Presence of Process Variation
abstract
Current integration scales allow designing chip multiprocessors (CMP) where cores are interconnected by means of a network-on-chip (NoC). Unfortunately, the small feature size of current integration scales cause some unpredictability in manufactured devices because of process variation. In NoCs,variability may affect links and routers causing that they do not match the parameters established at design time. In this paper we first analyze the way that manufacturing deviations affect the components of a NoC by applying a comprehensive and detailed variability model to 200 instances of an 8x8 mesh NoC synthesized using 45nm technology. A second contribution of this paper is showing that GALS-based NoCs present communication bottlenecks under process variation. To overcome this performance reduction we draft a novel approach, called performance domains, intended to reduce the negative impact of variability on application execution time. This mechanism is suitable when several applications are simultaneously running in the CMP chip.
Carles Hernández 0001, Antoni Roca 0001, Federico Silla, José Flich, José Duato
NOCS1
2009 A new mechanism to deal with process variability in NoC links
abstract
Associated with the ever growing integration scale of VLSI technologies is the increase in process variability, which makes silicon devices to become less predictable. In the context of network-on-chip (NoC), this variability affects the maximum frequency that could be sustained by each wire of the link that interconnects two cores in a CMP system. Reducing the clock frequency so that all wires can properly work is a trivial solution but, as variability increases, this approach causes an unacceptable performance penalty. In this paper, we propose a new technique to deal with the effects of variability on the links of the NoC that interconnects cores in a CMP system. This technique, called Phit Reduction (PR), retrieves most of the bandwidth still available in links containing wires that are not able to operate at the designed operating frequency. More precisely, our mechanism discards these slow wires and uses all the wires that can work at the design frequency. Two implementations are presented: Local Phit Reduction (LPR), oriented to fabrication processes with very high variability, which requires more hardware but provides higher performance; and Global Phit Reduction (GPR), that requires less additional hardware but is not able to extract all the available bandwidth. The performance evaluation presented in the paper confirms that LPR obtains good results both for low and high variability scenarios. Moreover, in most of our experiments LPR practically achieves the same performance than the ideal network. On the other hand, GPR is appropriate for systems where whithin-die variations are expected to be low.
Carles Hernández 0001, Federico Silla, Vicente Santonja, José Duato
IPDPS1