VLDB 2026 Research / reviewers in the wild / expert
Alessandro Biondi 0001
dblp:153/0102
· DBLP profile ↗
94ranked-venue papers
16as first author
53since 2021 · last 2026
0000-0002-6625-9336ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 46 · 7 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Time-Predictable Acceleration of Deep Neural Networks on FPGA SoCs with Multi-Core DPUs
Federico Aromolo, Niko Salamini, Jacopo Del Granchio, Alessandro Biondi 0001, Mauro Marinoni, Giorgio C. Buttazzo |
RTAS | 4 |
| 2026 | The use of the Simplex architecture to enhance safety in deep-learning-powered autonomous systemsabstractRecently, the outstanding performance reached by neural networks in many tasks has led to their deployment in autonomous systems, such as robots and vehicles. However, neural networks are not yet trustworthy, being prone to different types of misbehavior, such as anomalous samples, distribution shifts, adversarial attacks, and other threats. Furthermore, frameworks for accelerating the inference of neural networks typically run on rich operating systems that are less predictable in terms of timing behavior and present larger surfaces for cyber-attacks. To address these issues, this paper presents a software architecture for enhancing safety, security, and predictability levels of learning-based autonomous systems. It leverages two isolated execution domains, one dedicated to the execution of neural networks under a rich operating system, which is deemed not trustworthy, and one responsible for running safety-critical functions, possibly under a different operating system capable of handling real-time constraints. Both domains are hosted on the same computing platform and isolated through a type-1 real-time hypervisor enabling fast and predictable inter-domain communication to exchange real-time data. The two domains cooperate to provide a fail-safe mechanism based on a safety monitor, which oversees the state of the system and switches to a simpler but safer backup module, hosted in the safety-critical domain, whenever its behavior is considered untrustworthy. The effectiveness of the proposed architecture is illustrated by a set of experiments performed on two control systems: a Furuta pendulum and a rover. The results confirm the utility of the fall-back mechanism in preventing faults due to the learning component. Federico Nesti, Niko Salamini, Mauro Marinoni, Giorgiomaria Cicero, Gabriele Serra, Alessandro Biondi 0001, Giorgio C. Buttazzo |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Benchmarking the spatial robustness of DNNs via natural and adversarial localized corruptions
Giulia Marchiori Pietrosanti, Giulio Rossolini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
Pattern Recognit. | 3 |
| 2025 | Modeling the SL-LET Paradigm in AUTOSAR AdaptiveabstractThe AUTOSAR consortium proposed the AUTOSAR Adaptive standard to tackle the challenges introduced by the design of modern automotive systems. It consists of a service-oriented architecture (SoA) implemented in C++ and built on top of POSIX operating systems. However, unlike the previous AUTOSAR Classic specifications, this novel standard does not address non-functional requirements, including determinism, which is of key importance to guarantee the system's functional safety. This paper proposes a modeling extension to the AUTOSAR Adaptive standard aiming at guaranteeing a deterministic execution by leveraging the System-Level Logical Execution Time (SL-LET) paradigm, already used in the context of AUTOSAR Classic. A prototype implementation is also proposed, which is used to experimentally corroborate the feasibility of the proposed model extension with an evaluation based on a realistic automotive application built on the official AUTOSAR Adaptive Platform Demonstrator (APD). Davide Bellassai, Gerlando Sciangula, Claudio Scordino, Daniel Casini, Alessandro Biondi 0001 |
DATE | 5 |
| 2025 | Enabling Containerisation of Distributed Applications with Real-Time ConstraintsabstractContainerisation is becoming a cornerstone of modern distributed systems, thanks to their lightweight virtualisation, high portability, and seamless integration with orchestration tools such as Kubernetes. The usage of containers has also gained traction in real-time cyber-physical systems, such as software-defined vehicles, which are characterised by strict timing requirements to ensure safety and performance. Nevertheless, ensuring real-time execution of co-located containers is challenging because of mutual interference due to the sharing of the same processing hardware. Existing parallel computing frameworks such as Ray and its Kubernetes-enabled variant, KubeRay, excel in distributed computation but lack support for scheduling policies that allow guaranteeing real-time timing constraints and CPU resource isolation between containers, such as the SCHED_DEADLINE policy of Linux. To fill this gap, this paper extends Ray to support real-time containers that leverage SCHED_DEADLINE. To this end, we propose KubeDeadline, a novel, modular Kubernetes extension to support SCHED_DEADLINE. We evaluate our approach through extensive experiments, using synthetic workloads and a case study based on the MobileNet and EfficientNet deep neural networks. Our evaluation shows that KubeDeadline ensures deadline compliance in all synthetic workloads, adds minimal deployment overhead (in the order of milliseconds), and achieves lower worst-case response times, up to 4 times lower, than vanilla Kubernetes under background interference. Nasim Samimi, Luca Abeni, Daniel Casini, Mauro Marinoni, Twan Basten, Mitra Nasri, Marc Geilen, Alessandro Biondi 0001 |
ECRTS | 8 |
| 2025 | A Design Flow to Securely Isolate FPGA Bus Transactions in Heterogeneous SoCsabstractEmbedded computing systems are becoming increasingly complex. Modern system-on-chips come with heterogeneous designs that integrate diverse processing systems and a large variety of peripherals. When considering software with mixed and independent security and criticality levels, the heterogeneity of modern computing platforms poses considerable challenges in achieving strong isolation between execution domains. Tackling these challenges is even more difficult in platforms that integrate Field-Programmable Gate Array (FPGA) fabrics, which, due to their wide flexibility, introduce new security- and safety-related threats that can jeopardize isolation. As a matter of fact, if no proper countermeasures are in place, hardware accelerators (HAs) deployed on FPGA can be exploited to break the isolation capabilities implemented in a system by issuing dangerous bus transactions. This research proposes a design flow for heterogeneous platforms to strongly isolate bus transactions issued by HAs. The design flow is then specialized for the AMD Zynq UltraScale+ platform, leveraging the virtualizationrelated features of the Arm System Memory Management Unit (SMMU). The proposed solution jointly combines two new IPs for enforcing information transported by the AXI bus, a tool to verify the FPGA design, a principled configuration of the SMMU driver, and a secure boot flow. The proposal is evaluated with an industry-relevant use case related to embedded machine learning applied for the railway domain, in which isolation is established between two AMD Deep Learning Processor Units (DPU) and a set of FPGA HAs dedicated to a real-time critical application. Niko Salamini, Sara Alonso Salazar, Gabriele Serra, Giorgiomaria Cicero, Pietro Fara, Federico Aromolo, Alessandro Biondi 0001 |
RTAS | 7 |
| 2025 | Real-Time Multitasking of Deep Neural Networks With Nvidia TensorrtabstractGraphics processing units (GPUs) are often employed to accelerate the inference of deep neural networks (DNNs) in cyber-physical systems to implement advanced perception and control functionalities. Frameworks for GPU-accelerated DNN inference typically aim at maximizing the processing throughput rather than focusing on providing a predictable timing behavior, which is crucial for time-sensitive cyber-physical systems. This work proposes a framework for GPU-accelerated inference of DNNs on GPU-based embedded platforms in multitasking scenarios, which provides enhanced timing predictability using a design-time optimization procedure of the DNN workload and a specialized method to schedule the GPU acceleration requests of the DNNs at runtime based on fixed-priority limitedpreemptive scheduling. Fine-grained control of the inference is achieved by splitting the DNNs into smaller chunks, which are then scheduled using a specialized real-time scheduling mechanism. Experimental results on commercial embedded platforms report significant improvements in terms of schedulability. Federico Aromolo, Andrea Stevanato, Alessandro Biondi 0001, Giorgio C. Buttazzo |
RTSS | 3 |
| 2025 | Requirement-Based Analysis of Self-Suspending Tasks under EDFabstractWhile preemptive Earliest-Deadline-First (EDF) has been studied extensively in real-time systems, there are only few results when considering tasks with dynamic self-suspension behavior scheduled under EDF. Furthermore, all schedulability tests that have been developed in this context are based on analyzing specific intervals, hindering the performance of the analytical tightness of the result. In this work, we develop a schedulability test for EDF, built on a dynamic interval extension. That is, whenever the analysis cannot derive a decision to conclude the schedulability test, we iteratively extend the analysis interval to include additional carry-in jobs into the analysis. This is achieved by specifying execution-exceedance requirement for infeasibility of the system, i.e., by specifying how much workload must be accumulated within a certain time interval to achieve a deadline miss. Our approach outperforms all previous analyses and is the first to surpass the schedulability guarantees that can be provided for Deadline-Monotonic (DM) scheduling for dynamic self-suspending tasks, hence achieving a milestone in the analysis of EDF scheduling. Mario Günzel, Federico Aromolo, Alessandro Biondi 0001, Jian-Jia Chen |
RTSS | 3 |
| 2025 | Exploiting Inaccurate Branch History in Side-Channel Attacks
Yuhui Zhu, Alessandro Biondi 0001 |
USENIX Security Symposium | 2 |
| 2025 | AP-LET: Enabling deterministic Pub/Sub communication in AUTOSAR AdaptiveabstractThe automotive software industry is facing a paradigm shift driven by the need to develop more and more advanced functionality distributed on multiple electronic control units. The AUTOSAR Adaptive standard has been designed as a service-oriented architecture on top of a general-purpose operating system to tackle this paradigm shift. Nevertheless, it does not provide means to ensure deterministic communication, as required in safety-related components. This paper studies the integration of the System-Level Logical Execution Time (SL-LET) paradigm in AUTOSAR Adaptive. The key design challenges and requirements to support SL-LET in AUTOSAR Adaptive are described, highlighting how to overcome the considerable differences between the AUTOSAR Classic and Adaptive domains. Then, a meta-protocol named AP-LET is presented, together with two concrete instances: one based on high-priority tasks and another leveraging timestamps in the message payload to handle communications and ensure determinism. A complete implementation of both protocols is also described. AP-LET was finally evaluated with a realistic automotive application, showing its feasibility and effectiveness. Davide Bellassai, Claudio Scordino, Daniel Casini, Alessandro Biondi 0001 |
J. Syst. Archit. | 4 |
| 2025 | Time synchronization and performance analysis of the openSAFETY protocol via UDP over EthernetabstractThe growing demand for Ethernet-based Industrial Internet of Things (IIoT) is changing the shape of modern industrial systems and emphasizing the need for high-speed, reliable, scalable, and safe communication among industrial devices. Ethernet-based networks provide the basis for seamless device integration, real-time data exchange, and increased operational efficiency, making them the key to Industry 5.0 applications. As industrial automation becomes increasingly complex, the importance of functional safety grows exponentially. The openSAFETY protocol is a fieldbus-independent, scalable, and robust protocol for implementing functional safety. Our contribution is twofold. First, we analyze time synchronization in the openSAFETY to fully understand the interrelated timing parameters and give some practical guidelines to tune the safety application. We have proposed the parameter tuning approach, which is better in terms of performance and ensures continuous, safe operations. Second, we analyze the protocol’s performance via UDP over Ethernet under normal and degraded network conditions. We found the protocol resilient to network impairments under certain levels during the experiments. Under normal working conditions, the cycle time was successfully achieved in the microsecond range, even at full payload capacity. • Importance of functional safety and role of openSAFETY Protocol in industrial automation. • In-depth analysis of the time synchronization mechanism in the openSAFETY protocol. • Deriving the Configuration Parameters for tuning the Safety Application. • Performance analysis of the protocol under normal and degraded network conditions via UDP over Ethernet. Shoaib Zafar, Salvatore Sabina, Alessandro Biondi 0001, Giorgio C. Buttazzo |
J. Syst. Archit. | 3 |
| 2025 | AXI-REALM: Safe, Modular and Lightweight Traffic Monitoring and Regulation for Heterogeneous Mixed-Criticality SystemsabstractThe automotive industry is transitioning from federated, homogeneous, interconnected devices to integrated, heterogeneous, mixed-criticality systems (MCS). This leads to challenges in achieving timing predictability techniques due to access contention on shared resources, which can be mitigated using hardware-based spatial and temporal isolation techniques. Focusing on the interconnect as the point of access for shared resources, we propose AXI-REALM, a lightweight, modular, technology-independent, and open-source real-time extension to AXI4 interconnects. AXI-REALM uses a budget-based mechanism enforced on periodic time windows and transfer fragmentation to provide fair arbitration, coupled with execution predictability on real-time workloads. AXI-REALM features a comprehensive bandwidth and latency monitor at both the ingress and egress of the interconnect system. Latency information is also used to detect and reset malfunctioning subordinates, preventing missed deadlines. We provide a detailed cost assessment in a 12nm node and an end-to-end case study implementing AXI-REALM into an open-source MCS, incurring an area overhead of less than 2 %. When running a mixed-criticality workload, with a time-critical application sharing the interconnect with non-critical applications, we demonstrate that the critical application can achieve up to 68.2% of the isolated performance by enforcing fairness on the interconnect traffic through burst fragmentation, thus reducing the subordinate access latency by up to 24 times. Near-ideal performance, (above 95% of the isolated performance) can be achieved by distributing the available bandwidth in favor of the critical application. Thomas Benz, Alessandro Ottaviano, Chaoqun Liang, Robert Balas, Angelo Garofalo, Francesco Restuccia 0002, Alessandro Biondi 0001, Davide Rossi 0001, Luca Benini |
IEEE Trans. Computers | 7 |
| 2025 | Hazardous Behavior Analysis for Complex Pre-Existing Software in Safety-Critical Context: Approach and Application to the Linux Dynamic Memory AllocatorabstractAs system designers are transitioning to the usage of pre-existing software architectural elements to reduce time-to-market and costs, they face challenges in safety-critical applications. Pre-existing software may have been implemented without following any safety and/or quality standard, and its documentation may be incomplete or unclear. As recommended by part 6 of the functional safety standard ISO 26262 (product development at the software level), stringent rules and guidelines must be followed during the implementation of a new software. The qualification of a pre-existing software according to ISO 26262 may hence be very time consuming and expensive if not addressed in a structured way. This work presentsSTPA for Pre-existing Software(STPA-PES), a new software hazardous behavior analysis approach designed to identify criticalities or abnormal conditions in complex pre-existing software to be mitigated with adequate safety measures. The major challenge in such a process consists in the definition of a model that correctly represents the system under analysis when the architectural design of pre-existing software is not available. As a relevant example, the proposed approach is finally applied to dynamic memory allocator (DynMA) in Linux kernel to show how STPA-PES is able to derive safety requirements of this software architectural element. Raffaele Giannessi, Fabrizio Tronci, Alessandro Biondi 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | AXI-REALM: A Lightweight and Modular Interconnect Extension for Traffic Regulation and Monitoring of Heterogeneous Real-Time SoCsabstractThe increasing demand for heterogeneous functionality in the automotive industry and the evolution of chip manufac-turing processes have led to the transition from federated to integrated critical real-time embedded systems (CRTESs). This leads to higher integration challenges of conventional timing predictability techniques due to access contention on shared resources, which can be resolved by providing system-level observability and controllability in hardware. We focus on the interconnect as a shared resource and propose AXI-REALM, a lightweight, modular, and technology - independent real-time extension to industry-standard AXI4 interconnects, available open-source. AXI-REALM uses a credit-based mechanism to distribute and control the bandwidth in a multi-subordinate system on periodic time windows, proactively prevents denial of service from malicious actors in the system, and tracks each manager's access and interference statistics for optimal budget and period selection. We provide detailed performance and implementation cost assessment in a 12nm node and an end-to-end functional case study implementing AXI-REALM into an open-source Linux-capable RISC-V SoC. In a system with a general-purpose core and a hardware accelerator's DMA engine causing interference on the interconnect, AXI-REALM achieves fair bandwidth distribution among managers, allowing the core to recover 68.2 % of its performance compared to the case without contention. Moreover, near-ideal performance (above 95 %) can be achieved by distributing the available bandwidth in favor of the core, improving the worst-case memory access latency from 264 to below eight cycles. Our approach minimizes buffering compared to other solutions and introduces only 2.45 % area overhead compared to the original SoC. Thomas Benz, Alessandro Ottaviano, Robert Balas, Angelo Garofalo, Francesco Restuccia 0002, Alessandro Biondi 0001, Luca Benini |
DATE | 6 |
| 2024 | End-to-End Latency Optimization of Thread Chains Under the DDS Publish/Subscribe MiddlewareabstractModern autonomous systems integrate diverse soft-ware solutions to manage tightly communicating functionalities. These applications commonly communicate using frameworks implementing the publish/subscribe paradigm, such as the Data Distribution Service (DDS). However, these frameworks are real-ized with a multi-threaded software architecture and implement internal policies for message dispatching, posing additional chal-lenges for guaranteeing timing constraints. This work addresses the problem of optimizing a DDS-based interconnected real-time systems, proposing analysis-driven algorithms to set a vast range of parameters, ranging from classical thread priorities to other DDS-specific configurations. We evaluate our approaches on the Autoware Reference System, a realistic testbed from the Autoware autonomous driving framework. Gerlando Sciangula, Daniel Casini, Alessandro Biondi 0001, Claudio Scordino |
DATE | 3 |
| 2024 | Optimizing Per-Core Priorities to Minimize End-To-End Latencies
Francesco Paladino, Alessandro Biondi 0001, Enrico Bini, Paolo Pazzaglia |
ECRTS | 2 |
| 2024 | RT-Mimalloc: A New Look at Dynamic Memory Allocation for Real-Time SystemsabstractDynamic memory allocation is a pivotal feature of modern software systems but has mostly been scarcely used in real-time systems due to the limited time-predictability offered by dynamic memory allocators (DynMAs). While many general-purpose DynMAs have been proposed during the last decades, only a few efforts were devoted to the design of real-time DynMAs capable of providing bounded allocation times. Furthermore, the most notable of them dates back to almost 20 years ago. Motivated by this observation and the significant developments made in the field of general-purpose DynMAs in recent years, this work takes a new look at dynamic memory allocation for real-time systems. After analyzing and comparing modern DynMAs, we discuss how to modify the Mimalloc general-purpose DynMA into RT-Mimalloc, so that more predictable allocation times can be obtained. All the studies and evaluations performed in this work were based on both modern state-of-the-art benchmarks for memory allocation and synthetic workload to assess specific capabilities of the tested DynMAs. The evaluation showed that RT-Mimalloc is capable to improve the longest-observed allocation times of real-time DynMAs proposed in previous work while retaining most of the benefits of modern general-purpose DynMAs in terms of average-case performance. Raffaele Giannessi, Alessandro Biondi 0001, Alessandro Biasci |
RTAS | 2 |
| 2024 | In Search of Butterflies: Exceedance Analysis for Real-Time Systems under Transient OverloadabstractIn theory, real-time systems are provisioned based on provably sound worst-case execution times (WCETs), but in practice often only empirically derived, unsound execution-time estimates—i.e., nominal execution times (NETs)—are available since WCETs are difficult to obtain on modern hardware. NETs pose two significant challenges: First, since NETs may be exceeded at runtime, any response-time bounds derived from NETs are transitively unsound and may be violated. Second, even a minuscule NET violation can result in large, nonlinear response-time increases due to hard-to-predict, cascading scheduling effects. To explore the risk NET exceedance poses to a system’s temporal correctness, this paper provides the first general, systematic, and explainable methodology for exceedance analysis. The proposed approach supports fixed-priority (FP), earliest-deadline first (EDF), and first-in first-out (FIFO) scheduling on a uniprocessor or within a partitioned multiprocessor platform, and the full spectrum of preemption models from fully preemptive to fully non-preemptive workloads. Additionally, it produces explainable evidence in the form of tunable example traces that engineers can adjust to take system-specific expertise into account. The proposed methodology is evaluated with synthetic task sets and workloads based on an automotive benchmark, and in a case study applied to parts of the WATERS’17 industrial challenge. Matteo Zini, Filip Markovic 0001, Daniel Casini, Alessandro Biondi 0001, Björn B. Brandenburg |
RTSS | 4 |
| 2024 | Learning Memory-Contention Timing Models With Automated Platform ProfilingabstractCommercial off-the-shelf (COTS) multicore platforms are often used to enable the execution of mixed-criticality real-time applications. In these systems, the memory subsystem is one of the most notable sources of interference and unpredictability, with the memory controller (MC) being a key component orchestrating the data flow between processing units and main memory. The worst-case response times of real-time tasks is indeed particularly affected by memory contention and, in turn, by the MC behavior as well. This article presents FrATM2, a Framework to Automatically learn the Timing Models of the Memory subsystem. The framework automatically generates and executes micro-benchmarks on bare-metal hardware to profile the platform behavior in a large number of memory-contention scenarios. After aggregating and filtering the collected measurements, FrATM2 trains MC models to bound memory-related interference. The MC models can be used to enable response-time analysis. The framework was evaluated on an AMD/Xilinx Ultrascale+ SoC, collecting gigabytes of raw experimental data by testing tents of thousands of contention scenarios. Andrea Stevanato, Matteo Zini, Alessandro Biondi 0001, Bruno Morelli, Alessandro Biasci |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | CARLA-GeAR: A Dataset Generator for a Systematic Evaluation of Adversarial Robustness of Deep Learning Vision ModelsabstractAdversarial examples represent a serious threat for deep neural networks in several application domains and a huge amount of work has been produced to investigate them and mitigate their effects. Nevertheless, no much work has been devoted to the generation of datasets specifically designed to evaluate the adversarial robustness of neural models. This paper presents CARLA-GeAR, a tool for the automatic generation of photo-realistic synthetic datasets related to driving scenarios that can be used for a systematic evaluation of the adversarial robustness of neural models against physical adversarial patches, as well as for comparing the performance of different adversarial defense/detection methods. The tool is built on the CARLA simulator, using its Python API, and allows the generation of datasets for several vision tasks in the context of autonomous driving. The adversarial patches included in the generated datasets are attached to billboards or the back of a truck and are crafted by using state-of-the-art white-box attack strategies to maximize the prediction error of the model under test. Finally, the paper presents an experimental study to evaluate the performance of some defense methods against such attacks, showing how the datasets generated with CARLA-GeAR might be used in future work as a benchmark for adversarial defense in the real world. All the code and datasets used in this paper are available athttps://carlagear.retis.santannapisa.it. Federico Nesti, Giulio Rossolini, Gianluca D'Amico, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | On the Real-World Adversarial Robustness of Real-Time Semantic Segmentation Models for Autonomous DrivingabstractThe existence of real-world adversarial examples (RWAEs) (commonly in the form of patches) poses a serious threat for the use of deep learning models in safety-critical computer vision tasks such as visual perception in autonomous driving. This article presents an extensive evaluation of the robustness of semantic segmentation (SS) models when attacked with different types of adversarial patches, including digital, simulated, and physical ones. A novel loss function is proposed to improve the capabilities of attackers in inducing a misclassification of pixels. Also, a novel attack strategy is presented to improve the expectation over transformation (EOT) method for placing a patch in the scene. Finally, a state-of-the-art method for detecting adversarial patch is first extended to cope with SS models, then improved to obtain real-time performance, and eventually evaluated in real-world scenarios. Experimental results reveal that even though the adversarial effect is visible with both digital and real-world attacks, its impact is often spatially confined to areas of the image around the patch. This opens to further questions about the spatial robustness of real-time SS models. Giulio Rossolini, Federico Nesti, Gianluca D'Amico, Saasha Nair, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Robust-by-Design Classification via Unitary-Gradient Neural NetworksabstractThe use of neural networks in safety-critical systems requires safe and robust models, due to the existence of adversarial attacks. Knowing the minimal adversarial perturbation of any input x, or, equivalently, knowing the distance of x from the classification boundary, allows evaluating the classification robustness, providing certifiable predictions. Unfortunately, state-of-the-art techniques for computing such a distance are computationally expensive and hence not suited for online applications. This work proposes a novel family of classifiers, namely Signed Distance Classifiers (SDCs), that, from a theoretical perspective, directly output the exact distance of x from the classification boundary, rather than a probability score (e.g., SoftMax). SDCs represent a family of robust-by-design classifiers. To practically address the theoretical requirements of an SDC, a novel network architecture named Unitary-Gradient Neural Network is presented. Experimental results show that the proposed architecture approximates a signed distance classifier, hence allowing an online certifiable classification of x at the cost of a single inference. Fabio Brau, Giulio Rossolini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
AAAI | 3 |
| 2023 | Defending from Physically-Realizable Adversarial Attacks through Internal Over-Activation AnalysisabstractThis work presents Z-Mask, an effective and deterministic strategy to improve the adversarial robustness of convolutional networks against physically-realizable adversarial attacks. The presented defense relies on specific Z-score analysis performed on the internal network features to detect and mask the pixels corresponding to adversarial objects in the input image. To this end, spatially contiguous activations are examined in shallow and deep layers to suggest potential adversarial regions. Such proposals are then aggregated through a multi-thresholding mechanism. The effectiveness of Z-Mask is evaluated with an extensive set of experiments carried out on models for semantic segmentation and object detection. The evaluation is performed with both digital patches added to the input images and printed patches in the real world. The results confirm that Z-Mask outperforms the state-of-the-art methods in terms of detection accuracy and overall performance of the networks under attack. Furthermore, Z-Mask preserves its robustness against defense-aware attacks, making it suitable for safe and secure AI applications. Giulio Rossolini, Federico Nesti, Fabio Brau, Alessandro Biondi 0001, Giorgio C. Buttazzo |
AAAI | 4 |
| 2023 | Replication-Based Scheduling of Parallel Real-Time Tasks
Federico Aromolo, Geoffrey Nelissen, Alessandro Biondi 0001 |
ECRTS | 3 |
| 2023 | Bounding the Data-Delivery Latency of DDS Messages in Real-Time Applications
Gerlando Sciangula, Daniel Casini, Alessandro Biondi 0001, Claudio Scordino, Marco Di Natale |
ECRTS | 3 |
| 2023 | Virtualized DDS Communication for Multi-Domain Systems: Architecture and Performance Evaluation of Design AlternativesabstractModern applications for cyber-physical systems, such as autonomous driving, are more and more often characterized by the interconnection of software components with mixed levels of safety and security deployed on the same hardware platform. Virtualization by means of hypervisor technology is notably the most common approach to allow for the integration of such software components upon the same platform. The Data Distribution Service (DDS) is a publisher/subscriber middleware protocol that is establishing as a reference solution to put in communication distributed software components. Given these developments, the need for efficiently supporting DDS-based communication in virtualized systems is emerging in several industrial fields, especially when DDS communications interest in-platform software components. This paper presents the design of a virtualized DDS communication architecture for multidomain systems based on para-virtualization and hypervisor technology. The design is then specialized and implemented for the popular Xen hypervisor and Linux operating system, under which a set of implementation options are systematically studied and compared with a wide experimental evaluation. Andrea Stevanato, Alessandro Biondi 0001, Alessandro Biasci, Bruno Morelli |
RTAS | 2 |
| 2023 | Supporting logical execution time in multi-core POSIX systems
Davide Bellassai, Alessandro Biondi 0001, Alessandro Biasci, Bruno Morelli |
J. Syst. Archit. | 2 |
| 2023 | On the Minimal Adversarial Perturbation for Deep Neural Networks With Provable Estimation ErrorabstractAlthough Deep Neural Networks (DNNs) have shown incredible performance in perceptive and control tasks, several trustworthy issues are still open. One of the most discussed topics is the existence of adversarial perturbations, which has opened an interesting research line on provable techniques capable of quantifying the robustness of a given input. In this regard, the euclidean distance of the input from the classification boundary denotes a well-proved robustness assessment as the minimal affordable adversarial perturbation. Unfortunately, computing such a distance is highly complex due the non-convex nature of DNNs. Despite several methods have been proposed to address this issue, to the best of our knowledge, no provable results have been presented to estimate and bound the error committed. This paper addresses this issue by proposing two lightweight strategies to find the minimal adversarial perturbation. Differently from the state-of-the-art, the proposed approach allows formulating an error estimation theory of the approximate distance with respect to the theoretical one. Finally, a substantial set of experiments is reported to evaluate the performance of the algorithms and support the theoretical findings. The obtained results show that the proposed strategies approximate the theoretical distance for samples close to the classification boundary, leading to provable robustness guarantees against any adversarial attacks. Fabio Brau, Giulio Rossolini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Supporting AI-powered real-time cyber-physical systems on heterogeneous platforms via hypervisor technologyabstractAbstract The heavy use of machine learning algorithms in safety-critical systems poses serious questions related to safety, security, and predictability issues, requiring novel architectural approaches to guarantee such properties. This paper presents an architecture solution that leverages heterogeneous platforms and virtualization technologies to support AI-powered applications consisting of modules with mixed criticalities and safety requirements. The hypervisor exploits the security features of the Xilinx ZCU104 MPSoCs to create two isolated execution environments: a high performance domain running deep learning algorithms under the Linux operating system and a safety-critical domain running control and monitoring functions under the freeRTOS real-time operating system. The proposed approach is validated by a use case consisting of an unmanned aerial vehicle capable of tracking moving targets using a deep neural network accelerated on the FGPA available on the platform. Edoardo Cittadini, Mauro Marinoni, Alessandro Biondi 0001, Giorgiomaria Cicero, Giorgio C. Buttazzo |
Real Time Syst. | 3 |
| 2023 | Optimizing Inter-Core Communications Under the LET Paradigm using DMA EnginesabstractModern automotive applications are increasingly characterized by the need to transfer massive amounts of data in a predictable and deterministic way, possibly leveraging the Logical Execution Time (LET) paradigm. However, current proposals for LET communications are limited to core-commanded data transfers, which may result in large delays for data-intensive systems. To address this issue, we explore the use of Direct Memory Access (DMA) to handle LET communication with improved parallelism. Each DMA transfer operates on a contiguous memory area, thus calling for an optimized memory mapping to maximize performance. Modern DMA engines offer also advanced configurations, such as linked-lists of data transfers, which may provide more flexibility at the expenses of an increased (initial) programming overhead. Leveraging all such features of DMA engines, we propose a set of designs and protocols for LET communications with trade-offs between latency and space requirements. For each option we present the formulation to compute the optimal scheduling and memory allocation solution as a mixed-integer linear programming problem. Experimental results show the feasibility of the approach and a comparison of the solutions obtained using the proposed methods, showing a considerable improvement in terms of data acquisition latency when compared to LET communication without DMA. Paolo Pazzaglia, Daniel Casini, Alessandro Biondi 0001, Marco Di Natale |
IEEE Trans. Computers | 3 |
| 2023 | Bounding Memory Access Times in Multi-Accelerator Architectures on FPGA SoCsabstractModern FPGA System-on-Chips (SoCs) embed large FPGA logics capable of hosting multiple hardware accelerators. Typically, hardware accelerators require direct access to the shared DRAM memory for reaching the high performance demanded by modern applications. In commercial FPGA SoCs, this goal is achieved by interconnecting the hardware accelerators on an interconnect based on AMBA AXI, which is the de-facto industrial standard for on-chip communications. The AXI standard provides great flexibility in the definition of the network topology. Nevertheless, such flexibility generates a significant unpredictability when attempting to bound the hardware accelerators’ response time when executing under contention. This work focus on bounding the worst-case memory access time of hardware accelerators deployed on commercial FPGA SoCs. We propose a modeling and analysis technique to bound the response time of the hardware accelerators and evaluate the schedulability of a system applicable to arbitrary AXI-based bus structures deployed on FPGA SoCs. Our results are validated on real execution traces collected on two popular FPGA SoCs belonging to the Xilinx ZYNQ-7000 and Zynq-Ultrascale+ families and by simulated results. Francesco Restuccia 0002, Marco Pagani, Alessandro Biondi 0001, Mauro Marinoni, Giorgio C. Buttazzo |
IEEE Trans. Computers | 3 |
| 2023 | Analyzing Arm's MPAM From the Perspective of Time PredictabilityabstractWith heterogeneous multi-core platforms being crucial to execute the highly demanding workloads of modern applications, memory-access predictability remains a key issue for the system's safety. Many solutions have been proposed over the years, but none has been applied on a large scale. Nowadays, we are in front of an unprecedented opportunity to have an impact on commercial platforms: the Memory System Resource Partitioning and Monitoring (MPAM) specification by Arm, which describes different memory-access regulation mechanisms, presenting a valuable industrial attempt to address this issue. However, several points of the specification are described at a high level only, leaving plenty of room for interpretation to hardware manufacturers. This paper takes a close look at the memory-access regulation mechanisms in the MPAM specification and provides some detailed instantiations of such mechanisms. A fine-grained memory contention analysis is presented for each of them to finally enable a comparison of their worst-case performance. Matteo Zini, Daniel Casini, Alessandro Biondi 0001 |
IEEE Trans. Computers | 3 |
| 2023 | Detecting Adversarial Examples by Input Transformations, Defense Perturbations, and VotingabstractOver the past few years, convolutional neural networks (CNNs) have proved to reach superhuman performance in visual recognition tasks. However, CNNs can easily be fooled by adversarial examples (AEs), i.e., maliciously crafted images that force the networks to predict an incorrect output while being extremely similar to those for which a correct output is predicted. Regular AEs are not robust to input image transformations, which can then be used to detect whether an AE is presented to the network. Nevertheless, it is still possible to generate AEs that are robust to such transformations. This article extensively explores the detection of AEs via image transformations and proposes a novel methodology, called defense perturbation, to detect robust AEs with the same input transformations the AEs are robust to. Such a defense perturbation is shown to be an effective counter-measure to robust AEs. Furthermore, multinetwork AEs are introduced. This kind of AEs can be used to simultaneously fool multiple networks, which is critical in systems that use network redundancy, such as those based on architectures with majority voting over multiple CNNs. An extensive set of experiments based on state-of-the-art CNNs trained on the Imagenet dataset is finally reported. Federico Nesti, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Increasing the Confidence of Deep Neural Networks by Coverage AnalysisabstractThe great performance of machine learning algorithms and deep neural networks in several perception and control tasks is pushing the industry to adopt such technologies in safety-critical applications, as autonomous robots and self-driving vehicles. At present, however, several issues need to be solved to make deep learning methods more trustworthy, predictable, safe, and secure against adversarial attacks. Although several methods have been proposed to improve the trustworthiness of deep neural networks, most of them are tailored for specific classes of adversarial examples, hence failing to detect other corner cases or unsafe inputs that heavily deviate from the training samples. This paper presents a lightweight monitoring architecture based on coverage paradigms to enhance the model robustness against different unsafe inputs. In particular, four coverage analysis methods are proposed and tested in the architecture for evaluating multiple detection logic. Experimental results show that the proposed approach is effective in detecting both powerful adversarial examples and out-of-distribution inputs, introducing limited extra-execution time and memory requirements. Giulio Rossolini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Software Eng. | 2 |
| 2022 | Placement of Chains of Real-Time Tasks on Heterogeneous Platforms under EDF SchedulingabstractWhen designing a real-time system, application architects are called to settle many non-trivial decisions that may severely influence the system's performance. With modern hardware platforms always being more and more complex and equipped with heterogeneous processor cores or even hardware accelerators such as TPUs, FPGAs, or GPUs, the complexities to be faced by application architects are exacerbated. Therefore, they are called to wisely allocate the computational resources provided by the hardware platform to application tasks in such a way to meet timing requirements and optimize other goals such as energy consumption. This paper proposes a mixed-integer linear programming formulation (MILP) to solve the task-to-heterogeneous-cores allocation problem while guaranteeing the schedulability of a real-time application running on the platform under partitioned Earliest Deadline First (EDF) scheduling. A new method to derive approximate worst-case response-time bounds is also presented and leveraged to setup the MILP formu-lation, which allows computing and minimizing the end-to-end latency of processing chains and considers energy requirements. The approach is evaluated on a task set based on the WATERS 2019 Industrial Challenge proposed by Bosch. Daniel Casini, Alessandro Biondi 0001 |
DSD | 2 |
| 2022 | Hardware Acceleration of Deep Neural Networks for Autonomous Driving on FPGA-based SoCabstractIn the last decade, enormous and renewed attention to Artificial Intelligence has emerged thanks to Deep Neural Networks (DNNs), which can achieve high performance in performing specific tasks at the cost of a high computational complexity. GPUs are commonly used to accelerate DNNs, but generally determine a very high power consumption and poor time predictability. For this reason, GPUs are becoming less attractive for resource-constrained, real-time systems, while there is a growing demand for specialized hardware accelerators that can better fit the requirements of embedded systems. Following this trend, this paper focuses on hardware acceleration for the DNNs used by Baidu Apollo, an open-source autonomous driving framework. As an experience report of performing R&D with industrial technologies, we discuss challenges faced in shifting from GPU-based to FPGA-based DNN acceleration when per-formed using the DPU core by Xilinx deployed on an Ultrascale+ SoC FPG A platform. Furthermore, it shows pros and cons of today's hardware accelerating tools. Experimental evaluations were conducted to evaluate the performance of FPGA-accelerated DNNs in terms of accuracy, throughput, and power consumption, in comparison with those achieved on embedded GPUs. Gerlando Sciangula, Francesco Restuccia 0002, Alessandro Biondi 0001, Giorgio C. Buttazzo |
DSD | 3 |
| 2022 | Response-Time Analysis for Self-Suspending Tasks Under EDF Scheduling
Federico Aromolo, Alessandro Biondi 0001, Geoffrey Nelissen |
ECRTS | 2 |
| 2022 | PAC-PL: Enabling Control-Flow Integrity with Pointer Authentication in FPGA SoC PlatformsabstractControl-flow integrity (CFI) is an effective technique to enhance the security of software systems. Processor designers recently started to provide hardware-based support to efficiently implement CFI, such as the pointer authentication (PA) feature provided by ARM starting from ARMv8.3-A processor architectures. These CFI mechanisms are also accompanied by support in the mainline codebase of popular compilers (such as GCC and LLVM) and the Linux operating system. As such, they are expected to establish as widespread security mechanisms. Nevertheless, many commercial chips still do not support hardware-assisted CFI, even some of the ones that just entered the market. This paper presents PAC-PL, a solution to enable hardware-assisted CFI on heterogeneous platforms that include a field-programmable gate array (FPGA) fabric, such as the Xilinx Ultrascale+ and Versal. PAC-PL comes with compiler-and OS-level support, is compatible with ARM’s PA, and enables advanced key management and attack detection strategies. A timing analysis for PAC-PL is also presented. PAC-PL was experimentally evaluated with state-of-the-art benchmarks in terms of run-time overhead, memory footprint, and FPGA resource consumption, resulting in a practical solution for implementing CFI. Gabriele Serra, Pietro Fara, Giorgiomaria Cicero, Francesco Restuccia 0002, Alessandro Biondi 0001 |
RTAS | 5 |
| 2022 | Evaluating the Robustness of Semantic Segmentation for Autonomous Driving against Real-World Adversarial Patch AttacksabstractDeep learning and convolutional neural networks allow achieving impressive performance in computer vision tasks, such as object detection and semantic segmentation (SS). However, recent studies have shown evident weaknesses of such models against adversarial perturbations. In a real-world scenario instead, like autonomous driving, more attention should be devoted to real-world adversarial examples (RWAEs), which are physical objects (e.g., billboards and printable patches) optimized to be adversarial to the entire perception pipeline. This paper presents an in-depth evaluation of the robustness of popular SS models by testing the effects of both digital and real-world adversarial patches. These patches are crafted with powerful attacks enriched with a novel loss function. Firstly, an investigation on the Cityscapes dataset is conducted by extending the Expectation Over Transformation (EOT) paradigm to cope with SS. Then, a novel attack optimization, called scene-specific attack, is proposed. Such an attack leverages the CARLA driving simulator to improve the transferability of the proposed EOT-based attack to a real 3D environment. Finally, a printed physical billboard containing an adversarial patch was tested in an outdoor driving scenario to assess the feasibility of the studied attacks in the real world. Exhaustive experiments revealed that the proposed attack formulations outperform previous work to craft both digital and real-world adversarial patches for SS. At the same time, the experimental results showed how these attacks are notably less effective in the real world, hence questioning the practical relevance of adversarial attacks to SS models for autonomous/assisted driving. Federico Nesti, Giulio Rossolini, Saasha Nair, Alessandro Biondi 0001, Giorgio C. Buttazzo |
WACV | 4 |
| 2022 | A Linux-based support for developing real-time applications on heterogeneous platforms with dynamic FPGA reconfiguration
Marco Pagani, Alessandro Biondi 0001, Mauro Marinoni, Lorenzo Molinari, Giuseppe Lipari, Giorgio C. Buttazzo |
Future Gener. Comput. Syst. | 2 |
| 2022 | Partitioning real-time workloads on multi-core virtual machines
Luca Abeni, Alessandro Biondi 0001, Enrico Bini |
J. Syst. Archit. | 2 |
| 2022 | Optimized partitioning and priority assignment of real-time applications on heterogeneous platforms with hardware acceleration
Daniel Casini, Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale |
J. Syst. Archit. | 3 |
| 2022 | ARTe: Providing real-time multitasking to Arduino
Francesco Restuccia 0002, Marco Pagani, Agostino Mascitti, Michael Barrow, Mauro Marinoni, Alessandro Biondi 0001, Giorgio C. Buttazzo, Ryan Kastner |
J. Syst. Softw. | 6 |
| 2022 | Profiling and controlling I/O-related memory contention in COTS heterogeneous platformsabstractAbstract Motivated by the increasing number of embedded applications that make use of traffic‐intensive I/O devices, this work studies the memory contention generated by I/O devices and investigates on the regulation of the bus traffic they generate by means of COTS regulators, namely the QoS‐400 by Arm. To this purpose, the behavior of the QoS‐400 regulators is analytically characterized and then, taking the Xilinx Ultrascale+ as a reference modern heterogeneous platform, a software infrastructure to control such regulators from Linux is proposed. As an experience report, this article presents the results of an extensive experimental evaluation, based on both benchmarks and microbenchmarks, aimed at validating the effectiveness of QoS‐400 regulators in predictably controlling I/O‐related memory traffic, as well as assessing the impact of the regulation on software applications and I/O devices themselves. Matteo Zini, Giorgiomaria Cicero, Daniel Casini, Alessandro Biondi 0001 |
Softw. Pract. Exp. | 4 |
| 2022 | An I/O Virtualization Framework With I/O-Related Memory Contention Control for Real-Time SystemsabstractModern applications are often characterized by a tight interaction with I/O devices. At the same time, many application domains are also facing a shift toward an integrated approach where multiple applications with mixed levels of safety and security need to co-exist on top of a shared hardware platform, which is typically managed by a hypervisor. This gives rise to the need for a predictable mechanism allowing multiple virtual machines to share I/O devices, while at the same time controlling contention delays when they access global memory. To deal with these shortcomings, this article proposes an I/O virtualization framework providing support for controlling the I/O-related memory contention by leveraging the ARM QoS-400 regulators. Extensive experiments are performed to compare the proposed solution with the Xen hypervisor, showing improvements up to$8\times $when controlling the I/O-related memory contention. Niccolò Borgioli, Matteo Zini, Daniel Casini, Giorgiomaria Cicero, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Optimal Memory Allocation and Scheduling for DMA Data Transfers under the LET ParadigmabstractThe Logical Execution Time (LET) paradigm is increasingly used to achieve predictable communications in modern multicore automotive applications. Direct Memory Access (DMA) engines can perform the data copies that are needed in a LET implementation on behalf of the cores with improved parallelism and reduced overheads. However, each DMA transfer operates on contiguous memory areas, and the performance is strongly dependent on the allocation in memory of the variables to be copied. This paper proposes a protocol to perform LET communications with a DMA and presents an optimal memory allocation scheme and scheduling using a mixed-integer linear programming formulation. Experimental results are reported to compare the performance of different communication approaches. Paolo Pazzaglia, Daniel Casini, Alessandro Biondi 0001, Marco Di Natale |
DAC | 3 |
| 2021 | Scheduling Replica Voting in Fixed-Priority Real-Time SystemsabstractReliability and safety are mandatory requirements for safety-critical embedded systems. The design of a fault-tolerant system is required in many fields (e.g., railway, automotive, avionics) and redundancy helps in achieving this goal. Redundant systems typically leverage voting techniques applied to the outputs produced by tasks to detect and even tolerate failures. This paper studies the integration of distributed voting protocols in fixed-priority real-time systems from a scheduling perspective. It analyzes two scheduling strategies for implementing voting. One is attractive and friendly for software developers and based on suspending the task execution until the replica provides the data to be voted. The other one is inspired by the Logical Execution Time (LET) paradigm and requires introducing additional tasks in the system to accomplish voting-related activities. Queuing and delays introduced by inter-replica communication interfaces are also analyzed. Experimental results are finally presented to compare the two strategies, showing that LET-inspired voting is much more predictable and hence more suitable than the other strategy for fixed-priority real-time systems. Pietro Fara, Gabriele Serra, Alessandro Biondi 0001, Ciro Donnarumma |
ECRTS | 3 |
| 2021 | Event-Driven Delay-Induced Tasks: Model, Analysis, and ApplicationsabstractParallel execution and hardware acceleration involving specialized devices such as GPUs and FPGAs are becoming increasingly relevant in the domain of embedded systems. Communication between jobs dispatched on different cores and hardware accelerators is most often implemented using asynchronous events. Modeling the timing behavior of such systems requires to account for the delays incurred by each task due to the additional time spent waiting for events. This paper presents the event-driven delay-induced (EDD) task model to explicitly deal with complex computing workloads that incur such kinds of delays. The EDD task model generalizes several state-of-the-art models, such as the DAG task model and the segmented self-suspending task model, and is particularly suited to analyze parallel tasks that issue asynchronous hardware acceleration requests. Two analysis techniques for EDD tasks executing on single core platforms are first provided. We then extend those approaches to analyze parallel real-time tasks under partitioned multicore scheduling by means of a model transformation. Experimental results are presented to compare the two analysis techniques for EDD tasks proposed in the paper. Finally, we compare the analysis of partitioned parallel tasks modeled with EDD tasks against federated scheduling. Federico Aromolo, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTAS | 2 |
| 2021 | Latency Analysis of I/O Virtualization Techniques in Hypervisor-Based Real-Time SystemsabstractNowadays, hypervisors are the standard solution to integrate different domains into a shared hardware platform, while providing safety, security, and predictability. To this end, a hypervisor virtualizes the physical platform and orchestrates the access to each component. When the system needs to comply with certification requirements for safety-critical systems, virtualization latencies need to be analytically bounded for providing off-line guarantees. This paper presents a detailed modeling of three I/O virtualization techniques, providing analytical bounds for each of them under different metrics. Experimental results compare the bounds for a case study and quantify the contribution due to different sources of delay. Daniel Casini, Alessandro Biondi 0001, Giorgiomaria Cicero, Giorgio C. Buttazzo |
RTAS | 2 |
| 2021 | A Multi-Domain Software Architecture for Safe and Secure Autonomous DrivingabstractThis work aims at making Apollo, a popular autonomous driving framework, safer and more secure by designing a multi-domain architecture, where its components are split between a feature-rich domain running Linux and a critical domain running a real-time operating system (RTOS). The two domains are isolated by a hypervisor. We implemented a prototype where the control component has been ported from Linux to the Erika automotive-grade RTOS, and we discuss a number of challenges that have been faced in moving the component to Erika. The proposed solution has been experimentally evaluated by measuring the latencies involving processing paths passing through the control component. Luca Belluardo, Andrea Stevanato, Daniel Casini, Giorgiomaria Cicero, Alessandro Biondi 0001, Giorgio C. Buttazzo |
RTCSA | 5 |
| 2021 | Time-Predictable Acceleration of Deep Neural Networks on FPGA SoC PlatformsabstractThis work focuses on the time-predictable execution of Deep Neural Networks (DNNs) accelerated on FPGA System-on-Chips (SoCs). The modern DPU accelerator by Xilinx is considered. An extensive profiling campaign targeting the Zynq Ultrascale+ platform has been performed to study the execution behavior of the DPU when accelerating a set of state-of-the-art DNNs for Advanced Driver Assistance Systems (ADAS). Based on the profiling, an execution model is proposed and then used to derive a response-time analysis. A custom FPGA module named DICTAT is also proposed to improve the predictability of the acceleration of DNNs and tighten the analytical bounds. A rich set of experimental results based on both analytical bounds and measurements from the target platform is finally presented to assess the effectiveness and the performance of the proposed approach on ADAS applications. Francesco Restuccia 0002, Alessandro Biondi 0001 |
RTSS | 2 |
| 2021 | Task Splitting and Load Balancing of Dynamic Real-Time Workloads for Semi-Partitioned EDFabstractMany real-time software systems, such as those commonly found in the context of multimedia, cloud computing, robotics, and real-time databases, are characterized by a dynamic workload, where applications can join and leave the system at runtime. Global schedulers can transparently support dynamic workload without requiring any off-line task-allocation phase, thus providing advantages to the system designer. Nevertheless, such schedulers exhibit poor worst-case performance when compared to semi-partitioned schedulers, which instead can achieve near-optimal schedulability performance when used in conjunction with smart task splitting and partitioning techniques, and they are also lighter in terms of run-time overhead. This article proposes an approach to efficiently schedule dynamic real-time workloads on multiprocessor systems by means of semi-partitioned scheduling. A linear-time approximation scheme for the C=D splitting algorithm under partitioned EDF scheduling is proposed. Then, a load-balancing algorithm is presented to admit new real-time workloads with a limited number of re-allocations. The article finally reports on a large-scale experimental study showing that (i) the linear-time approximation is characterized by a very limited utilization loss compared with the corresponding exact approach (that has a much higher complexity), and that (ii) the whole approach allows achieving considerable improvements with respect to global and partitioned EDF scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Computers | 2 |
| 2021 | Spatio-Temporal Optimization of Deep Neural Networks for Reconfigurable FPGA SoCsabstractThis article proposes a technique for optimizing the timing performance and the resource consumption of hardware accelerators for deep neural network (DNN) inference on FPGA-based system-on-chips (SoC). When required, the accelerators are decomposed into chunks, each exploiting at best the available FPGA area, and dynamic partial reconfiguration (DPR) is leveraged to schedule such chunks at run-time. To this end, the article presents accurate models of the resource consumption and timing of DNN accelerators provided by the Xilinx FINN framework. The models are then used to formulate an optimization problem that computes the optimal decomposition of DNN accelerators (and their configuration) by minimizing the inference time while ensuring area constraints on the FPGA. Experimental results on Zynq-7000 platforms demonstrate that the proposed technique provides consistent improvements with respect to both stock configurations of the accelerators and other configurations that can be obtained with a static FPGA allocation. Biruk B. Seyoum, Marco Pagani, Alessandro Biondi 0001, Sara Balleri, Giorgio C. Buttazzo |
IEEE Trans. Computers | 3 |
| 2020 | Predictable Memory-CPU Co-Scheduling with Support for Latency-Sensitive TasksabstractPredictable execution models have been proposed over the years to achieve contention-free execution of real-time tasks by preloading data into dedicated local memories. In this way, memory access delays can be hidden by delegating a DMA engine to perform memory transfers in parallel with processor execution. Nevertheless, state-of-the-art protocols introduce additional blocking due to priority inversion, which may severely penalize latency-sensitive applications and even worsen the system schedulability with respect to the use of classical scheduling schemes. This paper proposes a new protocol that allows hiding memory transfer delays while reducing priority inversion, thus favoring the schedulability of latency-sensitive tasks. The corresponding analysis is formulated as an optimization problem. Experimental results show the advantages of the proposed protocol against state-of-the-art solutions. Daniel Casini, Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale, Giorgio C. Buttazzo |
DAC | 3 |
| 2020 | AXI HyperConnect: A Predictable, Hypervisor-level Interconnect for Hardware Accelerators in FPGA SoCabstractFPGA-based system-on-chips (SoC) are powerful computing platforms to implement mixed-criticality systems that require both multiprocessing and hardware acceleration. Virtualization via hypervisor technologies is, de-facto, an effective technique to allow the co-existence of multiple execution domains with different criticality levels in isolation upon the same platform. Implementing such technologies on FPGA-based SoC poses new challenges: one of such is the isolation of hardware accelerators deployed on the FPGA fabric that belong to different domains but share common resources such as a memory bus. This paper proposes AXI HyperConnect, a hypervisor-level hardware component that allows interconnecting hardware accelerators to the same bus while ensuring isolation and predictability features. AXI HyperConnect has been implemented on modern FPGA-SoC by Xilinx and tested with real-world accelerators, including one for Deep Neural Network inference. Francesco Restuccia 0002, Alessandro Biondi 0001, Mauro Marinoni, Giorgiomaria Cicero, Giorgio C. Buttazzo |
DAC | 2 |
| 2020 | Modeling and Analysis of Bus Contention for Hardware Accelerators in FPGA SoCsabstractFPGA System-on-Chips (SoCs) are heterogeneous platforms that combine general-purpose processors with a field-programmable gate array (FPGA) fabric. The FPGA fabric is composed of a programmable logic in which hardware accelerators can be deployed to accelerate the execution of specific functionality. The main source of unpredictability when bounding the execution times of hardware accelerators pertains the access to the shared memories via the on-chip bus. This work is focused on bounding the worst-case bus contention experienced by the hardware accelerators deployed in the FPGA fabric. To this end, this work considers the AMBA AXI bus, which is the de-facto standard communication interface used in most the commercial off-the-shelf (COTS) FPGA SoCs, and presents an analysis technique to bound the response times of hardware accelerators implemented on such platforms. A fine-grained modeling of the AXI bus and AXI interconnects is first provided. Then, contention delays are studied under hierarchical bus infrastructures with arbitrary depths. Experimental results are finally presented to validate the proposed model with execution traces on two modern FPGA-based SoC produced by Xilinx (Zynq-7000 and Zynq-Ultrascale+ families) and to assess the performance of the proposed analysis. Francesco Restuccia 0002, Marco Pagani, Alessandro Biondi 0001, Mauro Marinoni, Giorgio C. Buttazzo |
ECRTS | 3 |
| 2020 | Safely Preventing Unbounded Delays During Bus Transactions in FPGA-based SoCabstractAdvanced eXtensible Interface (AXI) is an open-standard communication bus interface implemented in most commercial off-the-shelf FPGA System-on-Chips (SoC) to exchange data within the chip. Unfortunately, the AXI standard does not mandate any mechanism to detect possible misbehavior of the connected modules. This work shows that this lack of specification has a relevant impact on popular implementations of the AXI bus. In particular, it is shown how it is easily possible to inject arbitrarily-long delays on modern FPGA system-on-chips under the presence of misbehaving bus masters. To safely solve this issue, this paper presents a general timing analysis to bound the execution of periodically-invoked hardware accelerators in nominal conditions. This timing analysis is then used to conFigure a latency-free hardware module named AXI Stall Monitor (ASM), also proposed in this paper, capable of detecting and safely solving possible stalls during AXI bus transactions. The ASM leaves a quantified flexibility to the hardware accelerators when deviating from nominal conditions. The contribution is finally supported by a set of experiments on the Zynq-7000 and Zynq Ultrascale+SoCs by Xilinx. Francesco Restuccia 0002, Alessandro Biondi 0001, Mauro Marinoni, Giorgio C. Buttazzo |
FCCM | 2 |
| 2020 | The AMPERE Project: : A Model-driven development framework for highly Parallel and EneRgy-Efficient computation supporting multi-criteria optimizationabstractThe high-performance requirements needed to implement the most advanced functionalities of current and future Cyber-Physical Systems (CPSs) are challenging the development processes of CPSs. On one side, CPSs rely on model-driven engineering (MDE) to satisfy the non-functional constraints and to ensure a smooth and safe integration of new features. On the other side, the use of complex parallel and heterogeneous embedded processor architectures becomes mandatory to cope with the performance requirements. In this regard, parallel programming models, such as OpenMP or CUDA, are a fundamental brick to fully exploit the performance capabilities of these architectures. However, parallel programming models are not compatible with current MDE approaches, creating a gap between the MDE used to develop CPSs and the parallel programming models supported by novel and future embedded platforms.The AMPERE project will bridge this gap by implementing a novel software architecture for the development of advanced CPSs. To do so, the proposed software architecture will be capable of capturing the definition of the components and communications described in the MDE framework, together with the non-functional properties, and transform it into key parallel constructs present in current parallel models, which may require extensions. These features will allow for making an efficient use of underlying parallel and heterogeneous architectures, while ensuring compliance with non-functional requirements, including those on real-time performance of the system. Eduardo Quiñones, Sara Royuela, Claudio Scordino, Paolo Gai, Luís Miguel Pinho, Luís Nogueira, Jan Rollo, Tommaso Cucinotta, Alessandro Biondi 0001, Arne Hamann 0001, Dirk Ziegenbein, Hadi Saoud, Romain Soulat, Björn Forsberg, Luca Benini, Gianluca Mandò, Luigi Rucher |
ISORC | 9 |
| 2020 | A Holistic Memory Contention Analysis for Parallel Real-Time Tasks under Partitioned SchedulingabstractWhen adopting multi-core systems for safety-critical applications, certification requirements mandate bounding the delays incurred in accessing shared resources. This is the case of global memories, whose access is often regulated by memory controllers optimized for average-case performance and not designed to be predictable. As a consequence, worst-case bounds on memory access delays often result to be too pessimistic, drastically reducing the advantage of having multiple cores. This paper proposes a fine-grained analysis of the memory contention experienced by parallel tasks running on a multi-core platform. To this end, an optimization problem is formulated to bound the memory interference by leveraging a three-phase execution model and holistically considering multiple memory transactions issued during each phase. Experimental results show the advantage in adopting the proposed approach on both synthetic task sets and benchmarks. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTAS | 2 |
| 2020 | Integrating Online Safety-related Memory Tests in Multicore Real-Time SystemsabstractAlmost all functional safety standards that regulate safety-critical domains impose to periodically test hardware platforms at run-time. RAM memories are among the fundamental components of computing platforms and are notably subject to faults. Hence, they are also primary components to be tested. Unfortunately, RAM tests are destructive, require to be atomically executed, and are not cheap from a computational perspective. As such, if not properly managed, they can jeopardize the timing performance of a real-time system, especially when running upon a multicore platform.This paper proposes a software architecture to integrate online memory tests on multicore real-time systems. Furthermore, by jointly considering a task model and a safety model based on the EN50129 safety standard, it presents an approach to compute the optimal configuration of memory tests that preserves the system schedulability and guarantees a given tolerable functional failure rate (TFFR). Experimental results show that the proposed approach allows achieving a marginal impact on schedulability while preserving a TFFR that is compatible with the highest safety integrity level specified by the EN50129. Ciro Donnarumma, Alessandro Biondi 0001, Francesco De Rosa, Stefano Di Carlo |
RTSS | 2 |
| 2020 | Timing isolation and improved scheduling of deep neural networks for real-time systemsabstractSummary In recent years, the performance of deep neural networks (DNNs) is significantly improved, making them suitable for many application fields, such as autonomous driving, advanced robotics, and industrial control. Despite a lot of research being devoted to improving the accuracy of DNNs, only limited efforts have been spent to enhance their timing predictability, required in several real‐time applications. This paper proposes a software infrastructure based on the Linux operating system to integrate DNNs within a real‐time multicore system. It has been realized by modifying both the internal scheduler of the popular TensorFlow framework and the SCHED_DEADLINE scheduling class of Linux. The proposed infrastructure allows providing timing isolation of DNN inference tasks, hence improving the determinism of the temporal interference generated by TensorFlow. The proposal is finally evaluated with a case study derived from a state‐of‐the‐art benchmark inspired by an autonomous industrial system. Extensive experiments demonstrate the effectiveness of the proposed solution and show a significant reduction of both average and longest‐observed response times of TensorFlow tasks. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
Softw. Pract. Exp. | 2 |
| 2019 | Analyzing Parallel Real-Time Tasks Implemented with Thread PoolsabstractDespite several works in the literature targeted predictable execution models for parallel tasks, limited attention has been devoted to study how specific implementation techniques may affect their execution. This paper highlights some issues that can arise when executing parallel tasks with thread pools, which may lead to deadlocks and performance degradation when adopting blocking synchronization mechanisms. A new parallel task model, inspired to a realistic design found in popular software systems, is first presented to study this problem. Then, formal conditions to ensure the absence of deadlocks and schedulability analysis techniques are proposed under both global and partitioned scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
DAC | 2 |
| 2019 | Simple and General Methods for Fixed-Priority Schedulability in Optimization ProblemsabstractThis paper presents a set of sufficient-only, but accurate schedulability tests for fixed-priority scheduling. The tests apply to the general case of scheduling with constrained deadline where tasks can incur in blocking times, be subject to release jitters, activated with fixed offsets, or involved in transactions with other tasks. The proposed tests come in a linear closed-form with a number of conditions polynomial in the number of tasks. All tests are targeted for use when encoding schedulability constraints within Mixed-Integer Linear Programming for the purpose of optimizing real-time systems (e.g., to address task partitioning in a multicore system). The tests are evaluated with a large-scale experimental study based on synthetic workload, revealing a failure rate (with respect to the state-of-the-art reference tests) of less than 1% in average, and at most of 2% in a very small number of limit-case configurations. Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale |
DATE | 2 |
| 2019 | A Bandwidth Reservation Mechanism for AXI-Based Hardware Accelerators on FPGAsabstractHardware platforms for real-time embedded systems are evolving towards heterogeneous architectures comprising different types of processing cores and dedicated hardware accelerators, which can be implemented on silicon or dynamically deployed on FPGA fabric. Such accelerators typically access a shared memory to exchange a significant amount of data with other processing elements. Existing COTS solutions focus on maximizing the overall throughput of the system, rather than guaranteeing the timing constraints of individual hardware accelerators. This paper presents the AXI budgeting unit (ABU), a hardware-based solution to implement a bandwidth reservation mechanism on top of the AMBA AXI standard infrastructure for hardware accelerators deployed on FPGAs. An accurate and tractable model, as well as the corresponding analysis, are also proposed to bound the response time of hardware accelerators in the presence of ABUs, in order to verify whether they can complete before their deadlines. Finally, a set of experiments are reported to evaluate the proposed approach on a state-of-the-art platform, namely the Zynq-7020 by Xilinx. The resource consumption of the ABU has been quantified to be less than 1% of the total FPGA resources of the Zynq-7020. Marco Pagani, Enrico Rossi, Alessandro Biondi 0001, Mauro Marinoni, Giuseppe Lipari, Giorgio C. Buttazzo |
ECRTS | 3 |
| 2019 | Optimizing the Functional Deployment on Multicore Platforms with Logical Execution TimeabstractThe move to multicore systems requires methods and tools to support the designer in the partitioning of functions among the available cores and the definition of the task model. In this paper we present the formulation of a functional partitioning for real-time systems and we provide an optimization method for an efficient implementation of the Logical Execution Time (LET) paradigm, to enforce causality and determinism in the development of time-and safety-critical applications. A novel schedulability analysis for partitioned tasks executing according to the LET paradigm is also provided. Our methods are applied to the industry-size model of the WATERS challenge and compute solutions that easily outperform the initial solution provided. Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale |
RTSS | 2 |
| 2019 | Hierarchical scheduling of real-time tasks over Linux-based virtual machines
Luca Abeni, Alessandro Biondi 0001, Enrico Bini |
J. Syst. Softw. | 2 |
| 2019 | Handling Transients of Dynamic Real-Time Workload Under EDF SchedulingabstractReal-time dynamic workload consists of tasks that can arbitrarily join and leave the system at run-time. To avoid incurring deadline misses, tasks that request to join the system must pass an admission test, which has to cope with potential scheduling transients originated by the residual effect of the tasks that previously left the system. This phenomenon may require some tasks to suffer an admission delay before being accepted for execution. This paper focuses on uniprocessor earliest-deadline first (EDF) scheduling with constrained deadlines and explicitly considers methods for handling scheduling transients in the presence of dynamic real-time workload. A generalized analysis framework is first presented to overcome several limitations of the existing approaches (including the support for overlapping transients), and is then used to derive methods for computing bounds on the admission delays incurred by tasks. Building on such results, an on-line protocol is proposed to handle the admission control of a dynamic workload, which also comes with a variant that can execute in polynomial time to favor its practical application. Furthermore, the paper shows how the presented analysis can be used off-line for analyzing mode-changes among static task sets. Experimental results are finally presented to evaluate the proposed algorithms. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Computers | 2 |
| 2019 | Is Your Bus Arbiter Really Fair? Restoring Fairness in AXI Interconnects for FPGA SoCsabstractAMBA AXI is a popular bus protocol that is widely adopted as the medium to exchange data in field-programmable gate array system-on-chips (FPGA SoCs). The AXI protocol does not specify how conflicting transactions are arbitrated and hence the design of bus arbiters is left to the vendors that adopt AXI. Typically, a round-robin arbitration is implemented to ensure a fair access to the bus by the master nodes, as for the popular SoCs by Xilinx. This paper addresses a critical issue that can arise when adopting the AXI protocol under round-robin arbitration; specifically, in the presence of bus transactions with heterogeneous burst sizes. First, it is shown that a completely unfair bandwidth distribution can be achieved under some configurations, making possible to arbitrarily decrease the bus bandwidth of a target master node. This issue poses serious performance, safety, and security concerns. Second, a low-latency (one clock cycle) module named AXI burst equalizer (ABE) is proposed to restore fairness. Our investigations and proposals are supported by implementations and tests upon three modern SoCs. Experimental results are reported to confirm the existence of the issue and assess the effectiveness of the ABE with bus traffic generators and hardware accelerators from the Xilinx’s IP library. Francesco Restuccia 0002, Marco Pagani, Alessandro Biondi 0001, Mauro Marinoni, Giorgio C. Buttazzo |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2019 | FLORA: FLoorplan Optimizer for Reconfigurable Areas in FPGAsabstractFloorplanning is a mandatory step in the design of hardware accelerators for FPGA platforms, especially when adopting dynamic partial reconfiguration (DPR). This paper presents FLORA, an automated floorplanner based on optimization via Mixed-Integer Linear Programming (MILP). The floorplanning problem is solved by means of a novel fine-grained modeling strategy of FPGA resources. Furthermore, differently from other proposals, our approach takes into account several realistic Partial Reconfiguration (PR) floorplanning constraints on FPGAs. FLORA was compared against state-of-the-art floorplanners by means of benchmark suites, showing that it is capable of providing better performance in terms of resource consumption, maximum inter-region, wire-length, and running time required to produce the solutions. Finally, FLORA was utilized to generate placements for a partially-reconfigurable video processing engine that was implemented on a Xilinx Zynq-7020. Biruk B. Seyoum, Alessandro Biondi 0001, Giorgio C. Buttazzo |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2018 | Beyond the Weakly Hard Model: Measuring the Performance Cost of Deadline MissesabstractMost works in schedulability analysis theory are based on the assumption that constraints on the performance of the application can be expressed by a very limited set of timing constraints (often simply hard deadlines) on a task model. This model is insufficient to represent a large number of systems in which deadlines can be missed, or in which late task responses affect the performance, but not the correctness of the application. For systems with a possible temporary overload, models like the m-K deadline have been proposed in the past. However, the m-K model has several limitations since it does not consider the state of the system and is largely unaware of the way in which the performance is affected by deadline misses (except for critical failures). In this paper, we present a state-based representation of the evolution of a system with respect to each deadline hit or miss event. Our representation is much more general (while hopefully concise enough) to represent the evolution in time of the performance of time-sensitive systems with possible time overloads. We provide the theoretical foundations for our model and also show an application to a simple system to give examples of the state representations and their use. Paolo Pazzaglia, Luigi Pannocchi, Alessandro Biondi 0001, Marco Di Natale |
ECRTS | 3 |
| 2018 | Achieving Predictable Multicore Execution of Automotive Applications Using the LET ParadigmabstractNext generation automotive applications require support for safe, predictable, and deterministic execution. The Logical Execution Time (LET) model has been introduced to improve the predictability and correctness of time-critical applications. The advent of multicore architectures, together with the need to ensure time predictability despite the complex memory hierarchy and the hardware resources shared by the cores, is an additional motivation for the use of the LET paradigm in conjunction with a suitable scheduling and memory access model. In this paper, we show how an implementation of the LET model on actual multicore platforms for automotive systems brings the potential to improve time determinism at the price of a modicum run-time overhead. Multiple implementation options are discussed using the automotive AUTOSAR model and operating system standard, and a realistic application defined by Bosch for the 2017 WATERS challenge. Experimental data of executions on the Infineon Aurix platform show the feasibility of the proposed approach. The paper also provides a discussion on further implementation optimizations and other issues related to the general problem of memory-aware analysis of automotive applications on multicores. Alessandro Biondi 0001, Marco Di Natale |
RTAS | 1 |
| 2018 | Memory Feasibility Analysis of Parallel Tasks Running on Scratchpad-Based ArchitecturesabstractThis work proposes solutions for bounding the worst-case memory space requirement for parallel tasks running on multicore platforms with scratchpad memories. It introduces a feasibility test that verifies whether memories are large enough to contain the maximum memory backlog that may be generated by the system. Both closed-form bounds and more accurate algorithmic techniques are proposed. It is shown how one can use max-plus algebra and solutions to the max-flow cut problem to efficiently solve the memory feasibility problem. Experimental results are presented to evaluate the efficiency of the proposed feasibility analysis techniques on synthetic workload and state-of-the-art benchmarks. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTSS | 2 |
| 2018 | Partitioned Fixed-Priority Scheduling of Parallel Tasks Without PreemptionsabstractThe study of parallel task models executed with predictable scheduling approaches is a fundamental problem for real-time multiprocessor systems. Nevertheless, to date, limited efforts have been spent in analyzing the combination of partitioned scheduling and non-preemptive execution, which is arguably one of the most predictable schemes that can be envisaged to handle parallel tasks. This paper fills this gap by proposing an analysis for sporadic DAG tasks under partitioned fixed-priority scheduling where the computations corresponding to the nodes of the DAG are non-preemptively executed. The analysis has been achieved by means of segmented self-suspending tasks with nonpreemptable segments, for which a new fine-grained analysis is also proposed. The latter is shown to analytically dominate state-of-the-art approaches. A partitioning algorithm for DAG tasks is finally proposed. By means of experimental results, the proposed analysis has been compared against a previouslyproposed analysis for DAG tasks with non-preemptable nodes managed by global fixed-priority scheduling. The comparison revealed important improvements in terms of schedulability performance. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTSS | 2 |
| 2018 | The SRP Resource Sharing Protocol for Self-Suspending TasksabstractMotivated by the increasingly wide adoption of realtime workload with self-suspending behaviors, and the relevance of mechanisms to handle mutually-exclusive shared resources, this paper takes a new look at locking protocols for self-suspending tasks under uniprocessor fixed-priority scheduling. Pitfalls when integrating the widely-adopted Stack Resource Policy (SRP) with self-suspending tasks are firstly illustrated, and then a new finegrained SRP analysis is presented. Next, a new locking protocol, named SRP-SS, is proposed to overcome the limitations of the original SRP. The SRP-SS is a generalization of the SRP to cope with the specificities of self-suspending tasks. It therefore reduces to the SRP under some configurations and hence theoretically dominates the SRP. It also ensures backward compatibility for applications developed specifically for the SRP. The SRP-SS comes with its own schedulability analysis and configuration algorithm. The performances of the SRP and SRP-SS are finally studied by means of large-scale schedulability experiments. Geoffrey Nelissen, Alessandro Biondi 0001 |
RTSS | 2 |
| 2018 | A survey of schedulability analysis techniques for rate-dependent tasks
Timo Feld, Alessandro Biondi 0001, Robert I. Davis 0001, Giorgio C. Buttazzo, Frank Slomka |
J. Syst. Softw. | 2 |
| 2018 | A design flow for supporting component-based software development in multiprocessor real-time systems
Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
Real Time Syst. | 1 |
| 2018 | On the ineffectiveness of 1/m-based interference bounds in the analysis of global EDF and FIFO scheduling
Alessandro Biondi 0001, Youcheng Sun |
Real Time Syst. | 1 |
| 2018 | Response-Time Analysis of Engine Control Applications Under Fixed-Priority SchedulingabstractEngine control systems include computational activities that are triggered at predetermined angular values of the crankshaft, and therefore generate a workload that tends to increase with the engine speed. To cope with overload conditions, a common practice adopted by the automotive industry is to design such angular tasks with a set of modes that switch at given rotation speeds to adapt the computational demand. This paper presents an exact response time analysis for engine control applications consisting of periodic and engine-triggered tasks scheduled by fixed priority. The proposed analysis explicitly takes into account the physical constraints of the considered systems and is based on the derivation of dominant speeds, which are particular engine speeds that are proved to determine the worst-case behavior of engine-triggered tasks from a timing perspective. Experimental results are finally reported to validate the proposed approach and compare it against an existing sufficient test. Alessandro Biondi 0001, Marco Di Natale, Giorgio C. Buttazzo |
IEEE Trans. Computers | 1 |
| 2018 | Selecting the Transition Speeds of Engine Control Tasks to Optimize the PerformanceabstractEngine control applications include functions that need to be executed at specific rotation angles of the crankshaft. The tasks performing these functions are activated at variable rates and are programmed to be adaptive with respect to the rotation speed of the engine to avoid overloading the CPU. Simplified control implementations are used at high speeds; for example, reducing the number of fuel injections or the complexity of the computations. Such different control implementations define execution modes with different execution times for different ranges of the rotation speed. The selection of the switching speeds for the operating modes of such tasks is an optimization problem, consisting in determining the optimal transition speeds that maximize the engine performance while guaranteeing schedulability. This article presents three methods for tackling such an optimization problem under a set of assumptions about the performance metrics: two heuristics and a branch and bound method that guarantees finding the optimal solution within a given speed granularity. In addition, a simple method to compute a performance upper bound is presented. The approach and the hypothesis are validated using a Simulink model of the engine and the computational tasks, considering the engine efficiency and the production of pollutants (NO 2 ) as metrics of interest. Simulation experiments show that the performance of proposed heuristics is quite close to that of the upper bound and the optimum within a finite granularity. Alessandro Biondi 0001, Marco Di Natale, Giorgio C. Buttazzo, Paolo Pazzaglia |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2018 | Modeling and Analysis of Engine Control Tasks Under Dynamic Priority SchedulingabstractIn automotive systems, engine control applications include computational activities that are triggered by specific rotation angles of the crankshaft, causing their activation rate to be proportional to the engine speed. In order to avoid overloads at high engine speeds, these tasks are implemented to adapt their functionality based on the angular velocity of the engine. This paper proposes a task model for expressing a number of realistic features of such engine control tasks and presents a real-time schedulability analysis for applications consisting of multiple engine control tasks and classical periodic/sporadic tasks scheduled by the earliest deadline first algorithm. Differently from other efforts spent in analyzing engine-control applications, the presented approach is focused on simplicity, providing linear-time and quadratic-time schedulability tests based on utilization bounds. Experimental results are finally presented to assess the performance of the presented analysis techniques. Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Semi-Partitioned Scheduling of Dynamic Real-Time Workload: A Practical Approach Based on Analysis-Driven Load BalancingabstractRecent work showed that semi-partitioned scheduling can achieve near-optimal schedulability performance, is simpler to implement compared to global scheduling, and less heavier in terms of runtime overhead, thus resulting in an excellent choice for implementing real-world systems. However, semi-partitioned scheduling typically leverages an off-line design to allocate tasks across the available processors, which requires a-priori knowledge of the workload. Conversely, several simple global schedulers, as global earliest-deadline first (G-EDF), can transparently support dynamic workload without requiring a task-allocation phase. Nonetheless, such schedulers exhibit poor worst-case performance. This work proposes a semi-partitioned approach to efficiently schedule dynamic real-time workload on a multiprocessor system. A linear-time approximation for the C=D splitting scheme under partitioned EDF scheduling is first presented to reduce the complexity of online scheduling decisions. Then, a load-balancing algorithm is proposed for admitting new real-time workload in the system with limited workload re-allocation. A large-scale experimental study shows that the linear-time approximation has a very limited utilization loss compared to the exact technique and the proposed approach achieves very high schedulability performance, with a consistent improvement on G-EDF and pure partitioned EDF scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
ECRTS | 2 |
| 2017 | Real-Time Analysis and Design of a Dual Protocol Support for Bluetooth LE DevicesabstractModern distributed embedded systems frequently involve wireless communication nodes where messages have to be delivered within given timing constraints. This goal can be achieved by adopting a suitable real-time communication protocol. In addition, connecting such systems with mobile devices is also desirable for performing configuration, monitoring, and maintenance activities. The Bluetooth low energy (BLE) protocol would be an attractive solution for this purpose, because it is supported by consumer devices, such as tablets and smart phones, for implementing personal area networks with reduced energy consumption. Unfortunately, however, it cannot guarantee a bounded delay for managing real-time traffic. Modern BLE radio transceivers allow partitioning the network bandwidth between the BLE protocol and another user-defined protocol running on top of the raw radio. This paper exploits this feature to provide an analysis and a design methodology to guarantee the feasibility of a real-time custom protocol that shares the radio with the BLE. Experimental results on a Nordic reference platform show the feasibility of the dual-protocol approach and its capability to support a custom real-time protocol on the raw radio with a bounded overhead. Mauro Marinoni, Alessandro Biondi 0001, Pasquale Buonocunto, Gianluca Franchino, Daniel Cesarini, Giorgio C. Buttazzo |
IEEE Trans. Ind. Informatics | 2 |
| 2016 | Real-time analysis of engine control applications with speed estimation
Alessandro Biondi 0001, Giorgio C. Buttazzo |
DATE | 1 |
| 2016 | Lightweight Real-Time Synchronization under P-EDF on Symmetric and Asymmetric MultiprocessorsabstractThis paper revisits lightweight synchronization under partitioned earliest-deadline first (P-EDF) scheduling. Four different lightweight synchronization mechanisms - namely preemptive and non-preemptive lock-free synchronization, as well as preemptive and non-preemptive FIFO spin locks - are studied by developing a new inflation-free schedulability test, jointly with matching bounds on worst-case synchronization delays. The synchronization approaches are compared in terms of schedulability in a large-scale empirical study considering both symmetric and asymmetric multiprocessors. While non-preemptive FIFO spin locks were found to generally perform best, lock-free synchronization was observed to offer significant advantages on asymmetric platforms. Alessandro Biondi 0001, Björn B. Brandenburg |
ECRTS | 1 |
| 2016 | OSEK-Like Kernel Support for Engine Control Applications under EDF SchedulingabstractEngine control applications typically include computational activities consisting of periodic tasks, activated by timers, and engine-triggered tasks, activated at specific angular positions of the crankshaft. Such tasks are typically managed by a OSEK-compliant real-time kernel using a fixed-priority scheduler, as specified in the AUTOSAR standard adopted by most automotive industries. Recent theoretical results, however, have highlighted significant limitations of fixed-priority scheduling in managing engine-triggered tasks that could be solved by a dynamic scheduling policy. To address this issue, this paper proposes a new kernel implementation within the ERIKA Enterprise operating system, providing EDF scheduling for both periodic and engine-triggered tasks. The proposed kernel has been conceived to have an API similar to the AUTOSAR/OSEK standard one, limiting the effort needed to use the new kernel with an existing legacy application. The proposed kernel implementation is discussed and evaluated in terms of run-time overhead and footprint. In addition, a simulation framework is presented, showing a powerful environment for studying the execution of tasks under the proposed kernel. Vincenzo Apuzzo, Alessandro Biondi 0001, Giorgio C. Buttazzo |
RTAS | 2 |
| 2016 | A Framework for Supporting Real-Time Applications on Dynamic Reconfigurable FPGAsabstractComputing platforms are evolving towards heterogeneous architectures including processors of different types and field programmable gate arrays (FPGAs), used as hardware accelerators for speeding up specific functions. The increasing capacity and performance of modern FPGAs, with their partial reconfiguration capabilities, have made them attractive in several application domains, including space applications.This paper proposes a framework for supporting the development of safety-critical real-time systems that exploit hardware accelerators developed through FPGAs with dynamic partial reconfiguration capabilities.A model is first presented and then used to derive a response-time analysis to verify the schedulability of a real-time task set under given constraints and assumptions. Although the analysis is based on a generic model, the proposed framework has been conceived to account for several real-world constraints present on today's platforms and has been practically validated on the Zynq platform, showing that it can actually be supported by state-of-the-art technologies. Finally, a number of experiments are reported to evaluate the worst-case performance of the proposed approach on synthetic workload. Alessandro Biondi 0001, Alessio Balsini, Marco Pagani, Enrico Rossi, Mauro Marinoni, Giorgio C. Buttazzo |
RTSS | 1 |
| 2016 | A Blocking Bound for Nested FIFO Spin LocksabstractBounding worst-case blocking delays due to lock contention is a fundamental problem in the analysis of multiprocessor real-time systems. However, virtually all fine-grained (i.e., non-asymptotic) analyses published to date make a simplifying (but impractical) assumption: critical sections must not be nested. This paper overcomes this fundamental limitation and presents the first fine-grained blocking bound for nested non-preemptive FIFO spin locks under partitioned fixed-priority scheduling. To this end, a new analysis method is introduced, based on a graph abstraction that reflects all possible resource conflicts and transitive delays. Alessandro Biondi 0001, Björn B. Brandenburg, Alexander Wieder |
RTSS | 1 |
| 2016 | Schedulability Analysis of Hierarchical Real-Time Systems under Shared ResourcesabstractSharing resources in hierarchical real-time systems implemented with reservation servers requires the adoption of special budget management protocols that preserve the bandwidth allocated to a specific component. In addition, blocking times must be accurately estimated to guarantee both the global feasibility of all the servers and the local schedulability of applications running on each component. This paper presents two new local schedulability tests to verify the schedulability of real-time applications running on reservation servers under fixed priority and EDF local schedulers. Reservation servers are implemented with the BROE algorithm. A simple extension to the SRP protocol is also proposed to reduce the blocking time of the server when accessing global resources shared among components. The performance of the new schedulability tests are compared with other solutions proposed in the literature, showing the effectiveness of the proposed improvements. Finally, an implementation of the main protocols on a lightweight RTOS is described, highlighting the main practical issues that have been encountered. Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
IEEE Trans. Computers | 1 |
| 2015 | Engine control: task modeling and analysis
Alessandro Biondi 0001, Giorgio C. Buttazzo |
DATE | 1 |
| 2015 | Supporting Component-Based Development in Partitioned Multiprocessor Real-Time SystemsabstractThe fast evolution of multicore systems, combined with the need of sharing the same platform for independently developed software, demands for new methodologies and algorithms that allow resource partitioning, while guaranteeing the isolation of concurrent applications. Unfortunately, a major problem that can break the isolation property of concurrent partitions is resource sharing. Although a number of resource access protocols exist for hierarchical uniprocessor systems, no protocols are available today for managing hierarchical partitions implemented on top a multiporcessor platform under partitioned scheduling. This paper presents a framework to support component based design on a multiprocessor platform and proposes a novel reservation server mechanism, called M-BROE, to handle shared resources in multiprocessor systems in the presence of resource reservation scheduling mechanisms. Alessandro Biondi 0001, Giorgio C. Buttazzo, Marko Bertogna |
ECRTS | 1 |
| 2015 | Feasibility Analysis of Engine Control Tasks under EDF SchedulingabstractEngine control applications include software tasks that are triggered at predetermined angular values of the crankshaft, thus generating a computational workload that varies with the engine speed. To avoid overloads at high rotation speeds, these tasks are implemented to self adapt and reduce their computational demand by switching mode at given rotation speeds. For this reason, they are referred to as adaptive variable rate (AVR) tasks. Although a few works have been proposed in the literature to model and analyze the schedulability of such a peculiar type of tasks, an exact analysis of engine control applications has been derived only for fixed priority systems, under a set of simplifying assumptions. The major problem of scheduling AVR tasks with fixed priorities, however, is that, due to engine accelerations, the interarrival period of an AVR task is subject to large variations, therefore there will be several speeds at which any fixed priority assignment is far from being optimal, significantly penalizing the schedulability of the system. This paper proposes for the first time an exact feasibility test under the Earliest Deadline First scheduling algorithm for tasks sets including regular periodic tasks and AVR tasks triggered by a common rotation source. In addition, a set of simulation results are reported to evaluate the schedulability gain achieved in this context by EDF over fixed priority scheduling. Alessandro Biondi 0001, Giorgio C. Buttazzo, Stefano Simoncelli |
ECRTS | 1 |
| 2015 | Dual-protocol support for Bluetooth LE devicesabstractLow energy consumption is one of the primary issues that have to be addressed in body area networks to prevent frequent battery recharges in the nodes. Such networks are being increasingly used to acquire sensory data that need to be processed in real-time. The Bluetooth Low Energy (BLE) protocol is an attractive solution for implementing personal area networks with reduced energy consumption, also because it is supported by consumer devices such as tablets and smart phones; however, it cannot guarantee a bounded delay for managing real-time traffic. This paper overcomes such a limitation by presenting a bandwidth sharing mechanism that allows partitioning the available network bandwidth between the BLE and another user-defined protocol built on top of the raw radio transceiver. Experimental results are also reported to characterize the timing behavior of the dual protocol on a specific platform. Mauro Marinoni, Gianluca Franchino, Daniel Cesarini, Alessandro Biondi 0001, Pasquale Buonocunto, Giorgio C. Buttazzo |
INDIN | 4 |
| 2014 | Optimal Design for Reservation Servers under Shared ResourcesabstractModularity and hierarchical-based design are crucial features that need to be supported in complex embedded systems characterized by multiple applications with timing requirements.Resource reservation is a powerful scheduling mechanism for achieving such goals and providing temporal isolation among different real-time applications. When different applications share mutually exclusive resources, a precise feasibility analysis can still be performed in isolation, using specific resource access protocols, taking into account only the application features and the reservation parameters. This paper presents a methodology for selecting the parameters of each reservation in order to guarantee the feasibility of the served applications and minimize the required bandwidth. Alessandro Biondi 0001, Alessandra Melani, Marko Bertogna, Giorgio C. Buttazzo |
ECRTS | 1 |
| 2014 | Exact Interference of Adaptive Variable-Rate Tasks under Fixed-Priority SchedulingabstractEngine control applications require the execution of tasks activated in relation to specific system variables, such as the crankshaft rotation angle. To prevent possible overload conditions at high rotation speeds, such tasks are designed to vary their functionality (hence their computational requirements) for different speed ranges. Modeling and analyzing such a type of tasks poses new research challenges in the schedulability analysis that are now being addressed in the real-time literature. This paper advances the state of the art by presenting a method for computing the exact worst-case interference of such adaptive variable-rate tasks under fixed priority scheduling, enabling a tight analysis and design of engine control applications. Alessandro Biondi 0001, Alessandra Melani, Mauro Marinoni, Marco Di Natale, Giorgio C. Buttazzo |
ECRTS | 1 |