EDBT 2026 Demo / reviewers in the wild / expert
Daniel Casini
dblp:201/8076
· DBLP profile ↗
40ranked-venue papers
18as first author
30since 2021 · last 2026
0000-0003-4719-3631ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 12 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRAM Bank-Aware Memory Allocation for Embedded Real-Time Virtualization
Eduardo Bischoff Grasel, Giovani Gracioli, Daniel Casini |
ISORC | 3 |
| 2026 | Shape-Aware Analysis of End-to-End Latency Under LET
Mario Günzel, Matthias Becker 0004, Daniel Casini |
RTAS | 3 |
| 2025 | Modeling the SL-LET Paradigm in AUTOSAR AdaptiveabstractThe AUTOSAR consortium proposed the AUTOSAR Adaptive standard to tackle the challenges introduced by the design of modern automotive systems. It consists of a service-oriented architecture (SoA) implemented in C++ and built on top of POSIX operating systems. However, unlike the previous AUTOSAR Classic specifications, this novel standard does not address non-functional requirements, including determinism, which is of key importance to guarantee the system's functional safety. This paper proposes a modeling extension to the AUTOSAR Adaptive standard aiming at guaranteeing a deterministic execution by leveraging the System-Level Logical Execution Time (SL-LET) paradigm, already used in the context of AUTOSAR Classic. A prototype implementation is also proposed, which is used to experimentally corroborate the feasibility of the proposed model extension with an evaluation based on a realistic automotive application built on the official AUTOSAR Adaptive Platform Demonstrator (APD). Davide Bellassai, Gerlando Sciangula, Claudio Scordino, Daniel Casini, Alessandro Biondi 0001 |
DATE | 4 |
| 2025 | Enabling Containerisation of Distributed Applications with Real-Time ConstraintsabstractContainerisation is becoming a cornerstone of modern distributed systems, thanks to their lightweight virtualisation, high portability, and seamless integration with orchestration tools such as Kubernetes. The usage of containers has also gained traction in real-time cyber-physical systems, such as software-defined vehicles, which are characterised by strict timing requirements to ensure safety and performance. Nevertheless, ensuring real-time execution of co-located containers is challenging because of mutual interference due to the sharing of the same processing hardware. Existing parallel computing frameworks such as Ray and its Kubernetes-enabled variant, KubeRay, excel in distributed computation but lack support for scheduling policies that allow guaranteeing real-time timing constraints and CPU resource isolation between containers, such as the SCHED_DEADLINE policy of Linux. To fill this gap, this paper extends Ray to support real-time containers that leverage SCHED_DEADLINE. To this end, we propose KubeDeadline, a novel, modular Kubernetes extension to support SCHED_DEADLINE. We evaluate our approach through extensive experiments, using synthetic workloads and a case study based on the MobileNet and EfficientNet deep neural networks. Our evaluation shows that KubeDeadline ensures deadline compliance in all synthetic workloads, adds minimal deployment overhead (in the order of milliseconds), and achieves lower worst-case response times, up to 4 times lower, than vanilla Kubernetes under background interference. Nasim Samimi, Luca Abeni, Daniel Casini, Mauro Marinoni, Twan Basten, Mitra Nasri, Marc Geilen, Alessandro Biondi 0001 |
ECRTS | 3 |
| 2025 | AP-LET: Enabling deterministic Pub/Sub communication in AUTOSAR AdaptiveabstractThe automotive software industry is facing a paradigm shift driven by the need to develop more and more advanced functionality distributed on multiple electronic control units. The AUTOSAR Adaptive standard has been designed as a service-oriented architecture on top of a general-purpose operating system to tackle this paradigm shift. Nevertheless, it does not provide means to ensure deterministic communication, as required in safety-related components. This paper studies the integration of the System-Level Logical Execution Time (SL-LET) paradigm in AUTOSAR Adaptive. The key design challenges and requirements to support SL-LET in AUTOSAR Adaptive are described, highlighting how to overcome the considerable differences between the AUTOSAR Classic and Adaptive domains. Then, a meta-protocol named AP-LET is presented, together with two concrete instances: one based on high-priority tasks and another leveraging timestamps in the message payload to handle communications and ensure determinism. A complete implementation of both protocols is also described. AP-LET was finally evaluated with a realistic automotive application, showing its feasibility and effectiveness. Davide Bellassai, Claudio Scordino, Daniel Casini, Alessandro Biondi 0001 |
J. Syst. Archit. | 3 |
| 2025 | Managing real-time constraints through monitoring and analysis-driven edge orchestrationabstractEmerging real-time applications are increasingly moving to distributed heterogeneous platforms , under the promise of more powerful and flexible resource capabilities. This shift inevitably brings new challenges. The design space to deploy chains of threads is more complex, and sound estimates of worst-case execution times are harder to obtain. Additionally, the environment is more dynamic, requiring additional runtime flexibility on the part of the application itself. In this paper, we present an optimization-based approach to this problem. First, we present a model and real-time analysis for modern distributed edge applications. Second, we propose a design-time optimization problem to show how to set the main parameters characterizing such applications from a time-predictability perspective. Then, we present an orchestration and runtime decision-making mechanism that monitors execution times and allows for runtime reconfigurations , spanning from graceful degradation policies to re-distributions of workload. A prototypical implementation of the proposed approach based on the QNX RTOS and its evaluation on a realistic case study based on an edge-based valet parking application conclude the paper. Daniel Casini, Paolo Pazzaglia, Matthias Becker 0004 |
J. Syst. Archit. | 1 |
| 2025 | To MILP or not to MILP? On AI techniques for the design and optimization of real-time systemsabstractAbstract Artificial intelligence (AI) is becoming increasingly relevant in many contexts. In embedded real-time systems, most of the previous research has focused on real-time guarantees for AI workloads (RT-for-AI). Instead, this position paper discusses the potential benefits and application cases of the complementary direction of using AI to optimize real-time systems themselves (AI-for-RT). It presents a vision where AI techniques, such as supervised and reinforcement learning, support system design and online configuration activities that are traditionally addressed using Mixed-Integer Linear Programming (MILP) or heuristic methods. The paper discusses scenarios where AI can potentially outperform classical techniques—such as recursive real-time analysis, systems with complex hardware/software interactions, and dynamic resource management—highlighting the promise of AI in both design-time and runtime real-time systems optimization. Solutions are left to future work: the goal is to populate the “Roadmap Towards Learning-Enabled and Learning-Assisted Real-Time Systems”, which is the target of this special issue. Daniel Casini |
Real Time Syst. | 1 |
| 2025 | Timerlat: Real-Time Linux Scheduling Latency Measurements, Tracing, and AnalysisabstractA trend in many embedded devices is the move from hardware-based to software-defined, such as software-defined networks and software-defined PLCs. This trend is motivated by multiple aspects, including the availability of complex software stacks and the consolidation of multiple devices into a single larger system. Due to its real-time capabilities and flexibility, Linux is the operating system of choice for many applications, including time-sensitive ones. However, assessing and debugging timing violations, especially those caused by scheduling latency, is challenging with the current state-of-the-art tools. This paper presentstimerlat, a tool that integrates scheduling latency measurements, tracing, and analysis in an easy-to-use interface. Its output includes an auto-analysis, providing insightful details on the composition of the scheduling latency. Experimental results are reported, evaluating the effectiveness of timerlat in assessing the latencies, considering different setups and workloads. Daniel Bristot de Oliveira, Daniel Casini, Juri Lelli, Tommaso Cucinotta |
IEEE Trans. Computers | 2 |
| 2024 | End-to-End Latency Optimization of Thread Chains Under the DDS Publish/Subscribe MiddlewareabstractModern autonomous systems integrate diverse soft-ware solutions to manage tightly communicating functionalities. These applications commonly communicate using frameworks implementing the publish/subscribe paradigm, such as the Data Distribution Service (DDS). However, these frameworks are real-ized with a multi-threaded software architecture and implement internal policies for message dispatching, posing additional chal-lenges for guaranteeing timing constraints. This work addresses the problem of optimizing a DDS-based interconnected real-time systems, proposing analysis-driven algorithms to set a vast range of parameters, ranging from classical thread priorities to other DDS-specific configurations. We evaluate our approaches on the Autoware Reference System, a realistic testbed from the Autoware autonomous driving framework. Gerlando Sciangula, Daniel Casini, Alessandro Biondi 0001, Claudio Scordino |
DATE | 2 |
| 2024 | In Search of Butterflies: Exceedance Analysis for Real-Time Systems under Transient OverloadabstractIn theory, real-time systems are provisioned based on provably sound worst-case execution times (WCETs), but in practice often only empirically derived, unsound execution-time estimates—i.e., nominal execution times (NETs)—are available since WCETs are difficult to obtain on modern hardware. NETs pose two significant challenges: First, since NETs may be exceeded at runtime, any response-time bounds derived from NETs are transitively unsound and may be violated. Second, even a minuscule NET violation can result in large, nonlinear response-time increases due to hard-to-predict, cascading scheduling effects. To explore the risk NET exceedance poses to a system’s temporal correctness, this paper provides the first general, systematic, and explainable methodology for exceedance analysis. The proposed approach supports fixed-priority (FP), earliest-deadline first (EDF), and first-in first-out (FIFO) scheduling on a uniprocessor or within a partitioned multiprocessor platform, and the full spectrum of preemption models from fully preemptive to fully non-preemptive workloads. Additionally, it produces explainable evidence in the form of tunable example traces that engineers can adjust to take system-specific expertise into account. The proposed methodology is evaluated with synthetic task sets and workloads based on an automotive benchmark, and in a case study applied to parts of the WATERS’17 industrial challenge. Matteo Zini, Filip Markovic 0001, Daniel Casini, Alessandro Biondi 0001, Björn B. Brandenburg |
RTSS | 3 |
| 2024 | The MATERIAL framework: Modeling and AuTomatic code Generation of Edge Real-TIme AppLications under the QNX RTOSabstractModern edge real-time automotive applications are becoming more complex, dynamic, and distributed, moving away from conventional static operating environments to support advanced driving assistance and autonomous driving functionalities. This shift necessitates formulating more complex task models to represent the evolving nature of these applications aptly. Modeling of real-time automotive systems is typically performed leveraging Architectural Languages (ALs) such as Amalthea, which are commonly used by the industry to describe the characteristics of processing platforms, operating systems, and tasks. However, these architectural languages are originally derived for classical automotive applications and need to evolve to meet the needs of next-generation applications. This paper proposes an automatic framework for the modeling and automatic code generation of dynamic automotive applications under the QNX RTOS. To this end, we extend Amalthea to describe chains of communicating tasks with multiple operating modes and to consider the QNX’s reservation-based scheduler, called APS, which allows providing temporal isolation between applications co-located on the same hardware platform. Finally, an evaluation is presented to compare different implementation alternatives under QNX that are automatically generated by our code generation framework. Matthias Becker 0004, Daniel Casini |
J. Syst. Archit. | 2 |
| 2024 | Introduction to the Special Issue on Real-Time Computing in the IoT-to-Edge-to-Cloud ContinuumabstractSpecial Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ... Daniel Casini, Dakshina Dasari, Matthias Becker 0004, Giorgio C. Buttazzo |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | Enhancing the Availability of Web Services in the IoT-to-Edge-to-Cloud Compute Continuum: A WordPress Case StudyabstractThe IoT-to-Edge-to-Cloud compute continuum presents vast opportunities for innovative applications, including crowdsensing, which leverages interconnected devices to gather real-time data. In domains like autonomous driving, crowdsensing enables traffic information sharing through web services. In this context, web services, like those based on Content Management Systems (CMS), are often used by drivers and passengers to share data about user experience, traffic congestion, and high-definition maps. However, ensuring high availability becomes crucial to maintain accessibility and reliability against usage peaks. This paper proposes a modern WordPress deployment approach that takes advantage of cloud-based to realize a cost-effective horizontal scalable architecture, leveraging Amazon AWS. The architecture suggested was implemented to test the effectiveness and released as a set of architecture-ready-to-use templates. Experimental results are provided to measure per-request response times under different autoscaling policies and bootstrap times. Gabriele Serra, Pietro Fara, Daniel Casini |
DSD | 3 |
| 2023 | Bounding the Data-Delivery Latency of DDS Messages in Real-Time Applications
Gerlando Sciangula, Daniel Casini, Alessandro Biondi 0001, Claudio Scordino, Marco Di Natale |
ECRTS | 2 |
| 2023 | On the QNX IPC: Assessing Predictability for Local and Distributed Real-Time SystemsabstractWith the advent of massively distributed applications such as those required by the IoT-to-Edge-to-Cloud compute continuum (i.e., automotive, smart agriculture, smart manufacturing, and more), real-time communication mechanisms allowing physically distributed nodes to seamlessly communicate as if they were running on the same host acquired noteworthy importance. To this end, the synchronous inter-process communication (IPC) mechanism provided by the QNX operating system (OS) is a promising candidate, as it allows using the application programming interface for communicating both on a single- and multi-node setting. Furthermore, it provides priority and partition inheritance mechanisms to improve predictability when working with the Adaptive Partitioning Scheduler (APS), a reservationbased scheduler provided by the QNX OS. This paper explores the behavior of the QNX synchronous message-passing (SyncMP) IPC with an extensive set of experiments, using them to formalize its behavior and model it from a real-time perspective. Then, it provides a response-time analysis for client-server applications based on the QNX SyncMP building upon self-suspending task theory. Finally, we evaluate the analysis on an application based on the WATERS 2019 Challenge by Bosch. Matthias Becker 0004, Dakshina Dasari, Daniel Casini |
RTAS | 3 |
| 2023 | Guest editorial: special issue on predictable machine learning
Daniel Casini, Giorgio C. Buttazzo |
Real Time Syst. | 1 |
| 2023 | Operating System Noise in the Linux KernelabstractAs modern network infrastructure moves from hardware-based to software-based using Network Function Virtualization, a new set of requirements is raised for operating system developers. By using the real-time kernel options and advanced CPU isolation features common to the HPC use-cases, Linux is becoming a central building block for this new architecture that aims to enable a new set of low latency networked services. Tuning Linux for these applications is not an easy task, as it requires a deep understanding of the Linux execution model and the mix of user-space tooling and tracing features. This paper discusses the internal aspects of Linux that influence the Operating System Noise from a timing perspective. It also presents Linux'sosnoisetracer, an in-kernel tracer that enables the measurement of the Operating System Noise as observed by a workload, and the tracing of the sources of the noise, in an integrated manner, facilitating the analysis and debugging of the system. Finally, this paper presents a series of experiments demonstrating both Linux's ability to deliver low OS noise (in the single-digit$\mu$s order), and the ability of the proposed tool to provide precise information about root-cause of timing-related OS noise problems. Daniel Bristot de Oliveira, Daniel Casini, Tommaso Cucinotta |
IEEE Trans. Computers | 2 |
| 2023 | Optimizing Inter-Core Communications Under the LET Paradigm using DMA EnginesabstractModern automotive applications are increasingly characterized by the need to transfer massive amounts of data in a predictable and deterministic way, possibly leveraging the Logical Execution Time (LET) paradigm. However, current proposals for LET communications are limited to core-commanded data transfers, which may result in large delays for data-intensive systems. To address this issue, we explore the use of Direct Memory Access (DMA) to handle LET communication with improved parallelism. Each DMA transfer operates on a contiguous memory area, thus calling for an optimized memory mapping to maximize performance. Modern DMA engines offer also advanced configurations, such as linked-lists of data transfers, which may provide more flexibility at the expenses of an increased (initial) programming overhead. Leveraging all such features of DMA engines, we propose a set of designs and protocols for LET communications with trade-offs between latency and space requirements. For each option we present the formulation to compute the optimal scheduling and memory allocation solution as a mixed-integer linear programming problem. Experimental results show the feasibility of the approach and a comparison of the solutions obtained using the proposed methods, showing a considerable improvement in terms of data acquisition latency when compared to LET communication without DMA. Paolo Pazzaglia, Daniel Casini, Alessandro Biondi 0001, Marco Di Natale |
IEEE Trans. Computers | 2 |
| 2023 | Analyzing Arm's MPAM From the Perspective of Time PredictabilityabstractWith heterogeneous multi-core platforms being crucial to execute the highly demanding workloads of modern applications, memory-access predictability remains a key issue for the system's safety. Many solutions have been proposed over the years, but none has been applied on a large scale. Nowadays, we are in front of an unprecedented opportunity to have an impact on commercial platforms: the Memory System Resource Partitioning and Monitoring (MPAM) specification by Arm, which describes different memory-access regulation mechanisms, presenting a valuable industrial attempt to address this issue. However, several points of the specification are described at a high level only, leaving plenty of room for interpretation to hardware manufacturers. This paper takes a close look at the memory-access regulation mechanisms in the MPAM specification and provides some detailed instantiations of such mechanisms. A fine-grained memory contention analysis is presented for each of them to finally enable a comparison of their worst-case performance. Matteo Zini, Daniel Casini, Alessandro Biondi 0001 |
IEEE Trans. Computers | 2 |
| 2022 | Placement of Chains of Real-Time Tasks on Heterogeneous Platforms under EDF SchedulingabstractWhen designing a real-time system, application architects are called to settle many non-trivial decisions that may severely influence the system's performance. With modern hardware platforms always being more and more complex and equipped with heterogeneous processor cores or even hardware accelerators such as TPUs, FPGAs, or GPUs, the complexities to be faced by application architects are exacerbated. Therefore, they are called to wisely allocate the computational resources provided by the hardware platform to application tasks in such a way to meet timing requirements and optimize other goals such as energy consumption. This paper proposes a mixed-integer linear programming formulation (MILP) to solve the task-to-heterogeneous-cores allocation problem while guaranteeing the schedulability of a real-time application running on the platform under partitioned Earliest Deadline First (EDF) scheduling. A new method to derive approximate worst-case response-time bounds is also presented and leveraged to setup the MILP formu-lation, which allows computing and minimizing the end-to-end latency of processing chains and considers energy requirements. The approach is evaluated on a task set based on the WATERS 2019 Industrial Challenge proposed by Bosch. Daniel Casini, Alessandro Biondi 0001 |
DSD | 1 |
| 2022 | End-to-End Analysis of Event Chains under the QNX Adaptive Partitioning SchedulerabstractModern autonomous cars run classic AUTOSAR applications alongside advanced driving assistance systems on a single-vehicle computer. Ensuring safety and predictability in such a complex system is challenging and requires temporal isolation between the various components. A promising solution is the POSIX-compliant QNX operating system: it meets the automotive standards for functional safety at the highest level (ISO 26262 ASIL-D) and provides temporal isolation through the Adaptive Partitioning Scheduler (APS), a resource reservation algorithm that guarantees processor bandwidth to groups of threads. These guarantees make it an ideal platform for composing diverse and complex applications on centralized vehicle computers. However, so far, there is no precise description or analysis of the APS reservation mechanism in real-time literature. In this paper, we provide the first description of the behavior of the APS from a real-time point of view and validate the results by running experiments on a real QNX platform. Based on the derived scheduler rules, we develop a response-time analysis to bound the end-to-end latency of event chains under APS. Finally, we evaluate different design strategies on a case study based on a real autonomous construction vehicle. Dakshina Dasari, Matthias Becker 0004, Daniel Casini, Tobias Stark |
RTAS | 3 |
| 2022 | A Theoretical Approach to Determine the Optimal Size of a Thread Pool for Real-Time SystemsabstractParallel workloads most commonly execute onto pools of thread, allowing to dispatch and run individual nodes (e.g., implemented as C++ functions) at the user-space level. This is relevant in industrial cyber-physical systems, cloud, and edge computing, especially in systems leveraging deep neural networks (e.g., TensorFlow), where the computations are inherently parallel. When using thread pools, it is common to implement fork-join parallelism using blocking synchronization mechanisms provided by the operating system (such as condition variables), with the side effect of temporarily reducing the number of worker threads. Consequently, the served tasks may suffer from additional delays, thus potentially harming timing guarantees if such effects are not properly considered. Prior works studied such phenomena, providing methods to guarantee the timing behavior. However, the challenges introduced by thread pools with blocking synchronization cause current analyses to incur a notable pessimism. This paper tackles the problem from a different angle, proposing solutions to determine the optimal size of a thread pool in such a way as to avoid the undesired effects that arise from blocking synchronization. Daniel Casini |
RTSS | 1 |
| 2022 | Optimized partitioning and priority assignment of real-time applications on heterogeneous platforms with hardware acceleration
Daniel Casini, Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale |
J. Syst. Archit. | 1 |
| 2022 | Profiling and controlling I/O-related memory contention in COTS heterogeneous platformsabstractAbstract Motivated by the increasing number of embedded applications that make use of traffic‐intensive I/O devices, this work studies the memory contention generated by I/O devices and investigates on the regulation of the bus traffic they generate by means of COTS regulators, namely the QoS‐400 by Arm. To this purpose, the behavior of the QoS‐400 regulators is analytically characterized and then, taking the Xilinx Ultrascale+ as a reference modern heterogeneous platform, a software infrastructure to control such regulators from Linux is proposed. As an experience report, this article presents the results of an extensive experimental evaluation, based on both benchmarks and microbenchmarks, aimed at validating the effectiveness of QoS‐400 regulators in predictably controlling I/O‐related memory traffic, as well as assessing the impact of the regulation on software applications and I/O devices themselves. Matteo Zini, Giorgiomaria Cicero, Daniel Casini, Alessandro Biondi 0001 |
Softw. Pract. Exp. | 3 |
| 2022 | An I/O Virtualization Framework With I/O-Related Memory Contention Control for Real-Time SystemsabstractModern applications are often characterized by a tight interaction with I/O devices. At the same time, many application domains are also facing a shift toward an integrated approach where multiple applications with mixed levels of safety and security need to co-exist on top of a shared hardware platform, which is typically managed by a hypervisor. This gives rise to the need for a predictable mechanism allowing multiple virtual machines to share I/O devices, while at the same time controlling contention delays when they access global memory. To deal with these shortcomings, this article proposes an I/O virtualization framework providing support for controlling the I/O-related memory contention by leveraging the ARM QoS-400 regulators. Extensive experiments are performed to compare the proposed solution with the Xen hypervisor, showing improvements up to$8\times $when controlling the I/O-related memory contention. Niccolò Borgioli, Matteo Zini, Daniel Casini, Giorgiomaria Cicero, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | Optimal Memory Allocation and Scheduling for DMA Data Transfers under the LET ParadigmabstractThe Logical Execution Time (LET) paradigm is increasingly used to achieve predictable communications in modern multicore automotive applications. Direct Memory Access (DMA) engines can perform the data copies that are needed in a LET implementation on behalf of the cores with improved parallelism and reduced overheads. However, each DMA transfer operates on contiguous memory areas, and the performance is strongly dependent on the allocation in memory of the variables to be copied. This paper proposes a protocol to perform LET communications with a DMA and presents an optimal memory allocation scheme and scheduling using a mixed-integer linear programming formulation. Experimental results are reported to compare the performance of different communication approaches. Paolo Pazzaglia, Daniel Casini, Alessandro Biondi 0001, Marco Di Natale |
DAC | 2 |
| 2021 | Latency Analysis of I/O Virtualization Techniques in Hypervisor-Based Real-Time SystemsabstractNowadays, hypervisors are the standard solution to integrate different domains into a shared hardware platform, while providing safety, security, and predictability. To this end, a hypervisor virtualizes the physical platform and orchestrates the access to each component. When the system needs to comply with certification requirements for safety-critical systems, virtualization latencies need to be analytically bounded for providing off-line guarantees. This paper presents a detailed modeling of three I/O virtualization techniques, providing analytical bounds for each of them under different metrics. Experimental results compare the bounds for a case study and quantify the contribution due to different sources of delay. Daniel Casini, Alessandro Biondi 0001, Giorgiomaria Cicero, Giorgio C. Buttazzo |
RTAS | 1 |
| 2021 | A Multi-Domain Software Architecture for Safe and Secure Autonomous DrivingabstractThis work aims at making Apollo, a popular autonomous driving framework, safer and more secure by designing a multi-domain architecture, where its components are split between a feature-rich domain running Linux and a critical domain running a real-time operating system (RTOS). The two domains are isolated by a hypervisor. We implemented a prototype where the control component has been ported from Linux to the Erika automotive-grade RTOS, and we discuss a number of challenges that have been faced in moving the component to Erika. The proposed solution has been experimentally evaluated by measuring the latencies involving processing paths passing through the control component. Luca Belluardo, Andrea Stevanato, Daniel Casini, Giorgiomaria Cicero, Alessandro Biondi 0001, Giorgio C. Buttazzo |
RTCSA | 3 |
| 2021 | A ROS 2 Response-Time Analysis Exploiting Starvation Freedom and Execution-Time VarianceabstractRobots are commonly subject to real-time constraints. To ensure that such constraints are met, recent work has analyzed the response times of processing chains under ROS 2, a popular robotics framework. However, prior work supports only scalar worst-case execution time bounds and does not exploit that the ROS 2 scheduling mechanism is starvation-free.This paper proposes a novel response-time analysis for ROS 2 processing chains that accounts for both the high execution-time variance typically encountered in robotics workloads and the starvation freedom of the default ROS 2 callback scheduler. Experimental results from both synthetic callback graphs and a real ROS 2 workload empirically show the proposed analysis to be much more accurate (often by a factor of 2× or more). Tobias Stark, Daniel Casini, Sergey Bozhko, Björn B. Brandenburg |
RTSS | 2 |
| 2021 | Task Splitting and Load Balancing of Dynamic Real-Time Workloads for Semi-Partitioned EDFabstractMany real-time software systems, such as those commonly found in the context of multimedia, cloud computing, robotics, and real-time databases, are characterized by a dynamic workload, where applications can join and leave the system at runtime. Global schedulers can transparently support dynamic workload without requiring any off-line task-allocation phase, thus providing advantages to the system designer. Nevertheless, such schedulers exhibit poor worst-case performance when compared to semi-partitioned schedulers, which instead can achieve near-optimal schedulability performance when used in conjunction with smart task splitting and partitioning techniques, and they are also lighter in terms of run-time overhead. This article proposes an approach to efficiently schedule dynamic real-time workloads on multiprocessor systems by means of semi-partitioned scheduling. A linear-time approximation scheme for the C=D splitting algorithm under partitioned EDF scheduling is proposed. Then, a load-balancing algorithm is presented to admit new real-time workloads with a limited number of re-allocations. The article finally reports on a large-scale experimental study showing that (i) the linear-time approximation is characterized by a very limited utilization loss compared with the corresponding exact approach (that has a much higher complexity), and that (ii) the whole approach allows achieving considerable improvements with respect to global and partitioned EDF scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Computers | 1 |
| 2020 | Predictable Memory-CPU Co-Scheduling with Support for Latency-Sensitive TasksabstractPredictable execution models have been proposed over the years to achieve contention-free execution of real-time tasks by preloading data into dedicated local memories. In this way, memory access delays can be hidden by delegating a DMA engine to perform memory transfers in parallel with processor execution. Nevertheless, state-of-the-art protocols introduce additional blocking due to priority inversion, which may severely penalize latency-sensitive applications and even worsen the system schedulability with respect to the use of classical scheduling schemes. This paper proposes a new protocol that allows hiding memory transfer delays while reducing priority inversion, thus favoring the schedulability of latency-sensitive tasks. The corresponding analysis is formulated as an optimization problem. Experimental results show the advantages of the proposed protocol against state-of-the-art solutions. Daniel Casini, Paolo Pazzaglia, Alessandro Biondi 0001, Marco Di Natale, Giorgio C. Buttazzo |
DAC | 1 |
| 2020 | Demystifying the Real-Time Linux Scheduling LatencyabstractLinux has become a viable operating system for many real-time workloads. However, the black-box approach adopted by cyclictest, the tool used to evaluate the main real-time metric of the kernel, the scheduling latency, along with the absence of a theoretically-sound description of the in-kernel behavior, sheds some doubts about Linux meriting the real-time adjective. Aiming at clarifying the PREEMPT_RT Linux scheduling latency, this paper leverages the Thread Synchronization Model of Linux to derive a set of properties and rules defining the Linux kernel behavior from a scheduling perspective. These rules are then leveraged to derive a sound bound to the scheduling latency, considering all the sources of delays occurring in all possible sequences of synchronization events in the kernel. This paper also presents a tracing method, efficient in time and memory overheads, to observe the kernel events needed to define the variables used in the analysis. This results in an easy-to-use tool for deriving reliable scheduling latency bounds that can be used in practice. Finally, an experimental analysis compares the cyclictest and the proposed tool, showing that the proposed method can find sound bounds faster with acceptable overheads. Daniel Bristot de Oliveira, Daniel Casini, Rômulo Silva de Oliveira, Tommaso Cucinotta |
ECRTS | 2 |
| 2020 | A Holistic Memory Contention Analysis for Parallel Real-Time Tasks under Partitioned SchedulingabstractWhen adopting multi-core systems for safety-critical applications, certification requirements mandate bounding the delays incurred in accessing shared resources. This is the case of global memories, whose access is often regulated by memory controllers optimized for average-case performance and not designed to be predictable. As a consequence, worst-case bounds on memory access delays often result to be too pessimistic, drastically reducing the advantage of having multiple cores. This paper proposes a fine-grained analysis of the memory contention experienced by parallel tasks running on a multi-core platform. To this end, an optimization problem is formulated to bound the memory interference by leveraging a three-phase execution model and holistically considering multiple memory transactions issued during each phase. Experimental results show the advantage in adopting the proposed approach on both synthetic task sets and benchmarks. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTAS | 1 |
| 2020 | Timing isolation and improved scheduling of deep neural networks for real-time systemsabstractSummary In recent years, the performance of deep neural networks (DNNs) is significantly improved, making them suitable for many application fields, such as autonomous driving, advanced robotics, and industrial control. Despite a lot of research being devoted to improving the accuracy of DNNs, only limited efforts have been spent to enhance their timing predictability, required in several real‐time applications. This paper proposes a software infrastructure based on the Linux operating system to integrate DNNs within a real‐time multicore system. It has been realized by modifying both the internal scheduler of the popular TensorFlow framework and the SCHED_DEADLINE scheduling class of Linux. The proposed infrastructure allows providing timing isolation of DNN inference tasks, hence improving the determinism of the temporal interference generated by TensorFlow. The proposal is finally evaluated with a case study derived from a state‐of‐the‐art benchmark inspired by an autonomous industrial system. Extensive experiments demonstrate the effectiveness of the proposed solution and show a significant reduction of both average and longest‐observed response times of TensorFlow tasks. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
Softw. Pract. Exp. | 1 |
| 2019 | Analyzing Parallel Real-Time Tasks Implemented with Thread PoolsabstractDespite several works in the literature targeted predictable execution models for parallel tasks, limited attention has been devoted to study how specific implementation techniques may affect their execution. This paper highlights some issues that can arise when executing parallel tasks with thread pools, which may lead to deadlocks and performance degradation when adopting blocking synchronization mechanisms. A new parallel task model, inspired to a realistic design found in popular software systems, is first presented to study this problem. Then, formal conditions to ensure the absence of deadlocks and schedulability analysis techniques are proposed under both global and partitioned scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
DAC | 1 |
| 2019 | Response-Time Analysis of ROS 2 Processing Chains Under Reservation-Based SchedulingabstractBounding the end-to-end latency of processing chains in distributed real-time systems is a well-studied problem, relevant in multiple industrial fields, such as automotive systems and robotics. Nonetheless, to date, only little attention has been given to the study of the impact that specific frameworks and implementation choices have on real-time performance. This paper proposes a scheduling model and a response-time analysis for ROS 2 (specifically, version "Crystal Clemmys" released in December 2018), a popular framework for the rapid prototyping, development, and deployment of robotics applications with thousands of professional users around the world. The purpose of this paper is threefold. Firstly, it is aimed at providing to robotic engineers a practical analysis to bound the worst-case response times of their applications. Secondly, it shines a light on current ROS 2 implementation choices from a real-time perspective. Finally, it presents a realistic real-time scheduling model, which provides an opportunity for future impact on the robotics industry. Daniel Casini, Tobias Stark, Ingo Lütkebohle, Björn B. Brandenburg |
ECRTS | 1 |
| 2019 | Handling Transients of Dynamic Real-Time Workload Under EDF SchedulingabstractReal-time dynamic workload consists of tasks that can arbitrarily join and leave the system at run-time. To avoid incurring deadline misses, tasks that request to join the system must pass an admission test, which has to cope with potential scheduling transients originated by the residual effect of the tasks that previously left the system. This phenomenon may require some tasks to suffer an admission delay before being accepted for execution. This paper focuses on uniprocessor earliest-deadline first (EDF) scheduling with constrained deadlines and explicitly considers methods for handling scheduling transients in the presence of dynamic real-time workload. A generalized analysis framework is first presented to overcome several limitations of the existing approaches (including the support for overlapping transients), and is then used to derive methods for computing bounds on the admission delays incurred by tasks. Building on such results, an on-line protocol is proposed to handle the admission control of a dynamic workload, which also comes with a variant that can execute in polynomial time to favor its practical application. Furthermore, the paper shows how the presented analysis can be used off-line for analyzing mode-changes among static task sets. Experimental results are finally presented to evaluate the proposed algorithms. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
IEEE Trans. Computers | 1 |
| 2018 | Memory Feasibility Analysis of Parallel Tasks Running on Scratchpad-Based ArchitecturesabstractThis work proposes solutions for bounding the worst-case memory space requirement for parallel tasks running on multicore platforms with scratchpad memories. It introduces a feasibility test that verifies whether memories are large enough to contain the maximum memory backlog that may be generated by the system. Both closed-form bounds and more accurate algorithmic techniques are proposed. It is shown how one can use max-plus algebra and solutions to the max-flow cut problem to efficiently solve the memory feasibility problem. Experimental results are presented to evaluate the efficiency of the proposed feasibility analysis techniques on synthetic workload and state-of-the-art benchmarks. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTSS | 1 |
| 2018 | Partitioned Fixed-Priority Scheduling of Parallel Tasks Without PreemptionsabstractThe study of parallel task models executed with predictable scheduling approaches is a fundamental problem for real-time multiprocessor systems. Nevertheless, to date, limited efforts have been spent in analyzing the combination of partitioned scheduling and non-preemptive execution, which is arguably one of the most predictable schemes that can be envisaged to handle parallel tasks. This paper fills this gap by proposing an analysis for sporadic DAG tasks under partitioned fixed-priority scheduling where the computations corresponding to the nodes of the DAG are non-preemptively executed. The analysis has been achieved by means of segmented self-suspending tasks with nonpreemptable segments, for which a new fine-grained analysis is also proposed. The latter is shown to analytically dominate state-of-the-art approaches. A partitioning algorithm for DAG tasks is finally proposed. By means of experimental results, the proposed analysis has been compared against a previouslyproposed analysis for DAG tasks with non-preemptable nodes managed by global fixed-priority scheduling. The comparison revealed important improvements in terms of schedulability performance. Daniel Casini, Alessandro Biondi 0001, Geoffrey Nelissen, Giorgio C. Buttazzo |
RTSS | 1 |
| 2017 | Semi-Partitioned Scheduling of Dynamic Real-Time Workload: A Practical Approach Based on Analysis-Driven Load BalancingabstractRecent work showed that semi-partitioned scheduling can achieve near-optimal schedulability performance, is simpler to implement compared to global scheduling, and less heavier in terms of runtime overhead, thus resulting in an excellent choice for implementing real-world systems. However, semi-partitioned scheduling typically leverages an off-line design to allocate tasks across the available processors, which requires a-priori knowledge of the workload. Conversely, several simple global schedulers, as global earliest-deadline first (G-EDF), can transparently support dynamic workload without requiring a task-allocation phase. Nonetheless, such schedulers exhibit poor worst-case performance. This work proposes a semi-partitioned approach to efficiently schedule dynamic real-time workload on a multiprocessor system. A linear-time approximation for the C=D splitting scheme under partitioned EDF scheduling is first presented to reduce the complexity of online scheduling decisions. Then, a load-balancing algorithm is proposed for admitting new real-time workload in the system with limited workload re-allocation. A large-scale experimental study shows that the linear-time approximation has a very limited utilization loss compared to the exact technique and the proposed approach achieves very high schedulability performance, with a consistent improvement on G-EDF and pure partitioned EDF scheduling. Daniel Casini, Alessandro Biondi 0001, Giorgio C. Buttazzo |
ECRTS | 1 |